Digital Digging with Henk van Ess

Digital Digging with Henk van Ess

How to fix public data

Inspired by Netflix's Freefall — the full manual on how I built aircraftdefects.com with GLM-5.3 Flash

Henk van Ess's avatar
Henk van Ess
Aug 28, 2026
∙ Paid

In the Netflix documentary Freefall I watched the father of a crash victim try to read the database of aircraft defect reports kept by the FAA, the US aviation regulator. It took him hours. With my new tool it takes seconds.

I entered the tool in Cerebral Valley‘s GLM-5.3 Flash hackathon. The Foundation for Aviation Safety wrote to me. Ed, from the Foundation, after a first look: “It looks impressive. I’m hoping to dig into it further.”

They asked for a walkthrough. That is the test of a tool like this: whether the people who work on the problem every day want to use it.

It is not only for them. Anyone about to board can type a flight number and a date and read what mechanics wrote about that aircraft, and how often it came back for the same thing. A journalist gets leads: the same part written up on several aircraft of one airline on one day, or a box ticked that contradicts the write-up under it. A researcher gets 1.76 million reports readable at once, and can ask them questions in plain words. This 28 page guide is about how you can fix public data for about $1 in just 2 days time.

Relatives of victims were attempting to find answers to their questions partly via a governmental website. Row after row of capital letters. Zone numbers. Single character codes. It’s called Service Difficulty Report and is maintained by the Federal Aviation Administration. When technicians in the United States find an issue with a plane (e.g., a crack in a component; rust under a piece of equipment; or a failed seal), they file it. These reports are then made available to everyone. You don’t need to sign in, pay a fee, or submit a request for records. There are 1,757,828 of these reports since 1995 through today. They relate to 54,634 separate planes.

I was shocked.

The reports are so bureaucratic that almost everything has been replaced by codes, and the codes conflict. A report will never tell you there was an issue with the landing gear. It will say ZONE 700. An emergency landing? That is just an A. But the same letter elsewhere on the form means the report came from an airline. For researchers, the database is a nightmare. I saw the father of one of the victims searching it, trying to make sense of it.

It was something public, but the government presented it as random unlabeled ingredients on a table. Over 1.7 million records were hidden behind bureaucratic codes. I want to make this completely transparent, to help families and researchers.

The FAA documents never tell you that “the crew made an emergency landing.” Instead, it will simply say “A.”

An airline’s name is made up of letters such as CALA. The technician’s finding is a completely different code. The code used to identify how the technician found the issue initially is a third code.

When I attempted to conduct a search using a legitimate query, I encountered timeout and latency issues. Typical IT issues with software that has gone unmaintained for a considerable period of time. Creating links between two pieces of data were difficult.

But I have questions I want answered. What types of defects were reported yesterday? Which companies reported the most? At what time did they report them? Who were the individuals involved? What location on the airplane was affected? What type of defect was reported?

What action did the crew take in response to the reported defect? Without knowing the “when,” “what,” “where” and “who,” I cannot begin to determine “why.” Why does a single operator send the same airplane out for repairs 60 times in a year? How can that occur?

In the past (read 2025), I would download all the data into a spreadsheet, create a pivot table, rotated it one direction, rotated it another direction, and hoped the spreadsheet did not crash due to its size. In 2026, I type what I desire in my own language and AI creates it for me if I am precise enough. I build about two original, highly hyper-focused research tools a week for myself and my clients.

Imagination is the limit to creating things, not your ability to remember formulas or knowing how to code. If you can express what you desire, you can create it. How does that work? How did I build a full fledged tool?

A black box with white letters

A majority of users who use AI will probably be doing so by way of a chat box in their browser. If they paste a document into the chat box, the AI will read it. If they paste forty documents into the chat box, some will drop off. If they ask the AI to create a complex tool, they will receive code that they must save, install and run themselves.

But if the same AI is running within a terminal (Windows, Mac, no download needed) it can access your file system, download anything that is required for execution, execute it, observe any errors and correct those errors. The latter portion is the key to this process. After completing its task, the AI observes the outcome of its work and attempts to complete the same task again. You act as an intermediary to convey information: do this, do that. If the terminal and AI is installed on a server, you can post the output directly onto the web. So out of a window that looks like it came from 1989 (it does) you get a working tool on your own machine, or a website anyone can visit. The limit is no longer what you can build. Can you describe what you want?

Let’s start building the tool in the Terminal app. For a researcher who cannot code, the barrier was never imagination. It was having to learn Python, to structure a project, to build in fail-safes, to know how not to leak your own material. Some of that still matters. The part where you personally type the lines of code does not.

I used the brand new version of the model, GLM-5.3 Flash. There was already a first version, built with Claude Code, but it only tested one thing: could the database be made accessible at all? The answer was yes.

It answers back in the same green characters. Yes, that is possible. Should I show you how?

It warned me about two of my questions. The who means airlines, and the airlines are hiding behind odd-looking codes across 26 years of history, some of them obsolete and some not. The what means every engineer writing up a repair in a private thicket of abbreviations, which will all have to be looked up. The where is the awkward one, because a report will often describe a repair without ever saying where on the aircraft it happened. The location is implicit. The mechanic knew. The AI found a list of FAA codes, I checked it, and granted it to rebuild the public database in human language.

During a hackathon of GLM-5.3 Flash , I decided to go way further. I found the information still overwhelming. How do I change that?

The first picture that arrived in my head was a schematic aircraft, seen from the side, sitting at the top. Can AI build that for me? And how do I explain what I want to see? I asked:

User's avatar

Continue reading this post for free, courtesy of Henk van Ess.

Or purchase a paid subscription.
© 2026 𝚑𝚎𝚗𝚔 𝚟𝚊𝚗 𝚎𝚜𝚜 · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture