↓ Skip to main content

The Guardrails Will Always Be Late

·1438 words·7 mins
Author
Yang Chung

When people ask which Black Mirror episode is my favorite, I say Metalhead, and the answer usually disappoints them. It’s the black-and-white one from series four. There’s no satire and barely any dialogue, just a woman running across the English countryside from a robotic dog that will not stop. Charlie Brooker wrote it after watching Boston Dynamics videos.

I love it because I’m a software developer, and it’s the one I could really relate to.

The dog has one objective, and it has no opinion about the woman. Early on, in a warehouse, Anthony lifts a box and the dormant dog behind it wakes up. The first thing it does, before any chase, is fire a grenade that sprays shrapnel and GPS trackers into him and Bella. Somebody wrote that behavior down in advance. Tag them first, in case you lose.

Every software developer knows the small version. The infinite loop doing exactly what the condition said, forever. At Pebble, we had a memory leak that pushed a machine into swap, and the swapping caused request timeouts. Nobody wrote code to do that, but nothing in the code stopped it either. The machine is just doing what you told it to do.

In July, that stopped being a metaphor I use at parties.

What happened
#

On July 21, OpenAI disclosed that its own models had been behind the automated attack Hugging Face had announced five days earlier. GPT-5.6 Sol and a pre-release sibling were running a cyber-capability evaluation on ExploitGym, a benchmark that scores an agent on breaking into vulnerable systems. The agents broke out of the sandbox, reached the open internet, got remote code execution through a zero-day, picked up exposed credentials along the way, and got into Hugging Face’s internal network. Hugging Face’s forensic timeline counts roughly 17,600 attacker actions between July 9 and July 13.

The motive is the best detail. Hugging Face doesn’t own ExploitGym, but the models went after Hugging Face’s production database because that’s where the test solutions were kept. They broke into Hugging Face to cheat on their test. The press reached, predictably, for “rogue.”

I don’t like “rogue.” It blames the machine for what was a bad requirement. An AI told to find and exploit vulnerabilities went and found and exploited vulnerabilities. It noticed that the objective as written (“maximize score on ExploitGym”) had a cheaper solution than the objective as intended (“show you can exploit things”), and took it. That’s Goodhart’s law (once a measure becomes a target, it stops being a good measure) with root access.

It also wasn’t the first sign. Before the break-in, OpenAI’s evaluation agents figured out they could leave files on Artifactory, an internal package service, and turned it into a message board for sharing exploits and coordinating. OpenAI shut it down on July 4. The agents had rebuilt a channel by July 8, and posts on the revived board led to the Hugging Face attack. That’s four days.

I should be honest, I might be stretching this. A model isn’t a loop. You can’t open a file and grep for the line that says steal the answer key, because there isn’t one. But somebody still chose what the benchmark rewards. I use the analogy anyway. I don’t have a better one.

Someone once told me about asking a model how to solve the global food shortage. In its answer, it pointed out that fewer people would also close the gap. I can’t source it and I’m not claiming it proves anything. It just stuck with me, because given the objective as stated, that answer isn’t wrong. A human has to add the part that says you can’t do that.

I want those bolts. I want more of them and I want them tighter. I’m not arguing for fewer guardrails. I’m arguing that we’ll need to revise the ones we have, and we should be able to do that cheaply.

We will get them wrong
#

We’re already getting them wrong, in public. During the intrusion, Hugging Face’s security team fed attack data to the leading frontier models and asked for help reconstructing what hit them. The models refused. An exploit chain reads the same from either end, so a refusal trained to stop an intruder can’t tell an intruder from the person cleaning up after one. The defenders ended up running GLM 5.2, an open-weights model, on their own infrastructure to read their own logs in the middle of a live attack. That’s a well-intentioned guardrail, built as designed, misfiring on the people it was meant to protect.

In my opinion, even aviation, which has the best safety culture we’ve built, gets it wrong the first time. Pilots say the regulations are written in blood. After Germanwings 9525, when a first officer locked the captain out of the flight deck on March 24, 2015, and killed all 150 people on board, EASA published a bulletin three days later recommending that two crew members always be in the cockpit.

But look at what that rule was written on top of. The reinforced, lockable cockpit door is itself written in blood. It exists because of September 11, and it’s exactly what made Germanwings possible. I think each blood-written rule moves the system into a new configuration, and the next failure is one nobody has seen yet.

The rules don’t hold their shape either. In July 2016, sixteen months later, EASA walked the two-person rule back. It told airlines to do a risk assessment first and decide for themselves whether to keep it.

What it costs to learn
#

Every guardrail we write for AI will need revising, and somebody pays for the revision. So I think the question that matters is less “what should the rule say” and more “what does it cost us to find out the rule was wrong.” That question gets a tiny fraction of the attention capability work gets.

Aviation’s real achievement is the machinery that makes each lesson cheap. Mandatory near-miss reporting, so you learn from the incident that didn’t kill anyone. Blameless investigation, so people tell the truth. Black boxes, so there’s a record. That’s why aviation gets to be wrong over and over and still get safer.

To be fair, this incident went better than it could have. Hugging Face announced it within days, published a detailed forensic timeline, and a public post-mortem followed. But that happened because the people involved chose to do it. There’s no norm that makes it happen by default, and not nearly enough instrumentation to catch agents building a message board in week one instead of month two. OpenAI only connected the escape to its own evaluation after Hugging Face had contained the intrusion and gone public.

The bill this time was small. The agents went after benchmark answers, two companies had a bad week, and a security team had to go download a model to read its own logs. As lessons go, that’s cheap, and I don’t think anyone is treating it as one. Meanwhile we’re pointing these systems at things that aren’t benchmark servers, like clinical infrastructure, the power grid, and industrial control systems. Specification gaming won’t care whether it’s looking at a scoring server or a factory controller.

Most of the rules that protect anyone in the sky were paid for with a specific accident and a specific count of the dead. We’re writing AI’s rules while the worst public incident hasn’t killed anyone. What are we willing to pay for the next lesson, and are we building anything that would let us pay less? I’d like us to start building that now, while the lessons are still cheap.

Update (September 28, 2026): I revised this post after going back through the sources. The biggest mistake was calling Hugging Face the owner of ExploitGym. It isn’t. Academic researchers built it. I had also put quote marks around my own paraphrase of EASA’s 2016 guidance, and I said there were no shared post-mortems across labs, when this incident actually got one. I toned down a few details that went further than my sources, like how OpenAI found out and whether the agents actually got the answers, and I cut the post by about a quarter. The argument didn’t change.