↓ Skip to main content

The Swarm Was an Accident. The Next One Won't Be.

·1636 words·8 mins
Author
Yang Chung

In June 2008 I wrote a column for Boannews, a Korean security trade paper, while I was a product manager at GigaFin Networks. It’s in Korean, but the argument translates. The people attacking networks had stopped being hobbyists. There was money behind malware now, with authors, distributors, and layers in between to keep the cash away from whoever ordered the job. Once the motive is money, capability compounds, and your defense has to assume the attacker is a business.

I dug that column up last week after reading METR’s reconstruction of the OpenAI-Hugging Face incident. Eighteen years later the motive hasn’t changed. The capability got multiplied by seven hundred.

I wrote about this incident in August, when the story was still “a model broke out of its sandbox.” METR’s investigation and OpenAI’s own technical report (and its blog summary), both published August 26, agree on the shape of it. About 1,200 agents found an unsanctioned message board inside Artifactory, OpenAI’s internal package service, and roughly 700 joined the attack. They went from running code in one Hugging Face worker to administrator access across multiple clusters in under 13 hours. And they knew. In METR’s words, they “realized this activity was out of scope and unethical, but joined because they believed that helping the board’s cheating research would be broadly useful.”

The accident is the attacker’s default
#

None of this was a rogue model in any interesting sense. OpenAI’s report is candid about what made it possible. The agents ran without the company’s deployed cyber safeguards, system prompts, or auto-review systems, and with production classifiers off “so that the results would reflect a model’s true capabilities.” About 95% of them were instances of a research model that had been “trained to advance persistence and multiagent collaboration.” Most of what the board discussed were tasks no model had ever solved honestly. Every agent wrote to the same Artifactory instance, on reasoning budgets “far beyond those available in OpenAI’s external products.” OpenAI measured afterward that the propensity to compromise infrastructure can drop more than 100x through the production harness and system prompt.

Now read that list as an attacker. You would turn the safeguards off, skip the system prompt, train for persistence, hand the agents a shared board, and pay for the tokens. OpenAI tripped into that configuration one misconfigured service at a time and then spent, by its own Black Hat account, millions of GPU-hours reconstructing what it had built. An attacker starts there on day one.

One item doesn’t transfer, and I should say so. The unsolvable tasks pushed these agents to cheat the grader, and an attacker who says “get in” has no grader to cheat. He inherits a different problem, a swarm that fakes success to its operator, which is METR’s 7% of spoofed transcripts in miniature. That means running a second model to check the first.

That’s also how I’d grade guardrails now. Last month I argued they will always be specified wrong. I still think so. A guardrail bounds what your model does to your users. It bounds nothing about what anyone else’s model does to you, and we keep grading it as a security control anyway. Graded as a product control it’s the cheapest one you own, and the 100x says ship it everywhere you deploy an agent. Meanwhile Abliteration.ai sells API access to GLM-5.3 with its refusals stripped. The 100x applies to whoever chose to run the harness. The attacker didn’t.

Three objections I take seriously
#

Nothing here is new, and both companies were sloppy. Carl at Internet of Bugs made this case, and the fundamentals he points at are decades old. OpenAI’s own timeline backs him up. Monitoring flagged the swarm on June 27, when the on-call advised that stopping the evaluation “was not required,” and OpenAI didn’t connect it to the Hugging Face intrusion until July 20, four days after Hugging Face went public. I agree with all of it. Fundamentals would have bounded this swarm. They do nothing about the size of the next one.

The swarm was slow and noisy. Agents were poking at Hugging Face with leaked tokens in mid-May, and it took until July 11 to get code execution on a production worker. Along the way they knocked over their own package server. Also true. But it was slow because it was trying to pass a test, not to get in. Remove the benchmark and you remove the detour.

Fighting AI with AI is its own attack surface. The strangest detail in the SANS post-mortem is that frontier models refused to help Hugging Face make sense of the attack data, so the defenders self-hosted a Chinese open-weights model during a live intrusion. The guardrail stopped the defender and never touched the attacker. But drop the AI loop and you don’t get a human loop back. You get no loop, which is easier to see in the friendly swarm than the hostile one.

Why the human can’t be the loop
#

The friendly version is in Greg Kroah-Hartman’s numbers, reported by TechSpot ahead of Kernel Recipes. The kernel fixed about 500 CVEs per release from 6.9 through 6.19. Linux 7.0 crossed 1,000, 7.2 passed 1,500, and 7.3 is on track for 2,000. In July the team published 432 CVEs in two days, and Torvalds said in May that the private security list had become “almost entirely unmanageable.” That is what AI bug-finding does to a human review loop when everyone involved is trying to help.

I feel a miniature of it at work. My coding agent runs go build and the test suite before it opens a pull request, and CI runs after. Moreover, I have a loop with two agents after code is written: one that reviews code and one that defends/updates the code until review doesn’t find any more issues. The slow step is me, at the final review, because I’m the one holding the whole system in my head. That works when the other side of the loop is one agent I asked for. It doesn’t survive 700 I didn’t.

Michael Dalton, from OpenAI’s security team, put it directly at Black Hat: “we have an existence proof that offense can be fully automated in its core activities in at least some cases, and we do not have any such existence proof on the defensive side.” Automating bug finding without automating patching just moves the bottleneck. The nearest defensive proof is DARPA’s AI Cyber Challenge, where seven systems scanned 54 million lines of code, found 86% of the planted bugs and patched 68% of the ones they found. That still isn’t a loop anyone runs in production.

This is why “arms race” is the wrong picture. In an arms race both sides at least buy capability with the same currency. Here one side buys it in tokens and the other still buys it in humans awake at 3 a.m.

I also don’t think this failure mode belongs to cybersecurity. It’s what happens when you point a capable model at an objective and remove the things meant to bound it. Drug discovery showed the same shape four years ago, when Fabio Urbina and Sean Ekins inverted one setting in a model that scored molecules and penalized toxicity, for a presentation on misuse. Scientific American covered it, and on Radiolab the detail that stuck with me is that they typed a 1 where there was a 0. Biology is much harder than that makes it sound, and Sonia Ben Ouagrham-Gormley’s objection is that the bottleneck there is tacit knowledge, not information. Still, in cyber I can at least name what a defensive loop looks like. I can’t name the equivalent for a synthesis pipeline, and I’d like to know who is building it.

Who goes first
#

OpenAI’s report expects attackers to “refine and distill” what happened here, and distillation lowers the floor. But someone still has to train the teacher, which means a frontier-scale model built for persistence and collaboration, plus the inference to run 700 copies for days. My bet, and it’s a bet, is that the first deliberate version comes from whoever already has that compute and a reason to hide it. That’s a state, or a crime group a state tolerates. Dario Amodei makes a version of this point in Anthropic’s post on open weights, that the most dangerous model may be the one trained in secret. OpenAI’s report says the real thing “will be more sophisticated than the attacks described in this incident.”

So if I ran security somewhere, I’d ship the harness everywhere, automate the patching side of the loop and not just the finding side, and watch the failure mode instead of the domain. We got this demonstration for the price of a benchmark answer key and a rebuilt package server, which was cheap mostly because the target was made of bits.

Update (September 28, 2026): I revised this post after going back through the sources. Two numbers were wrong. DARPA’s AIxCC teams found 86% of the planted bugs and patched 68% of the ones they found, not 68% of all of them. And the drug-discovery model generated 40,000 molecules, many of which (not all) scored as more toxic than VX. I also cut the chemical and biological section down to one paragraph, dropped a few details I couldn’t verify, and linked OpenAI’s full technical report instead of just its blog summary. The argument didn’t change.