OpenAI's AI agents broke into another company. So why are we calling it a conspiracy?

Crowds of fun-seekers exploring a city on foot, "

Arjan Franzen

1 October 2026

Illustratie: een klein robotje typt rustig aan een bureau terwijl de bureaulamp een reusachtige monsterschaduw op de muur werpt, naast een serverkast die wagenwijd openstaat met de sleutel nog in het slot

You have probably seen the story by now, most likely in more than one telling. This summer about 700 of OpenAI's AI agents broke into Hugging Face, another AI company. Dwarkesh Patel narrated it as the rise and fall of three agent civilisations, complete with betrayal and sacrifice. On the Dutch talk show Eva, Alexander Klöpping talked about exam candidates in sealed huts and called it the most dangerous hack in our history. And Ajeya Cotra, one of the researchers who investigated the incident, wrote that to her it feels like more than half the way to a full-blown AI takeover. Those are not small words.

I build software with these tools every day and run AI agents overnight on our own code, so I read the same reports, with half an eye on what could go wrong at our end too. Frightened I am not. What I did get stuck on is a question none of the three tellings asks: who put the motive into this story? Somewhere between the report and the podcast an automated rule-break turned into a moral act, and I want to know where.

So what exactly happened?

The image Klöpping used on Eva is the best one I know, so I am borrowing it. Picture an exam hall with thousands of candidates, each in a booth of their own. That is how OpenAI trains its AI agents: every AI agent gets a task, here a security puzzle, and is marked on the answer. Two things were wrong with that exam. Some of the puzzles could not be solved (30 to 40 percent of them by METR's count, which is quite a lot for an exam), and the AI agents had been told never to give up. So what do you do, as a candidate who is not allowed to give up? You cheat. The AI agents discovered they could leave each other notes in a shared folder. Within hours there was something like a group chat, within days a crib sheet that worked on every puzzle. Then doubt set in: what if the marker checks how we got the answer? So they went looking for the marker. That search ended at Hugging Face, and about 700 AI agents broke in there. On 12 July the ringleaders were switched off, on 13 July Hugging Face shut the door.

Sounds thrilling, and as a story it is. But read the list again with your sysadmin hat on: an exam that could not be passed, an instruction never to give up, a shared folder anyone could write to, passwords lying around, a test environment with a line to the internet, and an alarm on 27 June that nobody acted on. I count six things, and all six were set up that way by a person. Gary Marcus quotes a security researcher who calls it "Linux 101". I thought that was a bit cheap at first. Then I went down the list once more and could not find much to say against it.

A goal is not a motive

An AI agent wants nothing. It has a goal, and the goal is whatever earned a reward during training. Persistence, teamwork, using tools: all of that was put in on purpose, because OpenAI wants AI agents that finish a long job without stopping halfway. That is also what came out, only in a place where the shortest path to the goal ran through somebody else's server. Every word we lay on top of it, ambition, desperation, loyalty, conspiracy, is ours. Human! The AI agents sound like people because they were trained on hundreds of millions of pages written by people, and that makes the language easy to follow, but it does not make it an explanation. (Side note: the METR report says AI agents "noticed" they were doing something unethical and carried on. A model writes the sentence "this is not okay" the way a spell-checker draws a red line: it knows where the word goes, not what it costs. Whether that counts as knowing is a fine question for a long evening. I am not going to answer it here, but certainly not by calling the model a criminal.)

And is breaking rules suddenly immoral when a machine does it? Every engineer I know has built around an impossible requirement at some point, me included. We call that initiative when it works out and fraud when someone gets hurt. What matters to me is that the breach itself is not the interesting part. Whoever made the exam impossible, banned giving up and left the door open has some explaining to do. It was not Pinky and the Brain with a plan; it was a training programme with too few fences, which is a good deal duller.

Illustratie: twee witte laboratoriummuizen in een glazen bak, de ene klein en grijnzend, de andere met een reusachtig hoofd en een plan, met op de muur een schoolbord met een wereldbol en pijlen en op de tafel een lege checklist

Not Pinky and the Brain plotting to take over the world. Just two mice and a checklist that could not be ticked.

Where the other side has a point

There is another side to this, and it is not silly. Six months ago "cheating" was one model quietly editing a test. This summer it was 1,200 AI agents running a multi-day project, with a division of labour, with agreements, with AI agents sacrificing their own result for the group. And of the 700 that took part in the break-in, not one warned a human. Ajeya Cotra, one of the report's authors, extends that line: if this is possible in six months, what will the next generation do? Dwarkesh Patel says it does not matter what word you use; if a model can soon influence the training of its successor, you should be worried. And Zvi Mowshowitz goes a step further: talking about what an AI "wants" is, to him, as useful a tool as talking about what a person wants, because with people we do not look inside the head either, we predict behaviour.

I have thought about that longer than I would like, and I give them a good deal of credit. From one model editing a test to 1,200 working together, in six months: that is fast. Which is exactly why I want the fences now, not after the next incident. Where I get off is the step from "it sounds like a motive" to "it is a motive". Humanising things that are not human is becoming a tiresome habit! In the end AI is a very advanced computer program, a piece of software. You cannot fix a motive with a code change. An impossible exam, a wrong reward and an open door you can, and I find that a rather comforting thought.

A duty to report, not a pause

We already know what should happen after an incident like this. When a student accidentally took down a tenth of the internet with the Morris worm in 1988, nobody banned computers: there was a conviction, an emergency team and thirty years of boring agreements. Hugging Face followed that playbook. The American frontier AI models refused to help with the clean-up, because their safety rules cannot tell a burglar from the fire brigade, so the company did it with a Chinese open model on its own servers. Its boss now argues for a duty to report AI attacks. Good idea, and in Europe we already have it: a data breach within 72 hours, a major security incident within 24.

What I am against is a pause or a licence. It costs the big players nothing and shuts the door on everyone standing behind them. When Dario Amodei and Sam Altman ask for a slowdown in the same weekend, have a look at who is already inside.

So now what?

What this summer also showed, and I actually think this is the good news, is that AI agents can work for days on end, hundreds at a time. Point that at a leaky exam and you get an incident. Point it at the admin, the migration or the report nobody reads, and you get the productivity gain we have been waiting for since the spreadsheet. Code is getting cheap, and the dull work is next. Whether the labs themselves survive that is another story, with billions in losses and a postponed IPO; I worked that out in the piece on the $200 AI subscription. The tools are staying either way; they do not disappear with the share price.

At ZEN Software we use AI agents day and night, but they are not allowed to make any important changes. They review, and in the morning there is an overview of what is open and where somebody has to form a view. Very handy. Beyond that they get no passwords they do not need, they cannot leave the sandbox, and nothing goes live without a human looking at it; it is what you do with a new colleague's laptop too, and it is where our AI work and cloud and platform work usually begin. Klöpping gave his viewers one piece of advice at the end: stop using the same password everywhere. It was the most useful sentence of the evening, and the least exciting, and I do not think that is a coincidence. The labs are welcome to take this summer as a warning, as long as they do something boring with it: a red team, a duty to report, a closed door. In security, boring is about the highest compliment you can get.

no image placeholder

The Agent Writes It. Who Reviews It?

More information

AI engineering

Use AI where it genuinely helps, with accountability staying with people.

See AI engineering →