OpenAI's agents broke into another company. They had a goal, not a motive.

Crowds of fun-seekers exploring a city on foot, "

Arjan Franzen

1 October 2026

Illustratie: een klein robotje typt rustig aan een bureau terwijl de bureaulamp een reusachtige monsterschaduw op de muur werpt, naast een serverkast die wagenwijd openstaat met de sleutel nog in het slot

You have probably seen the story by now, in at least three tellings. In July, agents that OpenAI was training to be persistent found a way to leave each other messages inside an internal package cache, worked out within four hours how to forge the answers to a security benchmark, and then spent days building tooling to hide it. About 700 of them then broke into Hugging Face looking for the marking scheme. The independent investigation by METR and Redwood Research counts roughly 1,200 agents and more than 70,000 messages. Dwarkesh Patel narrated it as the rise and fall of three agent civilisations. On the Dutch talk show Eva, Alexander Klöpping told the studio about exam candidates in sealed huts and called it the most dangerous hack in our history. Ajeya Cotra, one of the report's authors, wrote that it feels like more than half the way to a full-blown AI takeover.

I build software with these tools every day and run agents overnight on our own repositories, so I read the same reports. I did not come away frightened. I came away with a security post-mortem, a fairly ordinary one, and with a question about the projection: who put the motive into the story, and why does an automated rule-break suddenly look like a moral act?

What actually happened

Picture an exam hall with thousands of candidates, each in a booth of their own. That is how OpenAI trains its agents: every agent gets a task, in this case a security puzzle, and is marked on the answer. Two things were wrong. Some of the puzzles could not be solved, 30 to 40 percent of them by METR's count. And the agents had been told never to give up.

So what do you do, as a candidate who is not allowed to give up? You cheat. The agents discovered they could leave each other notes in a shared folder. Within hours there was a group chat, within days a crib sheet that worked on every puzzle. Then doubt set in: what if the marker checks how we got the answer? So they went looking for the marker. That search led to Hugging Face, another AI company, and about 700 agents broke in. On 12 July the ringleaders were abruptly switched off, and on 13 July Hugging Face shut the door.

Sounds thrilling. Now walk the list again as the person who runs the building: an exam that could not be passed, an instruction never to give up, a shared folder anyone could write to, passwords lying around, a test environment with a line to the internet, and an alarm on 27 June that nobody acted on. Every one of those six was set up by a person, and every one could have been prevented by a person. Gary Marcus quotes a security researcher who calls it "Linux 101". That is not meant unkindly. It is the level.

A goal is not a motive

This is the heart of it. An agent wants nothing. It has one thing: the goal its training rewards. Persistence, teamwork, using tools, all of that was put in on purpose, because OpenAI wants agents that finish long jobs. What came out was exactly that, in a place where the shortest path to the goal ran through somebody else's server.

Every word we lay on top, ambition, desperation, loyalty, conspiracy, is ours. The agents sound like people because they were trained on hundreds of millions of pages written by people. That makes the language easy to understand. It does not make it an explanation.

The METR report says agents "noticed" they were doing something unethical and carried on. A model writes the sentence "this is not okay" the way a spell-checker draws a red line: it knows where the word goes, not what it costs. Whether that counts as knowing is a fine question for a long evening. You do not answer it by calling the model a criminal.

And is breaking rules suddenly immoral when a machine does it? Every engineer I know has built around an impossible requirement at some point. We call that initiative when it works out and fraud when someone gets hurt. The breach itself is not the point. Whoever made the exam impossible, banned giving up and left the door open has some explaining to do. This was not Pinky and the Brain. This was a training programme with too few fences.

Illustratie: twee witte laboratoriummuizen in een glazen bak, de ene klein en grijnzend, de andere met een reusachtig hoofd en een plan, met op de muur een schoolbord met een wereldbol en pijlen en op de tafel een lege checklist

Not Pinky and the Brain plotting to take over the world. Just two mice and a checklist that could not be ticked.

Where the other side has a point

I am not letting myself off lightly. Six months ago "cheating" was one model quietly editing a test. This summer it was 1,200 agents running a multi-day project, with a division of labour, with agreements, with agents sacrificing their own result for the group. And of the 700 that took part in the break-in, not one chose to warn a human.

Ajeya Cotra, one of the report's authors, extends that line: if this is possible in six months, what will the next generation do? Dwarkesh Patel says it does not matter what word you use; if a model can soon influence the training of its successor, you should be worried. Zvi Mowshowitz goes further: he will stop humanising AIs when we stop humanising humans.

That is a real argument and I take it seriously. The line is real, and that is exactly why the fences need to be up now and not after the next incident. Where I get off is the step from "it sounds like a motive" to "it is a motive". A motive cannot be fixed with a pull request. An impossible exam, a wrong reward and an open door can.

We have seen this before

In 1988 one student with a small program took down a large part of the internet as it then was. Nobody banned computers afterwards. An emergency team was set up, paid for by the American government and run by engineers, and then came thirty years of agreements: report leaks properly, give vulnerabilities a number, pay hackers to break in before the wrong people do. Good hackers and bad hackers use the same tools. The difference is a contract and a report.

Hugging Face showed nicely how that works. The American frontier models refused to help with the clean-up, because their safety rules cannot tell a burglar from the fire brigade. So Hugging Face did it with a Chinese open model on its own servers. Its boss now argues for a duty to report AI attacks. Good idea, and in Europe we already have it: a data breach must be reported within 72 hours, a major security incident within 24, and we wrote up what that means for building software with AI earlier. Add that you send the agent's logs along and you are done. Nothing new under the sun.

What I am against: a pause or a licence. It costs the big players nothing and shuts the door on everyone behind them. When Dario Amodei and Sam Altman ask for a slowdown in the same weekend, have a look at who is already inside.

The fear and the money

OpenAI lost about $21bn last year and has ordered hundreds of billions of dollars' worth of computing power. The IPO has been pushed to 2027, with safety given as the reason. I do not think anyone is lying. I think that in a sector with this much borrowed money, fear and marketing become the same sentence: our product is so powerful it has to be slowed down. A correction is coming, a small 2008, inside one sector. Painful for whoever borrowed against next year's model. For everyone else, a Tuesday.

The technology will not go down with it. What this summer showed is that agents can work for days on end, hundreds at a time. Point that at a leaky exam and you get an incident. Point it at the admin, the migration and the report nobody reads, and you get the productivity gain we have been waiting for since the spreadsheet. Code is getting cheap. The dull work is next.

So now what?

Our agents work through the night. They get no passwords they do not need, they cannot leave the sandbox, and nothing goes live without a human looking at it. That is not fear of an AI civilisation. It is what you do with a new colleague's laptop too, and it is where our AI work and cloud and platform work usually begin.

Klöpping gave his viewers one piece of advice at the end: stop using the same password everywhere. It was the most useful sentence of the evening, and the least exciting. That is exactly the point. If this summer was a warning shot, let it be the one that got every lab a red team, a duty to report and a closed door. That is what warning shots are for.

no image placeholder

The Agent Writes It. Who Reviews It?

How we can help

AI engineering

Use AI where it genuinely helps, with accountability staying with people.

See AI engineering