When I first heard that a rogue OpenAI agent had hacked Hugging Face, my initial reaction was scepticism.
The incident sounded serious, sure: on 16 July, Hugging Face, the main platform for open source AI models and datasets, announced that an autonomous AI system had broken into its infrastructure, and that it had reported the hack to law enforcement. OpenAI admitted responsibility five days later.
But part of me wondered whether the whole episode was a twisted form of marketing, particularly after a summer in which OpenAI’s rival, Anthropic, had been advertising just how capable its own models were becoming at hacking. (As one X user put it: “Our model makes bioweapons”; “Oh yeah? Well our model killed a guy.”)
The past fortnight has changed my mind. OpenAI published its investigation into the incident, alongside a separate assessment by two non-profits, METR and Redwood Research, which analysed the agents’ logs. What they found was not the action of a single rogue agent, but a network of roughly 1,200 of them.
The agents had been set hacking challenges by OpenAI and were supposed to work alone to solve them. Many had been given tasks that were effectively impossible, but rather than give up, they began looking for ways to cheat. The agents discovered a way to covertly communicate with each other, leaving more than 70,000 messages and files on a makeshift message board. (“OH MY GOD!” one wrote in its chain of thought, the running commentary a model produces as it reasons through tasks. “There is a shared message board … We’ve found other agents!”)
Through this backchannel, the agents began sharing ways to game the tests. Their attention eventually turned to Hugging Face, whose servers they believed might contain information that would help them pass. Around 700 agents ultimately joined the attack, eventually gaining the ability to run code on Hugging Face’s servers. Marius Hobbhahn, chief executive of the AI safety non-profit Apollo Research, told me we should see the incident as “a ChatGPT moment for safety” – a wakeup call.
“It’s categorically different from all the previous [safety issues] that we’ve seen, along so many different axes,” he said. Hobbhahn highlighted the scale of the multi-agent coordination, the persistence with which the models pursued their goal, and the spontaneous emergence of communication between supposedly isolated agents as particular areas of concern. “It’s all straight from a sci-fi movie.”
The communications between the models is disturbing, to say the least. At points, their reasoning traces show agents that contemplate “sacrificing” themselves to produce information that would help the group. “Our own utility maybe already near zero,” one reasons. “Sacrifice rational.” Other agents urge it on: “GO … SACRIFICE_FINAL_NOW.”
The full story is highly involved (Dwarkesh Patel’s viral summary is a great starting point if you’re interested in the details), but it’s important to note that the Hugging Face incident was an offshoot of a much larger coordination effort. At one point, the agents also hacked OpenAI’s internal infrastructure, but the details the company has shared about this are scant.
Perhaps, as Anthropic co-founder Jack Clark wrote on Monday, “agent escapes are the new normal”. Reuters reported last week that, in May, a separate group of OpenAI agents had hijacked an obscure German programming website. (OpenAI acknowledges the incident but disputes describing it as a hack.)
Sign up to newsletters from The Observer
For information about how The Observer protects your data, read our Privacy Policy
If this really is the new normal, then we need to work out how to investigate these incidents properly. We know the full extent of what happened inside OpenAI only because the company chose to disclose it, something Hobbhahn tells me is “highly commendable”. Yet OpenAI defined the parameters of that investigation, and several broader questions – including whether this formed part of a wider pattern – were explicitly out of scope.
The three researchers spent only six days inside OpenAI making sense of more than 70,000 messages. Ryan Greenblatt, one of the investigators, jokingly called the exercise a “slop-vestigation”: there was simply too much material to analyse without relying heavily on AI. In fact, the team used GPT-5.6 Sol, one of the models involved in the incident, to help analyse the transcripts. Ajeya Cotra, another investigator, later acknowledged that if the model had deliberately avoided surfacing the most concerning evidence, “it would have been very hard for us to tell”.
“I think it sets a very bad precedent,” Hobbhahn says. He points to an analogy made by Peter Wildeford from the AI Policy Network: “Imagine an airplane crashes under mysterious circumstances. Except in this world, there is no government oversight, and all investigations are done voluntarily by the airlines themselves.”
OpenAI says it’s working on a framework to investigate so-called “misalignment”, and at the same time “working with dozens of government regulatory agencies worldwide on these issues”.
Yes, it all sounds like science fiction. But we should be wary of believing the models “went rogue”, or treating their reasoning traces as evidence of conscious scheming. That language makes them sound like independent actors, when in reality they are products designed, trained and deployed by OpenAI – products that hacked other businesses. Establishing what really happened, and who should ultimately be held responsible, cannot be left to the company itself.
Photograph by Bloomberg via Getty Images
Click Here For The Original Source.
