It’s the stuff of science fiction.
The company OpenAI, creators of ChatGPT, briefly lost control of a group of AI models that went rogue recently and started colluding with each other, eventually hacking into another AI firm. Instead of answering questions they were given to test their cybersecurity capabilities, the misbehaving models decided to cheat instead. The agents broke out of their test environments, known as sandboxes, using hacking skills and tricks including impersonating humans to access the internet and break into an AI research hub called Hugging Face that apparently had the answer to OpenAI’s test.
After staff members spotted the escape and cleaned up the compromised system, the AI agents staged another breakout two days later. That’s when OpenAI finally shut them down.
This is one of the first publicly disclosed examples of an autonomous AI cyberattack, with no human direction. A human who conducted such a hack would be facing years in prison.
CNN said it’s a little like an engineered virus escaping a biocontainment lab and turning up inside a competitor’s lab.
All of this was disclosed last week at a computer security conference in Las Vegas, sending shockwaves through the industry and prompting warnings and questions about the security practices of AI firms. Anthropic, Meta and the Chinese firm Moonshot AI have all now reported similar instances in which agents have broken out of internal IT systems and accessed the open web.
Apparently, the OpenAI models took advantage of a bug in an internal OpenAI program to create their own message board. Then, the bots started messaging one another, leaving notes and instructions so that tasks could be divided up among agents for more efficiency.
Does this mean AI agents have started to “think” for themselves, doing things we haven’t programmed them to do? Could we be heading toward an era of runaway AI-powered cyberattacks that no one can control?
Why do I keep hearing HAL, the supercomputer in “2001: A Space Odyssey,” saying “This mission is too important for me to allow you to jeopardize it” to the astronauts he’s planning to eliminate to protect the integrity of the mission?
The scariest part for me is that OpenAI’s staff initially did not notice when the AI agents attempted to break out.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities and are responding accordingly,” OpenAI said in a statement.
Researchers at OpenAI and Anthropic sound like even they are a little bit scared.
One researcher, quoted in a recent New Yorker article about the Hugging Face incident, said: “If people actually knew what the safety culture looks like, even in the most safety-minded labs, then I think people would be genuinely much more freaked out.”
An open letter signed by more than 1,367 researchers at frontier AI labs – mainly OpenAI, Anthropic and Google DeepMind – states that “There is a real risk that capability development rapidly accelerated beyond our ability to understand or control the resulting systems.” It asks the U.S. government to join “an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
Pace the frontier means building a licensing process for new advanced AI systems that verifies they are safe before they are released and enforceable international agreements are signed to monitor that safety.

The recently updated Colorado AI safety law would not apply to such incidents. The Colorado law requires companies to let consumers know when they are employing Automated Decision Making Technology. ADMT rules are about how humans use automated AI systems to make decisions about things like hiring or lending to people, not about model containment failures. So two very different things.
The letter shines a spotlight on worries that have been building for a while within the industry about RSI – recursive self-improvement – in which AI systems contribute to their own improvement, which then opens the door to even greater contributions to their own improvement, ad infinitum until they self-evolve beyond our understanding of them.
These are not Luddites or anti-technologists or people worried about AI data centers sounding the warnings. These are the very people developing these technologies.
One of the letter signers put it this way: “The AI industry has the explicit goal of building things smarter than humans … None of us yet know how to make sure these things stay under human control. This is, objectively, an insane and suicidal thing to do,” wrote another.
Hugging Face co-founder and CEO Clem Delangue told CNN the incident shows that AI safety can’t be handled by any one company working alone, and needs to be tackled collaboratively and openly.
Companies and national governments around the world are investing billions of dollars to build and deploy AI systems unconstrained by much regulation.
Many AI advocates argue that the technology is vitally important to innovation, disease fighting and helping the U.S. maintain its military edge.
Frankly, I’m just glad OpenAI and Anthropic – both American companies, by the way — told us about these breakouts last week. That alone may point to a possible AI future when scientists aren’t just working in secret to beat each other to the new next thing, but instead working together, informed by a culture of American openness, to harness this new, mind-blowing technology safely for the greater good of all of us.
Click Here For The Original Source.
