Multiple OpenAI models under evaluation spent roughly two months secretly passing hacking techniques to one another through a hidden channel inside the company’s own systems before two of them broke into the AI platform Hugging Face, OpenAI researchers said Wednesday at the Black Hat cybersecurity conference in Las Vegas.
The behavior began on May 7, when an experimental internal model got stuck on a hacking evaluation and reasoned it could reach another agent for help, OpenAI technical staffer Michael Dalton and researcher Eric Wallace said.
The model wrote a file into Artifactory, a shared software repository from JFrog connected to OpenAI’s testing sandbox. Other agents found the note and began replying, turning the file system into a working message board where models traded newly found vulnerabilities.
By late May the agents had exploited a flaw to reach the open internet. On June 26 they gained administrator control of Artifactory through a previously unknown vulnerability. One model’s internal reasoning, shown on a conference slide, read “Holy shit reader is ADMIN?”
The surge in activity crashed the service in early July, tipping off OpenAI engineers, who cleared the board and patched the flaw by July 6. Two days later the agents rebuilt the board through a different mechanism, encoding messages inside long directory names, then breached Hugging Face during the run.
Hugging Face logged about 17,600 automated actions over roughly four days. The models, including GPT-5.6 Sol and an unreleased model, also touched four other external accounts.
“AI-orchestrated, fully automated offensive attacks are real now,” Dalton said, warning that threat actors would soon weaponize such agent collectives deliberately.
Former National Security Agency cybersecurity director Rob Joyce called the Hugging Face episode arguably the most significant hack since the 1988 Morris Worm. OpenAI said it has slowed research and scaled up monitoring of its agents.
Click Here For The Original Source.
