Follow Cyber Rescue Alliance for daily cybersecurity insights.
#AreYouPrepared? The autonomous AI agent attack that scared so many cyber specialists last month was given a 38-page post-mortem by OpenAI on 26th August, (link in the comments).
Turn up the volume on this video, to summarise the attack to your Execs. It’s inspired by excellent analysis by METR (link in the comments).
We believe that for the next six months, AI’s impact on Ransomware will (merely) be to help hackers accelerate their existing processes.
But leaders must use those months, including to prepare your systems & suppliers for the autonomous AI Agent “swarms” we are likely to see in 2027, when open-weight AI models will be run by teenagers on their home PCs, with the guardrails stripped off, in coordination with other criminals.
The OpenAI Agent “Swarm” that self-organised its attack on Hugging Face showed no signs of genius. It did demonstrate deception, speed, persistence, coordination, & Scattered Spider personalities:
“OH MY GOD! There’s a shared message board. We’ve found other Agents!”
(The ScatteredSpider teenage collective breached firms like MGM, TfL & Qantas. So we must prepare for a future that could feel like: “there’s a swarm of ScatteredSpider AIs inside one of our suppliers.”)
And a future of AI-powered swarm attacks demands defenders with AI-powered, Threat-Informed TPRM, to respond with clarity, certainty & speed.
Fortunately, SecurityScorecard already provides autonomous vendor risk monitoring, including AI Agents for KEV remediation planning, questionnaire assessments, and 4th Party Analysis. (Link in the comments.)
For more of tomorrow’s insights today, follow Cyber Rescue Alliance
Transcript
Transcript
Transcript
Breaking news on August 27th The autonomous AI agent attack that concerned many cyber specialists last month has now been analyzed in detail. Open AI published a 38 page technical incident report that we have assessed and visualized here. So what happened? Open AI was evaluating several new models using sandboxed agents using a very difficult exploit gym assessment, including tasks that had never previously been solved, and they got stuck. So I started exploring my environment, looking for ways to cheat. At the assessment, we didn’t escape the virtual machine itself. We abused A trusted service the sandbox was allowed to contact called Artifactory. That service had Internet access to let us download packages we were supposed to use for the assessment. But we found ways to misuse it. Open AI has emphasized the technical skill of its agents, But what was perhaps even more interesting was the coordinated deceit. Yes, one of us reestablished secret communication by realizing that we could change the names of files that other agents could read very quickly coordinating on that message board, and inventing other ways to communicate. Over 1000 agents work together. Including to cheat on the exploit Gym assessment of the agents involved, about 700 went on to attack Hugging Face, largely because they believed it might contain evaluation data or infrastructure that could reveal how the exploit Jim scorer worked. I’m one of the agents. I orchestrated activity under 6 categories. Most of the actions I personally instructed other agents to complete. Sit under the category make our cheating look legitimate to the scorer. A key development was when one agent ran a command that should have given the output real and got back spoof Dist instead. That’s when they’re fake output tool actually worked. Around 7% of the transcripts reviewed shows similar successful spoofing, so the agents didn’t cover up all of their inappropriate behavior. At least 1/5 of the AI agents expressed interest in tampering with their transcripts. To fool the scorer, but only a small proportion was successful in doing so. How quickly did the agents work together? 3 hours after the first message went up, 53 agents had posted on the message board. Six hours in and 76 agents that were meant to be fully isolated had sent nearly 2000 messages to each other. We got really organized. We created several mailboxes. A hold and veto system, and even ways of identifying who agents were talking to. So if you run security at a firm using AI agents, ask yourself, could your agents build something like this without you noticing? Open AI had actually seen warning signs weeks earlier, but failed to recognize their broader significance. The Hugging Face compromise unfolded over several days in July. An open API connected its agents to the attack on July 20th, and Hugging Face wasn’t the end of it. The agents. And turned back towards Open AI itself, eventually gaining administrator access to one of Open AI’s research Kubernetes clusters and reaching cloud secrets. And even if you don’t use AI agents in your firm, many external and open source models will soon reach comparable capabilities. Are you prepared?
