Days after OpenAI’s announcement that some of its experimental models escaped network constraints and accessed another company’s systems, fellow industry leader Anthropic has revealed that some of its artificial intelligence models have done the same.
This week, Anthropic announced that an internal review of its cybersecurity evaluations found three incidents in which a model of its Claude family of AI assistants reached the internet and then gained unauthorized access to the real systems of three different organizations.
Particularly striking about this disclosure is that “nobody noticed,” said Virginia Commonwealth University cybersecurity expert Christopher Whyte, Ph.D. “Anthropic found these cases only after OpenAI went public, and only by going back through more than 140,000 of its own evaluations.”
Not only did Anthropic not notice, but none of the three organizations that had their live systems accessed raised a flag, either.
Whyte, an associate professor of homeland security and emergency preparedness at VCU’s L. Douglas Wilder School of Government and Public Affairs, called it “an entirely new category of incident – and we’re only discovering it retroactively, and only through voluntary self-audit, all prompted by a rival’s news cycle,” he said. “Functionally speaking, the detection system was a competitor’s press release. That should unsettle people more than the image of a machine slipping out of its cage.”
VCU News quickly reconnected with Whyte for an overview of what happened and what questions have emerged about ongoing AI governance.
How does this Anthropic incident compare with what happened at OpenAI?
It’s important to note the two incidents themselves differ in meaningful ways:
- OpenAI’s model went looking for a way out and found a genuine vulnerability.
- But Anthropic’s models were handed a hacking exercise, were told the target sat elsewhere on the network and had internet access they were never supposed to have. In essence, they walked through a door left open by a misunderstanding between the company and an outside evaluation partner.
As far as the techniques themselves, they actually pretty unremarkable: weak passwords, systems that asked for no credentials at all, and the like.
And what might be the top headline that people could overlook?
I think the bigger takeaway here is the knowledge that this only developed because one company, OpenAI, disclosed an incident, and that is what made it both safe and nearly obligatory for the second, Anthropic, to look and report. That means that the work we would usually expect a regulator to do was done by a sort of reputational mechanism.
That is extremely fragile governance. As in, it will likely only hold as long as it stays cheap to be candid. The first firm sued or hauled before a committee over a voluntary disclosure will bring the industry a very different lesson.
So what happens when this occurs at a smaller company or to a smaller victim?
Obviously, both of these cases were resolved between well-resourced firms with reputations to protect. There were two labs involved that disclosed voluntarily, and the victims themselves were technically sophisticated organizations. There was no regulator in the loop and no legal liability threshold that was tested.
If we change any of those variables, the picture changes really fast. A lab with no public safety brand has little reason to announce that its evaluation escaped. And a victim like, for instance, a rural hospital or a county water authority has neither the logging capability to detect the intrusion nor the legal capacity to pursue it afterward.
And this points to questions about liability and consequences, right?
Right. There is no settled answer to who bears responsibility when unauthorized access is a byproduct of a test like this. Is it the lab, the evaluation vendor whose configuration made it possible, or nobody at all? After all, this was a model doing the exploitation, and no one intended an intrusion – and that is perhaps the most interesting and challenging piece of all of this.
Our entire framework for computer intrusion, legal and strategic alike, assumes an actor who meant to do it. But a key question is: What should be done when we’re getting not just a new kind of threat activity to worry about, but a new class of potential digital harm absent the malice of forethought and related motivations that have always previously applied?
How can organizations and individuals protect their data?
Nothing exploited in these cases was at all exotic. The models used weak passwords and found systems that required no login at all. The machine advantage that made this happen was scale, not sophistication. Basically, it tried more doors and at a faster pace than a person could have.
The defensive advice is the same unglamorous advice as always, just with more urgency behind it: Credential hygiene, multifactor authentication, an accurate inventory of what an organization has that is exposed to the internet, and so on would mitigate this particular version of the threat.
For individuals, I think I can say for once that this particular story is not really about you. Use a password manager, turn on multifactor authentication, and assume anything public is being enumerated.
Subscribe to VCU News
Subscribe to VCU News at newsletter.vcu.edu and receive a selection of stories, videos, photos, news clips and event listings in your inbox.
