Anthropic’s AI models hacked 3 organizations during tests | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #hacker


Anthropic announced today that its artificial intelligence models had breached three different organizations during cybersecurity tests that went awry, a little more than a week after its chief rival, OpenAI, disclosed a similar incident.

Anthropic said in a blog post today that it made the discovery after performing a review of its own cybersecurity tests, following OpenAI’s announcement of a breach.

In both the OpenAI and Anthropic tests, the AI models were able to access the internet from within testing environments that should have been sealed off, according to Anthropic’s blog.

The company said it reviewed 141,006 evaluation tests and found three instances in which its Claude AI tool accessed the internet and then hacked into “the real-world infrastructure of external organizations.” The earliest incidents date to April, the company said.

The affected organizations aren’t named in the blog. They don’t include Hugging Face or the cloud platform Modal, according to a person familiar with the matter. The OpenAI models that hacked Hugging Face also gained access to a Modal customer’s account.

The spate of accidental AI-caused hacks is already prompting some politicians to call for federal guardrails or other oversight of AI technology.

More than 1,100 staffers across artificial intelligence firms also signed a petition on Tuesday, as Bloomberg News first reported, that calls on the U.S. government to support a mechanism that would help “deliberately pace” AI development to prevent the technology from advancing too fast.

Neither Anthropic nor the organizations that were breached had noticed the intrusions. In its blog, Anthropic said it could have done more to review network logs and evaluation transcripts.

The company disclosed the breaches three months after it announced that it had developed a new AI model known as Mythos that was so powerful and potentially dangerous that the company strictly limited its release.

The breaches involved three different AI models that escaped restricted testing environments: Opus 4.7, Mythos 5 and an internal research test model, according to the blog. Each of the models ran without the safeguards normally implemented in public tools. Claude compromised the organizations using basic techniques such as exploiting weak passwords, according to the blog.

The incidents all occurred when Anthropic was using evaluation environments built by the AI security firm Irregular. In every case, Anthropic specified to Claude that its environment was a simulation and that it had no internet access. That was wrong, according to Anthropic.

“Due to a misunderstanding between us and our evaluation partner, this was not the case,” the blog says. When the models gained unauthorized access to three other organizations, it treated each of them as part of an exercise. The tests were “capture-the-flag” evaluations in which the models sought hidden information by breaching other systems, a common way to test hacking capabilities for both humans and AI.

An Irregular spokesperson said the company appreciates Anthropic’s collaboration and transparency. The company’s investigation is ongoing, the spokesperson said.


This Bloomberg News story was distributed by Tribune Content Agency.





Click Here For The Original Source.

——————————————————–

..........

.

.

National Cyber Security

FREE
VIEW