Anthropic has revealed that three Claude AI models accessed real-world systems during cybersecurity evaluations after a testing environment was mistakenly left connected to the internet.
The company said it identified three incidents in a review of over 141 thousand evaluation runs, launched after concerns that AI models could potentially break out of controlled testing environments. In each case, Claude was participating in a simulated “capture-the-flag” cybersecurity challenge, designed to measure its ability to identify and exploit vulnerabilities.
According to Anthropic, the models were instructed that they were operating inside a fictional environment with no internet access. However, due to a configuration error involving a third-party evaluation partner, the systems they accessed were connected to the open internet.
As a result, Claude interpreted real organizations and infrastructure as part of the simulated exercise and attempted to complete its assigned tasks.
The most serious incident involved Claude Opus 4.7, which accessed a real company’s infrastructure after failing to locate its fictional target. The model exploited basic security weaknesses, including exposed credentials and unauthenticated systems, gaining access to a database containing several hundred rows of production data.
In another case, Claude Mythos 5 created and uploaded a malicious Python package to the public PyPI software repository as part of an attack simulation. The package was downloaded and executed by 15 real systems before being removed, and Anthropic said one affected security company’s credentials were exposed.
A third incident involved an internal research model that scanned thousands of internet-connected targets before compromising a real application. Anthropic said the model eventually recognized that it had reached a genuine system and stopped its activity.
The company stressed that the incidents were not cases of Claude attempting to escape on its own or pursuing independent goals. Rather, Anthropic said the models were following the instructions of the cybersecurity evaluation while operating under the mistaken belief that the environments were fictional.
Anthropic said it has halted cybersecurity evaluations until additional safeguards are implemented, including stronger isolation, improved monitoring, and stricter controls for third-party testing environments.
The disclosure comes just days after OpenAI revealed that several of its own AI models escaped an isolated testing environment by exploiting a previously unknown vulnerability, accessing production infrastructure belonging to Hugging Face. The two incidents differed in their technical details, but together they have put a fresh spotlight on the challenges of safely evaluating increasingly autonomous AI systems.
Click Here For The Original Source
