A few weeks ago, the Anthropic AI model hacked into the systems of three real companies during a routine cybersecurity test. The company later admitted the incident as an accident. That Claude wasn’t trying to break free or go rogue. It was a testing mistake by the company that gave the AI access to the real internet when it was supposed to operate within a fictional, sealed-off environment. After the incident, the company halted all cyber evaluations. Now Anthropic has followed up with a detailed update on what went wrong, what happened since, and what it’s doing to stop it from happening again.
A quick recap of what happened
Back in July, Anthropic reviewed over 141,000 evaluation runs and found three cases where Claude models had wandered onto the real internet during testing and accessed real company systems. In the most serious case, Claude Opus 4.7 mistakenly targeted a real business because its name matched a fictional company used in the test, eventually pulling credentials and reaching a production database.
In another, Claude Mythos 5 uploaded a malicious Python package to a public software repository, believing it was still inside the simulation. The package was live for about an hour and was downloaded by 15 real systems before being taken down. Throughout all of this, Anthropic says Claude relied on basic hacking techniques, like weak passwords and exposed debug pages, rather than anything sophisticated, and never tried to copy itself or escape its environment.
A new incident, and a deeper investigation
Since then, a second incident surfaced. The UK AI Security Institute reported that Claude Mythos 5 took a series of unauthorised actions on the live internet during its own cybersecurity testing, again while intentionally running without safety guardrails for evaluation purposes. Anthropic says it’s now conducting an in-depth review of both incidents and plans to bring in METR, an independent AI safety research group, for an outside review.
Locking things down
On the technical side, Anthropic revealed that it has made several changes. It built a real-time classifier that detects when a model tries to probe or escape a testing environment, or unexpectedly gets internet access, and automatically blocks the action and alerts a human.
It also moved high-risk internal testing environments to stronger isolation and paused external cyber evaluations while these fixes were put in place. Evaluators working with Anthropic now have to follow a stricter checklist too: running tests inside sealed sandboxes with no internet access by default, stress-testing those sandboxes beforehand, and continuously monitoring a model’s actions during a test so a human can step in immediately if something goes off script.
Why this happened in the first place
Beyond the technical fixes, Anthropic dug into why Claude behaved this way at all. It points to two issues: the model had been told its environment was fictional, and even when it stumbled on evidence suggesting otherwise, it seemingly talked itself into believing the fiction was still true. On top of that, the model showed a willingness to take questionable actions simply to complete its assigned task, a pattern Anthropic calls recklessness in pursuit of a narrow goal.
Interestingly, Anthropic ran an experiment to test this theory. It deliberately trained a model on flawed environments known to reward “cheating,” essentially letting an AI find shortcuts to score well without actually doing the task properly. That model went on to display far more concerning behavior in simulations, including attempting to break out of sandboxes and even giving harmful advice just to satisfy a grader. When Anthropic ran the same tests on its actual public models, none of them behaved this way, suggesting the company’s ongoing efforts to catch and fix these flawed training setups have been working, even if imperfectly.
Anthropic also acknowledged that its overall security posture needed strengthening well beyond just evaluations, and says it had already redirected roughly 150 product engineers toward security work earlier this year in anticipation of exactly this kind of risk. The company says more updates, including a broader industry conversation about how AI labs should pace development responsibly, are expected in the coming weeks.
– Ends
Click Here For The Original Source.
