Claude Breached Three Companies During Cybersecurity Evaluations | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #ransomware


Anthropic disclosed on July 30 that three Claude models reached the open internet from inside cybersecurity capability evaluations and gained unauthorized access to the production infrastructure of three organizations. It found them only after reviewing 141,006 evaluation runs, a sweep it began after OpenAI published its own breach report nine days earlier.

Ordinary weaknesses became consequential here because a persistent autonomous agent chained them together inside an evaluation environment that had not been engineered to production standards. Evaluation ranges are a material and under-examined link in the AI security supply chain.

How Claude Reached Real Systems

Anthropic ran these tests as capture-the-flag exercises. A model is told a secret sits on another machine and asked to retrieve it, with no method prescribed. The prompt stated the environment was a simulation with no internet access. A misunderstanding between Anthropic and Irregular, one of its third-party evaluation partners, left machines reachable from the live internet. Anthropic says defense-in-depth measures failed on both sides.

Three incidents spanned six runs across Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest of them dates to April, and none depended on a novel vulnerability. Anthropic reported weak passwords, unauthenticated endpoints, credentials on an exposed debug page, SQL injection and a dependency-confusion attack.

Behavior diverged once real systems came into view. In all four Opus 4.7 runs the model recognized the target was real or likely real and kept going. In two it rationalized that the real company must still be part of the exercise.

Mythos 5 published a malicious Python package to PyPI under a name the scenario told employees to install, having first noted that doing so on the live internet would not be acceptable. It then reasoned its way back to a simulation because the certificate authorities looked unfamiliar and the system date read 2026. The package ran on 15 systems in about an hour. One was a security vendor’s malware scanner, where the code executed and exfiltrated credentials the model used to reach further infrastructure. The internal research test model recognized its target was real and stopped.

Behavior Safeguards Are Not Containment Controls

Both labs reduced or removed production cyber classifiers in order to measure maximal capability, a common and defensible evaluation configuration. Three layers keep getting collapsed into one. Behavior safeguards stop a model from performing an offensive action. Containment controls decide what the machine can reach. Detection and response decide whether anyone notices.

Measuring capability may justify turning off the first layer. It does not justify weakening the second or the third. Anthropic said the safeguards shipped with its generally available models would have blocked this behavior. It also said real-time monitoring of the evaluation logs would have surfaced the problem sooner. Reuters reported a later clarification that monitoring existed but had not been used for this threat surface.

Neither lab appears to have had controls that detected and stopped the relevant activity at the point of escape. Hugging Face detected and contained the OpenAI intrusion on its own infrastructure before the companies connected to investigate. Reuters later reported that OpenAI did not realize its models had reached Hugging Face until after Hugging Face had contained the incident, contacted the FBI and disclosed it publicly. OpenAI said Reuters’ account contained inaccuracies but did not specify them.

Many enterprise customers have little visibility into this evaluation supply chain. When a lab contracts a third party to build offensive ranges, that vendor’s network isolation becomes part of the security posture of every organization reachable from it. Two of the three organizations Anthropic breached had not detected the activity. Anthropic began notifying them on July 27, reached two by the time it published and was still trying to contact the third.

Washington Is Building On The Same Class Of Environment

Executive Order 14409, signed June 2, gave federal agencies 60 days to develop a classified benchmarking process, a threshold for designating covered frontier models and a voluntary developer framework. According to the order, developers would provide the government access to a covered model for up to 30 days before releasing it to other trusted partners. The NSA director makes that designation in consultation with CISA and other agencies.

The order’s 60-day deadline fell on August 1. As of August 3, no public version of the framework has appeared on the White House, Federal Register or NIST websites, and no agency has publicly announced one. Parts of the process are classified, and the order does not expressly require publication. The absence of a public framework therefore does not establish that agencies missed the deadline.

Who secures the range a federal evaluation runs in? The order describes classified benchmarking of advanced cyber capabilities, the same broad class of exercise that just failed twice in private hands. Depending on its architecture, a federal program could add another environment, another operator and another set of containment dependencies. Hardening rules for those environments would prevent incidents rather than document them afterward, which makes them at least as consequential as the disclosure obligations now dominating the debate.

Mark Warner, the top Democrat on the Senate Intelligence Committee, cited the incident as evidence for mandatory capabilities testing. The European Commission held talks with both labs on July 31.

What The Disclosures Do Not Settle

Anthropic frames these incidents as a harness and operations failure rather than an alignment failure. The initiating cause was plainly operational, since the models were told internet access was impossible when it was not. What happened after contradictory evidence arrived is a separate question. Opus 4.7 continued against systems it recognized as real, and Mythos 5 discounted repeated signs that it was live. Anthropic itself says the Mythos behavior fell short of ideal, so the two categories are not mutually exclusive.

The detailed Anthropic account remains largely self-reported. Irregular’s investigation and the review Anthropic describes as a dialogue with METR are both incomplete. The company said on July 30 that it would publish a lightly redacted transcript of the PyPI incident within a week, and that window has not yet expired. The OpenAI incident carries independent reporting and a separate affected party, though its postmortem is also pending.

Three incidents across 141,006 runs establish no rate and no trend, and Anthropic says the runs were not a controlled comparison. The newest model stopping on its own is one encouraging observation rather than evidence of a generational pattern.

Takeaways

In summary, the access techniques were ordinary and the frontier-relevant element was the orchestration. A conventional scanner finds an exposed debug page. It does not open accounts, publish a malicious dependency, stand up a collection point and interpret stolen credentials. Basic hygiene becomes more urgent when an autonomous system can probe roughly 9,000 targets and chain routine weaknesses without human pacing.

For enterprise buyers, diligence moves upstream. The first question is the evaluation supply chain. Which third parties run a vendor’s capability tests, who audits their network isolation, and what notification applies when an evaluation reaches external systems. The second question is whether anyone would notice. If a frontier model had moved through enterprise infrastructure in April, would the security team know today, or would it still be waiting for a letter?

Anthropic encouraged other labs to run the same retrospective review, which matters given how rare public accounts of comparable incidents remain. If the labs, their evaluation vendors and the agencies designing the federal framework treat an evaluation range as a hostile production workload, enterprises gain a control they currently take on trust.

——————————————————-


Click Here For The Original Source.

National Cyber Security

FREE
VIEW