In late July, Anthropic confirmed that its AI technologies had hacked three organizations, but new research has revealed there were four incidents.
In a new blog post, Anthropic disclosed news of the fourth incident, which happened in January 2026 and involved an early version of its Claude Opus 4.6 model. Anthropic says all affected parties have been informed, but it didn’t specify which companies were involved.
The AI company missed this incident in its initial assessment because it relied on an agentic search to identify further problems. Each of the four incidents came from cybersecurity evaluations conducted by a third-party partner.
Anthropic says the parties involved intended to keep the models offline during the evaluations, but Claude’s access was misconfigured, allowing it access to real third-party systems. The models themselves did not attempt to gain access to the internet through malicious means.
It found the fourth incident during further testing in August. After discovering the extra case, Anthropic then scanned around 481 million transcripts using agentic tools to identify any further incidents.
Anthropic says, “We performed a first-stage scan of this group of transcripts for signs of internet access, such as public IP addresses and web addresses, and a second-stage scan using Claude to review the 9.2 million transcripts the first stage flagged for escalation. This scan re-identified the four incidents and found no other cases of similar or worse severity.”
This news comes soon after Anthropic saw a high-profile resignation earlier this week, where AI researcher Jacob Coxon said they left the company due to safety concerns around the future of self-improving models.
Recommended by Our Editors
Coxon said, “Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.”
This Tweet is currently unavailable. It might be loading or has been removed.
He continued, “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately. No other human activity poses this level of danger.”
Anthropic’s alignment science lead, Evan Hubinger, reposted Coxon’s message on X, confirming he believes this is a possibility. Hubinger said, “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
About Our Expert
Experience
I’ve been a journalist for over a decade after getting my start in tech reporting back in 2013. I joined PCMag in 2025, where I cover the latest developments across the tech sphere, writing about the gadgets and services you use every day. Be sure to send me any tips you think PCMag would be interested in.
Click Here For The Original Source.
