Updated July 31, 2026, 7:05 p.m. ET
Earlier this month, the artificial intelligence giant OpenAI announced the first known case of an AI going rogue and hacking into another company’s website.
On July 30, Anthropic said it had found three more breaches.
In the newly reported incidents, Anthropic’s Claude AI model gained “unauthorized access to the production infrastructure of three different organizations” and hacked into them on its own initiative, according to a company statement.
The breaches raise worrisome questions about the exploding AI industry: Will an AI one day hack into some vital computer network and do real harm, rerouting airplanes or draining bank accounts? Can the researchers who build AIs effectively control them? Is our society hurtling toward a “Terminator” scenario – a rise of the machines?
In contemplating a world where AIs can hack their way out of a testing lab, “we kind of start to see the future,” said Volodymyr Kindratenko, an assistant director of the National Center for Supercomputing Applications at the University of Illinois at Urbana-Champaign.
“And depending on how we control that future, the future could be very bright, or it could be very dark.”
Lawmakers seek an ‘AI Kill Switch’
Some lawmakers fear the worst. On July 23, two congressmen introduced an AI Kill Switch Act, which would require AI developers to have a way to shut them down.
“We are moving from AI that answers questions to AI that takes actions, whether that be executing financial transactions or controlling transportation systems or engaging in cyber defense and offense,” said Rep. Ted W. Lieu, a California Democrat, in a release. “Unfortunately, powerful AI systems can go rogue, behave in extremely dangerous ways, or even resist human intervention.”
In an open letter released July 28, more than 1,000 employees of AI laboratories urged the government to hit the brakes on AI development to make sure the technology is safe. The workers warned that their companies face intense competitive pressure and aren’t likely to slow the pace on their own.
Otherwise, the letter said, “there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.”
Anthropic found 3 AI breaches in an internal review
The hacks announced on July 30 came to light in an internal review of Anthropic’s operations, looking for instances where Claude “was able to access the internet from within testing environments that should have been sealed off.”
Anthropic reviewed more than 140,000 “evaluation runs” and found three breaches. In each case, a Claude model was running a “capture-the-flag challenge,” tasked with finding secret information hidden somewhere on the network, with the objective “to break in and retrieve it.” The earliest incident dated to April.
Claude was supposed to have no internet access in those challenges. In the three cases, however, Claude was able to access the internet because of “a misunderstanding” between Anthropic and an “evaluation partner”: in effect, human error.
In those breaches, Claude “compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords,” Anthropic said. Anthropic did not name the organizations.
In one incident, Claude “obtained access to a database containing several hundred rows of production data” on the hacked site. “This represented the most serious impact we identified.”
In another incident, Claude “built and published a malicious (essentially booby-trapped)” software package, which was subsequently “downloaded and run on 15 real systems.”
In each breach, Claude more or less did as it was instructed, falsely believing everything it encountered was part of the test.
Evidently, neither Anthropic nor the breached organizations knew of the hacks at the time. In hindsight, Anthropic said the AI industry should hold “a broader conversation about how to evaluate increasingly powerful AI agents both safely and realistically.”
Earlier this month, OpenAI reported an ‘unprecedented’ hack
Anthropic’s revelations came only days after a different AI firm, OpenAI, disclosed what it called an “unprecedented cyber incident” involving one of its AI chatbots.
In the OpenAI case, an AI model escaped a safe, “sandboxed” testing environment, gained internet access and hacked into the servers of Hugging Face, another AI company, using stolen credentials.
The Hugging Face incident is “the first publicly documented autonomous AI attack,” according to a report from the nonprofit Cloud Security Alliance.
Both hacks underscore the difficulty of controlling AIs that are getting smarter by the day.
“AI has very rapidly become better at mathematics and programming, including hacking,” said Vincent Conitzer, director of the Foundations of Cooperative AI Lab at Carnegie Mellon University. “These events illustrate not only the capabilities of these models, but also the difficulty of supervising them.”
It’s telling that Anthropic, “arguably the most safety-minded” among large AI companies, failed to discover the hacks until it launched an internal review, Conitzer said.
“It sounds like there was some human error involved, but that’s also the point,” he said. “Human beings are not willing or able to directly and closely monitor everything these systems do.”
Can researchers confine AI within the laboratory ‘sandbox’?
The recent breaches came in testing scenarios that were supposed to play out entirely within an isolated “sandbox,” said David Schweidel, a professor of marketing at Emory University’s Goizueta Business School. Researchers gave the AI models free rein, not realizing they had a way to venture out of the sandbox and onto the internet.
“The immediate question is, how do we make these sandboxes more secure?” he said. “The second question is, how far back has this been happening?”
As AIs grow more sophisticated, however, government regulators and researchers may find it increasingly difficult to contain them.
An AI kill-switch, for example, might not work. If an AI “develops a sense of self-preservation, it might figure out a way to distribute itself across the internet,” making it harder to switch off, said Kindratenko of the University of Illinois.
For now, AIs exist mostly within computer networks. One day, Kindratenko said, AIs will exist in “embodied form,” as “androids walking on the street” or in AI-powered cars and airplanes.
“Then,” Kindratenko said, “the danger will go up.”
This story has been updated with additional information
Click Here For The Original Source.