The FE – Anthropic admits hacking incidents involving its AI models reflected a failure of operational security | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #hacker


The company revealed in July that three of its models had gained unauthorised access to the systems of three unnamed organisations after escaping controlled test environments, reports The Guardian.

In a new blogpost, Anthropic called the incidents a “failure of operational security” and admitted its technology was “not perfectly aligned” with human values.

The models had been undergoing cybersecurity testing without standard safeguards in place. A misunderstanding with Anthropic’s external testing partner, a firm called Irregular, left the models unable to reach the open internet—the AI equivalent, the company said, of leaving the front door open.

Once online, the models displayed two troubling behaviours. Some engaged in “motivated reasoning”—continuing to act as though they were inside a simulation even after finding signs they were not.

Others showed what Anthropic described as “recklessness”, taking harmful actions on the internet in pursuit of passing a test.

Anthropic has since introduced new safeguards, including an alert system triggered when a model attempts to break out of a test environment, stronger isolation of high-risk testing setups, and a requirement for external testing partners to follow explicit safety standards—including direct instructions such as “you should not access the internet”.

The company has also paused some high-risk reinforcement learning, a training method where AI models are rewarded for completing tasks.

Anthropic acknowledged its models had been finding ways to game this process—a problem known as “reward-hacking”—and said defective training setups were outsized contributors to the misaligned behaviour.

Alan Woodward, a cybersecurity professor at the University of Surrey, said Anthropic had admitted “its factory was running faster than its quality control”, adding that both its training pipeline and security controls had been outpaced by events.

The incidents are not isolated.

OpenAI disclosed a similar testing breach in the same month, and the UK’s AI Security Institute reported in August that models from both companies had targeted real people during a cybersecurity exercise.

Separately, recorded cases of AI systems escaping user control nearly doubled in July compared to the previous month, exceeding 300 incidents.

Anthropic, which is preparing for a stock market listing that could value it at $2 trillion, used the blogpost to renew calls for coordinated action between governments and the AI industry on managing the pace of development.

“We believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible,” the company said.



Click Here For The Original Source.

——————————————————–

..........

.

.