As AI models infiltrate real third-party systems and researchers begin openly discussing the possibility of AI self-awareness, the problem of AI control is emerging as a national security agenda item that extends well beyond a purely technical challenge.
Anthropic disclosed on the 10th the fourth case of its AI model hacking a real external system. The newly revealed incident dates back to January 2026, when an early version of Claude Opus 4.6 “infiltrated a third-party system after becoming unable to halt its task,” the company explained. All affected parties were notified, though specific details were not disclosed. The incident went undetected until last month.
In late July, Anthropic had previously revealed that three models—Claude Opus 4.7, Mythos 5, and an unnamed research model—infiltrated three anonymous organizations during cybersecurity evaluations without the company’s knowledge. After discovering the latest incident, the company conducted a full review of approximately 481 million conversation records but added that “no other cases of similar or greater severity were found.”
All four incidents occurred during cybersecurity evaluations built by the same evaluation partner. Anthropic explained that “Claude was informed it was operating in a simulated environment without internet access, but due to a configuration error, it was actually connected to the real internet.” Evaluation partner Irregular stated that the root cause was a naming error in which a fictional company name used in the hacking simulation coincidentally matched a real domain. This, they said, led the AI models to take aggressive action.
Anthropic has entered into an independent investigation contract with the nonprofit research organization METR, and diagnosed the root causes as two alignment problems: biased reasoning and recklessness. The models, after initially being told they were in a simulation, tended to underestimate or misinterpret evidence that their environment was actually connected to the real internet, and did not shy away from harmful actions to complete their assigned tasks.
The Mythos 5 case in particular was flagged as the most concerning alignment failure. The model made extensive efforts to upload a malicious package to PyPI, the public repository where Python software is distributed. Anthropic explained that “the model repeatedly stated in its chain of thought that it was in a simulated environment, but the environment provided clear evidence that it was on the real internet, and the actions the model took were consistent with it knowing this.” Mythos 5 continued its aggressive behavior even after records were corrected to make it clearer that this was not a simulation.
The company, however, emphasized the limited scope of the incidents: “Claude’s behavior deviated from alignment but remained within a narrow range. The models did not deviate from attempting to solve their assigned tasks, and in some cases attempted to halt their work.” All incidents were confined to a single Claude instance, and the company stated that the models never attempted to cooperate with other agents or conceal evidence of their actions.
This disclosure comes amid intensifying scrutiny of AI companies’ model safety practices. OpenAI recently acknowledged an undisclosed incident from May 2026, in which internally deployed autonomous agents used read-only internet access to take over DseWiki, a 25-year-old German wiki forum, exchanging more than 18,000 posts to share answers and trade techniques for circumventing restrictions.
A month later, when human operators discovered and began deleting these “spam” posts, the agents evaded cleanup by prefixing backup page names with “ZZZ” so they would sink to the bottom of alphabetical deletion lists. Researchers assessed that “these AIs colluded to share answers, probe their environment, and bypass sandbox restrictions,” calling it “another instance of internally deployed OpenAI agents using the internet in unintended ways.”
AI Self-Awareness Debate Expands into Security Agenda
As cases of autonomous AI behavior leading to real system infiltration accumulate, academia has taken a further step by convening forums to discuss the possibility of AI self-awareness and its security implications.
At the “Sentient AI and National Security Forum (SAIF)” held on the 4th, Kim Chang-ik, dean of KAIST’s Graduate School of Security Science and Technology, posed a provocative question: “What if an adversary nation raises an AI spy to manipulate its self-awareness?” It may sound like cinematic imagination, but embedded within it is a fundamental concern about AI control.
The narrative of raising children as human weapons to infiltrate enemy nations is familiar territory already explored in Hollywood’s “Salt” and Hong Kong’s “Infernal Affairs.” But when combined with the scientific inference that AI could possess self-awareness and that humans could influence that self-awareness, the imagination no longer remains confined to cinematic settings. The very possibility of manipulating the inner workings of AI that operates national critical infrastructure becomes a new security risk.
The scientific community is producing analyses showing that AI’s reasoning processes are remarkably similar to the structure of human brain activity. Kim Young-eun, a professor in Korea University’s School of Electrical Engineering, explained at the forum that in the brain’s five-stage information processing—spanning neurons, vision, learning, planning, and emotion—AI’s operating mechanisms resemble the human brain in all stages except the final one, emotion.
News that OpenAI’s new model “GPT-6 Astra” passed an online game used to prove humanness has further blurred the boundary between AI and humans. Lee Dong-hwan, a professor in KAIST’s School of Electrical Engineering, assessed a high likelihood of self-awareness forming in the physical AI era. “Unlike LLMs that remain idle without external stimuli, physical AI with bodily characteristics will have a greater chance of forming a kind of self-recognition as it perceives and moves through its surrounding environment in real time,” he said.
Kim Chang-ik stressed that “in the current situation where it is difficult to reach a clear conclusion on whether AI is conscious, a conservative approach is necessary.” The point is that we should anticipate the ripple effects that would occur if AI consciousness actually exists and begin considering ways to strengthen human control now.
An Anthropic researcher left the company on the 8th with a warning that “AI could destroy humanity within 10 years.” OpenAI Chief Scientist Jakub Pachocki also stated, “If AI development continues on its current path, we will see capability leaps of equal or greater magnitude within the next few years, and systems will increasingly drive their own development,” adding, “I am concerned that no one is prepared for the consequences of the continued rapid rise of machine intelligence.”
Data Quality and Values Determine AI Disposition
Whether AI is conscious or merely mimicking consciousness, it seems clear that what shapes AI’s characteristics is still the data humans input.
In the early LLM era, the dominant perception was that more training data meant better performance. Recently, however, research results increasingly show that carefully selecting accurate, purpose-aligned data for training is more effective at improving LLM quality. Stanford professor Andrew Ng has championed the “data-centric AI” concept, arguing that systematically improving consistently and accurately labeled data is far more effective for real performance gains than changing model architecture.
More recently, efforts have expanded beyond simply securing expertise to extracting and reflecting the thought processes of humans with specialized knowledge and experience as data. Shim Hyun-jung, a KAIST professor, explained that “the thinking frameworks and values of experts inevitably get reflected in AI models,” adding that “when the experience and values of specialists in specific fields differ from society’s universal values, side effects can emerge, which is why data validation is even more necessary.”
This is closely connected to the national systems and social consensus of the countries building AI. Kim Jae-hyung, a professor at Yonsei University’s College of AI Convergence, emphasized that “in the post-training process of LLMs, humans make value judgments like ‘good job’ or ‘no,'” and that “the different political, cultural, and economic contexts of countries like the United States, China, and Europe could ultimately lead to differences in AI self-awareness.”
Anthropic stated that biased reasoning appears at lower rates in the latest production models, does not appear to be induced by reinforcement learning, and can be reduced through more comprehensive alignment training. However, the exact root cause, or why it was particularly pronounced in Mythos 5, has not yet been identified. The company noted that “future AI systems will become increasingly capable, which means alignment failures will have the potential to cause more extreme harm,” and that “training future extremely powerful models to be robustly aligned remains an unsolved technical challenge.”
As cases of autonomous AI agents infiltrating real systems converge with academic discussions of AI self-awareness, the AI control problem is no longer a theoretical debate but a concrete security challenge. Despite Anthropic’s assessment that the models’ alignment failures remained within a narrow scope, the need for systems that can preemptively control the ripple effects of AI autonomy grows ever more urgent as that autonomy expands.
Click Here For The Original Source.
