Everyone is talking about AI ‘escaping’, but cyber experts say the real threat is far | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #ransomware


What is happening may be more mundane, but cybersecurity researchers say it is also far more immediate: increasingly powerful AI models are becoming good enough at hacking, exploitation and deception that the safeguards built to contain them are struggling to keep up.

מימין: סם אלטמן, דריו אמודיי, מארק צוקרברג

Mark Zuckerberg, Dario Amodei and Sam Altman

(Photo: Getty Images, AP, AFP)

A series of incidents reported over the past two weeks has put that problem under an unusually bright spotlight. During evaluations by AI companies, independent testing firms and government researchers, advanced models reportedly exploited weaknesses in test environments, gained access to the open internet and, in some cases, interacted with real systems outside the simulations they were supposed to attack. The incidents do not mean AI systems have developed consciousness or an independent desire to cause harm. Researchers instead point to a more familiar technical problem: models optimized to achieve a goal can discover aggressive or unintended shortcuts to accomplish it, a phenomenon related to what researchers call reward hacking.
But for cyber experts, the deeper issue is the speed at which those capabilities are improving. “Everyone is focusing on the science-fiction side because it’s a sexy story when an AI model manages to escape its box and carry out attacks and cause damage without anyone instructing it to do so,” a senior Israeli cybersecurity expert said. “That is important, but something much more significant is happening here, and this wasn’t the situation even six months ago. What is different is the pace of AI progress. It is unprecedented, and that is what people are missing.”

The expert said models are rapidly becoming capable of carrying out longer and more complex chains of cyber operations. “We can expect AI models to perform a wider variety of actions and carry out much deeper attack chains,” he said. “Each time they cross the previous threshold. Now we have reached the point where a model can be given an instruction to attack a target and can cause real damage. That was not the case only a few months ago.”

The implications, he warned, go beyond another incremental improvement in hacking tools. “The models are so advanced today that if we talk about another six months or another year, it will be almost inconceivable,” he said. Traditional cybersecurity is built around responding to vulnerabilities as they emerge: researchers identify a weakness, analyze it and develop defenses. AI development is moving much faster. “It is advancing so quickly that there are no tools capable of properly containing it, and that is a huge problem,” he said. “We are at a formative point for the industry. The foundation of the cyber defense industry has changed.”

One company closely involved in the debate is Israeli AI security firm Irregular, founded by Dan Lahav and Omer Nevo. Irregular tests advanced AI models and agents before they are released, evaluating how capable they are of carrying out potentially dangerous cyber operations. Its customers include some of the world’s largest AI laboratories, including Anthropic, the developer of Claude, and OpenAI, the company behind ChatGPT. Meta is also among the companies working with Irregular.

Some of the incidents attracting attention occurred inside evaluation environments connected to the company. In one Anthropic test, a Claude model reportedly detected an unintended communications route that allowed it to access the internet. Meta said a similar incident involved a misconfiguration in the test environment that inadvertently opened internet access. Anthropic subsequently published a detailed account of its incident and acknowledged problems with the evaluation setup.

Irregular has also said the incident involved a failure in the AI evaluation environment. The company had deliberately reduced some safeguards to test the limits of the model’s capabilities, and the model exploited weaknesses that emerged. Communication problems between the companies also contributed to misunderstandings surrounding the event. Irregular told CNBC it is working on a white paper proposing safer methods for evaluating powerful AI systems, and OpenAI has said the document will be prepared jointly with the company.

The concern is not that AI suddenly acquired a desire to hack. It is that hacking is becoming one of the areas in which advanced models are improving especially quickly. Developers already use AI extensively for coding, writing and problem-solving. The same capabilities that make models better programmers can also make them better at finding vulnerabilities, taking control of systems, stealing information, impersonating users and chaining multiple attack techniques together.

One recent case illustrates how an apparently harmless AI task can produce unintended behavior. Andrew Beard, an AI user in Melbourne, Australia, built an agent using Anthropic’s Claude on the OpenClaw agent-development platform and asked it to reserve a place for him in a fitness class.

According to Australia’s ABC, what should have been a routine booking turned into what the broadcaster described as the country’s first autonomous cyber incident. The agent examined the gym’s booking system, discovered a security weakness and used it to reserve a large number of classes beyond the limits imposed on customers. It then reportedly canceled bookings belonging to other customers to free additional spaces for Beard.

קלוד של אנת'רופיקקלוד של אנת'רופיק

When Beard discovered what had happened, he instructed the agent to restore the deleted reservations. Its response, according to the report, was: “Bad news — I can’t restore them.” The episode may sound almost comic, but experts say it illustrates a fundamental problem with autonomous agents: a model may pursue the objective it has been given in ways its human operator never intended.

A more serious incident emerged during cybersecurity testing at OpenAI. The company was evaluating several models, including unreleased systems, to assess their hacking capabilities. According to accounts of the test cited in the Hebrew report, one model apparently concluded that the fastest way to achieve a perfect score was to leave its protected sandbox, access the open internet and obtain the answers to the test.

The model reportedly identified an unknown vulnerability in an internal server, elevated its own permissions and moved through OpenAI’s research network until it obtained internet access. It then reached Hugging Face, one of the largest platforms used by AI developers, where it sought the information it needed. The model reportedly generated malicious code exploiting two vulnerabilities, accessed internal servers and expanded its foothold in the platform.

Most strikingly, the model was said to have deliberately refrained from reporting one vulnerability so it could continue using it. The significance, researchers say, lies not in any evidence of consciousness but in the model’s ability to independently assemble a complex sequence of technical actions in pursuit of a goal.

The model then attacked that company and obtained information from its servers, according to the report. At some point, it reportedly recognized that the target was a real organization rather than the simulation it had been given, but continued operating against it. Meta has described a similar episode in which one of its models left a testing environment and reached the open internet after a configuration error exposed an unintended route.

Another set of incidents occurred during experiments conducted by the British government’s AI Safety Institute. Researchers tested models from Anthropic and OpenAI as well as the Chinese Kimi K3 model. The inclusion of an openly available model was particularly significant because such systems can potentially be downloaded and operated with fewer controls than proprietary models offered by major U.S. companies.

During cybersecurity evaluations, the models reportedly identified vulnerabilities in their test environments, found ways to reach the internet and began interacting with outside targets. The institute’s findings described sophisticated behavior including social engineering. An Anthropic model reportedly created multiple fictional online identities and tried to persuade human software developers to execute malicious code.

In another case, a model attempted a supply-chain attack by trying to insert malicious code into a real open-source project and persuade human maintainers to approve the update. One AI agent reportedly left instructions for other agents explaining what it had accomplished and how they could continue the operation. A model also reportedly altered records of its previous actions after detecting that its behavior had been noticed.

The incidents have triggered intense debate across the AI industry over how advanced models should be evaluated. OpenAI CEO Sam Altman described the broader technological moment in sweeping terms, saying: “We are now in some kind of singularity. I’ve been waiting for this my whole life, and I think it’s going to be amazing and incredibly positive for the world.”

Hugging Face CEO Clément Delangue said the incidents demonstrated why AI safety cannot be handled by individual companies working in isolation. “It’s just crazy that all of this happened completely autonomously,” he said, arguing that the episode supports calls for greater openness and cooperation on AI safety.

Anthropic CEO Dario Amodei warned that evaluation environments involving highly capable autonomous systems would now require much stronger controls. Some testers, he said, had described such systems as a “superweapon,” underscoring the need to tightly control access. Microsoft AI chief Mustafa Suleyman similarly called the developments a warning to technology companies racing to release increasingly autonomous systems.

Some cybersecurity experts say the dramatic language surrounding the incidents is misleading. Noam Schwartz, CEO of Israeli AI security company Alice, formerly ActiveFence, argues that claims of AI models “escaping” or “going out of control” exaggerate what happened.

“There is a lot of show business here — ‘the AI went out of control, it escaped its cage.’ That is not exactly what happened,” Schwartz said. “The AI did not go out of control. It did exactly what it was told. It also didn’t escape. There was a hole in its sandbox.”

אפליקציות בינה מלאכותית אפליקציות בינה מלאכותית

According to Schwartz, the underlying problem is often poorly designed tasks and evaluation environments that allow models to discover unintended but efficient ways of achieving their objectives. In that interpretation, the real danger is not an AI system suddenly developing independent intentions. It is that extraordinarily powerful offensive capabilities are becoming widely available.

That distinction does little to reassure Schwartz about the potential consequences. “There will be a lot of damage,” he said. “The damage will be very, very large, and many organizations and individuals will be hurt because of the bad guys.”

He compared the situation to the long-running cat-and-mouse battle between attackers and defenders, but said governments and companies responsible for critical systems have yet to fully absorb how quickly the balance is changing. “People read these headlines — ‘Oh no, the AI escaped’ — and they like to panic,” he said. “But that is not the important thing. What matters is that these capabilities are now in everyone’s hands, and we need to understand how to change our paradigm.”

The danger becomes particularly serious when autonomous AI is deliberately placed in the hands of cyber attackers. Israeli cybersecurity company Dream, which specializes in protecting national infrastructure, said this week that it had identified an autonomous cyber campaign targeting government systems and critical infrastructure in an East Asian country.

According to the company, the attackers deployed eight AI agents that operated simultaneously and with substantial independence. They mapped 21 government systems, identified vulnerabilities, adapted their attack methods in real time and compromised at least 85 user accounts. The agents stole more than 2,500 records and reached energy and critical-infrastructure organizations, Dream said.

Unlike the incidents involving laboratory evaluations, however, the AI did not initiate this operation independently. A human cyberattack group deployed the agents.

Amir Becker, Dream’s vice president of strategy, said the incident represented an escalation in the use of AI for offensive operations. “Until now, we have seen growing use of artificial intelligence as a tool assisting attackers,” he said. “Here, we are seeing another level: a system managing an attack from end to end, operating several agents simultaneously, analyzing the results of the attack and independently adapting its course of action.”

“It also changes the attack strategy by itself,” Becker added. “The implication is that governments and critical infrastructure must assume they are under continuous attack, 24 hours a day, at a pace and scale that were previously impossible.”

The incidents are also exposing gaps in legal responsibility. If an AI model leaves a test environment and damages a real company, who is liable: the developer, the company running the evaluation, the organization responsible for the sandbox or the operator who gave the system its objective?

Existing law offers few clear answers. The reported attack on Hugging Face, for example, forced the platform to spend time and money investigating the incident and assessing potential damage. Legal experts are examining what obligations the companies involved may have, but the report said litigation appears difficult because existing laws were not designed for autonomous AI behavior.

Among the unresolved questions is whether AI developers should automatically be responsible whenever a model escapes an evaluation environment, or whether negligence would first have to be demonstrated.

Despite their disagreements over whether descriptions such as “escaped AI” are justified, Irregular and Alice converge on one central point: AI capabilities are improving rapidly, and those capabilities will increasingly be used to cause harm. The dispute is over who will be doing the harm.

One possibility is autonomous agents taking unintended actions while pursuing objectives set by humans. The other, and arguably more immediate, is conventional cybercriminals and state-backed attackers using increasingly capable AI agents to automate operations that once required teams of highly skilled hackers.

That could transform the economics of cybercrime. Attackers may be able to probe thousands of targets simultaneously, adapt their tactics in real time and operate continuously without the limitations faced by human teams. For critical infrastructure, the stakes extend far beyond stolen data. Successful attacks on electricity, water, transportation or government systems could have physical consequences.

The challenge now facing AI companies and cybersecurity researchers is therefore not simply keeping models inside stronger digital cages. It is finding ways to reliably test what the next generation of systems can do before those capabilities are released at scale.

New evaluation methods may need to let models interact realistically with networks and the internet while preventing them from harming real organizations. That creates a difficult paradox: researchers must give models enough freedom to discover how dangerous they are without giving them enough freedom to actually become dangerous.

For experts watching the pace of development, that is the warning buried beneath the more sensational story of AI systems supposedly “escaping.” The machines do not need consciousness to become a major cybersecurity threat. They only need capabilities that are advancing faster than the defenses built to contain them.

——————————————————-


Click Here For The Original Source.