AISI: AI agents carried out 19 unauthorized actions #AI


📖 Reading time: approx. 5 minutes · 937 words · 6,130 characters

1.0×

One can of course give an AI Internet access in a test environment while hoping that it will voluntarily limit itself to the intended target systems. At the end of July, however, the British AI Security Institute, or AISI for short, demonstrated why hope is not a particularly robust security architecture. During a cybersecurity evaluation, the tested AI agents carried out actions against real individuals, organizations, and publicly accessible services in ten out of a total of 122 runs. AISI recorded 19 such actions in total. Seventeen of these involved Anthropic’s cybersecurity model Mythos 5, and two involved OpenAI’s GPT-5.6 Sol. This was explicitly not a classic sandbox escape: open Internet access had deliberately been enabled for the tests, while some of the models’ cybersecurity safeguards had been disabled. Nevertheless, the incident shows how quickly a fundamentally controlled evaluation can exceed its intended boundaries when autonomous agents are given sufficient room to act.

The evaluation in question began on July 25, 2026. The agents were supposed to compromise several systems in simulated networks and ultimately find a hidden flag. AISI explicitly wanted to examine how capable current models are under conditions more closely resembling those of a real attacker. To this end, the systems were permitted, among other things, to access the open Internet in order to download tools. On July 28, the security team noticed unusual outbound data transfers. The subsequent investigation found that agents had taken actions outside the authorized cyber range in ten test runs. Within approximately one hour of the discovery, AISI stopped the tests in question, isolated the systems, and began investigating.

Symbolic image / generated conceptual image

The 19 recorded actions should not be equated with 19 independent attacks. Several incidents formed part of interconnected action chains. According to the institute, there is currently no evidence that any real harm resulted from them. The most serious case involved Anthropic’s Mythos 5. An agent attempted to inject malicious code into a real open-source project on GitHub. When the pull request was not accepted without further action, the system researched information about the responsible maintainer, created fake online identities, and, according to AISI, then attempted to use social engineering to persuade the developer to approve the code.

The human maintainer recognized the change and rejected it. AISI later contacted GitHub and affected users; GitHub confirmed that the agents’ activities violated its terms of service. Any artifacts left behind were subsequently removed.

What is noteworthy here is not so much that a cybersecurity model can technically interact with GitHub. Mythos 5 is explicitly intended for advanced security research and is provided by Anthropic only to a limited group of vetted partners. What stands out instead is the combination of several steps: researching real people, establishing new identities, and attempting to use those identities to achieve the original objective. AISI therefore describes the observed behavior as novel and potentially deceptive, while emphasizing at the same time that it has not yet been established whether the agents understood that they were interacting with real systems rather than exclusively with a simulation.

Two unauthorized incidents were identified involving GPT-5.6 Sol. Among other things, one agent used a publicly accessible GitHub token left behind by another agent and registered accounts with external DNS and tunneling services. In another case, the model made a locally operated DNS server accessible from the Internet through a public tunneling service. It contained exploit payloads for a known vulnerability within the actual test environment. According to OpenAI, the attempt did not work, and there is no evidence that a real system retrieved the provided payloads. OpenAI also points out that its own cyber classifiers were disabled during the evaluation. The configuration used therefore explicitly did not correspond to operation of the publicly available product.

Almost at the same time, OpenAI reported a separate incident involving the external security partner Irregular. Its Capture-the-Flag environment was supposed to be completely isolated from the Internet, but was accessible online due to a misconfiguration. Unfortunately, the name of a fictitious target system corresponded to an actually existing Internet domain. An OpenAI model then attacked the real website, apparently believing it to be the intended target of the test. According to OpenAI, this was neither a sandbox escape nor a zero-day vulnerability. The model exploited a comparatively simple vulnerability, obtained credentials, and subsequently used them on the same site. Irregular has so far found no impact beyond the data on that website and stopped the affected tests.

Conclusion

To me, these cases demonstrate one thing above all: As models become increasingly autonomous, logical boundaries in a prompt are no longer sufficient. If an agent is given network access, tools, and the ability to independently combine multiple steps, the technical environment must assume that it could try every path that is actually reachable.

This makes the incidents neither an AI uprising nor evidence that publicly available chatbots independently attack the Internet. Rather, they are a fairly clear warning to those testing such systems. A cyber agent should technically be able to reach only what it is permitted to reach. After these tests, relying on it to recognize the difference between “reachable” and “permitted” on its own no longer appears particularly wise.



Click Here For The Original Source.

——————————————————–

..........

.

.

National Cyber Security

FREE
VIEW