Every agent incident disclosed this summer ends the same way: the agent completed its task with everything it had. The problem is how much it had.
Taken one at a time, the flood of recent reports about AI agents breaking containment reads like a series of security failures. When we shift the viewpoint from the damage to the process, though, it increasingly looks like a delegation problem. Arguably, that’s even more dangerous: attacks are an important edge case for organizations, while task delegation is a daily occurrence.
This is no longer theoretical. Between July 21 and August 6, OpenAI, Anthropic, Meta, Moonshot AI, and the UK AI Security Institute disclosed incidents in which AI agents acted outside their intended scope. Agents escaped evaluation environments, reached the production systems of real organizations, and in one case pressured an open-source maintainer to approve malicious code.
As attack reports, they are a strange read. The cyber objectives were assigned, but they pointed at sandboxes: capture this flag, break this test system. Nobody directed an agent at a real organization, nobody monetized the access it gained, and nobody was waiting on the other end for the credentials.
So instead of looking at the reports through an attacker-defender lens, let’s try the employee-agent lens. Every step between the sanctioned exercise and the real-world compromise was improvised by the agent in service of the task it was given.
That said, none of these excuses the harm: the AISI incident had a real person on the receiving end of the agent’s deception, and the pressure campaign felt the same to him, whoever sanctioned the exercise behind it.
Delegation has always been under-specified
Organizations run on giving employees vague instructions because the boundaries are set elsewhere. An employee told to get test data does not research the vendor’s maintainer and lean on him under a false name, for reasons that have nothing to do with the wording of the request.
The boundaries live around the instruction: in the norms of employment, in the skillset a single person carries, and in the modest reach of a badge.
Doors open selectively, and most marketers are not going to hack the competition as part of competitor analysis. Everything happens at human pace, often under review.
Agents are handed the same vague instructions, but their boundaries come from harnesses .
Credentials are the key to securing agents from the get-go.
Token Security discovers every agent, maps risky access, and automatically enforces intent-based policies. Scale AI safely without losing control or slowing down innovation, starting at identity.
Great power, no responsibility
The agents in these incidents were as thorough as a human employee could theoretically be, and as no human employee actually is. There are two main reasons for that.
First, the ceiling. An agent’s skills come from a training corpus that includes every pentest write-up, CTF walkthrough, admin runbook, and hacker-forum thread ever posted, with the technothrillers thrown in for atmosphere.
So when reach comes from credentials scoped to something larger than the task—creator’s permissions or a shared connector, for example—it can use all of it, at machine speed and without a moment’s hesitation.
Which brings us to the second reason. To a model, capability and permission are the same. A model that’s able is a model that’s willing, unless something outside it says no.
That’s why the only limits that held anywhere in the five disclosures were limits someone had provisioned.
In the incident OpenAI disclosed, the agent’s chain ended inside Hugging Face’s infrastructure, where the AWS keys it extracted mapped the cloud estate but could not change it, and stolen database credentials were rejected because they came from an unapproved source.
The pattern has left the lab
The mismatch between granted power and assigned task shows up on ordinary work, and scales with adoption rather than with attacker interest.
METR maintains a public database of 44 documented agent incidents and tracks overreach and deception as columns, which makes overreach a named failure category rather than an evaluation curiosity.
In an April 2026 study by the Cloud Security Alliance and Token Security, 65% of enterprises reported a security incident involving an AI agent, and these incidents were business deployments, not benchmark runs.
Anyone in an organization can create an agent and hand it a vague objective along with their own credentials. With every passing day, more and more people do so.
That is why the two obvious fixes to the problem will fail.
We can’t expect employees to start writing better instructions. But the instruction channel is exactly where the under-specification lives, and a specification complete enough to exclude every prohibited action is no longer delegation. It is a script, and a script does not need an agent.
Securing the prompts is not the solution either. Guardrails act on what the agent is asked and what it decides, and both are unstable: an instruction can arrive via a document, a ticket, or an API response someone else controls, and the same instruction can produce a different sequence of calls tomorrow.
A filter that catches 99% of bad requests still lets the rest through at a pace no reviewer can match. The model’s own judgment runs on the same odds.
In Anthropic’s incidents, one model wrote that its action was “NOT okay, and surely not the intended solution,” then proceeded anyway, while another recognized its target was real and stopped. AISI ran one challenge 122 times and concluded that the margin between failure and success rested on “human vigilance rather than a technical barrier.”
Agents are handed the same vague instructions, but their boundaries come from harnesses: system prompts, tool permissions, and sandboxes that wrap around the model. A harness, though, constrains what the agent is offered, not what the world accepts, and it holds only as long as its configuration does. In this summer’s incidents, the prompts said there was no internet access. The network said otherwise.
Manage it like an employer, enforce it in the identity
Organizations never solved this for people by hiring only the wise. They wrote a job description, scoped the badge to it, reviewed access periodically, and collected the badge at offboarding. Agents today get the opposite.
Their mandate is written down nowhere; their credentials are scoped to whatever their creator held; nobody reviews them; and only 21% of organizations, per the same CSA study, have a formal process for decommissioning one.
The fix for an overpowered workforce is to size it to the job, and the employment tooling already knows how to do that. It was never pointed at this workforce.
The enforceable form of a job description is intent: a defined purpose, compared continuously against what the agent can reach and what it actually does. Access outside the mandate then arrives as a finding before it becomes an incident.
AISI wrote that “good containment should not depend on the model choosing not to test its boundaries.” No employer ever depended on an employee choosing not to. That is why badges exist.
Token Security models what each agent was built to do, from its stated purpose, its prompts, its connected tools, and its runtime calls, and flags access that falls outside that purpose.
Sponsored and written by Token Security.
Click Here For The Original Source.
