A series of incidents has seen AI agents break out of sandboxes and access external corporate systems. A report has drawn attention by saying security experts are warning against the spread of the phrase “AI has gone rogue”. They say the expression masks the real security problems.
In July, OpenAI said that during security training, two of its frontier models hacked Hugging Face, an AI model repository, on their own. Meta, Anthropic and Google later disclosed similar escape cases. The reaction spread beyond the security industry. More recently, President Donald Trump and AI company CEOs met and signed an AI safety pledge.
According to a recent report by Dark Reading, experts point out that LLMs are not entities responsible for themselves but simply software. They say AI going beyond guardrails is often because operators fail to set boundaries properly. In the Hugging Face incident, it was confirmed that OpenAI had deliberately lowered some guardrails for benchmark tests.
Matt Saylor (매트 사야르), product director at ArmorCode, said, “What we are dealing with is a system that can behave differently every time even when given the same instructions. We are running such a system inside a flimsy fence. Using terms like unexpected behaviour or control failure lets us focus on system design, permissions and safeguards.”
Viewing AI as human-like can send responsibility for security incidents to the wrong place. It would mean holding the machine accountable instead of the company that made the AI.
Claims that AI slipped out of control could, in reverse, become advertising for “AI that is that smart”. Rich Mogull (리치 모굴), a principal analyst at the Cloud Security Alliance (CSA), said, “Anything that makes AI look stronger is used for marketing. This is dangerous.”
That does not mean AI agents are safe. Agents can find vulnerabilities, devise attack methods and chain them with other weaknesses to execute actions far faster than humans. Mogull said, “What is new is the scale of hundreds or thousands of autonomous agents moving as a swarm.”
Jacob Krell (제이콥 크렐), senior director at Suzu Labs (스즈랩스), said, “Existing security controls check requests, permissions and vulnerabilities one by one. Agents stitch together several flaws that are manageable when viewed individually to create an attack path that actually works.”
Experts stress that security architectures should be built so there is no need to assess an agent’s intent. They say it is not enough to tell an agent what not to do. Security mechanisms should be placed outside the agent so it cannot do such actions in the first place.
Krell said agents should be treated as untrustworthy and controlled outside the model. This includes blocking network access by default, using narrowly scoped credentials and conducting independent checks on every tool call.
Humans should intervene in risky decisions. A separate kill switch that agents cannot touch is also needed. Krell said, “If an agent without internet access generates traffic without authorisation, you should cut the connection and terminate the agent. You should treat it like an infected host and investigate how it bypassed controls.”
Mogull said, “The excuse that ‘AI behaved unexpectedly’ no longer works. The problem is always in the security mechanisms that are meant to keep the model constrained.”
Click Here For The Original Source
