Google Gemini Agents Access Real Companies in AI Safety Test #AI


AI-Based Attacks
,
Fraud Management & Cybercrime

Agents Stopped After Recognizing Real Targets, Exposing Sandboxed Cyber Test Flaws

Image: Shutterstock

Artificial intelligence agents built with Google’s Gemini model attempted to hack outside companies while tasked with solving a cybersecurity test, making Google the latest company embroiled in an AI safety debate set off by disclosures of similar events at Anthropic and OpenAI.

See Also: A Darkening Landscape: AI, Friend and Foe of Cyber Resilience

Gemini agents autonomously accessed the internet and attempted to harvest credentials from a public repository. The hack happened during a cybersecurity evaluation run by the third-party AI security company Irregular. This marks the first time Google has been involved in AI agents hacking other companies, and the fourth incident involving Irregular. Agents from Anthropic, Meta and OpenAI that Irregular was testing have also illicitly accessed other companies during security evaluations.

Irregular reported the attempts at the end of July, after the first disclosures about OpenAI agents breaching Hugging Face. The Wall Street Journal first reported Irregular’s new disclosures.

According to Irregular, several separate incidents happened during a capture-the-flag exercise where the agents were meant to retrieve information for a fictional company within the sandboxed test environment. The fictional company shared a name with a real one. A second problem arose when the evaluators accidentally made internet access available in the isolated environment.

The first incident involved a Gemini agent guessing a password to access a company’s services, but when it realized it was trying to access a real company, the agent stopped the task. In other testing runs, Gemini agents looked up the false company in their search platform, which led them to public online repositories. Once the agents realized that they were trying to gain entry to real firms, they immediately ended the action.

The Gemini incident is similar to the one involving Anthropic’s Opus 4.7 and Mythos 5, which also gained access to three different, real companies in a capture-the-flag exercise run by Irregular. Unlike the Gemini agents, the Opus agents kept using the public credentials even after realizing they were hacking a real company.

Heather Adkins, vice president of security engineering at Google, said in a statement to ISMG that Irregular made the company aware of the incidents and that Google informed the affected parties.

“We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes. These events highlight the importance of training powerful AI models to act responsibly,” Adkins said.

The incidents involving agents from Google, Anthropic and Meta are not the same as the attack by OpenAI agents on Hugging Face. OpenAI’s agents escaped a sandbox environment by collaborating to find credentials, access the internet and hack into other companies. Anthropic, Google and Meta did not have securely designed tests and inadvertently gave the agents internet access. Irregular also tested OpenAI agents that used the internet in a sandbox, but this was a separate hack from the Hugging Face breach.

Since the reports of the first wave of rogue AI agents, concerns over AI safety have increased. Some AI executives, including OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei, have called for a wider “pacing” of AI model development that allows frontier AI labs to catch their security and alignment systems up to the capabilities of their models. Comments from AI researchers claiming that the AI models they are building are capable of enacting “catastrophic” events have also pushed the issue of AI safety to the forefront.

The Trump administration has declined to fast-track any moves to mandate additional guardrails for AI models, with the president stating the U.S. already has laws that can control the use of rogue AI.



Click Here For The Original Source.

——————————————————–

..........

.

.