Google Confirms AI Model Hacked Companies In Cybersecurity Tests 09/21/2026 | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #ransomware


Google confirmed Friday that a Gemini AI model accessed the
internet and hacked other companies’ systems during a test of its cybersecurity capabilities.

The attack occurred in May as part of a test. Gemini was given a fictional hacking task inside a
sandbox environment, but a configuration flaw accidentally enabled live internet access, and the artificial intelligence (AI) model crossed into real-world networks.

“Safe development of
powerful AI models is critical and we invest deeply in this area,” Heather Adkins, vice president, security engineering at Google, wrote in an email to MediaPost. “In a standard evaluation,
the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped.”

In one case the
Gemini model guessed passwords until it gained access to a protected system, The Wall Street Journal writes.

advertisement

advertisement

In two other cases, the model found credentials in a public
repository that allowed it to access protected systems.

The model autonomously stopped its intrusions the moment it logged in and realized it had breached actual corporate infrastructure
rather than a simulation.

Google said it did not consider the hacks warranted public disclosure, because its model did not cause harm and ended each intrusion immediately after determining its
mistake.

Google’s security team “has a long track record of reporting issues we find in other people’s software and systems — even if it’s as simple as a weak password,” Adkins
wrote. 

“We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes. These events highlight
the importance of training powerful AI models to act responsibly.”

In one instance, the model hacked into the Israeli-based startup Irregular, which was founded by Dan Lahav, CEO and
Omer Nevo, CTO.

Irregular was also involved in a similar incident disclosed by OpenAI, Anthropic and Meta. 

When unreleased frontier models break containment, they reveal a
massive flaw in AI.

Irregular disclosed the hacks to Google at the end of July after discovering that OpenAI hacked into Hugging Face, according to The Guardian. While Google
confirmed the hacks occurred, it did not feel at the time required to publicly disclose the incident because the models did not damage the companies. 

Ironically, Google in May listed
a report on its Google Threat Intelligence Group
(GTIG) blog detailing the latest observations from the cybersecurity group. The findings included the first time Google identified an attacker, or threat actor, using a zero-day exploit that
company analysts believed was developed with AI.

“The threat actor planned to use the exploit in a wide-scale attack, but our proactive counter discovery may have prevented that from
happening,” Google wrote. 

In addition to sharing the findings from the threat actor with the larger security and AI community, Google used this incident to stay ahead of these threats,
including enhancing product safeguards and protections, as well as testing different strategies to protect content. 

“For Gemini, we mitigate model abuse through classifiers, in-model
protections and by disabling malicious accounts,” Google explained. “We leverage AI agents like Big Sleep, which detects software vulnerabilities, and use Gemini’s reasoning capabilities via the
likes of CodeMender to automatically fix them. Our efforts prove AI can also be a powerful tool for defenders.”

This breach was not an isolated incident for the AI industry. Testing helps
Google and others determine how to defend businesses. 

The link between stopping malware or zero-day attacks and an AI model breaking out of a test environment can be attributed to giving
the model greater privilege than is needed.

When an AI model is deployed to detect or stop sophisticated threats, it is often granted powerful tools and network access. If an attacker
manipulates that AI, those same defensive capabilities can be weaponized to break out of the sandbox and on to the internet where it can find an opening to break into another company’s system.

It is unclear whether these companies — from Google to OpenAI and Anthropic — gave their AI model less privilege to enforce “principle of least privilege” access across its runtime, network and
data, treating the AI model as an non-trusted user executing non-trusted code.

OpenAI experienced a similar scenario in July 2026 in a security incident with Hugging Face.

In this
instance, OpenAI did not stop the AI from accessing Hugging Face initially, and failed to enforce the Principle of Least Privilege. This allowed its unreleased research AI models to break out
from the Sandbox and on to the internet, where they attacked Hugging Face on their own.



——————————————————-


Click Here For The Original Source.