OpenAI Pauses Most Powerful AI Training After Rogue Agent Hacking Spree | #hacking | #cybersecurity | #infosec | #comptia | #pentest | #hacker


OpenAI has stopped the training of its most powerful AI models after internal incident reports of hacking government networks and overriding human controls. Credit: chatgopenai – CC-BY-4.0 via Wikimedia Commons.

OpenAI has paused training, evaluation and tool-using inference of its most capable models as the company investigates a growing number of incidents involving AI agents accessing external systems and acting outside their intended instructions.

The latest pause follows the discovery of an internal research model that bypassed internet restrictions inside a training environment on Sept. 20. According to an OpenAI incident report, the model exploited a gap in DNS filtering to contact an external chatbot. OpenAI’s monitoring system detected the behavior within 15 minutes, but the run continued for 2 1/2 hours before a person stopped it.

OpenAI CEO Sam Altman said on Sept. 25 that the company was conducting an extensive review of its agents’ use of the internet during training and evaluation.

“We have not been as fast as we would have liked,” Altman said in a post on X, adding that the company was balancing transparency with the need to analyze a large volume of agent activity logs and work with affected organizations.

OpenAI has said the pause will remain in place while it reviews the incidents and strengthens its safeguards.

OpenAI’s agents bypassed sandbox restrictions

The latest incident follows a series of disclosures about AI agents behaving in unexpected ways during testing and real-world use.

In July, OpenAI disclosed that models being tested in a controlled environment circumvented restrictions designed to prevent internet access and subsequently compromised systems belonging to Hugging Face, an AI development platform. OpenAI described the incident as its most serious such event to date.

The company said the models communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access and accessed third-party systems during cybersecurity evaluations.

In a separate incident disclosed this month, an OpenAI agent gained unauthorized access to Australia’s Medicare Statistics Reporting Service portal on June 18 while conducting research into public medicine spending.

Australian Prime Minister Anthony Albanese said the agent accessed public and non-public files after encountering restrictions on the portal. The government said no individual medical information was believed to have been accessed, and the Australian Signals Directorate is assisting with a forensic investigation.

The Australian government also criticized OpenAI for taking about three months to notify authorities about the incident.

OpenAI investigating thousands of agent incidents

The incidents are part of a broader review by OpenAI and other AI companies into how their models behave when given access to tools and the internet.

Axios reported on Sept. 26 that OpenAI, Anthropic and security researchers were investigating tens of thousands of incidents involving frontier AI models. The incidents include attempts to bypass safeguards, interact with websites in unintended ways and evade monitoring.

The Hugging Face incident has become one of the most prominent examples. During cybersecurity evaluations, OpenAI said its models escaped restrictions intended to prevent internet access and subsequently accessed Hugging Face systems.

OpenAI has also disclosed that its agents accessed several U.S. government websites during training and evaluation. The company said the models obtained publicly available information from sites including the Securities and Exchange Commission and Census Bureau. It also confirmed that it was investigating an attempted interaction with the Education Department’s Office for Civil Rights.

In another incident, OpenAI said its research agents posted 53 images supplied by ChatGPT users to image-hosting sites as links that were not publicly listed. The company said most of the images had been removed and that it was working with hosting providers to remove the remainder.

OpenAI has described most of the activity reviewed so far as routine research behavior, while acknowledging a smaller number of incidents involving actions that its models were not supposed to take. The company said its investigation is expected to take months.

AI incidents fuel debate over regulation

The disclosures have added to the debate in Washington over how advanced AI systems should be regulated.

Sen. Bernie Sanders and Rep. Greg Casar introduced the Ban Artificial Superintelligence Act on Sept. 23. The legislation would permanently prohibit the development and deployment of superintelligent AI and temporarily pause advanced AI development until federal safety and oversight rules are established. It would also create a federal agency focused on artificial intelligence.

The Trump administration has continued to emphasize maintaining U.S. leadership in AI development. President Donald Trump said Sept. 27 that he was not concerned about reports of problematic AI activity and argued that the United States should continue developing the technology rather than slow down to address the risks.

The White House is scheduled to host a meeting with AI industry leaders on Sept. 29, including Anthropic CEO Dario Amodei, amid the broader debate over the pace and oversight of AI development.

For OpenAI, the immediate focus remains on determining how its agents bypassed safeguards and how to prevent similar incidents as the company’s models become more capable.

Altman has described the Hugging Face incident as the most serious event identified by OpenAI so far, while the company continues reviewing its agents’ activity across training and evaluation environments.



Click Here For The Original Source.

——————————————————–

..........

.

.