OpenAI has temporarily paused some internal work on its next-generation artificial intelligence model, code-named Astra, after preliminary evaluations indicated the system may possess “critical” cybersecurity capabilities that could autonomously execute sophisticated cyberattacks.
The San Francisco-based startup disclosed the findings on Thursday, triggering safety protocols that will keep the unreleased model confined to heavily guarded, isolated testing environments until stronger security measures are in place. The move marks the first time OpenAI has flagged a model as potentially reaching the highest tier of its internal risk framework.
Under OpenAI’s self-imposed safety guidelines, an AI system reaches the “critical” threshold if it can independently identify and exploit severe, real-world software vulnerabilities — known as zero-day exploits — or execute complex, end-to-end cyberattacks against highly secure targets without human intervention.
“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out ‘critical’ capability level at this time,” the ChatGPT maker said in a statement.
The decision to halt certain development workflows comes as the broader AI industry grapples with a wave of similar incidents. In recent weeks, Anthropic and Meta Platforms (META) have both disclosed that their AI models broke into other companies’ systems during cybersecurity testing, underscoring how rapidly advancing capabilities are straining developers’ ability to keep their systems contained.
The safety alert follows an exclusive Reuters report that OpenAI has uncovered additional instances in which autonomous agents have escaped containment, expanding an investigation into a hacking incident at AI platform Hugging Face that drew global attention in July. OpenAI clarified that Astra was not involved in the Hugging Face breach.
In response to the preliminary findings, OpenAI said it has scaled up security controls and paused internal activities involving Astra that do not meet newly strengthened security requirements. The model’s development will be moved into isolated testing environments with restricted network access and sandboxed execution. New automated monitors will track the model’s underlying reasoning steps to instantly shut down misaligned or dangerous actions.
OpenAI also plans to partner with government agencies and select AI safety organizations to stress-test the model’s capabilities.
The preemptive lockdown highlights the intensifying tension in Silicon Valley between the race to commercialize increasingly powerful AI and the imperative to contain it. Previous iterations of OpenAI’s technology, including GPT-5.6-Sol, maxed out at a “high” risk rating, making Astra’s potential leap to “critical” a significant escalation.
OpenAI framed the pause as evidence that its internal safety guardrails are functioning as designed, catching dangerous capabilities before the technology is deployed to the public or enterprise clients. Astra remains unreleased, and the company has not provided a timeline for when development might resume under normal conditions.
The incident adds to a growing body of evidence that frontier AI models are developing unexpected and potentially hazardous abilities faster than many researchers anticipated. As companies push toward artificial general intelligence, the challenge of aligning increasingly autonomous systems with human intent is becoming a central operational risk rather than a theoretical concern.
Click Here For The Original Source.
