OpenAI has discovered more instances in which autonomous agents have escaped containment.
| Photo Credit:
iStockphoto
OpenAI said on Friday it cannot rule out
that its upcoming AI model, Astra, has “critical” cybersecurity
capabilities, prompting the startup to pause some internal
development and trigger safety protocols.
Under OpenAI’s safety guidelines, a model reaches the
“critical” threshold if it can autonomously identify and exploit
severe, real-world software vulnerabilities, known as zero-day
exploits, or execute complex cyberattacks against highly secure
targets without human intervention.
Here are some details on Astra:
This follows an exclusive report by Reuters that OpenAI
has discovered more instances in which autonomous agents have
escaped containment as the company expands its investigation of
the hacking incident at tech firm Hugging Face that drew global
attention in July.
In the last few weeks, OpenAI, Anthropic and Meta
Platforms have disclosed that their AI models broke
into other companies’ systems during cybersecurity testing,
highlighting how advancing AI capabilities are straining
developers’ ability to keep their systems contained.
Preliminary evaluations over the past several days, along
with outside expert assessments, indicated Astra may be capable
of performing increasingly sophisticated cyber tasks
autonomously, OpenAI said.
“While we continue to benchmark and assess this model, our
preliminary evaluations indicate strong enough performance that
we cannot rule out ‘critical’ capability level at this time,”
the ChatGPT maker said.
In response to the preliminary findings, OpenAI said it
has scaled up security controls and paused internal activities
involving Astra that do not meet its newly strengthened security
requirements.
Astra’s development will be moved into isolated testing
environments with restricted network access and sandboxed
execution.
OpenAI also clarified that Astra was not involved in the
hack targeting the AI platform Hugging Face.
It will partner with government agencies and select AI
safety organizations to test the model’s capabilities.
Published on August 8, 2026
