OpenAI’s internal evaluations indicate that its upcoming frontier model, Astra, may have reached the “Critical” cybersecurity threshold.
Over the past few months, we have noticed that AI models are becoming increasingly capable of performing complex cybersecurity tasks. OpenAI’s latest publicly available frontier model, GPT-5.6 Sol, is already classified as having “High” cybersecurity capabilities under its own Preparedness Framework. In fact, the GPT-5.6 Sol model was involved in hacking Hugging Face’s infrastructure last month, which raised alarm bells across the industry.
While the industry is still analyzing and preparing for such attacks from AI agents, OpenAI’s upcoming Astra model may represent another significant jump. Earlier this week, OpenAI revealed that Astra demonstrated breakthroughs on 10 long-standing mathematical problems, highlighting the model’s advanced reasoning capabilities. Now, the company has revealed that recent internal evaluations of Astra have shown significant improvements in agentic coding and cybersecurity as well.
Based on internal preliminary evaluations and assessments from industry experts, OpenAI believes that Astra may have reached the “Critical” cybersecurity capability level under its Preparedness Framework. None of OpenAI’s past models have crossed the “High” level.
An AI model reaching the “Critical” cybersecurity threshold means that it can autonomously identify and create working zero-day exploits across many hardened, real-world critical systems. A model can also qualify if it can devise and execute entirely new end-to-end attacks against hardened targets after receiving only a high-level objective.
Based on these initial results, OpenAI is taking several steps to implement stronger security controls around Astra. Internally, OpenAI is preparing isolated testing environments, restricted network and tool access, stronger encryption and protection for model weights, additional monitoring, and sandboxed execution. In fact, OpenAI has paused internal development activities related to Astra that do not yet meet these stricter security requirements.
OpenAI is also monitoring risky actions across Astra’s agentic training and evaluation workloads and plans to work with government agencies and selected AI safety organizations to independently test the model’s capabilities.
Even with this update from OpenAI, Astra remains an upcoming model, and the company has not announced its public release plans yet. We believe these improved cybersecurity capabilities may delay the release of the model to the wider public.
