OpenAI says upcoming Astra model may approach its Critical cybersecurity capability threshold
OpenAI says preliminary internal evaluations of its upcoming Astra model are strong enough that it cannot rule out the Critical cybersecurity capability level under its Preparedness Framework, prompting stronger development and testing controls.
What OpenAI announced
On 7 August 2026, OpenAI said preliminary internal evaluations of an upcoming model called Astra showed major gains in agentic coding and cybersecurity. The company did not say Astra has definitively been classified at the Critical level. Instead, it said the evidence is strong enough that it cannot currently rule out the Critical cybersecurity threshold defined in its Preparedness Framework.
That distinction matters. The announcement is a precautionary capability assessment while benchmarking and expert review continue, not a final declaration that the model can reliably perform every task associated with the Critical category.
What “Critical” means in OpenAI’s framework
OpenAI’s updated Preparedness Framework separates advanced-risk capability levels into High and Critical thresholds. In its 7 August statement, OpenAI described the Critical cybersecurity threshold as involving capabilities such as independently finding and developing functional zero-day exploits across many hardened real-world critical systems, or devising and executing novel end-to-end attack strategies against hardened targets from a high-level goal.
OpenAI said earlier frontier systems, including GPT-5.6 Sol, had been assessed at the High rather than Critical cybersecurity threshold. Astra’s preliminary results therefore represent a potentially important capability transition if later evaluation confirms them.
Controls OpenAI says it is applying
OpenAI said it has increased robustness testing and strengthened the security environment around Astra. Measures described in the company’s announcement include isolated testing environments, tighter network and tool access, stronger model-weight protection and encryption, additional monitoring and detection, and sandboxed execution.
The company also said it has paused internal Astra activities that do not yet meet the strengthened security-control requirements and plans to work with relevant government agencies and selected AI-safety organizations on further testing. OpenAI also said third-party testing partners will receive recommended security controls for higher-risk evaluations and workloads.
Why the update is significant
Cybersecurity is a dual-use capability area: improvements can help defenders discover and remediate vulnerabilities faster, but the same underlying capabilities may also increase offensive potential if misused. OpenAI’s Preparedness Framework is intended to tie model-development and deployment decisions to measured capability thresholds and the strength of corresponding safeguards.
In its April 2025 Preparedness Framework update, OpenAI said systems reaching the High threshold require safeguards that sufficiently minimize severe risk before deployment, while systems reaching the Critical threshold require adequate safeguards during development as well. OpenAI’s May 2026 Frontier Governance Framework further described how the company maps preparedness practices to emerging regulatory requirements covering areas such as cyber offense, CBRN risk, harmful manipulation and loss of control.
What to watch next
The key next step is whether subsequent benchmarking, expert assessment and external testing confirm Astra at the Critical cybersecurity capability level or place it below that threshold. A final evaluation would be more informative than the current precautionary statement.
For security teams and AI governance practitioners, the announcement is also worth watching because it provides a concrete example of how a frontier developer says it changes internal controls when a model may cross into a higher-risk capability tier. Any future OpportunityAtlas coverage should distinguish between confirmed capability ratings, preliminary assessments and deployment status rather than treating them as equivalent.
This article is built from the source material below. Open the originals for full context and the latest updates.