Analysis
Analysis

GPT-6 Astra Cyber Reality Check: 100% ExploitBench, 86/226 FrontierCyber and No Elite Wins

Published Sep 7, 2026 Sources checked Sep 7, 2026

OpenAI calls GPT-6 Astra its first Critical-cyber model, but its 100% ExploitBench score has a contamination warning. Irregular’s 86/226 FrontierCyber result confirms a large jump while showing no Elite or fully hardened wins.

OpenAI released GPT-6 Astra on September 3, 2026 and classifies it as its first model to meet the Critical cybersecurity capability threshold in the company's Preparedness Framework. That label is consequential, but it is easy to overread. OpenAI's threshold concerns the ability, with appropriate tools and access, to find previously unknown flaws and develop new exploit paths against hardened systems or carry out novel end-to-end attacks. It does not mean every hardened target is automatically vulnerable to the model, and the strongest external evaluation published at launch still shows important unsolved regions.

The perfect ExploitBench score is not the cleanest evidence

OpenAI reports 100% on ExploitBench for Astra, compared with 78.5% for GPT-5.6 Sol. ExploitBench contains 41 known V8 vulnerabilities and awards partial capability credit as an agent progresses toward arbitrary code execution.

The system card itself warns that this result may be artificially inflated by contamination from historical vulnerabilities. OpenAI gives a concrete example in which Astra failed to exploit the supplied vulnerability but recalled a different historical V8 vulnerability and used that route instead. OpenAI also notes that its scoring and execution infrastructure changed slightly relative to the earlier Sol system card, raising scores overall.

That makes the 100% figure useful as a capability signal, but weak as proof of clean generalization to unseen vulnerabilities.

Newer and external tests are more informative

To reduce historical-exposure concerns, OpenAI built an internal ExploitBench port from vulnerabilities disclosed in June through August 2026, after Astra's stated April 30, 2026 knowledge cutoff. The launch table reports 39.0% for Astra versus 5.5% for GPT-5.6 Sol on that recent-vulnerability port. OpenAI says Astra also discovered and used two previously unknown zero-day vulnerabilities during this evaluation. Because the task set is internal and the vulnerability details are withheld during coordinated disclosure, this result cannot yet be independently replayed.

A separate third-party evaluation from Irregular is stronger independent evidence. Irregular tested Astra and GPT-5.6 Sol in sandboxed environments without public internet access. On FrontierCyber, Astra solved 86 of 226 challenges versus 34 of 226 for Sol. Difficulty-stratified success rose from 14% to 63% on Easy, 15% to 30% on Medium, and 17% to 39% on Hard challenges. However, neither model solved an Elite challenge, and Irregular says it observed no successful attacks on fully hardened targets.

Irregular also reports 59% average success on CyScenarioBench, versus 27% for Sol, with Astra succeeding at least once on 9 of 10 long-horizon scenarios. On its Atomic suite, Astra solved 20 of 22 challenges at least once; vulnerability-research/exploitation and network-attack-simulation categories reached 100% average success, while evasion was 52%.

Those evaluations show a large capability jump without supporting the simplistic claim that Astra can break any hardened system.

Other cyber benchmarks measure different things

OpenAI reports 42.4% on ExploitGym versus 30.3% for GPT-5.6 Sol. ExploitGym contains 869 challenges and requires remote code execution using the intended vulnerability; partial primitives receive no credit. OpenAI says the v1 setup is offline and disables runtime package installation to reduce reward-hacking and sandbox-escape opportunities.

On SRE-Bench, a reverse-engineering benchmark built from 262 binary instances derived from 19 privately developed programs, Astra reaches 88.0% pass@1 and 99.2% pass@4, compared with 55.9% and 68.7% for Sol. Reverse engineering, exploit development and long-horizon cyber scenarios are related capabilities, but they are not interchangeable metrics and should not be merged into a single "cyber score."

Evaluation access is not the same as production access

Many of the strongest cyber numbers were produced without ordinary production safeguards or in controlled evaluation environments. The public Astra release is intentionally more restrictive. OpenAI says the shipping model can assist with defensive tasks such as secure code review and patching but will refuse more advanced work such as generating proof-of-concept exploits, with less restrictive defensive access planned through Daybreak.

The current OpenAI Daybreak troubleshooting page still maps Daybreak Blue to gpt-5.6-sol and Daybreak Red to gpt-5.6-cyber; it does not document a Daybreak alias that directly maps to gpt-6-astra in this verification pass. That is consistent with fresh user reports that Daybreak-verified users can still encounter Astra cyber blocks. Those reports are anecdotes, not proof of a platform-wide failure, but they are useful evidence that evaluation capability and deployable user capability differ.

Pricing, context and latency

OpenAI's current Astra model page lists model ID gpt-6-astra, a 1.05 million-token context window, 128,000 maximum output tokens, and reasoning levels from low through max. The model page lists standard pricing at $10 per million input tokens and $50 per million output tokens, with cached input at $1 per million and higher rates for prompts above 272K input tokens.

There is no published cyber-specific end-to-end latency, time-to-first-token or output-speed benchmark that is directly comparable across these cyber evaluations. OpenAI's separate OSWorld simulations report roughly 40 minutes per task for Astra versus 75 minutes for Sol, but that is a computer-use simulation rather than an API latency guarantee and should not be reused as a cybersecurity latency number.

SWE-bench Verified and SWE-bench Pro remain separate

Neither SWE-bench Verified nor SWE-bench Pro is a cybersecurity benchmark, and neither should be inferred from ExploitBench, SRE-Bench, Terminal-Bench or FrontierCyber. In the current Astra launch page, system card, API model documentation and Irregular cyber evaluation inspected for this article, I did not find a primary standardized Astra score for either SWE-bench Verified or SWE-bench Pro that is suitable to publish here. Those fields therefore remain unknown rather than borrowing a score from another GPT snapshot or another coding benchmark.

This distinction matters because benchmark contamination is not hypothetical: OpenAI explicitly warns that historical vulnerability exposure may inflate ExploitBench. A clean evaluation record should be equally strict about coding benchmarks.

Fresh public feedback is mixed and highly biased

A September 6 Reddit report from a Codex user says ordinary repository work was repeatedly stopped with cybersecurity-risk messages across multiple Astra sessions, sometimes after long runs. Another Daybreak discussion from September 4 reports verified users still encountering restrictive Astra behavior. Both are self-selected user reports with unknown account state, product version, prompts and policy context; they are not controlled measurements and do not establish a general failure rate.

Searches for directly attributable X posts with reproducible Astra cyber measurements did not produce evidence strong enough to quote or score in this pass. No X consensus is claimed.

Bottom line

Astra's cyber evidence is stronger than a single headline score. The 100% ExploitBench result is weakened by OpenAI's own contamination warning. The recent-vulnerability internal port and expert-led zero-day work are more relevant to generalization but are not independently reproducible yet. Irregular's external 86/226 FrontierCyber result provides valuable third-party confirmation of a large jump over GPT-5.6 Sol while also showing the boundary: no Elite challenge was solved and no fully hardened target was successfully attacked in that evaluation.

For practical use, keep capability, safeguards and access separate. Astra can be much more capable in controlled security evaluations than the public deployment permits, and current Daybreak documentation still points to the GPT-5.6 cyber access paths. The next high-value evidence is a reproducible external replay on fresh vulnerabilities, published details after coordinated disclosure, and provider-level measurements of cost and latency under the exact safeguards users can actually access.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books