Analysis
Analysis

GPT-6 Astra Benchmark Update: AA v4.3 Ties Fable 5.1 at 53; Terminal-Bench 59.1, SWE-bench Still Unpublished

Published Sep 8, 2026 Sources checked Sep 8, 2026

Artificial Analysis v4.3 now ties GPT-6 Astra and Claude Fable 5.1 at 53. Astra scores 59.1% on a 66-task Terminal-Bench 4.0 run, while no authoritative exact SWE-bench Verified or Pro score was accepted.

Why this is a material update, not another launch recap

GPT-6 Astra already has launch-week coverage, but the benchmark picture changed materially on September 7, 2026. Artificial Analysis moved its Intelligence Index to v4.3, replacing Terminal-Bench 2.1 with Terminal-Bench 4.0 and replacing τ³-Banking with a held-out AutomationBench-AA set. Under that new index, GPT-6 Astra (max) and Claude Fable 5.1 (max with fallback) both score 53.

That is not the same experiment as the launch-era Artificial Analysis v4.1.1 row reproduced on OpenAI's September 3 launch page, where Astra was shown at 61.2 and Fable 5.1 at 65.7. The index changed twice after launch. A lower v4.3 number should therefore not be reported as an eight-point Astra regression, and Fable's move from 65.7 to 53 should not be treated as a thirteen-point model collapse. The benchmark composition, task privacy and difficulty changed.

This update focuses on the newest version-pinned evidence and on what still remains unknown.

Primary and independent sources:

Artificial Analysis v4.3: 53 is a tie, but the component scores are not identical

Artificial Analysis says Intelligence Index v4.3 contains ten evaluations across four categories: Agents 30%, Coding 20%, General 30%, and Scientific Reasoning 20%. Private questions or answers now account for 45% of the index weight.

The headline result is a tie:

  • GPT-6 Astra (max): 53
  • Claude Fable 5.1 (max with fallback): 53
  • Claude Opus 5 (max): 51
  • Claude Fable 5 (with fallback): 50
  • Muse Spark 1.3 (max): 48
  • GPT-5.6 Sol (max): 47

A tie on the composite does not mean the systems behave identically. Artificial Analysis says Fable 5.1 scores higher on AA-Briefcase and SciCode, while Astra scores higher on Terminal-Bench 4.0 and AutomationBench-AA. The useful production question is therefore not "which model is number one?" but "which component resembles my workload, and what does that success cost?"

The v4.3 article also reports a $3.26 average cost per Intelligence Index task for Astra max versus $7.63 for Fable 5.1 max with fallback under Artificial Analysis's accounting. This is evaluation-specific cost, not a universal application bill.

Terminal-Bench 4.0: 59.1% independently measured, 66 tasks × 3 attempts

Artificial Analysis's new Terminal-Bench 4.0 result is unusually well scoped. It says it runs all 66 tasks three times and reports average pass@1. The benchmark covers terminal-based work spanning software, machine learning, science, operations, security, hardware and media.

In that implementation:

  • GPT-6 Astra (max): 59.1%
  • Claude Fable 5.1 (max with fallback): 52.0%
  • Claude Opus 5 (max): 49.0%
  • GPT-5.6 Sol (max): 39.9%

These are independently benchmarked by Artificial Analysis. They should still be labeled as Artificial Analysis's implementation, because an agent benchmark measures more than raw model weights: environment, time limits, tool surface, reasoning configuration, retries and verifier behavior can all change the result.

OpenAI's launch page separately reports 57.9% for Astra on Terminal-Bench 4.0. The 57.9% vendor result and the 59.1% independent result are close, but they are not automatically interchangeable. Unless the full harness, task revision, attempt policy and endpoint configuration are pinned as identical, the safest reporting is to keep both rows and their provenance.

AutomationBench-AA adds 657 held-out workflows and guardrail-sensitive scoring

The other large v4.3 change is AutomationBench-AA. Artificial Analysis says it evaluates a held-out 657-task set, using benchmark version v1.0.6, across Finance, HR, Marketing, Operations, Sales and Support.

The evaluator reports two different measures:

  1. Score — the average share of task objectives completed, with any guardrail violation reducing that task to zero.
  2. Tasks Completed — the share of workflows where every objective was completed with no guardrail violation.

Astra max scores 68.5% on the first measure and fully completes 41.6% of workflows. Those numbers should not be merged into one "68.5% completion rate." Partial objective credit and full workflow completion answer different questions.

This is especially relevant for production agents. A model that completes most subtasks but violates a policy boundary can still be unusable for finance, HR or customer-support automation. Conversely, a strict all-or-nothing completion metric can hide useful partial progress. Reporting both measures gives a better picture.

Coding Agent Index v1.4: Astra 67 in Codex versus Fable 5.1 70 in Claude Code

Artificial Analysis's separate Coding Agent Index v1.4 should not be mixed with Intelligence Index v4.3 or Terminal-Bench 4.0.

The current coding-agent methodology covers 326 tasks across:

  • DeepSWE: 113 tasks
  • Terminal-Bench v2.1: 89 tasks
  • SWE-Atlas-QnA: 124 tasks

Each task is attempted three times; each component reports task-normalized average pass@1, and the three benchmark components receive equal weight in the composite.

Under the current public comparison:

  • GPT-6 Astra (max) in Codex: Coding Agent Index 67
  • Claude Fable 5.1 (max with fallback) in Claude Code: 70

The component rows make the difference clearer:

  • DeepSWE: Astra/Codex 67%, Fable/Claude Code 66%
  • Terminal-Bench v2.1: Astra/Codex 83%, Fable/Claude Code 89%
  • SWE-Atlas-QnA: Astra/Codex 51%, Fable/Claude Code 56%

Artificial Analysis also reports pooled efficiency measures of roughly $4.72 per task, 26.8 minutes per task and 4.0M tokens per task for Astra/Codex, compared with $9.18, 24.0 minutes and 7.1M tokens for Fable 5.1/Claude Code.

This is valuable evidence, but it is a model-plus-agent-stack comparison, not a clean foundation-model-only head-to-head. Codex and Claude Code have different orchestration, prompts, tool behavior, context management and implementation details. The same model can score differently in another harness.

Methodology:

OpenAI's launch coding table remains vendor evidence, not the independent v4.3 table

OpenAI's September 3 launch page reports these Astra coding results:

  • Terminal-Bench 4.0: 57.9%
  • DeepSWE v1.1: 74.1%
  • FrontierCode 1.1 Extended: 64.5%
  • FrontierCode 1.1 Main: 53.3%
  • Internal Database Migration Tasks: 63.9%
  • Artificial Analysis Coding Agent Index v1.4: 67.0

The page also reproduces the then-current Artificial Analysis Intelligence Index v4.1.1 score of 61.2.

These numbers are useful launch evidence, but the date and benchmark version are part of the result. The independently maintained v4.3 composite now says 53, and Artificial Analysis's current Terminal-Bench 4.0 implementation says 59.1. A current comparison should not copy the launch table and call it today's independent leaderboard.

SWE-bench Verified: no authoritative Astra score found

SWE-bench Verified and SWE-bench Pro remain separate benchmarks.

For GPT-6 Astra, the primary OpenAI launch page reviewed for this update contains no SWE-bench entry at all. That omission matters because some secondary pages now circulate exact Astra numbers labeled "SWE-bench Verified." I did not find a primary OpenAI, official SWE-bench, or version-pinned independent run manifest supporting an exact Astra Verified score in this bounded review.

That cell should therefore remain unpublished / not verified, not filled with a number copied from a secondary leaderboard.

OpenAI has also publicly argued since February 2026 that SWE-bench Verified is increasingly contaminated for frontier-model evaluation. That policy context can explain why OpenAI may prefer newer coding evaluations, but it does not authorize inventing a missing Astra score.

SWE-bench Pro: still no version-pinned Astra result accepted

The same rule applies to SWE-bench Pro. I did not accept an exact Astra Pro score in this run because I could not trace one to a current primary or independent version-pinned evaluation with enough methodology to identify the suite revision, task count, scaffold, retries, model endpoint and date.

Artificial Analysis Coding Agent Index v1.4 does not solve this gap: its current components are DeepSWE, Terminal-Bench v2.1 and SWE-Atlas-QnA. The methodology's version history says SWE-Bench-Pro-Hard-AA was removed from the index in version 1.1.

Likewise, DeepSWE, FrontierCode and Terminal-Bench are different evaluations. Their scores must not be relabeled as SWE-bench Pro.

The correct benchmark table is therefore:

Evaluation Astra result used here Evidence label
SWE-bench Verified Not established No accepted exact current run
SWE-bench Pro Not established No accepted exact current run
DeepSWE v1.1 74.1% OpenAI/vendor launch result
Terminal-Bench 4.0 59.1% Artificial Analysis independent implementation
Coding Agent Index v1.4 67 Astra in Codex; model+agent-stack result

Current API facts: 1.05M context, 128K output, and a long-context price step

OpenAI's current API model page identifies the exact API model as gpt-6-astra and lists:

  • 1,050,000-token context window
  • 128,000 maximum output tokens
  • knowledge cutoff: April 30, 2026
  • reasoning efforts: low, medium, high, xhigh and max
  • text and image input; text output
  • no audio or video input/output for the model
  • web search, file search, image generation, code interpreter, hosted shell, apply-patch, skills, computer use and MCP support through the Responses API tool surface
  • fine-tuning: not supported

Standard short-context token rates are $10/M input, $1/M cached input, $12.50/M cache writes and $50/M output.

There is a major practical caveat to the 1.05M context headline: OpenAI says requests with more than 272K input tokens are charged at 2× input/cache rates and 1.5× output rates for the full request. Batch and Flex are listed at 50% of Standard rates, while Fast mode is 2× the applicable rate.

A production team should therefore not treat "1M context" as merely a capacity feature. Crossing the long-context threshold changes unit economics for the entire request.

Official API source:

Independent speed and latency: max effort is not a normal chat-latency number

At verification time on September 8, Artificial Analysis's live Astra max page reported about 61.7 output tokens/second and 328.62 seconds time to first token on the first-party OpenAI API.

That TTFT requires careful interpretation. It is for the max reasoning configuration in Artificial Analysis's standardized measurement environment. It includes the delay before the first visible answer token and can include substantial internal reasoning. It should not be reported as "GPT-6 Astra takes 5.5 minutes to respond" for every request.

Artificial Analysis's release comparison shows much lower latency at lower reasoning efforts; the low-effort Astra variant is listed around a few seconds TTFT. Provider load, prompt shape, reasoning effort, region and processing mode can all move these numbers.

The useful production measurement is a distribution under your own workload: P50/P95 time to first useful token, end-to-end task time, cost per successful task and retry rate—measured separately for each reasoning effort.

Early practitioner feedback: useful anecdotes, not a population benchmark

Two recent OpenAI Developer Community posts illustrate why launch-week feedback must be labeled carefully.

A Berlin teacher posted on September 7 that they used Codex with Astra over September 5–7 to turn an existing teaching-site prototype into a production website, describing work across implementation, design iteration, debugging, automated tests and deployment. That is a concrete positive build report, but it has no controlled comparison or standardized task set.

A separate user posted on September 7 that a roughly 40-minute existing-repository task at medium reasoning changed their weekly Codex allowance estimate from about 87% remaining to about 47%, and called the consumption impractical for their workflow. That is a concrete negative cost/quota anecdote, but allowance percentage is not the same as metered API tokens or dollars, and plan quotas can change independently of API pricing.

Sources:

These reports disagree in emphasis—one highlights end-to-end productivity, the other resource consumption—and both are self-selected. They should guide what to measure, not establish a consensus.

A bounded search also surfaced claims mirrored from X about repository coding tests, but I did not obtain a stable direct X post with a sufficiently reproducible run manifest for this update. Those claims were not promoted into benchmark evidence, and no X consensus is inferred.

Availability is improving, but benchmark access and product-surface access are different questions

The API documentation currently exposes the model page, pricing and rate-limit tiers, with the free API tier marked unsupported. OpenAI's launch page still describes ChatGPT availability as a rollout across paid plans and notes that Enterprise administrators can control access.

Public forum threads continue to show account- and surface-specific differences between Chat, Work and Codex. Those reports are useful evidence of rollout friction, but they should not be mixed into capability rankings. A model can be highly capable in an independent evaluation and still be temporarily unavailable, capacity-limited or surfaced differently in a subscription product.

For production procurement, verify the exact endpoint and account access rather than assuming a launch announcement guarantees immediate availability in every interface.

Practical verdict

The strongest new evidence since Astra's launch is not a new "winner" label. It is a clearer decomposition of the benchmark story.

Broad composite: Artificial Analysis v4.3 now places Astra max and Fable 5.1 max-with-fallback in a tie at 53.

Terminal work: Artificial Analysis measures Astra max at 59.1% on Terminal-Bench 4.0, using all 66 tasks three times.

Business-agent workflows: Astra max scores 68.5% on AutomationBench-AA objective completion and 41.6% on fully completed workflows under the 657-task held-out evaluation.

Coding agents: Astra in Codex scores 67 on Coding Agent Index v1.4, while Fable 5.1 in Claude Code scores 70. Because the harnesses differ, this is a system comparison, not a pure model ranking.

SWE-bench: no accepted exact Astra Verified or Pro score was found. Those cells remain blank.

Economics: $10/$50 is only the short-context Standard headline. Cache writes, long-context requests above 272K, Fast mode, tool calls, reasoning effort and retries can change real task cost substantially.

The next high-value evidence would be a version-pinned Astra SWE-bench Pro run under a clearly specified harness, plus matched Astra/Fable runs using the same agent scaffold and tool policy, and repeated real-repository tests that report success, regression rate, wall-clock time, tokens and dollars together. Until then, the fair conclusion is not that one model universally beats the other; it is that Astra's independent v4.3 evidence is now much stronger and more current than the launch-week leaderboard snapshot, while important coding-benchmark gaps remain.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books