Terminal-Bench 4.0: GPT-6 Astra Leads by 0.30 Points, but the Top Systems Overlap
A current, source-backed reading of Terminal-Bench 4.0: GPT-6 Astra tops the latest 18-system snapshot at 58.18%, but the 0.30-point gap sits well inside published uncertainty and the rows use different agent harnesses.
The leaderboard changed after GPT-6 Astra arrived
Terminal-Bench 4.0 is a benchmark for difficult computer work performed through a terminal. The official site describes the metric as resolution rate and plots 95% confidence intervals around each system result. It is important to call them system results: a row combines a model with an agent harness such as Codex, Claude Code, mini-SWE-agent or another scaffold.
The latest public 18-system snapshot reviewed on September 5, 2026 places GPT-6 Astra in Codex at max reasoning at 58.18%, followed by Claude Fable 5.1 in Claude Code at 57.88%. GPT-6 Astra at xhigh and high effort are also listed at 57.88%. That is a 0.30 percentage-point headline lead, not evidence of a decisive model-to-model win.
The official Terminal-Bench page is dynamic, so static mirrors can briefly lag. BenchLM's current snapshot shows the 18-row board with Astra first at 58.18%. Snorkel AI's dedicated Terminal-Bench 4.0 page still exposes an earlier snapshot with Fable 5.1 first at 57.9%, while other current Snorkel benchmark pages already cross-link a Terminal-Bench 4.0 card showing Astra at about 58.2%. This is a useful reminder to attach a date to leaderboard claims rather than treating a live benchmark as a timeless table.
Sources: Terminal-Bench, BenchLM Terminal-Bench 4.0, and Snorkel AI Terminal-Bench 4.0.
Why 58.18% versus 57.88% is not a clean ranking
Terminal-Bench explicitly says the whiskers on its leaderboard are 95% confidence intervals. A current independent read of the published board reports a ±2.79-point interval for Astra max, ±3.83 for Fable 5.1, ±2.71 for Astra xhigh and ±2.98 for Astra high. Those uncertainty bands are far larger than the 0.30-point difference between first and the next three rows.
So the defensible conclusion is narrower than “Astra beats Fable.” Astra has the highest point estimate in the current snapshot, while the leading systems' uncertainty intervals overlap substantially. Additional trials, a changed task set or a future benchmark revision could easily change the order.
There is a second confound: Astra was run through Codex, while Fable 5.1 was run through Claude Code. Terminal-Bench therefore measures a model-plus-agent stack. Tool policy, prompting, context management, retry behavior and other harness choices can contribute to the score.
Source: Terminal-Bench current leaderboard and current board analysis.
What Terminal-Bench 4.0 actually changed
Version 4.0 is not a cosmetic rename of 3.0. The maintainers recalibrated time, CPU and memory, moved all tasks to a flat eight-hour agent timeout, fixed task instructions/environments/verifiers, and removed tasks that no longer provided a clean frontier signal.
The official release says eight tasks were removed: two for saturation, two for refusals, two because public solutions existed, and two for unresolved quality or platform-compatibility problems. The maintainers say 19 tasks were fixed. The GitHub v4.0.0 release lists the eight removed tasks and the modified task entries. Snorkel describes 20 revisions because its count includes the global resource-level revision alongside task-specific changes; that counting difference should not be mistaken for a factual contradiction.
The maintainers also adopted semantic versioning for a continuous benchmark. They explain that resource changes and changes to the task set are breaking changes that require trials to be rerun. That means a 4.0 score should not be compared directly with a 3.0 or 2.x score as if only the model changed.
Sources: Terminal-Bench 4.0 release note and v4.0.0 GitHub release.
Sample size and harness
The current 4.0 snapshot contains 66 tasks, with five trials per task, giving 330 task trials for a complete system run. BenchLM documents the five-trial setup and eight-hour timeout for the current 18-row snapshot. The benchmark's purpose is broader than code completion: tasks exercise autonomous terminal work across software, systems and other professional computer workflows.
This sample is large enough to be more informative than a single demo, but it is not large enough to make a 0.30-point difference meaningful when the published uncertainty is several percentage points. Repeated trials also do not remove harness effects: the same base model can move when effort, agent software, tools or resources change.
Source: BenchLM Terminal-Bench 4.0 methodology and Terminal-Bench 4.0 methodology.
Cost is useful evidence, but it is not a production quote
Terminal-Bench publishes cost and token totals alongside resolution rate. A current board read reports total evaluation cost of about $3,267 for Astra max and $6,244 for Fable 5.1 max. Astra high is reported at about $2,269 while matching Fable's 57.88% point estimate in this snapshot.
Those numbers are informative for this benchmark run, but they are not universal “price per task” figures. The systems used different agent harnesses, cache behavior, reasoning settings and token volumes. Production workloads can have very different prompt reuse, tool-call patterns, output lengths and failure/retry rates. A derived cost-per-solved-trial ratio can help compare this particular board, but it should be labeled as arithmetic derived from the published run totals rather than as a vendor price.
Source: Terminal-Bench leaderboard and cost analysis derived from the board.
Maintainer discussion on X: useful context, not independent model feedback
On August 29, 2026, Terminal-Bench maintainer Ryan Marten posted on X that the team removed eight tasks across four reasons—saturation, refusals, public solutions, and unresolved quality/platform issues—and fixed 19 tasks after user reports and leaderboard-run findings. In the same thread he said 4.0 had fewer agent timeouts and errors than 3.0 and explained that the move to 4.0, rather than 3.1, reflects semantic versioning because resources and the task set changed.
That is primary maintainer commentary about benchmark construction. It should not be treated as independent evidence that any model is better, and it is not a sample of community sentiment.
Post: Ryan Marten on X, August 29, 2026.
Do not translate this into a SWE-bench result
Terminal-Bench 4.0 is not SWE-bench Verified or SWE-bench Pro. The benchmarks test different tasks, use different evaluation infrastructure, and have different known limitations. A 58% Terminal-Bench result cannot be converted into or directly ranked against a SWE-bench percentage.
For this intelligence series, SWE-bench Verified remains a separate 500-task human-validated subset with documented contamination concerns at the frontier, while SWE-bench Pro is a separate 1,865-task benchmark with its own task-quality audit concerns. Those caveats do not invalidate Terminal-Bench, and Terminal-Bench's score does not resolve the caveats in either SWE-bench family.
Practical reading for developers and AI buyers
The current evidence supports four practical conclusions:
- Astra has the highest current Terminal-Bench 4.0 point estimate, but the lead over Fable 5.1 and two lower Astra effort settings is tiny relative to the published uncertainty.
- Agent harness matters. Codex versus Claude Code is part of the experiment, so the table is not a clean isolated-base-model comparison.
- High effort may be more attractive than max for Astra on this snapshot. It posts 57.88% versus 58.18% at max while the total benchmark run cost is lower, and the 0.30-point difference is inside the uncertainty band.
- Version numbers matter. Do not use old Terminal-Bench 2.x or 3.0 scores to claim a generation-over-generation improvement without controlling for the changed task set and resource protocol.
Confidence is high on the 4.0 methodology and release changes because they come from the benchmark maintainers and release repository. Confidence is medium-high on the September 3–4 leaderboard ordering because multiple current mirrors agree while the official table is dynamically rendered and some partner pages lag. Confidence is low on extrapolating benchmark-run cost to production economics.
The most useful next evidence would be additional same-version trials, same-harness comparisons where possible, transparent latency and failure-rate reporting, and future 4.1 verifier updates without conflating them with 5.0 task-set changes.
This article is built from the source material below. Open the originals for full context and the latest updates.