Analysis
Analysis

Intelligence Index v4.3 Reality Check: 45% Private Weight, 66-Task Terminal-Bench 4 and 657-Task AutomationBench-AA

Published Sep 8, 2026 Sources checked Sep 8, 2026

Artificial Analysis v4.3 raises private-evaluation weight to 45%, upgrades Terminal-Bench to 66-task v4.0 and adds 657-task AutomationBench-AA. Astra and Fable 5.1 tied at 53 in the dated launch snapshot.

What changed on September 7

Artificial Analysis released Intelligence Index v4.3 on September 7, 2026, replacing two components of its composite model-evaluation index: Terminal-Bench 2.1 moves to Terminal-Bench 4.0, and the earlier τ³-Banking component is replaced by AutomationBench-AA.

The release matters because this is not merely another leaderboard refresh. It changes the tasks underlying the score. A v4.3 score therefore should not be read as a direct trend line from v4.2 or older Index versions. Some movement can come from model quality, but some can come from the benchmark mix, task difficulty, scoring rules and updated agent environments.

Artificial Analysis says v4.3 keeps the same category weights as v4.2 — Agents 30%, Coding 20%, General 30% and Scientific Reasoning 20% — while increasing the share of Index weight assigned to evaluations with private questions or answers from 40% to 45%.

Source: Artificial Analysis — Intelligence Index v4.3 announcement

The dated launch snapshot: Astra and Fable 5.1 both scored 53

In the September 7 launch article, GPT-6 Astra at max effort and Claude Fable 5.1 at max effort with the evaluator's default fallback both score 53 on Intelligence Index v4.3. Claude Opus 5 follows at 51, Claude Fable 5 at 50, Muse Spark 1.3 at 48 and GPT-5.6 Sol at 47.

That 53/53 result should be treated as a dated v4.3 launch snapshot, not a permanent model rating. Artificial Analysis model pages are live surfaces and may change as evaluations are rerun or additional result processing lands. For that reason this article does not silently replace the launch article's numbers with later dynamic page values.

The overall tie also hides different strengths. Artificial Analysis says Fable 5.1 is stronger on AA-Briefcase and SciCode, while Astra is stronger on Terminal-Bench 4.0 and AutomationBench-AA Score. A single composite number compresses those differences.

What the 100-point Index actually contains

Artificial Analysis publishes the v4.3 composition and weights:

Category Evaluation Private questions/answers Weight
Agents AA-Briefcase Yes 15%
Agents GDPval-AA v2 No 10%
Agents AutomationBench-AA Yes 5%
Coding Terminal-Bench 4.0 No 10%
Coding SciCode No 10%
General AA-Omniscience Accuracy Yes 10%
General AA-Omniscience Non-hallucination Yes 5%
General GDP.pdf No 10%
General AA-LCR v1.1 No 5%
Scientific reasoning Humanity's Last Exam No 10%
Scientific reasoning CritPt Yes 10%

That is why 45% private weight is more precise than saying "45% of the benchmarks are private." The percentage refers to Index weighting for evaluations with private questions or answers, not a simple count of benchmark names.

Private evaluation material has a real benefit: it reduces the opportunity for direct training contamination, benchmark-specific prompt tuning and public-solution memorization. The tradeoff is reduced outside reproducibility. Independent researchers cannot fully rerun a held-out task set they do not possess, so confidence rests more heavily on the evaluator's implementation, auditability and reporting.

Terminal-Bench 4.0: 66 tasks, three AA runs each

Artificial Analysis says it now runs all 66 Terminal-Bench 4.0 tasks three times and reports average pass@1. In its v4.3 run, GPT-6 Astra max scores 59.1%, Claude Fable 5.1 max with fallback scores 52.0%, and Claude Opus 5 max scores 49.0%.

Terminal-Bench 4 is materially different from Terminal-Bench 2.1. The v4 release raises task difficulty, recalibrates compute and time allowances, and improves task instructions, environments and verification. That makes the new score more relevant to current frontier agents, but it also means old Terminal-Bench 2.1 percentages are not directly comparable.

There is another important harness issue. Public Terminal-Bench leaderboards compare model-and-agent systems, not isolated base models. Harnesses such as Codex or Claude Code can change tool access, context management, retry behavior, environment interaction and token budgets. Artificial Analysis runs its own controlled implementation, so its 59.1 versus 52.0 comparison is an independent evaluator result — but it is still Artificial Analysis's harness, not a proof that the same gap will appear in every terminal agent.

Source: Artificial Analysis v4.3 release

AutomationBench-AA: 657 held-out business workflows

The other major change is AutomationBench-AA, Artificial Analysis's implementation of Zapier's AutomationBench.

Artificial Analysis says it uses the held-out test set of 657 tasks from benchmark version 1.0.6. The tasks span Finance, HR, Marketing, Operations, Sales and Support and operate across simulated business applications such as Gmail, Google Sheets, Slack, Salesforce, Zendesk, Jira and HubSpot.

This is a useful complement to coding benchmarks because the agent has to discover relevant APIs, manipulate application state and obey business rules across multiple systems.

Artificial Analysis changes the headline scoring from Zapier's official leaderboard. Its Score averages the share of objectives completed across all 657 workflows, but a guardrail violation makes that task score zero. It also reports Tasks Completed separately: the share of workflows where every objective was finished without a violation.

In the v4.3 launch run, GPT-6 Astra max scores 68.5% on the AA Score, ahead of Grok 4.6 high at 66.7% and GLM-5.3 max at 62.2%. Astra completes every objective without a guardrail violation on 41.6% of workflows. Claude Fable 5.1 max with fallback is reported at 32.1% Tasks Completed, while Claude Opus 5 max is at 28.3%.

Those figures must not be confused with Zapier's own official task_completed_correctly leaderboard. Zapier currently reports GPT-6 Astra max at 41.4% on AutomationBench 1.0.6, because its headline metric is strict full-task completion. The two leaderboards are related but use different headline metrics.

Sources: AutomationBench-AA methodology, Zapier AutomationBench leaderboard, AutomationBench v1.0.6 changelog

Why 68.5% and 41.4% are not contradictory

This is exactly the kind of benchmark comparison that can produce misleading headlines.

On AutomationBench-AA, a task with multiple objectives can earn partial objective credit when the agent completes some of the required work and does not violate a guardrail. Zapier's official leaderboard asks whether the entire task completed correctly. A model can therefore score well on objective completion while still failing the stricter all-or-nothing task.

The correct interpretation is:

  • 68.5% is Astra max's Artificial Analysis objective-based, guardrail-adjusted AutomationBench-AA Score in the v4.3 release.
  • 41.6% is the share of AA workflows Astra max completed in full without a guardrail violation.
  • 41.4% is Zapier's current official strict full-task score for Astra max on AutomationBench 1.0.6.

These numbers answer different questions. None should be substituted for another.

Pricing, latency and context are not baked into the 53-point intelligence score

Artificial Analysis separately tracks cost, speed and latency, but Intelligence Index v4.3 is a capability composite. A model does not receive extra intelligence points merely for being cheaper or faster.

Artificial Analysis says OpenAI's five reasoning-effort settings for GPT-6 Astra occupy much of the current Intelligence versus Cost per Task Pareto frontier, while Claude Fable 5.1 at higher effort also sits on that frontier. That is useful procurement evidence, but it remains workload-specific.

OpenAI's current first-party API page lists GPT-6 Astra at $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache-write tokens and $50 per million output tokens, with a 1,050,000-token context window and 128,000 maximum output tokens. Inputs above 272K tokens use OpenAI's documented long-context surcharge.

Artificial Analysis's live model page currently reports separate measured speed, latency and cost-per-Index-task statistics for Astra. Those measurements can move with provider routing and evaluation refreshes, so they should not be frozen into a model identity claim or used to reinterpret the dated 53-point v4.3 launch snapshot.

Sources: OpenAI GPT-6 Astra API specification, Artificial Analysis GPT-6 Astra page

SWE-bench Verified and SWE-bench Pro are separate — and neither is v4.3

SWE-bench Verified is not a component of Intelligence Index v4.3. SWE-bench Pro is not a component either.

The v4.3 Coding category is 10% Terminal-Bench 4.0 and 10% SciCode. Therefore a 53 Intelligence Index score does not imply any SWE-bench Verified or SWE-bench Pro score, and a Terminal-Bench result must not be relabeled as SWE-bench.

This article also does not transfer historical SWE-bench figures from earlier model checkpoints or different harnesses onto the exact v4.3 configurations. If a model's Verified or Pro result is discussed elsewhere, it needs its own checkpoint, task split, harness, retry policy, date and evidence trail.

That separation is especially important for agent models because the harness can materially change observed coding performance.

Independent evidence is stronger than a vendor chart, but not the same as multi-lab reproduction

Artificial Analysis is independent of OpenAI and Anthropic, so its runs are more useful for cross-vendor comparison than simply repeating each vendor's launch table. It also publishes methodology, task counts, evaluation weights and cost data.

But "independent evaluator" is not the same as "independently reproduced by multiple labs." In this review I did not find a same-day third-party lab rerunning the complete v4.3 Index with access to its private material, nor could such a lab fully reproduce private task sets without collaboration.

The evidence level is therefore:

independent evaluator result with disclosed methodology and partially private test material; not a public multi-lab reproduction.

That wording matters because the new 45% private weighting intentionally trades some public reproducibility for contamination resistance.

Public feedback: useful questions, not measured validation

I did not find a reliable same-day X post that independently reran Intelligence Index v4.3, so no X quote or consensus is invented here.

Public Reddit discussion around GPT-6 Astra is more skeptical than the leaderboard headlines. In a September 4 AI_Agents thread, users focused on whether strong benchmark scores translate into long-running production reliability, noting that a terminal benchmark score around the high-50s still means many tasks fail. A September 6 discussion similarly argued that Astra's biggest gains look agentic rather than like a uniform 2.5x jump in general reasoning, while highlighting its much higher list pricing versus GPT-5.6 Sol.

Those are self-selected anecdotes and interpretations, not controlled measurements. They are useful mainly because they point to the right production questions: completion reliability, human correction rate, end-to-end latency and cost per successful task.

Discussion sources: AI_Agents thread, September 4, ArtificialInteligence thread, September 6

Practical reading of v4.3

Intelligence Index v4.3 is a useful upgrade because it pushes the composite toward harder terminal work and broader cross-application automation while increasing contamination-resistant evaluation weight.

The safest interpretation is not "53 proves Astra and Fable 5.1 are identical." It is:

under Artificial Analysis's September 7 v4.3 methodology, the two models had the same rounded composite launch score while showing different benchmark strengths.

For teams choosing a model, the next step is workload matching. Terminal agents should inspect Terminal-Bench 4 behavior, business automation teams should distinguish AutomationBench-AA objective score from strict full-task completion, coding teams should keep SWE-bench Verified and Pro in their own evidence columns, and production buyers should measure cost, latency and success under the exact harness they intend to deploy.

The next evidence that would materially change this assessment is a stable later v4.3 snapshot with rerun details, a public explanation for any score revisions on dynamic model pages, and independent reproductions of the public benchmark components under matched scaffolds.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books