Analysis
Analysis

Artificial Analysis Intelligence Index v4.2 Reality Check: 40% Private Tests, Fable 57 vs Astra 55 and a Weighting Mismatch

Published Sep 6, 2026 Sources checked Sep 6, 2026

Artificial Analysis v4.2 doubles private held-out weighting to 40%. Fable 5.1 leads Astra 57–55, but old scores are not comparable and two official pages disagree on category weights.

What changed on September 4

Artificial Analysis released Intelligence Index v4.2 on September 4, 2026. This is a benchmark-suite revision, not a new model checkpoint. The update adds AA-Briefcase, adds Surge AI's GDP.pdf, removes the saturated GPQA Diamond evaluation, upgrades AA-LCR to v1.1, regrades SciCode, and changes the evaluation weights.

The most consequential design change is that 40% of the Index weighting now comes from private held-out test sets, double the share in v4.1. Artificial Analysis says the held-out material includes AA-Briefcase, AA-Omniscience and CritPt solutions, with a larger private share planned for v5. This can make direct benchmark gaming harder, but it also means outside researchers cannot independently inspect or reproduce every task in the composite score.

That trade-off should be explicit whenever the v4.2 leaderboard is used as evidence.

Fable 5.1 leads Astra 57 to 55 — but the label includes the evaluation configuration

The current v4.2 model pages put Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) at 57 and GPT-6 Astra (max) at 55. Artificial Analysis describes Fable 5.1 as the current leader and Astra as the runner-up.

Those labels matter. The Fable result is not described as a pure single-checkpoint run: Artificial Analysis keeps Default Fallback in the model configuration. In its September 1 pre-v4.2 evaluation, it said Anthropic's default server-side safety fallback routed safety-flagged requests to Claude Opus 4.8 or Opus 5 and accounted for about 4% of output tokens in that run. The current v4.2 model label still explicitly includes the fallback configuration, so readers should treat the score as a model-system result rather than silently dropping that qualifier.

OpenAI's exact public model identity is gpt-6-astra. Anthropic's public Fable family page confirms Claude Fable 5.1 availability and pricing. The comparison here is therefore between identifiable current products, but also between the exact serving/evaluation configurations Artificial Analysis tested.

Do not compare 57 and 55 directly with the old 66 and 61

Before v4.2, Artificial Analysis reported 66 for Fable 5.1 max with fallback on September 1 and 61 for GPT-6 Astra max on September 3 under Intelligence Index v4.1.1.

It would be wrong to say Fable "fell nine points" or Astra "fell six points" as if the model checkpoints regressed. Between those numbers and v4.2, Artificial Analysis changed the suite, grading infrastructure and weights. Version history says v4.1 emphasized Agents 34%, Coding 24%, Scientific Reasoning 24% and General 18%; v4.2 uses a different distribution and adds two evaluations while removing GPQA Diamond.

The old and new scores are useful as historical snapshots of what each benchmark version measured. They are not an apples-to-apples longitudinal model-quality series.

The current methodology says 30/20/20/30, while the benchmark FAQ says 25% each

There is also a live documentation mismatch worth preserving rather than smoothing over.

The versioned Intelligence Benchmarking methodology currently says the composite is weighted Agents 30%, Coding 20%, Scientific Reasoning 20%, General 30%. Its detailed table assigns: AA-Briefcase 15%, GDPval-AA v2 10%, tau3-Banking 5%, Terminal-Bench v2.1 10%, SciCode 10%, AA-Omniscience 15% split across accuracy and non-hallucination, GDP.pdf 10%, AA-LCR v1.1 5%, HLE 10% and CritPt 10%.

But the current Intelligence Index evaluation FAQ says four categories each contribute 25%.

These two official Artificial Analysis pages therefore conflict. Because the versioned methodology table supplies the per-evaluation weights and its version history explicitly records the v4.2 rebalance, 30/20/20/30 is the more specific current methodology description. Still, the 25%-each FAQ should be treated as a documentation inconsistency until Artificial Analysis reconciles it.

This matters for reproducibility: a composite score cannot be reconstructed confidently if its public weighting description is ambiguous.

What is actually inside v4.2

Artificial Analysis says v4.2 combines ten evaluations. Its current methodology provides useful task and repeat metadata:

  • AA-Briefcase: 91 tasks across four scenarios, one repeat, 15% weight, using an agentic file-output workflow.
  • GDPval-AA v2: 220 tasks, one repeat, 10%.
  • tau3-Banking: 97 tasks, five repeats, 5%.
  • Terminal-Bench v2.1: 89 tasks, three repeats, 10%.
  • SciCode: 288 test-set subproblems, three repeats, 10%.
  • AA-Omniscience: 6,000 questions, one repeat, 15% combined across accuracy and non-hallucination.
  • GDP.pdf: 100 tasks across ten domains, five repeats, 10%.
  • AA-LCR v1.1: 5%.
  • Humanity's Last Exam: 10%.
  • CritPt: 10%.

GDP.pdf itself uses 100 professional PDFs spanning ten domains and 4,592 pages, graded against 1,275 expert-authored atomic criteria. Its headline All-pass measure credits a task only if every criterion passes.

AA-Briefcase is deliberately closer to professional agent work: Artificial Analysis describes multi-week simulated projects with linked tasks and thousands of source files. Its implementation allows up to 500 turns per task in the Stirrup harness and operates in a sandbox without internet access.

The composite is therefore much broader than a single multiple-choice reasoning test, but its result remains dependent on the chosen task mix, graders, harnesses, weights and private material.

Private tests reduce gaming, but they also narrow independent reproducibility

The 40% private-held-out share addresses a real benchmark problem: public task sets can leak into training data, prompt engineering can overfit to known test distributions, and model developers can optimize against stable public leaderboards.

However, "private" is not the same as "independently reproducible." Outside groups cannot rerun undisclosed tasks exactly, inspect all failures, measure contamination directly, or verify whether the private set represents their own workloads.

That does not invalidate v4.2. It changes what the score can support. The Index is strongest as an independently operated comparative measurement under Artificial Analysis's controlled suite. It is weaker as a fully open scientific artifact that any lab can recreate end-to-end from public data.

Artificial Analysis estimates the composite Index has a 95% confidence interval narrower than ±1 point, based on experiments with more than ten repeats on certain models across the included datasets, while warning that individual evaluation intervals may be wider and saying more statistical detail will be disclosed later. A two-point Fable-versus-Astra gap is therefore notable in the current aggregate, but it should still be described as a narrow lead under this specific benchmark configuration—not proof that one model is universally superior.

Cost and token efficiency tell a different story from the rank

The current v4.2 model pages show both Fable 5.1 and GPT-6 Astra at $10 per million input tokens and $50 per million output tokens on their first-party APIs, consistent with their vendors' public pricing.

Yet Artificial Analysis measures very different benchmark costs and token use:

  • Fable 5.1 max with fallback: 57 Index, $6.12 per Intelligence Index task, about 160 million output tokens across the Index, and roughly 70.5 output tokens/second.
  • GPT-6 Astra max: 55 Index, $2.57 per Intelligence Index task, about 49 million output tokens, and roughly 71.3 output tokens/second.

So the list price alone does not determine cost per completed benchmark task. Under this suite and these settings, Astra is substantially more token-efficient and cheaper per weighted Index task while Fable holds the higher aggregate score.

Artificial Analysis also reports reasoning-inclusive first-answer latency of about 264.91 seconds for Fable 5.1 and 384.30 seconds for Astra max on the measured first-party routes. Its methodology explicitly says this latency includes reasoning-model "thinking" before the first answer token. These figures should therefore not be mistaken for simple network time-to-first-byte, nor treated as immutable model properties across providers, reasoning levels and workloads.

SWE-bench Verified and SWE-bench Pro are not part of this score

Intelligence Index v4.2 includes Terminal-Bench v2.1 and SciCode for coding. It does not include SWE-bench Verified or SWE-bench Pro.

That boundary is important. A 57 or 55 Intelligence Index score cannot be converted into a SWE-bench result, and a leaderboard lead here should not be described as a SWE-bench lead.

OpenAI separately said on February 23, 2026 that it had stopped reporting SWE-bench Verified because it found increasing contamination and recommended SWE-bench Pro instead. That is a separate methodological position, not evidence that Astra's v4.2 score equals a Pro result.

In this bounded review, no exact new primary-source SWE-bench Verified or SWE-bench Pro result with pinned current Fable 5.1/Astra model identity, scaffold, split and evaluation recipe was promoted from the v4.2 materials. Both remain separate columns in any serious coding comparison.

This is mainly an English text benchmark, not a universal multimodal ranking

Artificial Analysis's methodology states that the Intelligence Index is primarily text-based and English-language. It benchmarks image input, speech input and multilingual performance separately.

That means v4.2 should not be used as a universal claim about multimodal perception, voice, video, multilingual quality or every agent deployment. Teams choosing a model for those workloads should consult the corresponding specialized evaluations.

Public feedback: no reproducible X consensus was found

Artificial Analysis's own September 4 announcement and public model pages provide detailed first-hand benchmark evidence. A bounded search for attributable public X discussion did not surface a reproducible independent run that pinned the v4.2 benchmark version, exact model configuration and objective metrics strongly enough to promote here.

That absence is preferable to inventing a social-media consensus. Launch reactions can identify useful hypotheses, but they should not override the measured result without a controlled rerun.

Practical takeaway

Intelligence Index v4.2 is a meaningful benchmark refresh, especially because it adds longer professional tasks and raises the private held-out share to 40%. The current leaderboard's Fable 5.1 57 versus GPT-6 Astra 55 is useful evidence under Artificial Analysis's suite, while Astra's much lower measured cost per Index task shows why aggregate quality and operational efficiency should be tracked separately.

For model selection, record the benchmark version, exact model/effort/fallback configuration, harness, task set, grader, date, cost per successful task, output-token use and latency definition. Do not compare v4.2 scores directly with the old v4.1.1 values as if the benchmark were unchanged. Keep SWE-bench Verified and SWE-bench Pro separate. And until the official pages agree, preserve the current 30/20/20/30 methodology table versus 25/25/25/25 FAQ mismatch rather than silently choosing whichever weighting supports a preferred narrative.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books