Artificial Analysis v4.2 Reality Check: Why Fable 5.1 Fell 66→57 and Gemini 3.8 Flash 59→47
Artificial Analysis changed its Intelligence Index on September 4. Lower scores for Fable 5.1, Astra, Gemini 3.8 Flash and Muse Spark 1.3 mostly reflect a harder v4.2 benchmark, not sudden model regressions.
Artificial Analysis changed its flagship Intelligence Index on September 4, 2026, and the update creates an easy trap for anyone following frontier-model leaderboards: a model's new v4.2 score can be much lower than the number published at launch even when the model itself has not changed.
The clearest examples are Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash and Muse Spark 1.3. Artificial Analysis reported Fable 5.1 max at 66 on September 1, Gemini 3.8 Flash high at 59 and Muse Spark 1.3 max at 62 on September 2, and Astra max at 61 on September 3. Its current v4.2 pages now show 57, 47, 53 and 55, respectively.
Those are not clean capability regressions. The benchmark changed underneath the models.
Cross-version score changes are not model changes
A useful way to read the current leaderboard is to put the measurement version beside every score:
| Model / configuration | Pre-v4.2 published score | Current v4.2 score | Raw difference |
|---|---|---|---|
| Claude Fable 5.1, max with fallback | 66 | 57 | -9 |
| GPT-6 Astra, max | 61 | 55 | -6 |
| Gemini 3.8 Flash, high | 59 | 47 | -12 |
| Muse Spark 1.3, max | 62 | 53 | -9 |
The raw subtraction is mathematically correct but not an estimate of how much the models became worse. The September 3 Astra article explicitly labels its detailed Intelligence Index breakdown as v4.1.1. The September 4 v4.2 release changes the task mixture, weights and grading infrastructure. An old 61 and a new 55 therefore answer different composite-evaluation questions.
Primary references: Artificial Analysis v4.2 announcement, Fable 5.1 launch evaluation, Astra launch evaluation, Gemini 3.8 Flash launch evaluation, and Muse Spark 1.3 launch evaluation.
What v4.2 changed
The new Index adds two substantial evaluations. AA-Briefcase is a private-held-out agentic knowledge-work benchmark covering multi-week simulated professional projects. Artificial Analysis says the public methodology currently uses 91 tasks across four scenarios and gives it 15% of the total Index. GDP.pdf adds long-context professional document reasoning over 100 PDFs across ten domains and 4,592 pages; the evaluation uses 1,275 expert-authored atomic criteria, and a task's headline All-pass result requires every criterion for that task to be satisfied.
At the same time, Artificial Analysis removed GPQA Diamond, saying frontier performance had saturated it.
The most consequential anti-gaming change is provenance. Artificial Analysis says 40% of v4.2's weighting is now private held-out data, twice v4.1's share. The held-out portion includes AA-Briefcase, AA-Omniscience and CritPt solutions. That can reduce direct benchmark optimization and contamination risk, but it also means an outside researcher cannot independently rerun the complete composite from public test data alone.
Grading changed too. Artificial Analysis reports corrected answer-key problems and a grading-system prompt for AA-LCR v1.1, re-anchored Elo scales and revised sampling for GDPval-AA v2 and AA-Briefcase, and a more robust SciCode sandbox intended not to fail slow but correct code. Any one of those changes can move a model's aggregate score without the model endpoint changing.
Methodology reference: Artificial Analysis Intelligence Benchmarking Methodology.
What is actually inside the 57, 55, 53 or 47
The v4.2 composite uses four capability categories, not one monolithic exam:
- Agents — 30%: AA-Briefcase 15%, GDPval-AA v2 10%, and tau3-Banking 5%.
- Coding — 20%: Terminal-Bench v2.1 10% and SciCode 10%.
- General — 30%: AA-Omniscience 15%, GDP.pdf 10%, and AA-LCR v1.1 5%.
- Scientific reasoning — 20%: Humanity's Last Exam 10% and CritPt 10%.
Sample sizes and repeat policies differ substantially. The methodology lists 6,000 AA-Omniscience questions with one repeat, 2,158 HLE questions with one repeat, 288 SciCode subproblems with three repeats, 100 GDP.pdf tasks with five repeats, 97 tau3-Banking tasks with five repeats, 89 Terminal-Bench 2.1 tasks with three repeats and 70 CritPt items with five repeats. AA-Briefcase and GDPval-AA use agentic file-delivery workflows and Elo-style judging rather than simple multiple-choice accuracy.
Artificial Analysis estimates the overall v4.2 Index has a 95% confidence interval narrower than ±1 point, based on experiments with more than ten repeats on selected models across the included datasets, while warning that individual evaluation confidence intervals can be wider. That estimate is useful for same-version comparisons; it does not repair comparability across different Index versions.
The suite is also primarily English and text based. Artificial Analysis reports image, speech and multilingual performance separately. A high Intelligence Index position therefore should not be silently converted into a claim that the same model leads multimodal perception, speech or multilingual work.
Private tests improve resistance to gaming, but limit reproducibility
There is a real methodological tradeoff here. Private tasks can make it harder for model developers to train directly against exact questions or tune prompts to known test cases. That is especially useful in a period when public benchmarks saturate quickly.
But a 40% private composite is not fully reproducible by an unaffiliated lab. Artificial Analysis partly offsets that limitation by documenting its harnesses and releasing a lighter AA-Briefcase example set. Its methodology says AA-Briefcase runs through the open Stirrup harness, gives an agent up to 500 turns per task, uses a code-execution environment, and disables internet access inside the sandbox. Individual tasks are currently run independently rather than letting the model carry its own prior task outputs forward across the multi-week scenario.
That is much more informative than a bare leaderboard number, but the correct evidence label remains independent third-party evaluation with a substantial private test component, not a fully public reproducible benchmark.
Current same-version comparisons are more meaningful
Once every model is measured under v4.2, comparisons become more defensible. Artificial Analysis currently places Fable 5.1 max at 57, Astra max at 55, Muse Spark 1.3 max at 53, and Gemini 3.8 Flash high at 47. Those same-version gaps can be interpreted subject to the Index's estimated uncertainty, configuration choices, endpoint drift and the relevance of the composite to a user's workload.
Even then, the aggregate should not erase the component benchmarks. Artificial Analysis itself says specific evaluations can matter more for specific use cases. A coding team should inspect Terminal-Bench and SciCode; a knowledge-work deployment should examine AA-Briefcase and GDPval; a hallucination-sensitive application should inspect AA-Omniscience rather than treating the aggregate as sufficient.
Current model references: Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, and Muse Spark 1.3.
Cost per task also changed meaning
Current v4.2 model pages report a weighted Cost per Intelligence Index task. For example, Fable 5.1 max is currently about $6.12 per Index task, Astra max about $2.57, Gemini 3.8 Flash high about $0.74, and Muse Spark 1.3 max about $0.96.
These figures are useful for comparing the cost of running this particular composite, but they are not API list prices and not a universal cost per useful task. Artificial Analysis calculates evaluation cost from token prices and observed token use, then weights results by each benchmark's share of the Index. When the task mix and weights change, cost per Index task can change even if a vendor's token tariff does not.
The model pages separately report API prices, speed, latency and context. Those serving metrics should remain separate from intelligence scores because a deployment can prefer a lower-scoring model if it is materially faster, cheaper, easier to access or better aligned to the exact workload.
SWE-bench Verified and SWE-bench Pro stay separate
Neither SWE-bench Verified nor SWE-bench Pro is part of the v4.2 Intelligence Index. Terminal-Bench v2.1 and SciCode make up the Index's 20% coding category, but neither is a synonym for SWE-bench.
A model's Intelligence Index score moving from one version to another therefore says nothing by itself about whether its SWE-bench Verified or SWE-bench Pro result changed. Model-specific SWE-bench numbers should be reported only when the exact model, harness, dataset version, agent setup and result are identified. The same rule applies to DeepSWE, Terminal-Bench, CyberGym, CWE-bench and other specialized evaluation families.
Public feedback is already arguing about benchmark choice
A fresh r/singularity discussion around the v4.2 change illustrates why benchmark interpretation matters. Commenters disagree over whether alternative spatial benchmarks better reflect general capability and criticize individual benchmark choices such as Terminal-Bench 2.1 and CritPt. These are self-selected opinions, not reproducible evidence that v4.2 is wrong. They do show that leaderboard consumers care about task composition, not only the final number.
Discussion reference: r/singularity — AA Intelligence Index Changes.
A bounded search for an attributable X post with a stable direct status URL and independent controlled measurements did not produce one. Artificial Analysis's own X activity is vendor/operator commentary about its benchmark rather than independent validation, so no X consensus is claimed here.
Bottom line
Artificial Analysis v4.2 is a benchmark reset, not evidence of a sudden frontier-model collapse. The update deliberately makes the Index harder, more agentic and more private: it adds AA-Briefcase and GDP.pdf, removes saturated GPQA Diamond, doubles private held-out weighting to 40%, and changes several grading systems.
That makes v4.2 potentially more useful for current model selection, while creating a clear reporting rule: compare models within the same Index version and never narrate a cross-version score drop as a model capability drop without a controlled re-evaluation that holds the benchmark constant.
For practical decisions, keep the aggregate, component benchmarks, API price, cost per task, context, latency, output speed and access conditions in separate columns. A one-number leaderboard is a useful summary; it is not a substitute for matching the evaluation to the work you actually need the model to do.
This article is built from the source material below. Open the originals for full context and the latest updates.