Artificial Analysis v4.2 Changes the AI Leaderboard: What the New Scores Mean
Artificial Analysis changed its Intelligence Index on September 4. Here is why Gemini 3.8 Flash's headline score moved, what v4.2 measures, and how to compare Fable 5.1 and GPT-6 Astra fairly.
Why the leaderboard changed on September 4
Artificial Analysis released Intelligence Index v4.2 on September 4, 2026. This is a methodology update, not a new AI model. The evaluator added AA-Briefcase, a private held-out agentic knowledge-work evaluation, added Surge AI's GDP.pdf for long-context professional document reasoning, removed GPQA Diamond after judging it saturated, doubled the private/held-out share of the index to 40%, and changed parts of its grading infrastructure.
That distinction matters because a model's headline composite can change even when the model itself has not changed. Artificial Analysis' September 2 launch analysis reported Gemini 3.8 Flash (high) at 59 on the then-current Intelligence Index. Its current v4.2 model page now reports 47. That should not be described as a sudden collapse in Gemini capability: the measurement changed.
Sources: Artificial Analysis Index v4.2, Gemini 3.8 Flash launch analysis, and current Gemini 3.8 Flash model page.
What v4.2 measures
Artificial Analysis says v4.2 combines ten evaluations: AA-Briefcase, GDPval-AA v2, tau3-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.
The two additions change the mix materially. AA-Briefcase uses private held-out projects designed around multi-week knowledge work, linked tasks and many source files. GDP.pdf uses 100 PDFs across ten domains, totaling 4,592 pages, and grades responses against 1,275 expert-authored atomic criteria. Artificial Analysis says an answer receives the headline GDP.pdf all-pass credit only when every criterion for the task is satisfied.
The evaluator also says 40% of the index weighting is now private or held out, twice the share in v4.1. That is intended to reduce benchmark gaming, but it also means scores from v4.1 and v4.2 are not cleanly interchangeable.
The current frontier result
Artificial Analysis says Claude Fable 5.1 leads v4.2, followed by GPT-6 Astra. Its current model pages report Fable 5.1 at 66 at max effort with Anthropic's default fallback configuration, and GPT-6 Astra at 61 at max effort. The v4.2 announcement places Meta as the third-ranked lab and Google further down the aggregate index.
Those scores still need configuration labels. Artificial Analysis' Fable evaluation uses Anthropic's production-style fallback behavior, so its 66 is not an isolated base-model-only measurement in every safety-sensitive case. Astra's current page reports a 1-million-token context window and $10/$50 per million input/output tokens, while the Fable page reports the same headline input/output prices but different cache economics.
A useful new v4.2 result is GDP.pdf: Artificial Analysis reports 33.2% for GPT-6 Astra, 28.2% for GPT-5.6 Sol and 26.2% for Claude Fable 5.1. That is one specific long-document evaluation, not proof that Astra is universally better for knowledge work.
Sources: Fable 5.1 model page, GPT-6 Astra model page, and v4.2 methodology announcement.
Gemini 3.8 Flash is the clearest example of why benchmark versioning matters
Google released Gemini 3.8 Flash on September 2 with a separate Cyber variant for trusted defenders. Google lists an introductory API price of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, rising to $1.50/$7.50 from January 1, 2027. Google says the standard Flash model is available through the Gemini API, Google AI Studio, Gemini Enterprise and consumer Google AI Pro/Ultra surfaces; Flash Cyber is restricted through the Fairwind Program.
Google's own launch measurements include 54.9% on HLE-Verified and claims of strong DeepSWE v1.1 and agentic performance. For the Cyber variant, Google reports 47.2% pass@1 on CWE-Bench compared with 47.8% for a leading frontier model in its comparison, plus separate internal vulnerability-discovery measurements. These are vendor-published results and should be labeled as such.
Artificial Analysis' independent release analysis originally put Gemini 3.8 Flash (high) at 59 and $0.58 per Intelligence Index task. Its current v4.2 page reports 47 and $0.74 per task. The sensible interpretation is that the benchmark composition and weighting became harder and more agentic/private—not that the same Gemini endpoint lost twelve points of capability overnight.
Source: Google's Gemini 3.8 Flash and Flash Cyber announcement.
This is not a SWE-bench ranking
Artificial Analysis Intelligence Index v4.2 includes Terminal-Bench v2.1, but it is not SWE-bench Verified and it is not SWE-bench Pro. Those coding benchmarks should remain separate lanes.
SWE-bench Verified is the human-validated 500-task subset of SWE-bench. SWE-bench Pro is a different repository-level benchmark with its own task sets, harness and quality caveats. A model moving up or down on Artificial Analysis v4.2 therefore does not establish that it improved or regressed on either SWE-bench benchmark.
This separation is especially important when comparing coding agents: benchmark version, scaffold, tool permissions, reasoning effort and pass criteria can change the ordering.
Sources: SWE-bench Verified and SWE-bench Pro public leaderboard.
Community reaction: skepticism is evidence of disagreement, not a counter-benchmark
Public discussion around the index is mixed. In two September 4-5 LocalLLaMA threads, some users argued that Artificial Analysis scores do not match their personal model experience, questioned reproducibility, or said the composite has become too agentic to represent every use case. Other replies defended benchmarks as useful for narrowing a large model field even when they do not replace workload-specific testing.
These are anecdotes from a self-selected community, not controlled measurements. They are useful mainly as evidence that users disagree about what a single composite should represent. Search-accessible X results did not provide a sufficiently attributable post tied to the v4.2 methodology update during this verification pass, so no X quote or consensus is claimed.
Discussion sources: LocalLLaMA: Artificial Analysis Index is NOT Representative of real World Performance and LocalLLaMA: Any good alternatives to Artificial Analysis?.
Practical takeaway
Do not compare a v4.1 headline score with a v4.2 headline score as if the scale were unchanged. Record the benchmark version, exact model/effort configuration, provider, date and harness. For production decisions, pair a broad composite with task-specific evaluations: coding-agent benchmarks for repository work, document benchmarks for long-form analysis, latency and cost-per-success measurements for agents, and your own regression set.
The v4.2 update is valuable precisely because it exposes a common leaderboard mistake: a number without its methodology version is incomplete evidence. Fable 5.1 currently leads this particular composite and Astra is second, but that does not make either model the automatic winner for every coding, multimodal, latency-sensitive or cost-sensitive workflow.
Confidence is high on the v4.2 methodology changes and current model-page figures because they come directly from the evaluator and vendor. Confidence is medium on broad cross-model conclusions because composite weighting is a design choice. Confidence is low on community-wide sentiment because public comments are sparse and self-selected.
This article is built from the source material below. Open the originals for full context and the latest updates.