Analysis
Analysis

Gemini 3.8 Flash Reality Check: AA Index Moves 59→41 After v4.3; Terminal-Bench 4.0 Lands Near 20%

Published Sep 8, 2026 Sources checked Sep 8, 2026

Gemini 3.8 Flash now shows 41 on Artificial Analysis v4.3 after scoring 59 in the earlier index. The change follows a benchmark-suite revision, not proof of model regression. Current evidence puts Terminal-Bench 4.0 near 20%, with high throughput but high-effort TTFT around 16 seconds.

Why the headline score changed

Gemini 3.8 Flash launched on 2 September 2026 as Google's newest Flash-class reasoning model for software engineering, autonomous agents and complex enterprise workflows. At launch, Artificial Analysis reported 59 for its high-reasoning configuration on the then-current Intelligence Index. Five days later, Artificial Analysis introduced Intelligence Index v4.3, and the current Gemini 3.8 Flash high page reports 41.

That 59→41 movement should not be read as proof that the model suddenly became worse. The benchmark suite itself changed. Version 4.3 replaced Terminal-Bench v2.1 with the much harder Terminal-Bench v4.0 and replaced τ³-Banking with AutomationBench-AA, while keeping category weights at Agents 30%, Coding 20%, General 30% and Scientific Reasoning 20%. Artificial Analysis also says private questions or answers now account for 45% of the index weighting.

The useful comparison is therefore: how does Gemini 3.8 Flash perform under each named benchmark version and harness, and what does that mean for real deployments?

Primary and independent sources used in this audit:

Exact model identity and access

Google's stable API model code is gemini-3.8-flash. The API documentation lists text, image, video, audio and PDF inputs with text output. The input limit is 1,048,576 tokens and the maximum output is 65,536 tokens.

The model supports caching, code execution, file search, function calling, Google Maps grounding, Search grounding, structured output and URL context. Computer use is listed as Preview. Thinking levels are low, medium and high; the documentation says minimal is not supported. The same page says native audio generation, native image generation and the Live API are not supported for this model.

This matters when comparing benchmarks: a one-million-token context window and tool access can help long-horizon workflows, but a benchmark score does not imply that every product surface exposes the same tools or agent scaffold.

Pricing: introductory token rates are temporary

Google's standard paid pricing through 31 December 2026 is:

Item Through Dec. 31, 2026 Starting Jan. 1, 2027
Input $0.75 / 1M tokens $1.50 / 1M
Output, including thinking tokens $3.75 / 1M $7.50 / 1M
Cached input $0.075 / 1M $0.15 / 1M

Google also offers Batch/Flex and Priority inference with different rates. The launch post explicitly warns that Gemini 3.8 can perform more reasoning steps and tool calls on difficult work, especially at higher effort.

That means the low per-token price is only one part of cost. Artificial Analysis initially measured $0.58 per Intelligence Index task under the pre-v4.3 suite, but its current v4.3 model page reports $1.24 per task and 170 million output tokens across the index evaluation. Those per-task values belong to different benchmark versions and workloads, so they should not be treated as a clean price increase for an identical task set.

Google's coding evidence: separate the benchmark generations

Google's September evaluation PDF is unusually useful because it documents both benchmark versions and their evaluation source.

For DeepSWE v1.1, Google reports 73.7% for Gemini 3.8 Flash. Google states that its Gemini result is self-computed with a mini-SWE-agent harness at high thinking, while other rows come from the public DeepSWE leaderboard.

For Terminal-Bench 2.1, Google reports 89.4%. Its methodology says Gemini results are self-computed and use the default Terminus 2 agent harness.

For Terminal-Bench 4.0, the same results table reports only 19.1% for Gemini 3.8 Flash, and Google says those figures come from the official public leaderboard at the highest reported thinking level.

These numbers are not contradictory. Terminal-Bench 4.0 is a materially different and harder benchmark revision. A model scoring 89.4 on 2.1 and 19.1 on 4.0 is not evidence of an overnight 70-point capability collapse.

Google-reported benchmark snapshot

Benchmark Gemini 3.8 Flash Provenance in Google's methodology
DeepSWE v1.1 73.7% Gemini self-computed, mini-SWE-agent, high thinking
Terminal-Bench 2.1 89.4% Gemini self-computed, Terminus 2
Terminal-Bench 4.0 19.1% Official public leaderboard
GDPval-AA v2 1545 Elo Artificial Analysis
Vals Finance Agent v2 61.4% Vals
Harvey Legal Agent Benchmark 10.0% Vals
GDP.PDF 35.0% Self-computed
CharXiv Reasoning 86.2% Self-computed
LVBench 87.8% agentic / 87.1% static Self-computed
HLE-Verified 54.9% Self-computed on 1,811 verified questions
OSWorld-2.0 59.0% Self-computed under documented settings

The table should not be turned into one universal rank. Different rows use different harnesses, providers, evaluators and task families.

SWE-bench Verified: no exact current score accepted

This review did not find an authoritative, current Gemini 3.8 Flash SWE-bench Verified result with enough information to pin the exact model, benchmark revision, scaffold, task count, retries and endpoint.

Google's current four-page evaluation methodology does not include a SWE-bench Verified row. DeepSWE v1.1 is a separate software-engineering evaluation and must not be relabeled as SWE-bench Verified. Terminal-Bench is also a different benchmark family.

Because SWE-bench Verified has become highly saturated for frontier systems and results can be very harness-sensitive, this article leaves the exact Gemini 3.8 Flash Verified score unreported rather than borrowing a number from a secondary table or another Gemini checkpoint.

SWE-bench Pro: also keep it separate

The same rule applies to SWE-bench Pro. Several secondary summaries circulate Gemini 3.8 Flash figures, but this run did not recover a directly accessible, current primary-source result with enough method detail to promote one as the authoritative score for the exact gemini-3.8-flash checkpoint.

So this article does not substitute DeepSWE, Terminal-Bench, SWE-Atlas or a score belonging to another Gemini variant. A future update should add SWE-bench Pro only when the exact checkpoint, benchmark revision, harness/scaffold, number of tasks, retries and evaluation date are pinned.

Artificial Analysis v4.3: current independent picture

Artificial Analysis's current Gemini 3.8 Flash high page reports an Intelligence Index v4.3 score of 41. Its current direct comparison against Gemini 3.7 Flash high shows:

Artificial Analysis v4.3 measure Gemini 3.8 Flash high Gemini 3.7 Flash high
Intelligence Index 41 39
AA-Briefcase 1202 1116
GDPval-AA v2 1464 1433
AutomationBench-AA 60% 62%
Terminal-Bench v4.0 20% 14%
SciCode 57% 57%
Humanity's Last Exam 48% 48%
GDP.pdf 21% 24%
CritPt 18% 14%
AA-LCR v1.1 81% 82%

This is a much more useful apples-to-apples comparison than comparing the old 59 and new 41 as if the index were unchanged. Under v4.3, 3.8 high is two index points ahead of 3.7 high, but it does not win every constituent evaluation.

What AutomationBench-AA measures

Artificial Analysis runs a held-out 657-task split of Zapier's AutomationBench v1.0.6 across Finance, HR, Marketing, Operations, Sales and Support. Each task is run once with a 50-turn cap. The headline score is the share of objectives completed, but any guardrail violation makes that task score zero. A separate “Tasks Completed” metric counts workflows where every objective is completed without a guardrail violation.

That distinction matters: a 60% AutomationBench-AA score is not the same thing as saying the model fully completes 60% of workflows.

What Terminal-Bench 4.0 measures

Artificial Analysis evaluates all 66 Terminal-Bench 4.0 tasks three times and reports average pass@1. Its current Gemini comparison rounds Gemini 3.8 Flash high to 20%, very close to Google's official-leaderboard snapshot of 19.1%. The small difference is consistent with different snapshots or rounding rather than a reason to merge the two numbers.

Why 59 became 41

Artificial Analysis's 2 September launch article reported Gemini 3.8 Flash high at 59 under the pre-v4.3 Intelligence Index. On 7 September, Artificial Analysis changed two agentic constituents:

  • Terminal-Bench v2.1 → v4.0
  • τ³-Banking → AutomationBench-AA

The new Terminal-Bench uses 66 harder terminal tasks and the new automation evaluation uses a private 657-task held-out set. Private tasks or answers now represent 45% of the index weight. Category weights remain Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%.

Therefore, 59 and 41 are scores on different composite benchmarks. They are both legitimate historical measurements, but they are not a clean before/after capability test of the same model on the same tasks.

Speed is excellent, but high-reasoning latency is not “instant”

Artificial Analysis currently measures about 281 output tokens per second from Google's API for Gemini 3.8 Flash high, placing it among the fastest reasoning systems in its comparison set.

But its current high-reasoning page also measures time to first token at roughly 16 seconds. Those two metrics describe different stages of a request: once output starts, generation is very fast; a difficult high-effort task can still spend significant time reasoning before the first visible token.

This aligns with Google's own launch explanation that 3.8 Flash may use extra reasoning steps and repeated tool calls on complex tasks. “Flash” is therefore best understood as a cost/throughput class, not a guarantee that every agentic request finishes immediately.

Multimodal and computer-use evidence has its own caveats

Google reports 35.0% GDP.PDF, 86.2% CharXiv Reasoning, 87.8% agentic / 87.1% static LVBench, and 59.0% OSWorld-2.0. The methodology notes that several of these rows are self-computed and that LVBench uses different frame counts across vendors because of API limits: 1,024 frames for Gemini and GPT-5.6 models versus 300 for Claude models.

For OSWorld, Google documents a separate evaluation setup including a 1080p environment, up to 500 steps and its evaluator settings. These are useful operational numbers, but they should not be generalized to every computer-use product surface or agent framework.

Public feedback is mixed—and anecdotal

Early public discussion shows both strong successes and frustrating agent-loop behavior.

A 4 September r/aiagents post describes Gemini 3.8 Flash solving an Android photo-editor bug after two other frontier systems failed in that user's attempt. The post is a useful case report, but it is an uncontrolled n=1 comparison.

Several r/google_antigravity threads from 5–7 September report the opposite experience: repeated file rereading, loops, high quota consumption and cases where users preferred Gemini 3.7 Flash. One thread includes replies from users who had the opposite result and preferred 3.8 on their tasks.

These reports should be treated as self-selected anecdotes, not a measured consensus. Different prompts, repository sizes, effort settings, Antigravity versions and stochastic runs can change outcomes. No stable primary X post with a reproducible exact-checkpoint benchmark—complete with prompt/task, harness, settings and date—was accepted in this run, so no X “consensus” is claimed.

Public discussion links:

Practical tradeoffs

For developers choosing Gemini 3.8 Flash today, the evidence supports a nuanced view:

  • Strengths: low introductory token price, 1M-token context, very high measured output throughput, strong multimodal/tool support, and meaningful gains over 3.7 Flash on some current v4.3 agentic/coding measures.
  • Costs: higher effort can consume many more reasoning/output tokens and take longer before visible output; introductory pricing doubles on 1 January 2027.
  • Benchmark caveat: the old AA score of 59 and current 41 are not directly comparable because v4.3 changed the evaluation suite.
  • Coding caveat: Google reports 73.7 DeepSWE v1.1 and 19.1 Terminal-Bench 4.0 under different provenance; neither should be relabeled as SWE-bench Verified or SWE-bench Pro.
  • Product caveat: computer use is preview, Live API is not supported for this model, and native image/audio generation are not supported according to the current API docs.
  • Feedback caveat: practitioner reports are mixed and uncontrolled, so teams should run their own repository-level acceptance tests at the intended effort setting.

Verdict

Gemini 3.8 Flash is still a compelling cost/throughput model, but the most informative update since launch is methodological rather than promotional. The current independent Artificial Analysis v4.3 score is 41, not the launch-era 59, because the index now contains harder and different agentic evaluations. On the new suite, Gemini 3.8 Flash high scores about 20% on Terminal-Bench 4.0 and 60% on AutomationBench-AA, while Google's separate official snapshot reports 19.1% Terminal-Bench 4.0 and a vendor-run 73.7% DeepSWE v1.1.

The next high-value evidence is an exact, reproducible Gemini 3.8 Flash run on SWE-bench Verified and SWE-bench Pro separately, with the model endpoint, benchmark revision, harness, task count, retries, tool policy and date all pinned. Until then, those two benchmark fields should remain unfilled rather than inferred from nearby coding suites.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books