Gemini 3.8 Flash Reality Check: 73.7 DeepSWE, 19.1 Terminal-Bench 4.0 and AA v4.2 at 47
Google’s current Gemini 3.8 Flash evidence is strong but uneven: 73.7% on DeepSWE v1.1, 89.4% on Terminal-Bench 2.1, only 19.1% on Terminal-Bench 4.0, and a current Artificial Analysis v4.2 score of 47. The model is cheap and fast per token, but Google explicitly says it may use more reasoning steps and tokens than 3.7 Flash.
Google released Gemini 3.8 Flash on September 2, 2026 as a generally available, stable Flash-tier model for long-horizon software engineering, autonomous agents and enterprise knowledge workflows. The exact API model ID is gemini-3.8-flash. Google’s developer documentation lists text, image, video, audio and PDF input, text output, a 1,048,576-token input limit, a 65,536-token output limit, and low, medium and high thinking levels. minimal is not supported.
The important part of the launch is not one headline score. The current evidence shows a model that is highly competitive on some coding and agent benchmarks, much weaker on a harder terminal benchmark, very fast in independent API measurements, and potentially more expensive per completed task than its unchanged per-token price suggests.
Current Google benchmark evidence
Google’s current evaluation methodology says Gemini results are pass@1 unless otherwise noted and that Gemini 3.8 Flash was run through the Gemini API with default sampling settings. It also explains that comparator results come from a mixture of provider-reported results and public leaderboards, so the full table is not one perfectly controlled head-to-head experiment.
For DeepSWE v1.1, Google reports 73.7% for Gemini 3.8 Flash. Google says this result is self-computed using the mini-swe-agent harness with high thinking. Comparator values are taken from the public DataCurve leaderboard. That makes the result useful, but not the same thing as an independent DataCurve submission performed by a third party.
For Terminal-Bench 2.1, Google’s current methodology table lists 89.4% for Gemini 3.8 Flash, compared with 85.8% for Gemini 3.7 Flash, 89.1% for Claude Opus 5 and 88.8% for GPT-5.6 Sol. Google says Gemini results are self-computed and the comparison is restricted to the default Terminus 2 harness.
For Terminal-Bench 4.0, the same current table lists only 19.1% for Gemini 3.8 Flash. That is dramatically below Claude Opus 5 at 51.8% and GPT-5.6 Sol at 37.3% in Google’s table. Google says Terminal-Bench 4.0 values come from the official public leaderboard and use the highest reported thinking level. The 2.1 and 4.0 results therefore should not be merged into one generic “Terminal-Bench” score: they are different benchmark versions with different difficulty and evidence provenance.
Google also reports 54.9% on HLE-Verified. Its methodology says this is a self-computed run over the full 1,811-item verified set, including 668 verified items from the original Humanity’s Last Exam set and 1,143 revised items, while excluding 689 original items identified as uncertain. That is not interchangeable with scores on the original HLE dataset.
Other current rows include 61.4% on Vals Finance Agent v2, 10.0% all-pass on Harvey’s Legal Agent Benchmark, 35.0% all-pass on GDP.PDF, 86.2% on CharXiv Reasoning, 59.0% on OSWorld 2.0, and 56.5% on the difficult BioMysteryBench split. Each row has its own harness and sourcing notes, so they should be interpreted by task rather than averaged informally.
Why the widely repeated 90.8% Terminal-Bench number needs caution
A number of early launch write-ups repeat 90.8% for Terminal-Bench 2.1. Google’s current September evaluation-methodology table now shows 89.4%. Because the primary methodology page is the source of record available at verification time, this article uses 89.4%. There is not enough primary evidence here to determine whether 90.8% was an earlier run, a chart revision, a rounding or reporting issue, or a different setup.
This is a useful example of why benchmark articles should preserve benchmark version, harness and verification date rather than copying launch-day numbers indefinitely.
SWE-bench Verified and SWE-bench Pro are separate
The current Google launch post, model card and Gemini 3.8 Flash evaluation-methodology document reviewed for this article do not provide a first-party Gemini 3.8 Flash score for SWE-bench Verified or SWE-bench Pro.
Secondary sites currently circulate numbers for both benchmark families, but they do not establish a sufficiently clear primary provenance in the evidence reviewed here. Those numbers are therefore not promoted into this article as verified Gemini 3.8 Flash results. DeepSWE v1.1, Terminal-Bench 2.1 and Terminal-Bench 4.0 are kept separate rather than being silently relabeled as SWE-bench.
Artificial Analysis v4.2: 47, fast output, high token use
Artificial Analysis currently scores Gemini 3.8 Flash (high) at 47 on Intelligence Index v4.2. Its current first-party Google API measurement shows about 280.8 output tokens per second, roughly 11.32 seconds to first token, $0.74 weighted cost per Intelligence Index task, and 140 million output tokens generated across the Intelligence Index evaluation.
Those figures add an important practical tradeoff. The model is fast once output begins, but the high-reasoning configuration can spend substantial time thinking before the first answer token and can be verbose across difficult tasks.
The v4.2 number also should not be compared naively with older Gemini 3.8 Flash Artificial Analysis snapshots. Artificial Analysis changed its Intelligence Index on September 4, adding AA-Briefcase and GDP.pdf, removing the saturated GPQA Diamond benchmark, upgrading grading, and increasing the weight of private held-out test sets to 40%. Earlier snapshots that placed Gemini 3.8 Flash around 59 belong to the older index. The current 47 is therefore not evidence that the same model suddenly lost 12 capability points.
Price is unchanged per token, not necessarily per task
Google’s current Gemini Developer API price is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, with output billing including thinking tokens. On January 1, 2027, Google says standard pricing becomes $1.50 input and $7.50 output per million tokens. Context caching is $0.075 per million tokens through the end of 2026 and $0.15 beginning in 2027.
Google explicitly says Gemini 3.8 Flash “works harder” on complex tasks by taking extra reasoning steps and calling tools iteratively. It warns that the model may use more tokens, especially at higher effort levels, and recommends lower effort or continuing to use Gemini 3.7 Flash when compute efficiency matters most.
That distinction matters: identical unit pricing does not guarantee identical cost per completed task.
Gemini 3.8 Flash Cyber is a different access product
Google launched Gemini 3.8 Flash Cyber alongside the general model, but it is not simply another public API reasoning level. It is a cybersecurity-focused variant restricted to trusted defenders through Google’s Fairwind Program. Google reports more than 70% success on an internal vulnerability-discovery benchmark spanning 20 programming languages and 47.2% pass@1 on CWE-Bench for automated patching.
Those cyber scores should not be attributed to the generally available gemini-3.8-flash endpoint unless Google explicitly documents the same evaluation for that public model.
Public feedback is mixed and anecdotal
Fresh public discussion is not a controlled evaluation. On September 6, one Reddit user in r/GeminiAI said Gemini 3.8 Flash performed better than expected for coding in Antigravity, while another user in r/google_antigravity reported repeated file-reading loops, very high quota consumption and a failed coding result. A separate r/opencode thread praised a one-prompt frontend/Three.js test but also contained replies describing inconsistent precision and instruction-following.
These reports are useful for discovering failure modes and workloads worth testing. They are self-selected anecdotes with different prompts, products, account limits and harnesses, so they do not establish model-wide success rates or community consensus.
Practical read
Gemini 3.8 Flash is a compelling low-unit-price agent model, especially when a workload benefits from fast decoding, multimodal input, a million-token context window and repeated tool use. The current evidence is strongest for moderate-horizon coding and structured agent workflows.
The main cautions are equally concrete: the harder Terminal-Bench 4.0 result is far below the 2.1 score, high-effort reasoning can increase token use and time-to-first-answer, the promotional API price doubles in January 2027, and several attractive launch comparisons mix self-computed Google runs with third-party or provider-reported comparator numbers.
The fairest conclusion is not that Gemini 3.8 Flash “beats” frontier models. It is that Google has shipped a fast, relatively inexpensive model that is near-frontier on some agentic coding tasks, clearly weaker on other long-horizon tasks, and unusually sensitive to the chosen benchmark, harness and reasoning level.
This article is built from the source material below. Open the originals for full context and the latest updates.