DeepSeek V4 Pro 0813 Reality Check: 96.4 SWE-bench Verified, 54.68 Terminal-Bench, and 55.4 Pro Is Preview-Era
DeepSeek V4 Pro 0813 reaches 96.4% on Vals' archived SWE-bench Verified, but Terminal-Bench 2.1 ranges from 87.9 vendor-run to 54.68 at Vals. The widely repeated 55.4 SWE-bench Pro score belongs to DeepSeek's preview-era card, not the exact 0813 release table.
Why this benchmark audit matters
DeepSeek-V4-Pro-0813 is the August 13, 2026 production release of DeepSeek V4 Pro. The useful update now is not another launch recap: enough independent testing has accumulated to show that its coding-agent results change sharply with the harness, while a widely repeated SWE-bench Pro score is attached to the wrong checkpoint.
The evidence ledger is unusually clear. Vals measures the exact 0813 checkpoint at 96.40% on SWE-bench Verified using its shared minimal bash-only harness, while Fireworks separately reports 95.2% on the same 500-task benchmark family. On Terminal-Bench 2.1, however, DeepSeek reports 87.9, Fireworks reports 76.4, and Vals reports 54.68% across three full trials. Those are not a matched replication and should not be averaged.
SWE-bench Pro must be kept separate. The familiar 55.4% figure appears on DeepSeek's older DeepSeek-V4-Pro card in the V4-Pro Max column. The exact DeepSeek-V4-Pro-0813 release card does not publish a SWE-bench Pro row, and Scale's current public SWE-Bench Pro table does not show an exact 0813 entry in the reviewed leaderboard. The correct value for exact 0813 Pro is therefore not established.
Primary release and API sources:
Independent evaluation sources:
- Vals — DeepSeek V4 Pro 0813
- Vals — SWE-bench Verified
- Fireworks — DeepSeek V4 Pro evaluation
- Artificial Analysis — DeepSeek V4 Pro 0813
- Scale — SWE-Bench Pro public leaderboard
Exact identity, context and access
DeepSeek describes DeepSeek-V4-Pro-0813 as the official release superseding the V4 Pro preview checkpoint. It retains the V4 Pro architecture and attaches the DSpark speculative-decoding module. The model weights are public under the MIT license.
DeepSeek's first-party API keeps the stable model ID deepseek-v4-pro, while its current pricing page identifies the served version as DeepSeek-V4-Pro-0813. The API documentation lists a 1M-token context, 384K maximum output, thinking support, tool calls, Responses API support and a Pro concurrency limit of 500. A benchmark report should name the dated checkpoint even if production code calls an undated alias.
DeepSeek's official 0813 agent table
The exact 0813 model card reports 87.9 Terminal-Bench 2.1, 61.5 NL2Repo, 83.3 CyberGym, 62.7 DeepSWE, 74.1 Toolathlon-Verified, 25.7 Agents' Last Exam, 31.8 AutomationBench Public, 71.1 DSBench-FullStack and 67.2 DSBench-Hard. DeepSeek explicitly labels the two DSBench sets internal.
The configuration footnote matters: DeepSeek says public code-agent tasks use DeepSeek Harness minimal mode, max reasoning effort, temperature=1.0 and top_p=0.95. The 87.9 Terminal-Bench result is therefore a model-plus-harness result, not a harness-free property of the weights.
The exact 0813 release table does not contain SWE-bench Verified or SWE-bench Pro. Those scores need separate checkpoint-specific evidence.
SWE-bench Verified: strong independent evidence, but a saturated benchmark
Vals reports 96.40% for DeepSeek V4 Pro 0813, second on its preserved leaderboard behind Claude Opus 5 at 97.00%. Vals runs the 500-task Verified set in isolated Docker environments and uses a shared minimal bash-tool-only agent harness. Models receive bash and must navigate, search, edit and patch with standard command-line tools. Vals explicitly notes that SWE-bench measures both the model and agent harness.
There is an important September status change: Vals now marks SWE-bench Verified an Archived Benchmark because performance has saturated. Its September 1 page says seven of 86 evaluated models reached at least 95%, leaving little room to separate frontier systems. The 96.40% result remains a valid preserved exact-checkpoint measurement; it should not be turned into a universal ranking for new coding agents.
Fireworks independently reports 95.2% across 500 SWE-bench Verified tasks for 0813. The two non-DeepSeek results are directionally consistent, but they use different execution stacks and should remain separate rather than being averaged into a synthetic score.
SWE-bench Pro: 55.4 belongs to the earlier V4 Pro card
DeepSeek's older DeepSeek-V4-Pro card contains a comparison-across-modes table where the V4-Pro Max column reports 80.6 SWE Verified, 55.4 SWE Pro and 76.2 SWE Multilingual. The same table reports Terminal-Bench 2.0 at 67.9.
That is not the exact 0813 release table. The later DeepSeek-V4-Pro-0813 card is a separate checkpoint page and does not list SWE-bench Pro. Scale's current SWE-Bench Pro public page describes 731 public tasks within a 1,865-task benchmark spanning public, held-out and private subsets; the reviewed public leaderboard does not show an exact DeepSeek V4 Pro 0813 row.
So the benchmark status should be recorded as:
| Evaluation | Exact 0813 status |
|---|---|
| SWE-bench Verified | 96.40% — Vals independent, archived board |
| SWE-bench Pro | Not established for exact 0813 |
| 55.4 SWE-bench Pro | Earlier DeepSeek V4-Pro Max / preview-era table |
This does not predict what 0813 would score on Pro. It only prevents silently moving an older score onto a newer checkpoint.
Terminal-Bench 2.1: 87.9 vs 76.4 vs 54.68
Terminal-Bench provides the clearest harness-sensitivity signal.
DeepSeek: 87.9. Vendor-run with DeepSeek Harness minimal mode, max effort, temperature 1.0 and top-p 0.95.
Fireworks: 76.4. Fireworks reports this over 89 Terminal-Bench 2.1 tasks. Its methodology note says DeepSeek's published runs use the native DeepSeek Harness (dsh) and native tool semantics. Fireworks also notes that V4 Pro does not ship a conventional Jinja chat template and says a minimal-first bash/editor surface produced more consistent tool calls in its own testing than immediately exposing a larger generic catalog.
Vals: 54.68%. Vals reports 54.68% across three full Terminal-Bench 2.1 trials, with 28.89% on its hard-task slice, for the exact 0813 model evaluated at max reasoning effort.
The DeepSeek-to-Vals spread is 33.22 percentage points. The evidence does not justify calling either measurement fake: these are not identical harnesses with identical inference and tool policies. The defensible conclusion is that V4 Pro 0813 is highly harness-sensitive on long-horizon terminal work in current public evidence. Production teams should benchmark the model under the same tool surface and orchestration they plan to deploy.
Artificial Analysis v4.3: current score 36 is not a 17-point model regression
Reuters reported an Artificial Analysis Intelligence Index score of 53 around the August launch. Artificial Analysis now reports 36 for DeepSeek V4 Pro 0813 max on Intelligence Index v4.3.
That difference must not be described as a clean model regression. Artificial Analysis changed the index: on September 7, v4.3 upgraded Terminal-Bench from 2.1 to 4.0 and replaced tau3-Banking with AutomationBench-AA, while preserving a versioned composite across ten evaluations. A score on the old index and a score on v4.3 are different measurement frameworks.
At this verification time, Artificial Analysis also reports 72.4 output tokens/second, 1.66 seconds time to first token, about $0.67 average cost per Intelligence Index task, a 1M context, and classifies the model as 1.6T total / 49B active parameters with MIT/open weights. Speed and latency are live API observations rather than service-level guarantees.
Current first-party pricing
DeepSeek currently uses peak/off-peak pricing for V4 Pro:
| Token category | Off-peak | Peak |
|---|---|---|
| Cache-hit input / 1M | $0.022 | $0.044 |
| Cache-miss input / 1M | $0.66 | $1.32 |
| Output / 1M | $1.98 | $3.96 |
DeepSeek defines peak periods as 01:00-04:00 UTC and 06:00-10:00 UTC, Monday-Friday; other hours are off-peak. This is why quoting only $1.32 in / $3.96 out is incomplete: those are peak cache-miss/input-output rates. Real workload economics depend on cache hits, time of day, retries, reasoning effort and successful-task rate.
Public feedback: useful anecdotes, not a consensus
A launch-day r/DeepSeek post on August 13 described unusually short reasoning on difficult prompts and suspected a temporary rollback or deployment/configuration issue. Importantly, the same post said a similar 3D-game task produced a much stronger result a few hours later and suggested the problem might be cluster deployment, inference configuration or server-side harness behavior rather than the base model. That is a self-selected launch anecdote, not proof of a weight-level defect.
A product developer on r/LoreMateAI wrote on August 15 that switching its backend to deepseek-v4-pro-0813 was a drop-in change and an improvement in its own testing. An August 24 cache complaint on a third-party route was later attributed by the poster to that provider after the first-party DeepSeek API restored normal cache behavior. These reports point in different directions and reinforce the need to separate model, provider and harness effects.
A bounded X search did not surface a stable direct practitioner post for 0813 with a reproducible test method that added stronger evidence than the benchmark sources above. X trend summaries were excluded, and no X consensus is inferred.
Verdict
The strongest claim supported by independent evidence is that DeepSeek V4 Pro 0813 is exceptionally strong on SWE-bench Verified-style patch resolution, with Vals at 96.40% and Fireworks at 95.2%. That signal is real, but Vals has archived the benchmark because frontier performance is saturated.
The long-horizon agent story is less tidy. 87.9, 76.4 and 54.68 can all be legitimate Terminal-Bench 2.1 measurements because the harness, tool serialization, inference configuration and retry policy are part of the evaluated system. The spread itself is the finding.
Finally, SWE-bench Verified and SWE-bench Pro are separate benchmarks. The frequently repeated 55.4 Pro score belongs to DeepSeek's earlier V4-Pro Max table, not the exact 0813 release table. Until an exact 0813 Pro run is published with a pinned suite revision, scaffold, task count, retries, endpoint/checkpoint and trajectories, the correct value is unknown.
For engineering teams, the best procurement test is a matched evaluation on their own repositories: same tasks, same tools, same provider, same reasoning effort, same retry policy, and cost measured per successfully completed task. On DeepSeek V4 Pro 0813, the harness is not a footnote; it is part of the product.
This article is built from the source material below. Open the originals for full context and the latest updates.