DeepSeek Guide — whale logoDeepSeek GuideFAN SITE
BENCHMARKS5 MIN READ

DeepSeek V4 Pro Benchmarks: 0813 GA Scores & Independent Tests

UPDATED: AUG 16, 2026AUTHOR: INDEPENDENT FAN GUIDE
OVERVIEW

V4 Pro 0813 scores Terminal Bench 2.1 at 87.9 and tops Cybergym and AutomationBench. Full official table, independent tests, and the harness caveat.

01

The Headline Numbers

DeepSeek V4 Pro 0813 — the GA build released August 13, 2026 — scores **87.9 on Terminal Bench 2.1**, takes **first place on Cybergym (83.3)** and **AutomationBench (31.8)**, and beats its own preview by a wide margin across every agent benchmark[1][4].

  • Terminal Bench 2.1: 87.9 (preview was 72.1, Flash-0731 was 82.7)
  • Cybergym: 83.3 — first place in DeepSeek's published table
  • AutomationBench (Public): 31.8 — first place
  • DeepSWE: 62.7 (preview was 12.8)
  • Toolathlon-Verified: 74.1
  • Independent 8-task suite: 76.25% vs 24.8% for the preview
NOTE

All official numbers are vendor-reported on the Hugging Face model card, produced with DeepSeek's own agent configuration[4].

02

Official Benchmark Table (0813)

The complete published table from the DeepSeek-V4-Pro-0813 model card, with the nearest published competitors[4].

DSBench-FullStack (71.1) and DSBench-Hard (67.2) are internal test sets, marked with a dagger in the model card. Everything in this table except the internal rows uses the official DeepSeek Harness minimal mode with max reasoning effort, temperature 1.0, top_p 0.95[4].

The picture: V4 Pro 0813 leads the open-weight field on agentic coding, sits just below Opus-4.8 on Terminal Bench, and clears the preview-era Flash on every line — the GA build is a real generational step.

BenchmarkV4 Pro 0813Flash-0731Pro PreviewKimi K3Opus-4.8Fable-5
HLE (wo/w tools)42.7 / 60.037.8 / 51.537.7 / 48.240.5 / 54.743.5 / 56.049.8 / 57.9
Terminal Bench 2.187.982.772.181.088.385.0
NL2Repo61.554.238.548.969.7
Cybergym83.376.752.780.078.3
DeepSWE62.754.412.846.267.558.0
Toolathlon-Verified74.170.355.959.976.576.2
AutomationBench (Public)31.825.112.812.930.827.2
Agents' Last Exam25.725.216.523.827.625.7
03

The Harness Caveat: How These Scores Were Produced

The agent scores are not raw model numbers — every code-agent benchmark was run through the **DeepSeek Harness in minimal mode** at max reasoning effort[4]. That is a model-plus-harness score.

The same configuration note appeared in the July 31 Flash-0731 changelog, and it matters for comparisons: a different agent framework (Claude Code, OpenCode, Cline) can produce materially different completion rates on the same model. The DeepSeek Harness benchmarks page explains how the harness itself is scored and how to replicate the setup.

Practically, treat these tables as 'V4 Pro 0813 + Harness minimal' scores. If you run V4 Pro inside your own tooling, expect variance — that is normal, not a defect.

NOTE

DeepSeek's official statement for the preview said the same testing harness had not been published; the GA model card now links the open-source Harness, so the exact benchmark configuration is reproducible[4].

Sponsored
04

Independent Tests: MindStudio & Others

Independent testing shows the same direction but more conservative numbers. MindStudio's 8-task coding/reasoning suite scored V4 Pro 0813 at **61/80 (76.25%)**, roughly tied with Muse Spark 1.2 and below Kimi K3 and Opus 5 — but the same suite scored the preview at only 24.8%[6].

Reddit's r/LocalLLaMA thread on the 0813 numbers mostly reads them as 'sane' for an open flagship, with skepticism reserved for the vendor harness configuration and the internal DSBench rows[9].

There is no NIST-level independent audit of the 0813 build yet — the May CAISI evaluation covered the preview build only.

  • Independent 8-task suite: 76.25% (preview: 24.8%) — one of the largest single-version jumps recorded
  • Observed strengths: front-end generation, task planning, asking clarifying questions
  • Observed weaknesses: overthinking simple problems, rewriting code more than the task needs
  • Community verdict: for everyday workloads, V4 Flash often feels like the better tool — see the V4 Pro review
05

Preview vs GA: The Jump

The GA build's improvement over the preview is the story of this release: Terminal Bench 2.1 up 15.8 points (72.1 → 87.9), DeepSWE up from 12.8 to 62.7, AutomationBench up from 12.8 to 31.8[4][6].

The scale of the jump suggests the 0813 build is a materially retrained checkpoint rather than a configuration flip — consistent with the 'same architecture, new weights' pattern DeepSeek established with Flash-0731 on July 31[1].

If you benchmarked the preview in July and wrote it off, the GA numbers are worth re-checking.

BenchmarkPreview (Apr 24)0813 GA (Aug 13)Delta
Terminal Bench 2.172.187.9+15.8
DeepSWE12.862.7+49.9
Cybergym52.783.3+30.6
Toolathlon-Verified55.974.1+18.2
AutomationBench (Public)12.831.8+19.0
Sponsored
06

How to Read These Numbers

Three rules for using this table without fooling yourself[4][6][9].

The honest summary: V4 Pro 0813 is the strongest open-weight coding agent as of mid-August 2026 on vendor numbers, and a large improvement over the preview on independent numbers. It is not head-and-shoulders above the closed flagships — expect it to trade blows with Opus-4.8 and Kimi K3 depending on the task[4][6].

  • Compare like-for-like: same harness, same reasoning effort, same temperature
  • Use independent suites (MindStudio, your own) as the tiebreaker when vendor vs vendor tables disagree
  • Check the date: agent frameworks and model checkpoints move weekly, so any table older than two weeks is stale
NOTE

Benchmarks measure the model under one configuration. Your workload is the final test — run your own evals before committing production traffic.

Sponsored
Sponsored