Hero title card for the article 'DeepSeek V4 Pro Benchmarks': a scoreboard with a large gauge, headline 'DeepSeek V4 Pro — the benchmark scorecard', subtitle 'Official 0813 scores, independent checks, and what's still unreproduced', and chips reading 'TERMINAL-BENCH 2.1: 87.9', 'DEEPSWE: 62.7', 'AA INDEX: 53', '1M CONTEXT'. OrcaRouter logo composited bottom-right.
Guides & Insights

DeepSeek V4 Pro Benchmarks: The Full 0813 Scorecard, Every Independent Check, and What Nobody Has Verified Yet

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The one-line answer: DeepSeek V4 Pro posts a genuinely strong agentic scorecard on Deep​Seek's own numbers — Terminal-Bench 2.1 at 87.9, DeepSWE at 62.7, CyberGym at 83.3 — but nearly every headline figure is vendor-reported, and the independent picture is cooler. Artificial Analysis puts the official 0813 build at 53 on its Intelligence Index, level with GLM 5.2, four points behind GPT-5.6 Terra and seven behind Kimi K3; Vals AI ranks it 12th, below the previous-generation GPT-5.5. This article is the full aggregation nobody else has written: the complete 0813 scorecard, the extended coding set from the technical report, every independent check that exists, and an honest ledger of which numbers have been reproduced and which are still Deep​Seek's word alone. A week after the card, the serving picture has started to fill in too — on August 19, 2026, LMSYS and Ant Group published an H20 serving stack that reaches 271 output tokens/s at batch size 1, covered in its own section below and filed, like everything on the vendor side of this page, as self-reported until someone reproduces it.

The official 0813 scorecard — Deep​Seek's own numbers, all of them

Deep​Seek's official announcement (API docs, news260813, August 13, 2026) reports the production build against the April preview. Every figure in this section is vendor-reported — none has been independently reproduced as of August 16, 2026:

• Terminal-Bench 2.1 — 87.9 (preview: 72.1).

• DeepSWE — 62.7 (preview: 12.8) — a roughly 5x jump that is the single most striking claim on the card.

• CyberGym — 83.3 (preview: 52.7).

• Humanity's Last Exam — 42.7 without tools, 60.0 with tools (preview: 37.7 / 48.2).

• Toolathlon-Verified — 74.1 (preview: 55.9).

• AutomationBench (public) — 31.8 (preview: 12.8).

• DSBench-FullStack — 71.1; DSBench-Hard — 67.2 (preview: 41.8 / 31.1).

• NL2Repo — 61.5 (preview: 38.5).

• Agents' Last Exam — 25.7 (preview: 16.5).

The pattern to notice is where the gains concentrate: agentic and cybersecurity benchmarks. That is exactly the kind of scorecard a lab produces when it wants to advertise an agent model — and exactly the kind that needs a neutral rerun before you spend production money on it.

Scoreboard card for DeepSeek V4 Pro's official 0813 benchmarks: rows read Terminal-Bench 2.1 87.9 (preview 72.1), DeepSWE 62.7 (12.8), CyberGym 83.3 (52.7), HLE with tools 60.0 (48.2), DSBench-FullStack 71.1 (41.8), Toolathlon-Verified 74.1 (55.9), with a footer reading 'All figures DeepSeek-reported via API docs news260813, Aug 13 2026 — none independently reproduced.'

The extended coding set from Deep​Seek's technical report

Beyond the agentic card, the Deep​Seek technical report (as aggregated by BenchLM on August 16, 2026) adds a broader coding and long-context set. Also vendor-reported, and labeled "Provider exact" by BenchLM — no independent test:

• SWE-bench Verified — 80.6%; SWE-bench Pro — 55.4%.

• LiveCodeBench Pass@1-CoT — 93.5% — the best verified in BenchLM's catalog.

• Codeforces rating — 3206 — also the best verified in the catalog.

• BrowseComp — 83.4%; Terminal-Bench 2.0 — 67.9%.

• MRCR 1M — 83.5%; CorpusQA 1M — 62.0% — both catalog-best on 1M-token retrieval, which matters for a model that advertises a 1M-token context window.

• GPQA Diamond — 92.8, SciCode — 49.2, Long-Context Recall — 75.3 (listed on OrcaRouter's model page).

What independent evaluators actually measure

Three independent sources have run DeepSeek V4 Pro, and their verdicts cluster below the vendor card:

• Artificial Analysis Intelligence Index — 53 (SCMP's reporting, August 2026; OrcaRouter's model page shows 53.2, ranked 20th of 132). On par with GLM 5.2, four points behind GPT-5.6 Terra, seven behind Kimi K3.

• Artificial Analysis Coding — 68.8, ranked 24th of 130 (per OrcaRouter's model page).

• Vals AI Index — 12th, trailing the previous-generation GPT-5.5 and well behind Kimi K3 and Claude Opus 5 (SCMP). Vals AI's own testing flagged two concrete weak spots: completing tasks inside a sandboxed terminal environment, and building complex financial models in Excel spreadsheets.

• Vibe Code Bench v1.1 — 49.93 (Vals AI, the one "Benchmark exact" independent run in the set).

The honest ledger — what's been reproduced vs what's still just Deep​Seek's word

This is the section a benchmark article exists to provide, and the reason to bookmark this page over the launch coverage. Split every score into one of two buckets:

• Independently sourced: AA Intelligence Index 53 / 53.2, AA Coding 68.8, Vals AI rank 12th, Vibe Code Bench 49.93 — and a separately-sourced Terminal-Bench 2.1 run of 78.7 that sits alongside Deep​Seek's 87.9 on OrcaRouter's model page. That nine-point gap between the vendor's number and an independent run is the single most important datum on this page: it is the reproduction gap, quantified.

• Vendor-reported only, not independently reproduced: every figure in the official 0813 card above (Terminal-Bench 2.1 87.9, DeepSWE 62.7, CyberGym 83.3, HLE 42.7/60.0, Toolathlon 74.1, AutomationBench 31.8, DSBench 71.1/67.2, NL2Repo 61.5, Agents' Last Exam 25.7) and the whole extended coding set (SWE-bench Verified 80.6, LiveCodeBench 93.5, Codeforces 3206, BrowseComp 83.4, MRCR 83.5, CorpusQA 62.0, GPQA Diamond 92.8). Deep​Seek's own docs label the Terminal-Bench 2.1 figure a "provider run."

Read that way, the story is coherent: DeepSeek V4 Pro is a real mid-pack frontier model — AA 53, ranked in the teens to twenties — with a vendor scorecard that claims near-top agentic and coding performance. Both statements are true, and they are not in conflict; one is measured, the other is claimed.

Verification ledger card for DeepSeek V4 Pro: left column 'Independently sourced' lists AA Intelligence Index 53, AA Coding 68.8, Vals AI rank 12th, and a separate Terminal-Bench 2.1 run of 78.7; right column 'Vendor-reported only' lists Terminal-Bench 2.1 87.9, DeepSWE 62.7, CyberGym 83.3, LiveCodeBench 93.5, Codeforces 3206, SWE-bench Verified 80.6; footer reads 'Independent per Artificial Analysis, Vals AI, SCMP Aug 2026 · Vendor per DeepSeek API docs and technical report.'

What the scorecard costs

DeepSeek V4 Pro is the 1.6T-total / 49B-active MoE flagship with a 1M-token context window and 384K max output. On OrcaRouter it is priced at Deep​Seek's list price with zero markup: $0.435 per million input tokens and $0.87 per million output (cache read $0.060 per million), per the directory checked August 15, 2026 and confirmed on the live model page. At that rate the price-per-point math is the strongest argument for the model: it is roughly a tenth the output price of the models scoring near it on the independent index, so you are paying mid-tier prices for a scorecard that, if the vendor's claims hold, is near the top. One flag: Deep​Seek has announced a peak/off-peak pricing schedule (off-peak at 50% of peak) effective August 17, Beijing time, so the rate may move — the numbers here are the live list price as of this article.

Screenshot of the OrcaRouter model page for DeepSeek V4 Pro 0813: input price about $0.44 and output about $0.88 per million tokens, cache read $0.060, a 1M-token context window and 384K max output, live performance figures of 80.8 tokens per second output speed and 0.16 percent error rate, and an Artificial Analysis Intelligence score of 53.2 ranked 20th of 132.

Serving DeepSeek V4 Pro: 271 output tokens/s on H20

Benchmarks grade what the model knows; serving numbers grade how cheaply you can get it to think. On August 19, 2026, LMSYS — the organization behind Chatbot Arena — and Ant Group's open-source team published the first serious serving work on the 0813 build: a "scenario-specific serving stack" built on SGLang, the open-source engine LMSYS maintains. The headline figure is 271 output tokens/s at batch size 1 on an NVIDIA H20, which the team puts at 1.42× off a B300 — reached on a chip with no native FP4 Tensor Cores. These are LMSYS and Ant Group's own measurements, announced in their blog and on X; none has been independently reproduced.

The rest of the stack's numbers, all self-reported by LMSYS and Ant Group on their own stack:

• Decode latency — an "optimized DSpark" component cuts time-per-output-token (TPOT) by 74.8–78.0% at batch size 1.

• Prefill — a 1-million-token prefill completes in 43.7 seconds, a 36.5% geometric-mean prefill throughput gain.

• KV cache — a scheme the team calls "Humming MXFP4AFP8 + Online C128" expands full-token KV cache capacity 10.14×, the lever that makes 1M-token agent workloads practical on this class of silicon.

• Per-GPU decode — at 4K context, per-GPU decode throughput rises from 319.9 to 703.2 tokens/s/GPU, a 2.20× improvement.

Read the caveats before you size a fleet on it. Batch size 1 is the single-stream regime — what an interactive agent or a chat hits, not a high-concurrency API — and the 1.42×-off-B300 comparison is a ratio LMSYS states without publishing B300's absolute tokens/s, so the gap is asserted, not checkable. The result also measures the stack, not the chip: these gains come from SGLang-level optimization on hardware that would otherwise look underpowered for a 1.6T-total MoE. For context, OrcaRouter's live model page measures DeepSeek V4 Pro at 80.8 output tokens/s through a shared API endpoint — the H20 figure is a single-stream benchmark on a dedicated stack, a different number with a different meaning.

If your workload is served through an API, none of this changes how you call the model: DeepSeek V4 Pro is one OpenAI-compatible key on OrcaRouter at Deep​Seek's list price, with automatic failover if the provider stalls. The H20 numbers matter most if you are the team weighing a self-hosted H20 fleet against paying for the API — the gap the LMSYS and Ant Group work closes is a hardware-purchase decision, not a model-selection one.

When this recommendation is wrong

The honest cases where DeepSeek V4 Pro should not be your pick:

• You need independently-confirmed frontier reasoning. AA Index 53 is mid-pack — Kimi K3 (60) and GPT-5.6 Terra (57) are ahead and verified. If your decision hinges on the measured ceiling, the vendor card doesn't change it.

• You're betting production on DeepSWE 62.7 before someone reruns it. A 12.8→62.7 jump in one release is a claim, not a fact. Wait for a neutral reproduction — the 87.9-vs-78.7 Terminal-Bench gap shows what one might find.

• Your workload is sandboxed terminal automation or Excel financial modeling. Vals AI says these are exactly where the model struggles, independent of any vendor score.

• You're making a cybersecurity decision on CyberGym 83.3. That is Deep​Seek testing its own model on its own benchmark — a security vendor buying on this number is buying unaudited data.

• You're buying H20s on the 271 tokens/s number alone. That figure is a single-stack, self-reported benchmark at batch size 1 — promising, unreproduced, and one point on a cost curve that also includes utilization, power, and whatever serving stack you would actually run.

Bottom line

DeepSeek V4 Pro 0813 is a genuine frontier model with a vendor scorecard that outruns its independent record. The 0813 release turned a text preview into a serious agentic contender — we covered the release itself separately — and if you want to evaluate it today, it is one OpenAI-compatible key on OrcaRouter at Deep​Seek's list price, with automatic failover if the provider stalls. Just read the scorecard the way this article splits it: the independent checks (AA 53, Vals 12th, the 78.7 Terminal-Bench run) are what you can trust today; the 87.9s and 62.7s are what you should verify before you build on them. The same discipline now extends to serving: LMSYS and Ant Group's 271 output tokens/s on H20 is the best throughput picture we have, and it is still their word until a neutral run confirms it. Buy it for price-to-performance and 1M-token context, not for the unreproduced headline numbers.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily