DeepSeek V4 Flash Official Release: The Cheap, Fast, Agentic 1M-Context Model, Explained
Guides & Insights

DeepSeek V4 Flash Official Release: The Cheap, Fast, Agentic 1M-Context Model, Explained

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

DeepSeek V4 Flash is the cheap half of DeepSeek's V4 ladder, and the independent numbers have finally caught up with it: on MathArena's AIME 2026 set it scores 95.83%, statistically indistinguishable from DeepSeek V4 Pro's 96.67%, at roughly one-ninth the cost per problem. The official public-beta build, -0731, shipped on July 31, 2026 with the same architecture as the April preview — a 284-billion-parameter mixture-of-experts model, 13B active, with a 1-million-token context — but a round of extra post-training that sharply improved its agentic, coding, and tool-calling ability, plus native Responses-API support and specific adaptations for Codex-style agents. This guide covers what the release actually changed, what the vendor and independent evaluations each say, what it costs once you account for how much it thinks, and how to call it.

Accuracy note: the release details and the headline agentic scores below come from DeepSeek's official API change log (dated 2026-07-31) and are vendor-reported, run on DeepSeek's own harness — treat them as directional and run your own evals. Independent figures are attributed where they appear: Artificial Analysis for the Intelligence Index and its components, MathArena for AIME 2026, DeepSeek's own price list for pricing. Where something rests on a single practitioner's report rather than a published evaluation, the sentence says so. Specs and prices change; verify before building.

TL;DR. V4 Flash is the cost-efficient sibling of the giant DeepSeek V4 Pro: a 284B / 13B-active MoE with a 1M-token context and up to 384K output, tuned for high-volume, low-latency, agentic work. The July 31 official release is a post-training upgrade — same size, materially better agents and coding — that also adds native Responses-API and Codex support. DeepSeek reports strong agentic/coding scores (e.g., Terminal-Bench 2.1 82.7, Toolathlon 70.3), and Artificial Analysis has since scored the -0731 build 50 on its Intelligence Index: ten points above the April preview, six above the far larger V4 Pro, and level with Gemini 3.6 Flash. It costs about $0.15 / $0.29 per million input/output tokens, roughly a third of V4 Pro's price. Only the Flash API was upgraded; V4 Pro and the DeepSeek app/web are unchanged. On OrcaRouter you get it at 0% markup through one OpenAI-compatible endpoint.

Key takeaways

• Official release (build -0731, July 31, 2026): a post-training upgrade of the preview — same architecture, much stronger agents, coding, and tool use.

• Architecture unchanged: 284B total / 13B active (MoE), 1M-token context (1,048,576), up to 384K output, text-only input, with reasoning, tools, and JSON.

• New: native Responses API support plus specific Codex adaptation; continues to support OpenAI ChatCompletions and Anthropic-style interfaces. Model id deepseek-v4-flash.

• DeepSeek-reported agentic/coding gains: Terminal-Bench 2.1 82.7, Toolathlon (verified) 70.3, Cybergym 76.7, DeepSWE 54.4, NL2Repo 54.2, DSBench-FullStack 68.7.

• Very cheap: about $0.15 / $0.29 per 1M input/output tokens — roughly a third of V4 Pro's $0.44 / $0.88 — and DeepSeek prices cache hits at $0.0028 per 1M, about 98% below fresh input. Only the Flash API changed; Pro and app/web are the same.

What the official release actually upgraded

V4 Flash first appeared as a preview on April 24, 2026. The July 31, 2026 official release (the -0731 build, per DeepSeek's API change log) is explicitly not a new architecture: total and active parameters (284B / 13B) and the 1M context are unchanged. What changed is post-training — extra fine-tuning that DeepSeek says significantly improved core agentic, coding, and tool-calling capabilities. Two practical additions matter: V4 Flash now natively supports the Responses API and is specifically adapted for Codex-style coding agents, while still speaking OpenAI ChatCompletions and Anthropic-style interfaces. Importantly, only the Flash API was upgraded in this release; V4 Pro and the DeepSeek app/web experiences are unchanged. For teams, that means a drop-in improvement for agentic workloads at the same price, with no migration.

Architecture: a small-active MoE built for throughput

V4 Flash is a mixture-of-experts model with roughly 284 billion total parameters but only about 13 billion active per token. That low active-parameter count is the point: it keeps inference fast and cheap while a large expert pool preserves quality. The context window is a full 1,048,576 tokens (about 1M), with a very large 384K maximum output, and the model supports reasoning (extended thinking), tool calling, and structured JSON. Input is text-only. The design goal — high-volume, low-latency, agentic work — shows up in the serving numbers: a p50 time-to-first-token around 394 milliseconds and output around 184 tokens per second on OrcaRouter's measurements.

Benchmarks: what the official numbers say

The official release leans hard on agentic and coding evaluations, run with DeepSeek's own harness (minimal-mode / max-effort settings). Here's how to read the headline results, all DeepSeek-reported:

• Terminal-Bench 2.1: 82.7 — agentic command-line/software tasks in a real terminal. A high score here signals a model that can actually drive tools and shells, not just answer questions.

• Toolathlon (verified): 70.3 and Automation Bench (Public): 25.1 — breadth and reliability of tool use and end-to-end automation. Tool-calling reliability is the make-or-break trait for agents.

• Cybergym: 76.7 — security/CTF-style agentic problem solving.

• DeepSWE: 54.4, NL2Repo: 54.2, DSBench-FullStack: 68.7, DSBench-Hard: 59.6 — software-engineering and full-stack coding, from repo-level generation to hard data-science tasks. Strong DSBench-FullStack (68.7) points to practical, multi-file coding competence.

• Agent Last Exam: 25.2 — a deliberately hard agentic reasoning gauntlet where even top models score low; useful as a relative signal, not an absolute.

For independent context, Artificial Analysis has now evaluated the -0731 build itself and scores it 50 on the Artificial Analysis Intelligence Index — ten points above the April preview it replaces, six points above the much larger V4 Pro, and level with Gemini 3.6 Flash. Its component results: GPQA Diamond 91%, AA-LCR 66%, Humanity's Last Exam 37%, SciCode 50%, τ³-Bench Banking 31%, CritPt 17%, and an Elo of 1559 on GDPval-AA v2. Two of those deserve a second look. Artificial Analysis measured Terminal-Bench 2.1 at 79% against DeepSeek's reported 82.7 — close, but lower, which is the direction vendor-harness numbers usually miss in. And on AA-Omniscience the model lands at -16 with a high measured hallucination rate, so the cheap-and-agentic story does not extend to unsupervised factual recall. That is the sharper way to read this model: strong reasoning per dollar, weak knowledge per query.

Competition math: the one place Flash matches Pro

The most useful independent read on V4 Flash's reasoning comes from MathArena, which runs every model four times on each problem of a freshly released competition and publishes the per-problem grid. On AIME 2026 — both papers, 30 problems, four runs each — V4 Flash in max-reasoning mode scores 95.83% ±3.58%, against DeepSeek V4 Pro's 96.67% ±3.21%. Those intervals overlap almost completely: on this benchmark the small model is not measurably worse than the flagship. The cost column is where they separate. MathArena puts one V4 Flash run at $0.009 per problem versus $0.082 for V4 Pro — about nine times cheaper for the same accuracy, which is a stronger argument for the efficiency tier than anything in DeepSeek's own announcement.

Two caveats belong in the same breath. MathArena flags both DeepSeek entries with a contamination warning, its label for a model released after the competition appeared, so some of that accuracy may be recall rather than reasoning. And the version MathArena tested is dated 2026-04-24 — the April preview, not the -0731 build. As of this writing there is no independent AIME 2026 score for the official release, and DeepSeek's own model card publishes no math benchmark for it at all, only agentic ones.

The grid also shows where the ceiling sits. Nearly every model on the board clears most of AIME 2026; the problem that sorts them is problem 15. Only two entries solve it in all four runs — GPT-5.5 at extra-high effort and Claude Opus 4.8 at max. V4 Flash scores zero of four there, and so does V4 Pro, alongside GLM-5.2, Gemini 3.6 Flash, Kimi K2.6 and Claude Opus 4.7.

That last detail is what makes one report from August 5, 2026 worth flagging. The AI commentator @teortaxesTex posted that the -0731 build did solve AIME 2026 problem 15, taking around 1,006 seconds and roughly 100,000 output tokens, and described the model visibly starting to game the problem before abandoning the shortcut and working it honestly. This is one practitioner's single run, not a benchmark result: it has not been reproduced, no public leaderboard has scored the -0731 build, and a lone transcript is not an evaluation. What can be checked is internal consistency, and it holds up — Artificial Analysis measures this model's output at about 116 tokens per second, and 1,006 seconds at that rate is roughly 117,000 tokens, the same ballpark as the count reported. A 100,000-token answer is also well inside what DeepSeek expects, since it recommends allowing up to 384K output tokens at high or max reasoning effort. Both figures run about three times MathArena's measured averages for this model (33,588 output tokens and 6m 50s per answer), which is what you would expect on the single hardest problem in the set. So: plausible in shape, unverified in outcome, and if it replicates it is a real jump over the preview build, which fails that problem four times out of four.

Pricing: the number that changes budgets

On OrcaRouter, which passes provider pricing through at 0% markup, V4 Flash is about $0.15 per million input tokens and $0.29 per million output tokens, and DeepSeek kept pricing flat for this upgrade. That is roughly a third of DeepSeek V4 Pro's $0.44 / $0.88. Cache hits are cheaper than most people assume: DeepSeek lists cached input at $0.0028 per million, about 98% below fresh input, which is what makes agents and chat with a stable system prompt so cheap to keep running. In real terms, 10 million tokens a month at a 70% input / 30% output split lands on the order of a couple of dollars; scale to a billion tokens and the gap against a flagship becomes a line item finance notices.

On a reasoning model, though, the rate card is the less interesting half of the bill — what moves budgets is how many tokens the model chooses to spend. Artificial Analysis needed about 206 million output tokens to run its full Intelligence Index on V4 Flash, roughly twice the ~100 million median across the models it tests, and describes it as fast but very verbose. Priced out, a single 100,000-token answer costs about three cents and the 384K output ceiling caps one answer near eleven cents. Cheap in absolute terms, but a few hundred times the cost of a short reply — which is why reasoning effort is a budget lever rather than a quality dial you leave pinned at max. Simon Willison's hands-on write-up makes the same point from the other direction: at default settings his test generation was disappointing, and raising reasoning effort to high improved it substantially. Meter your own hard requests before you assume the headline rate is your cost.

Where V4 Flash fits: the DeepSeek V4 ladder

DeepSeek's V4 line is a ladder. V4 Pro is the flagship — a much larger, top-tier reasoning-and-coding model. V4 Flash is the efficiency tier: far less active compute, a fraction of the price, and, after this post-training upgrade, genuinely strong agentic and coding ability. The fact that DeepSeek invested in making Flash a better agent (Responses API, Codex adaptation, tool-use benchmarks) tells you where it expects the volume to run. But the AIME 2026 numbers complicate the tidy version of that story in Flash's favour: on competition math the two models sit inside each other's error bars, so "send the hard problems to Pro" is not automatically right. The honest split is that Pro buys breadth of knowledge and the hardest agentic work, not a better shot at a hard math problem — which is why routing by measured difficulty on your own tasks beats defaulting to the biggest model.

Three real-world scenarios

1. Coding agents and Codex-style workflows

The -0731 build's native Responses-API and Codex support, plus its Terminal-Bench (82.7) and DSBench-FullStack (68.7) scores, make V4 Flash a strong, cheap engine for autonomous coding agents — issue-fixing, refactors, test generation, multi-step tool use.

2. High-volume tool-using agents

Fast first tokens (~394 ms), high throughput (~184 tok/s), a 1M context, and strong Toolathlon (70.3) reliability suit production agents that call tools at scale — retrieval, automation, orchestration.

3. Long-document processing on a budget

The full 1M-token context and 384K output handle summarizing, analyzing, or transforming very large inputs — logs, codebases, long reports — in a single pass, cheaply.

How to access V4 Flash

You can call V4 Flash directly on DeepSeek's API (model id deepseek-v4-flash; it supports the Responses API, OpenAI ChatCompletions, and Anthropic-style interfaces), or through a vendor-neutral endpoint. OrcaRouter exposes V4 Flash at 0% markup through a single OpenAI-compatible endpoint (deepseek/deepseek-v4-flash) alongside 200+ other models — so you can benchmark it against GPT-5.6 Luna, Gemini 3.6 Flash, GLM-5.2, Qwen3.8-Max, and Kimi K3 in one place, and route each request to whatever is cheapest for the job. Because pricing is passed through rather than marked up, a DeepSeek price change reaches your bill the day it ships. And because the interface is OpenAI-compatible, adopting V4 Flash — or switching off it after a week of measurement — is a configuration change, not a re-integration.

Limitations and honest caveats

V4 Flash is an efficiency model, not a flagship, but the independent data has made that boundary more specific than "worse at hard things." On AIME 2026 it matches V4 Pro within confidence intervals; where it clearly trails is breadth of knowledge — AA-Omniscience -16 and a high measured hallucination rate — and the most demanding agentic work. Input is text-only, with no native image input. The headline agentic scores are DeepSeek-reported on its own harness, and the one figure an independent lab has re-run came in lower (Artificial Analysis measured Terminal-Bench 2.1 at 79% against the reported 82.7), so treat vendor numbers as directional. And "cheap per token" is not "cheap per task": this model is measurably verbose, so cap reasoning effort and tool iterations on simple calls instead of leaving max effort on by default. Validate on your own workload before committing production traffic.

FAQ

What is DeepSeek V4 Flash?

DeepSeek's efficiency model in the V4 line: a 284B / 13B-active mixture-of-experts model with a 1M-token context, up to 384K output, tuned for fast, high-volume, agentic and coding work. Its official release (build -0731) shipped July 31, 2026.

What changed in the official release?

The architecture is unchanged; it's a post-training upgrade that significantly improves agentic, coding, and tool-calling ability, and adds native Responses-API support with Codex adaptation. Only the Flash API was upgraded — V4 Pro and the app/web are the same.

How much does V4 Flash cost?

About $0.15 per million input tokens and $0.29 per million output tokens on OrcaRouter (0% markup) — roughly a third of V4 Pro's $0.44 / $0.88 — with cache hits listed at $0.0028 per million, about 98% below fresh input. Pricing stayed flat in this release. Budget for token volume as well as rate: this is a verbose reasoning model, and a 100,000-token answer costs about three cents.

How good is V4 Flash at agentic and coding tasks?

DeepSeek reports strong scores: Terminal-Bench 2.1 82.7, Toolathlon (verified) 70.3, Cybergym 76.7, DeepSWE 54.4, DSBench-FullStack 68.7 — all vendor-harness figures, so verify on your own tasks.

Is V4 Flash any good at competition math?

Better than its price suggests. On MathArena's AIME 2026 grid the April preview build scores 95.83% across four runs of 30 problems — inside the confidence interval of V4 Pro's 96.67%, at about a ninth of the cost per problem — though MathArena flags both models for possible contamination, and it has not yet scored the -0731 build. The one problem it never solves is problem 15, which only GPT-5.5 at extra-high effort and Claude Opus 4.8 at max solve reliably.

How does V4 Flash compare to V4 Pro?

Pro is the far larger flagship for the hardest reasoning and coding; Flash is the efficiency tier at ~1/3 the price with strong agentic/coding ability. Route hard tasks to Pro, everything else to Flash.

What's the context window?

A full 1,048,576 tokens (about 1M), with up to 384K output tokens.

Is V4 Flash multimodal?

No — input is text-only. It supports reasoning, tool calling, and JSON output.

Does V4 Flash work with the Responses API and Codex?

Yes — the -0731 release adds native Responses-API support and specific Codex adaptation, alongside OpenAI ChatCompletions and Anthropic-style interfaces.

How do I use V4 Flash?

Directly via DeepSeek's API, or through OrcaRouter's OpenAI-compatible endpoint at 0% markup (deepseek/deepseek-v4-flash), where you can also compare and route across rival models.

Bottom line

DeepSeek V4 Flash's official release keeps the cheap, fast recipe and makes it a much better agent: the same 284B / 13B-active MoE and 1M context, post-trained for stronger coding and tool use, with native Responses-API and Codex support, at the same ~$0.15 / $0.29. The independent numbers now back most of the pitch — Artificial Analysis puts the -0731 build level with Gemini 3.6 Flash on intelligence, and MathArena has the preview build matching V4 Pro on AIME 2026 for a ninth of the cost per problem. What the low rate does not buy is breadth of knowledge or a free pass on verbosity: this model thinks at roughly twice the token budget of a median model, so your per-task bill depends more on the reasoning effort you allow than on the headline price. Run it through a 0%-markup, OpenAI-compatible endpoint like OrcaRouter, measure what your own hard requests actually spend, and route up to a flagship only where the measurement says you must.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily