
DeepSeek V4.1 Flash vs DeepSeek V4 Flash: Same Flash Rate Card, a Re-Trained Model
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2123Intelligence49Coding
The week DeepSeek V4.1 Flash went from a rumored successor to the shipping Flash model has been short and loud. On September 8 the model surfaced as a two-day API test under the model id deepseek-v4.1-flash-expires-on-0910 — a build DeepSeek itself called an "intermediate version". On September 9 its open platform said the formal release, DeepSeek V4.1 Flash, would land around September 10 and that it "comprehensively surpasses" DeepSeek V4 Pro on capability, cost, speed and end-to-end time. On September 10 the release notice went with a new Flash-series rate card, cut at 12:00 Beijing time, that DeepSeek V4 Flash has not seen since an August price hike. Whatever else is true, the text workhorse DeepSeek V4 Flash is no longer the newest way to run a Flash workload, and this piece is about what that actually changes for the app you are running today.
Comparisons this early are lopsided, and it is worth saying so before the numbers start. DeepSeek V4 Flash has a published spec card, a month of production traffic behind the DeepSeek-V4-Flash-0731 update, open weights and independent measurements from Artificial Analysis. DeepSeek V4.1 Flash has an official announcement, two days of community speed tests, and no published card, no technical report and no third-party score yet. Vendor claims are labelled as vendor claims below; community figures are labelled community-measured.
The Flash tier just reset — three changes, one week
DeepSeek V4 Flash, the 284B-parameter mixture-of-experts model with 13B active, has been the sensible low-cost, high-throughput default for a large share of production text traffic since its July 31 refresh. Three announcements in three days moved the ground under it:
• September 8 — DeepSeek opened a time-boxed beta of V4.1 Flash to any API account, capped at 20 concurrent requests and scheduled to expire on September 10, under a model name that carried its own expiry date.
• September 9 — the open platform announced the formal release of DeepSeek V4.1 Flash for around September 10, describing a new model structure with native multimodal support, and claiming it beats DeepSeek V4 Pro on performance, cost, speed and total completion time.
• September 10 — the formal release brings a new Flash-series price list, effective 12:00 Beijing time, cutting the rates DeepSeek V4 Flash has charged since an August 17 increase that drew heavy developer criticism.
The last of those deserves its own beat because it is the rare price cut that is actually a price cut. On the new list, off-peak input with a cache hit drops from ¥0.05 to ¥0.02 per million tokens — a 60% cut — while cache-miss input goes from ¥1.5 to ¥1 and output from ¥4.5 to ¥4. Peak-hour rates are double the off-peak figures. In round US-dollar terms, that is roughly $0.003 per million cache-hit input, $0.14 cache-miss input and $0.56 output at off-peak, peak doubled. The 24-day gap between the August hike and this cut was not an accident of scheduling; the price reset is part of the V4.1 Flash launch.
Same tier, different model
The easy read is that V4.1 Flash is V4 Flash with a speed dial turned up. It is not. The 0731 refresh was a post-training update on the existing DeepSeek V4 Flash base — same architecture, same 284B-total / 13B-active mixture-of-experts layout. DeepSeek's description of V4.1 Flash points at something bigger: a new model structure, which several outlets read as a fresh pre-training run rather than a fine-tune. The distinction matters because a re-trained model does not inherit the old model's behaviour the way a refresh does — prompting, tool-calling and output style can all shift even where capability does not.
• Architecture — DeepSeek V4.1 Flash: new model structure, described by DeepSeek as a fresh build with native multimodal input. DeepSeek V4 Flash: the V4-Flash-0731 post-training refresh of a 284B-total / 13B-active MoE.
• Input — V4.1 Flash handles text and image input natively in the base model; DeepSeek V4 Flash is text-only, and images require the separate bolt-on build DeepSeek-V4-Flash-Vision-Exp.
• Context and output — DeepSeek V4 Flash publishes 1M context and 384K max output; DeepSeek's launch materials say V4.1 Flash keeps the million-token-scale context window, but no formal card has been published.
• Concurrency — DeepSeek V4 Flash lists a 2,500-concurrent ceiling; the V4.1 Flash beta was capped at 20 and DeepSeek's "high concurrency" claim for the release build was not yet backed by a published number at writing time.
• Weights — DeepSeek V4 Flash is open-weight under an MIT licence; nothing has been announced for V4.1 Flash.
• Independent standing — DeepSeek V4 Flash has been measured by Artificial Analysis; DeepSeek V4.1 Flash has no third-party score yet.
The text-versus-multimodal point deserves one extra sentence even in a text comparison, because it decides who V4.1 Flash is for. If your pipeline was running DeepSeek V4 Flash for text and DeepSeek-V4-Flash-Vision-Exp for images, V4.1 Flash is the first DeepSeek build that covers both paths in one model — one endpoint to maintain, no encoder bolted on the side.

Where they really diverge: throughput and what a finished task costs
At parity prices the two models separate on speed, and speed is where the numbers get loudest. Artificial Analysis measures DeepSeek V4 Flash 0731 at roughly 120–140 output tokens per second on DeepSeek's own API, depending on the configuration and measurement window. First-day testers of V4.1 Flash reported sustained generation between roughly 280 and 500 tokens per second depending on the task, with peaks above 500 in light-concurrency runs; the higher figures are community-measured, not vendor-published, and tokens-per-second is never an intrinsic property of a model. Still, the direction is consistent across every first-day report, and DeepSeek's own claim of faster generation and lower total completion time points the same way.
That speed gap changes the economics even with the two models on the same rate card, because you pay per token and tokens arrive faster. A long-context retrieval or code-generation job that took a minute on V4 Flash can finish in a fraction of the time on V4.1 Flash at the same per-token price — several first-day testers reported five-to-six-times-faster end-to-end on exactly those task classes. The beta's early "it spends money faster" complaints were the flip side of the same coin: at an unchanged per-token price, a model that emits tokens two to three times as fast drains a balance faster per minute. Whether the per-task bill actually falls depends on whether the new model finishes with fewer tokens and fewer retries, which is precisely what no one can yet prove.
On the price card itself, the two models sit side by side. Per million tokens off-peak on the September 10 list: ¥0.02 cache-hit input, ¥1 cache-miss input and ¥4 output, peak double — the same list that applied to the Flash tier before the cut had been ¥0.05, ¥1.5 and ¥4.5. For a cache-heavy workload the cut is the headline: ¥0.02 against ¥0.05 is a 60% reduction on the tokens that make up most of an autocomplete or agent-loop input stream. For a fresh-ingest job that misses cache, the saving is closer to a third. None of that makes V4.1 Flash cheaper per token than V4 Flash — they share the card — but the architecture's higher throughput is the differentiator, and it costs nothing extra.

The receipts are for V4 Flash; the claims are for V4.1 Flash
The single most important benchmark fact about this matchup right now is that there is no benchmark for the new model. DeepSeek V4.1 Flash's headline claim — that it comprehensively surpasses DeepSeek V4 Pro on capability, cost and speed — is vendor-reported and unreproduced. No parameter count, no eval table and no technical report has been published, and Artificial Analysis had not measured it as of writing. Against a family of models whose last flagship launch drew independent scrutiny, "trust us, it beats the Pro" is a claim a reader should hold at arm's length until a third party runs it.
DeepSeek V4 Flash, by contrast, has a real paper trail. Artificial Analysis counts the 0731 reasoning build among the leading models in its class, scores it well above the median for comparable open-weight models, and measures its output throughput on DeepSeek's own API in the ~120–140-token-per-second range cited above. Those are the numbers V4.1 Flash will be judged against when its own score lands, and they set a higher bar than the "fast cheap Flash" label implies. The model the new build is replacing is not weak; it is a genuinely strong open-weight workhorse. That is what makes the upgrade decision real rather than ceremonial.
Routing is doing the heavy lifting — inside DeepSeek, and for you
The most under-reported part of this launch is that DeepSeek is routing traffic between its own models. Its announcement is explicit: once V4.1 Flash is live and until V4.1 Pro ships, every API request aimed at deepseek-v4-pro will be served by DeepSeek V4.1 Flash and billed at the Flash-tier price. V4 Pro callers get the new architecture and a fraction of the per-token cost without changing a line of code. Note what does not get routed: deepseek-v4-flash calls. The old Flash model keeps serving whoever still points at it, which is exactly why this comparison exists — if you are on V4 Pro the switch happens for you, and if you are on V4 Flash it is a decision you have to make yourself.
A vendor quietly deciding that a cheaper model should serve the calls aimed at its premium one is a striking endorsement of V4.1 Flash — and it is also the same logic a routing layer applies across vendors. That is the gap a platform like OrcaRouter fills: one API for 200+ models, with provider list prices passed through at 0% markup, so DeepSeek's September 10 Flash cut is live on our side the same day. DeepSeek V4 Flash on OrcaRouter stays callable on the same key as every other model you run, which makes the safe way to adopt a day-one model genuinely cheap: point a slice of traffic at DeepSeek V4.1 Flash on DeepSeek's own API, keep DeepSeek V4 Flash as the automatic fallback on the same routing rule, and let the failover catch the edge cases while the new architecture earns its keep.
DeepSeek V4.1 Flash vs DeepSeek V4 Flash: which should you call?
If the honest answer were either model for everyone, there would be no article. The two now fit different risk profiles, and the split is unusually clean.

Move to DeepSeek V4.1 Flash if you are starting a new Flash workload today, if latency and throughput are the binding constraint, if you want image input without maintaining a second model, or if you are happy to let DeepSeek's own routing de-risk the Pro-tier claim. The model is on DeepSeek's first-party API under the deepseek-v4.1-flash name, priced on the same Flash card as the model it replaces, and every first-day signal says it is substantially faster.
Stay on DeepSeek V4 Flash if you are serving production traffic at the published 2,500-concurrency ceiling, if you rely on its open weights for self-hosting or compliance, or if your prompts, evals and tool-calling logic were tuned against 0731 behaviour and cannot absorb a re-trained model's quirks on day one. A fresh pre-training run is a behaviour change, not just a speed change, and the 0731 build is the version with the month of receipts.
For most teams the defensible default is the middle path: treat DeepSeek V4.1 Flash as the default for new workloads, keep DeepSeek V4 Flash as the failover for the workload that pays the bills, and let the first independent scorecard — plus a published spec card and a concurrency number — settle the argument that no two-day beta can. The Flash tier just got a new engine; the old one is not obsolete, it is the benchmark.
Curious what Pro-class traffic looks like while DeepSeek routes it to V4.1 Flash at Flash prices? DeepSeek V4 Pro on OrcaRouter — both on one key, billed at DeepSeek's list price.
Compared in this article4
Detected from this article · Benchmarks: Artificial Analysis · updated daily
