
DeepSeek V4.1 Flash Benchmarks: What 74.2 Does and Does Not Prove
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 984 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 110 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
DeepSeek V4.1 Flash scored 74.2 on DeepSWE v1.1, a tenth of a point above GPT-6 Astra, half a point above Gemini 3.8 Flash and Claude Opus 5, and the number has been quoted all week as a changing of the guard. It is worth being precise about what it actually is. The model itself is not new — DeepSeek V4.1 Flash went generally available on 2026-09-10, six days ago — so nothing here is a launch. What landed inside that window is the measurement layer: the first independent index entries, the first independent arena result, day-0 serving support across three stacks, and a vendor decision to retire its own flagship in favour of this one. This page is a roundup of what has actually been measured, with the source named for every figure, and a plain statement of what nobody has established yet.
The DeepSWE number, and the four models inside half a point of it
DeepSWE v1.1 is a long-horizon software engineering benchmark from Datacurve — 113 original tasks across 91 open-source repositories and five programming languages, built around multi-file changes and behavioural correctness. On DeepSeek's own model card, the vendor-reported table reads: DeepSeek V4.1 Flash 74.2, Claude Opus 5 74.0, GPT-5.6 Sol 73.0, Kimi K3 67.5, GLM-5.3 66.9, DeepSeek V4-Pro-0813 62.7, DeepSeek V4-Flash 54.4.
The independent leaderboard for the same benchmark does not reproduce that ordering, and the differences matter more than the headline. Datacurve's DeepSWE v1.1 Pass@1 leaderboard places Muse Spark 1.3 (FAIR) first at 75.40, ahead of DeepSeek V4.1 Flash at 74.20. Behind it: GPT-6 Astra at 74.12 (extra high) and 74.10 (max), Gemini 3.8 Flash at 73.83 (high), Claude Opus 5 at 73.65 (max).
So the 74.2 is a real entry on a real third-party leaderboard, and it is second, not first. The claim that circulated on release day — that this is state of the art on DeepSWE — does not survive contact with the leaderboard that carries the score. Muse Spark 1.3 is 1.2 points clear.
Now the part that gets skipped. Four frontier models sit inside 0.55 points of each other on this benchmark: 74.20, 74.12, 73.83, 73.65. That is a narrower band than the measurement noise. Published run-to-run variation on DeepSWE is on the order of 1.4 to 3.2 points — three to six times the size of the entire top-four spread. A 0.2-point lead over GPT-6 Astra is not a finding; it is what a tie looks like when you print it to two decimals.
Scaffold sensitivity is the second reason to hold the number loosely, and here the vendor's own documentation supplies the evidence. DeepSeek reports DeepSWE across eight scaffolds in the same release materials. The 74.2 was achieved on mini-SWE. The same model scored 72.6 on DSH Minimal, 70.5 on DSH Standard, 69.8 on Claude Code, 67.6 on DSH PTC, 66.2 on Pi, 65.6 on Codex and 65.5 on OpenCode. That is an 8.7-point spread on one model and one benchmark, driven purely by the harness wrapped around it. Anyone comparing a DeepSeek number produced on mini-SWE against a rival's number produced on a different agent loop is comparing scaffolds, not models.
Where the independent index actually places it
Artificial Analysis runs a fixed suite at declared effort settings, which makes its Intelligence Index the cleanest external check available. On the model page for DeepSeek V4.1 Flash, at reasoning max effort, the index reads 40 — ranked #6 of 113 models in that class, against a class median of 18. The closed rivals sit above it: Claude Opus 5 (max) 51, GPT-5.6 Sol (max) 47, GLM-5.3 (max) 45, Kimi K3 (max) 44, Gemini 3.8 Flash (high) 41. DeepSeek's own previous flagship, V4-Pro-0813, sits below at 36.
Read the two scoreboards side by side and the disagreement is structural, not a mistake. On DeepSeek's chosen benchmark, at the right scaffold, the model ties the frontier. On a fixed external suite averaged across reasoning, coding, maths and agentic work, it is eleven points behind Claude Opus 5 and one point behind Gemini 3.8 Flash. Both can be true. A vendor table is an argument; an index is an average.
The index page also carries two figures that matter for a routing decision and rarely make the headlines. DeepSeek V4.1 Flash runs at 214.4 output tokens per second at max effort — fast. It also generated 250M output tokens completing the Intelligence Index, against a median of 140M, which Artificial Analysis flags as markedly verbose. Verbosity is not a quality problem; it is a bill. A model that thinks in 79% more tokens than its peers gives back part of its price advantage on every long reasoning call.

ProgramBench: a single-model score and a multi-agent score are not the same measurement
This is the number most likely to be misread, so it is worth separating carefully.
• Single model, standard configuration — DeepSeek V4.1 Flash scores 20.3 on ProgramBench (Almost@1), per DeepSeek's own model card. That is below GPT-5.6 Sol at 23.0 and far below Claude Opus 5 at 37.0, and only slightly above GLM-5.3 at 19.0 and Kimi K3 at 17.5. Vendor-reported, and it does not flatter the model.
• Multi-agent team configuration — DeepSeek's tech report, DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression, describes a preliminary Agent Swarm RL recipe in which a lead agent coordinates teammates over a shared task board. On a high-confidence 172-task ProgramBench subset — tasks where the reference solution passes at 95% or better — the multi-agent configuration climbs from 13.59% at a one-hour deadline to a peak of 30.04% at eight hours, while the single-agent baseline moves from 12.79% to 20.39%.
The 30.04% is the source of the "30%" claim, and it is not a single-model score. It is a team of agents given eight hours, measured on a filtered subset, in the vendor's own preliminary evaluation. The comparison it invites — "well above single Sol, almost 2x Kimi K3" — is arithmetically fair against Sol's 23.0 and K3's 17.5, and it is also a comparison between a multi-agent configuration and single-model baselines on a different task set. The same report shows the effect on FrontierSWE v2: multi-agent reaching 32.90% at twenty hours against 28.20% single-agent. The gap is real, it narrows as the deadline extends, and the token cost of getting it is not small — one hands-on write-up reports more than 10x the tokens and runtimes approaching four hours.
The parameter count is not settled, so treat every size comparison as provisional
DeepSeek's release notes describe a "552B-parameter MoE". That is the vendor figure, and it is what most coverage repeats. It is also not the whole artifact.
DeepSeek's own model card documents a second component alongside the transformer backbone: "Engram conditional memory (196B parameters, sparsely accessed via token-based lookup)." Add the two and you get 748B. That is the arithmetic behind the community reading that circulated on r/LocalLLaMA within hours of release, and the model card's own safetensors index lists 763B parameters across its tensor types. The roughly 510 GB FP8 figure comes from the same lineage of analysis: what you actually download and hold in memory.
The reconciliation is that these count different things. 552B is the computational backbone — the parameters a token flows through. 196B is a lookup table consulted at specific layers that returns stored information rather than computing over it. 8B and 16B, the activation figures DeepSeek quotes, are the per-token slices active during prefill and decode respectively. All four numbers describe the same model.
None of that makes the comparison safe. Every claim of the form "this model matches a frontier model at a fraction of the size" is resting on a vendor headline that excludes a fifth of the artifact. Until someone publishes a reproducible parameter accounting, size-based arguments should carry the caveat inline.
ARC-AGI: no result for this model
There is no ARC-AGI-2 or ARC-AGI-1 score for DeepSeek V4.1 Flash. ARC Prize's verified results page carries no entry for this checkpoint, and DeepSeek's own benchmark table has no ARC row. The only ARC Prize-verified DeepSeek result is for the older V4 Flash 0731 checkpoint — 61.4% on ARC-AGI-2 semi-private at maximum reasoning effort and 89.0% on ARC-AGI-1, at $0.04 and $0.02 per task respectively. Those are the predecessor's numbers on a different model, and citing them as V4.1 Flash results would be wrong. If ARC-AGI matters to your evaluation, this model is currently unscored.
What else landed inside the window
Three things changed for a reader deciding whether to try this model, and none of them is the model's own release.
• Independent arena result. OpenDesign Arena ran thirteen models through identical design tasks — web apps, dashboards, mobile screens, landing pages. DeepSeek V4.1 Flash scored 81.2 out of 100 against GPT-6 Astra's 82.7, at $0.023 per finished design against $1.61, and finished in 5.3 minutes against 11.1. Delivery rate — outputs usable without revision — was 57.7% against Astra's 60%. This is one of the first genuinely independent measurements of the model, and it measures design output specifically, not general reasoning or coding.
• Day-0 serving support on three stacks. vLLM published a day-0 image (2026-09-09) validated on NVIDIA and AMD hardware. SGLang and Miles published day-0 inference and RL support on 2026-09-10, including an Engram host-offload path that raised KV cache capacity 36% on 4x GB300 and a bounded-replay optimisation that lifted prefill throughput 1.56x on 8x H200. Cambricon announced Day-0 adaptation on the vLLM stack the same day. For a model whose selling point is aggressive KV cache compression — a quarter of the previous generation's HBM footprint, an eighth of its SSD — being runnable on standard stacks from day one is the difference between a lab result and something you can deploy.
• The vendor retired its own flagship. DeepSeek is phasing out V4-Pro. From 04:00 UTC on 2026-09-14, any request naming “deepseek-v4-pro” routes to V4.1 Flash and is billed at Flash rates, continuing until a V4.1 Pro launches. DeepSeek V4.1 Pro is not released — it is a future model, and the redirect is a stopgap that keeps pinned integrations from breaking. Also retired: “deepseek-v4-flash” and “deepseek-v4-flash-vision-exp”, both temporarily routed to V4.1 Flash for compatibility. If you have a pinned model ID in production, that redirect is the one operational fact on this page you cannot ignore.
What the price does to the decision
Pricing is where this model stops being a benchmark story. Per DeepSeek's own published rates, effective 04:00 UTC 2026-09-10, with peak hours defined as 01:00–04:00 and 06:00–10:00 UTC on weekdays and everything else off-peak.
• Off-peak, per million tokens — $0.003 cache hit, $0.15 cache miss, $0.60 output.
• Peak, per million tokens — $0.006 cache hit, $0.30 cache miss, $1.20 output.
• Cached input, compared to a frontier closed model — three tenths of a cent per million against roughly fifty cents. For an agent looping over the same repository, that tier carries most of the tokens.
• Measured on our own catalogue — $0.15 in and $0.60 out per million, 784 ms median time to first token, 211 tokens per second over the last seven days.
Artificial Analysis lists the same model at $0.30 input and $1.20 output with a 98% cache discount and a blended $0.18 per million — that is the peak tier, and the two figures are the same price at different hours of the day. The distinction is worth internalising before you compare vendors: a model with a two-tier clock cannot be compared on a single headline rate.
That is also why pass-through pricing matters more than usual here. OrcaRouter routes DeepSeek V4.1 Flash alongside the rest of its catalogue at 0% markup — provider list price, unchanged, so DeepSeek's peak and off-peak tiers arrive on our side exactly as DeepSeek publishes them, and a vendor price change is live the same day. On a model whose rate doubles for four hours every weekday morning, a platform that flattens the tiers is hiding the only number you needed.


What nobody has established
The gap in this story is not a missing benchmark. It is that no credible head-to-head exists between DeepSeek V4.1 Flash and its named rivals on a shared harness, run by someone with no stake in the outcome. Every number on this page is one of three things: vendor-reported, produced by a third party on a different configuration, or an average across a fixed suite that was not designed to isolate this matchup. There is no run where DeepSeek V4.1 Flash and GPT-6 Astra were evaluated on the same tasks, same scaffold, same effort setting, published with variance.
Three specific things would change the picture, and none of them exists today.
• A scaffold-controlled DeepSWE comparison. The 8.7-point spread across DeepSeek's own scaffolds is larger than any gap between the top four models on the leaderboard. Until the frontier is measured on one harness, the ordering at the top of that table is not meaningful.
• An independent reproduction of the ProgramBench agent-team result. The 30.04% comes from the vendor's tech report on a filtered 172-task subset, in a preliminary configuration, with a token cost that has not been characterised for production budgets.
• ARC-AGI. No score exists for this checkpoint at all.
The practical consequence is a narrower claim than the headlines support. DeepSeek V4.1 Flash is measurably within noise of the frontier on one long-horizon coding benchmark, genuinely competitive on independent design tasks at roughly 1.4% of the cost, and roughly ten points behind on an aggregate index. That is enough to justify running it against your own workload and letting your evaluation settle the question. It is not enough to justify rewriting a production routing table on the strength of 0.2 points — and the fact that both of those statements are true is the whole finding.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
