
DeepSeek V4 Pro vs Qwen3.8 Max: The 3x Price Gap, Re-Run After the September Reprieve
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
DeepSeek V4 Pro and Qwen3.8 Max are the two most important open-or-cheap flagships of this generation, and the September events on the DeepSeek side changed the matchup more than either vendor's marketing has acknowledged. DeepSeek V4 Pro shipped on August 13, 2026, ten days after Alibaba's Qwen3.8 Max went generally available on August 3, and the pairing has always read as a one-sided contest: Qwen3.8 Max is the 2.4-trillion-parameter all-rounder at $2 per million input and $6 per million output tokens, while DeepSeek V4 Pro is the 1.6-trillion-parameter reasoning specialist at a fraction of that price, MIT-licensed, with public weights. The reason this comparison is worth re-running today is not the launch-week numbers but what happened a month later — DeepSeek spent the week of September 8 trying to retire V4 Pro, then reversed on September 11 and committed to keeping it online at unchanged billing. That reversal is the single most important fact about this matchup right now, because it decides whether the cheap half is a safe bet or a migration trap.
Independent measurements in this piece come from Artificial Analysis, captured today. Alibaba's and DeepSeek's own benchmark tables are labeled vendor-reported throughout — neither company's launch numbers have been fully reproduced by a third party, and DeepSeek's are now four weeks old against a moving scoreboard.
The September reprieve is the story, not the August launch
On September 8, DeepSeek announced that after September 14 every request to deepseek-v4-pro would be routed to V4.1 Flash and billed at Flash prices until a V4.1 Pro shipped. Anyone running Qwen3.8 Max versus DeepSeek V4 Pro as a serious comparison had to start planning for the cheap option to vanish. Then DeepSeek delayed the cutoff, and on September 11 it reversed entirely, stating it would continue providing API services for DeepSeek V4 Pro after September 14 with billing unchanged, and giving notice if that ever changes. The reversal was reported by Chinese state media and tech press the same day.
What this does to the matchup is remove the asymmetry that was quietly steering decisions. Through the first week of September, the honest answer to "should I build on the $0.66 output model or the $6 output model?" was being contaminated by survival risk: the price leader had a shutdown date. As of September 15, that contamination is gone. The cheap option is durable, and the comparison is finally about what it was always supposed to be about — intelligence, price, and speed.

Price: the same game, at a fraction of the output cost
On the current independent scoreboard the two sit four points apart, and the price gap between them is the largest in the open-weights tier. Qwen3.8 Max is $2.00 per million input and $6.00 per million output tokens on Alibaba's list price, with an 88% cache discount. DeepSeek V4 Pro is $0.66 per million input and $1.98 per million output off-peak, doubling to $1.32/$3.96 at peak hours, with a 97% cache discount.
• Price — DeepSeek V4 Pro $0.66/$1.98 per 1M off-peak (peak $1.32/$3.96) vs Qwen3.8 Max $2.00/$6.00 flat
• Intelligence Index — V4 Pro 36 (#7/113) vs Qwen3.8 Max 40 (#30/200)
• Output speed — V4 Pro ~81 tokens/sec vs Qwen3.8 Max ~41 tokens/sec
• Cost per task — V4 Pro $0.67 vs Qwen3.8 Max $2.67 (per AA index task)
• Context — both 1M tokens
• License — V4 Pro MIT open weights vs Qwen3.8 Max proprietary weights, commercial license
The output-price ratio is the whole argument: Qwen3.8 Max costs three times as much per input token and roughly three times per output token at V4 Pro's peak rate — and the gap on output widens to six times when V4 Pro is billed at its off-peak rate, which covers most of the week outside a narrow window. On a reasoning-heavy workload, where output tokens dominate the bill, that gap is the difference between a model you can leave running and one you watch in the dashboard. The cache lines widen it further: V4 Pro's 97% discount against Qwen's 88% means a long-context task with a reused prefix costs a fraction of its list price on the DeepSeek side.
Intelligence: four points apart, and the gap is the real product
The independent scoreboard makes this closer than the price ratio suggests. On the current Artificial Analysis Intelligence Index, Qwen3.8 Max scores 40 against DeepSeek V4 Pro's 36 — a four-point gap that shows up in reasoning-heavy work but disappears in token-heavy work. Qwen3.8 Max's edge is real and consistent: it ranks #30 of 200 overall, just inside the frontier, while V4 Pro ranks #7 of 113 in its open-weights class. On Alibaba's own benchmark table, Qwen3.8 Max claims a Terminal-Bench 2.1 of 86.6, GPQA Diamond 92.6, and a FrontierSWE 73.5 that nearly doubled its predecessor's 40.7 — all vendor-reported, none reproduced independently. DeepSeek's model card claims a Codeforces rating of 3,348 and a Terminal-Bench 2.1 of 87.9 at maximum reasoning effort — likewise vendor-reported and unverified by a third party.
The honest framing is that the two companies are arguing past each other. Qwen3.8 Max's vendor table is built around enterprise and scientific work — the 0902 refresh Alibaba shipped on September 2 was further post-trained on Coding and Cowork, which Qwen says strengthens exactly that profile. DeepSeek V4 Pro's own claims center on deep reasoning and agentic tool use, where its four-week-old Codeforces and Terminal-Bench numbers still look strong. Neither vendor's table has been reproduced by an independent evaluator, and on the one scoreboard that is independent, the four-point gap is the whole difference between the models.
Speed is where Qwen3.8 Max quietly loses ground
The dimension the spec sheet hides is output speed. Artificial Analysis measures DeepSeek V4 Pro at about 81 tokens per second against Qwen3.8 Max's roughly 41 — a two-times difference in how fast answers come back. On a chat workload that is a pause the user notices; on an agent loop with many sequential calls, it compounds into real wall-clock time. The price ratio already favors V4 Pro, and the speed gap makes the effective cost-per-answer gap even wider than the token math suggests.
There is a trade buried in that number, and it is Qwen3.8 Max's verbosity. Artificial Analysis notes the model is "notably slow and very verbose," and its cost per task — $2.67 against V4 Pro's $0.67 — reflects both the higher rate and the extra output tokens. A reasoning model that writes more per answer costs more per task even before you multiply by the rate card. If your workload is latency-sensitive or your agent loop makes dozens of calls, that is the number to run your own eval on.
Weights: MIT versus a commercial license
DeepSeek V4 Pro is MIT-licensed with public weights — you can download them, serve them yourself, fine-tune them, and keep your modifications closed. Qwen3.8 Max is the first Max-class Qwen ever to ship open weights (the 2.4T A95B checkpoint landed on August 12), but under a commercial license with scale-tier conditions, and Alibaba has not published the full inference stack the way DeepSeek has. For a company comparing the two, the deployment story is the tie-breaker: V4 Pro can be self-hosted as infrastructure; Qwen3.8 Max is primarily a service you rent. The weights difference matters most if your concern is vendor lock-in or if you want to escape the API price entirely.
Who should pick which
If your workload is reasoning-heavy, output-token-dominated, or latency-sensitive, DeepSeek V4 Pro is the answer on nearly every axis that shows up in the bill: roughly three to six times cheaper on output, about twice as fast, deeper cache discount, and MIT weights as an escape hatch. The four-point intelligence gap is real but narrow, and it rarely survives contact with the price difference on production traffic.
If your workload is enterprise-shaped — complex multi-step analysis, deep logical derivation, demanding agentic work, or any task where the marginal answer quality is worth three to six times the token cost — Qwen3.8 Max's extra four points on the independent index and its stronger enterprise benchmark claims are the reason to pay the premium. The 0902 refresh sharpens that profile. And if you need image or video input, the choice is not even close: Qwen3.8 Max accepts text, image, and video, while DeepSeek V4 Pro is text-only.
Both models are callable through OrcaRouter on a single key — DeepSeek V4 Pro and Qwen3.8 Max both route through the platform — and because OrcaRouter passes provider list prices through with no markup, the numbers in this article are the numbers you actually pay, with vendor price cuts landing the same day. The routing DSL lets you send the token-heavy or latency-sensitive traffic to DeepSeek V4 Pro and the multimodal or highest-stakes reasoning tasks to Qwen3.8 Max out of the same integration — which is exactly how most teams end up using this pairing: both, for what each does best.


The verdict
The September reprieve settled the question that was drowning this matchup: the cheap half is no longer a flight risk. DeepSeek V4 Pro is the cost-performance answer — three to six times cheaper on output, twice as fast, MIT-licensed, and now committed to staying online. Qwen3.8 Max is the capability answer — four independent index points higher, stronger enterprise claims, multimodal input, and a commercial license that keeps it out of your stack. Pick the four points when your evals prove they matter; pick the price when your bill is the eval that actually decides.
Send the token-heavy or latency-sensitive traffic to DeepSeek V4 Pro and the multimodal or highest-stakes reasoning to Qwen3.8 Max out of the same integration — which is exactly how most teams end up using this pairing: both, for what each does best.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
