
Qwen 4 Max vs DeepSeek V4 Pro: 95 billion active parameters against 49 billion
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
The Qwen 4 Max announcement at the Yunqi conference on September 22, 2026 gave the industry a name and a tier structure and nothing else — no weights, no price, no context window, no date. DeepSeek V4 Pro has been listed since April 24, 2026 as a 1.6-trillion-parameter mixture-of-experts with 49 billion parameters active per token, a 1-million-token context, and an off-peak price of $0.66 per million input tokens and $1.98 per million output tokens. The number worth putting side by side is not the total parameter count but the active count: the last published Qwen3.8-Max activates 95 billion parameters per token, roughly twice the 49 billion active parameters DeepSeek V4 Pro runs, and that one figure accounts for most of the three-fold price gap between the two labs before a single benchmark is consulted.
Two different bets on sparsity
Both labs build sparse mixture-of-experts models. They disagree, sharply, about how sparse to be. Qwen3.8-Max — the shipping Qwen flagship, and the only concrete baseline for what Qwen 4 Max will look like — is a 2.4-trillion-parameter model that activates 95 billion parameters per token, a ratio of roughly one part in twenty-five. DeepSeek V4 Pro is a 1.6-trillion-parameter model that activates 49 billion, a ratio of roughly one part in thirty-three, and it stores its expert weights in FP4 with the attention, router and norm layers left in FP8, which pushes per-token compute down again.
Those are not tuning differences. They are two different answers to the same question — what a frontier model should spend per token:
• Total parameters — Qwen 3.8 flagship 2.4T versus DeepSeek V4 Pro 1.6T
• Active per token — Qwen 3.8 flagship 95B versus DeepSeek V4 Pro 49B
• Activation ratio — roughly 1 in 25 versus roughly 1 in 33
• Input price, off-peak — Qwen3.8-Max $2.00 per 1M versus DeepSeek V4 Pro $0.66 per 1M
• Output price, off-peak — Qwen3.8-Max $6.00 per 1M versus DeepSeek V4 Pro $1.98 per 1M
• Flagship weights — Qwen 3.8 flagship under a custom Qwen licence versus DeepSeek V4 Pro published under MIT
• Context — Qwen3.8-Max 1M tokens versus DeepSeek V4 Pro 1M tokens
• Modality — Qwen3.8-Max text, image and video input versus DeepSeek V4 Pro text-only on its listing
• Observed latency — Qwen3.8-Max p50 time-to-first-token 3.29s versus DeepSeek V4 Pro 3.52s on the same catalogue
The sparsity difference and the price difference point the same way this time, which is not the direction the total parameter counts suggest. Qwen activates roughly twice as many parameters per token as DeepSeek and charges about three times as much, and sparsity alone does not close that gap. What closes the rest is modality: Qwen's flagship carries vision and video encoders, they are paid for on every request whether or not an image is in it, and that is why the input rate sits where it does. DeepSeek V4 Pro takes text and returns text, and it is priced like a model that does only that.
One qualifier belongs on the DeepSeek column before it is used for anything. DeepSeek prices on a peak and off-peak schedule: the $0.66 and $1.98 figures are the off-peak rates, and during the peak windows — 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays — the same model bills at $1.32 input and $3.96 output. Everything else, including all weekends and holidays, is off-peak. The rate you pay is therefore a function of when your traffic runs, and a cost model built on the off-peak number alone will be wrong for any workload with a weekday-morning component.

The Qwen 4 Max column is empty, and that is not a temporary condition
It is tempting to fill the comparison in with the assumption that Qwen 4 Max will land near the Qwen 3.8 numbers and price below them. Resist it. Alibaba described the Qwen 4 architecture as a new generation and said it is training; the only released artifact that describes that architecture is Qwen3.8-Flash-Next, open-sourced in late August 2026 and labelled on its own model card as a preview of the Qwen 4 design. It is a 125-billion-parameter main model with 51 billion parameters of N-gram embeddings activating around 6 billion per step, with QSA sparse attention, gated residuals and a 262,000-token native context extensible toward 1 million.
If that design scales to the Max tier, the active count should stay in the tens of billions rather than climbing toward DeepSeek's total — keeping per-step compute roughly flat as total parameters grow is the entire point of QSA and of N-gram embeddings that are looked up rather than densely computed. That is a direction rather than a specification, and it should not be printed as one.
What is knowable now is the operational asymmetry. A model that exists can be measured on your workload; a model that does not exist cannot be measured on anything. Qwen 4 Max is reachable, when it ships, through the vendor's own API and several third-party platforms — the same route every Alibaba flagship has taken. It is not on OrcaRouter today: the catalogue has zero entries for the Qwen 4 family. DeepSeek V4 Pro is on OrcaRouter now, and so is a dated snapshot of it.
Reproducibility is the argument nobody makes
DeepSeek ships dated snapshots and keeps them addressable. The catalogue carries both the floating DeepSeek V4 Pro identifier and a pinned 2026-08-13 variant — that date being the build DeepSeek itself designated as the general-availability release. That is unglamorous and it matters more than a benchmark delta for anyone running an evaluation harness or a compliance-controlled pipeline: when you pin the snapshot, your regression suite is measuring your code rather than the vendor's release cadence.
The Qwen 3.8 generation did the same thing — there is a dated Qwen3.8-Max snapshot on the same catalogue alongside the floating name — and there is no reason to think Alibaba will stop. But it is a property of shipped models, and it is worth stating plainly that on the day Qwen 4 Max is announced it will have no snapshot to pin, no stable identifier to test against, and no way to run a before-and-after on your own data. Whatever comparison you want to make, you will be making it later.

Running both without running two stacks
OrcaRouter routes DeepSeek V4 Pro as deepseek/deepseek-v4-pro at $0.66 input and $1.98 output per million tokens, which is DeepSeek's own off-peak list price passed through with 0% markup. That last part is worth more here than elsewhere in this batch, because $1.98 output is already close to the floor for a frontier-tier model: a platform margin of even ten percent would be a fifth of the difference between this model and a mid-tier alternative, and at 0% markup the rate you see on the page is the rate DeepSeek publishes — including the peak and off-peak split, which passes through on the same schedule rather than being averaged into a flat rate.
The same endpoint carries the Qwen 3.8 family — Qwen3.8-Max, Qwen3.8-Flash and Qwen3.8-27B — next to DeepSeek V4 Pro on one OpenAI-compatible API with close to 200 models, with automatic failover when a provider path degrades and a routing DSL for making the selection explicit rather than hard-coding it. If DeepSeek cuts the rate, the change lands on your key the same day, because there is no margin layer for a cut to get stuck behind. When a Qwen 4 tier ships, the same catalogue is where it would show up.
Where this leaves you
DeepSeek V4 Pro wins this comparison by forfeit, and it is worth being precise about why rather than dressing it up. It is cheaper on both axes by a factor of about three, it activates five times as many parameters per token, its context window matches, and it has a dated snapshot you can pin. Qwen 4 Max has a tier name and a conference slot.
The comparison that is actually live is DeepSeek V4 Pro against Qwen3.8-Max, and it is a real trade rather than a rout: DeepSeek is a text-only model at roughly a third of the price, with weights published under MIT and a leaner per-token activation profile, and Qwen's flagship takes image and video input at $2.00 and $6.00. Pick on modality and on measured cost per finished task, not on parameter counts, and re-run the comparison when Alibaba publishes a model card — at which point the interesting question will be whether Qwen 4 Max prices its multimodal capability above DeepSeek's text rate by less than it does today.

Compared in this article4
Detected from this article · Benchmarks: Artificial Analysis · updated daily
