
Qwen3.8-Max vs Qwen 3.8: One 2.4T Core, Rented or Run — After the 0902 Refresh
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
On September 2, 2026, Alibaba shipped a dated refresh of Qwen3.8-Max — the build it pins as Qwen3.8-Max-0902 — and an independent coding leaderboard put the new snapshot at #1 the same week. The downloadable version of that same 2.4-trillion-parameter core, the model most people mean when they write "Qwen 3.8" with no suffix, is still sitting at its August 12 checkpoint: text-only, reasoning locked on, and hundreds of gigabytes short of runnable on anything smaller than a GPU rack. That gap between the hosted flagship and the open weights did not exist in August. It is the whole comparison now, and it is why the same model keeps being sold to you under two names.
Qwen3.8-Max is Alibaba's hosted flagship: a 2.4-trillion-parameter sparse mixture-of-experts model, roughly 95 billion parameters active per token, live on a priced API since August 3 and carrying a 1-million-token multimodal context window. Qwen 3.8, taken as a standalone name, is the same generation's open release — most consequentially Qwen3.8-2.4T-A95B, the weights Alibaba published on August 12 under a custom license. They are not a cheaper sibling and a premium sibling. They are the same brain sold two ways, and a September post-training pass has started to pull the two deliveries apart. This piece sorts out which "Qwen 3.8" you are actually looking at, what the 0902 refresh changed, and which route — rent the API or run the weights — you should be on today.
The pairing that is not a rivalry
Before the dates and the rate cards, the architecture: Qwen3.8-Max and Qwen3.8-2.4T-A95B share a fine-grained MoE design with 512 experts per layer (10 routed plus one shared, per layer), 92 layers, and a hybrid stack in which 69 of the 92 layers use linear attention — Gated DeltaNet combined with gated attention — with full attention interleaved for the long-context work. That is why the model card can claim the million-token context on the API and why the open build reaches a native 262K that extends toward a million with the right serving configuration. The differences between the two names are not in the weights' architecture; they are at the edges, and the edges have moved since September 2.
• Parameters — Qwen3.8-Max: 2.4T total / ~95B active, MoE. Qwen3.8-2.4T-A95B: the same core, same active count.
• Delivery — Qwen3.8-Max: hosted API, live since August 3, refreshed as Qwen3.8-Max-0902 on September 2. Qwen3.8-2.4T-A95B: downloadable weights, first published August 12, no newer checkpoint since.
• Modality — Qwen3.8-Max: text, image, and video input. Qwen3.8-2.4T-A95B: text-only.
• Reasoning control — Qwen3.8-Max: thinking toggles, including a non-thinking mode. Qwen3.8-2.4T-A95B: thinking always on, with only a reasoning-effort dial (low / high / extra-high).
• Tooling — Qwen3.8-Max: native function calling, five built-in tools, and structured output. Qwen3.8-2.4T-A95B: no native tool layer; what you get depends on the inference framework you serve it with.
What the September 2 refresh changed
The 0902 build is a post-training pass, not a new architecture. Alibaba pointed it at the two things Qwen3.8-Max was already known for — coding and "cowork," its name for long-horizon agentic and office work — and the numbers that landed around it are the reason this refresh matters. On the independent Code Arena: WebDev leaderboard, Qwen3.8-Max-0902 debuted at 1,691 Elo, a 22-point jump over the previous build and enough to take first place on arrival. Alibaba also frames the build as the top-scoring model on that board's cost-performance Pareto frontier at a blended rate around $5 per million tokens — the $2-and-$6 list price weighted by output-heavy agentic traffic, roughly a quarter of what a comparable frontier call runs elsewhere. Those figures are a mix the reader should keep straight: the 1,691 Elo is an independent arena result; the Pareto-frontier framing is Alibaba's own characterization of that independent board.
Alibaba's internal evals for the 0902 pass, vendor-reported and unreproduced, show the same direction: Terminal-Bench 3.0 rising from 11.3 to 29.0, ProgramBench from 10.5 to 28.0, JobBench from 53.4 to 64.0, and WorkArena Elo from 1,348 to 1,468. Read those as "the refresh moved the agentic and coding axes," not as verified scores. On the general-intelligence side, Artificial Analysis' Intelligence Index has Qwen3.8-Max at 56 — competitive, but not the reason to reach for it; the coding and agentic results are.
The part most coverage has missed is that the 0902 post-training is, as of this writing, API-only. Alibaba has not published an open checkpoint of the refreshed model — the Qwen3.8-2.4T-A95B weights on Hugging Face still trace to the August 12 release. That means the hosted Qwen3.8-Max and the downloadable core are no longer the same model in the way they were in mid-August. The API has the September post-training; the weights do not. If your whole reason for running the open release was to reproduce exactly what the flagship does, that assumption quietly stopped being true on September 2.

The five places they genuinely diverge
Put the shared architecture aside and the practical differences line up cleanly. Modality is the first filter: if your workload includes images or video — screenshots, charts, documents with figures, video frames — only Qwen3.8-Max takes them, and its 1M-token window is native rather than an extension. The open weights are text-only. Reasoning control is the second: the hosted API can run with thinking off, which matters for latency-sensitive and high-volume extraction where a chain of thought is pure overhead; the open build cannot switch it off. Tooling is the third: function calling, the five built-in tools, and JSON mode are part of the hosted product, while the open build depends on whatever your serving layer provides.
Context is subtler than the spec sheet suggests. Both sides can reach roughly a million tokens, but the hosted model does it natively with multimodal input, while the open build's 262K native context has to be extended in configuration, and every token of that extended window is text. Output caps differ too — the API allows up to about 131K output tokens, and the open build is typically configured to a similar ceiling, but the practical ceiling on the open side is set by your serving stack, not by the model.
What each route actually costs
This is where the two deliveries stop being comparable at all. Qwen3.8-Max has a list price: $2.00 per million input tokens, $6.00 per million output, and $0.25 per million for cached input. Because OrcaRouter passes provider list prices through at 0% markup, that is the rate you pay here, and when Alibaba changes it, the new number is live the same day. The 0902 snapshot is pinned separately as qwen/qwen3.8-max-0902 for teams that want reproducible behaviour — same price, same capabilities, but a fixed checkpoint instead of the rolling model. On OrcaRouter's traffic, the 7-day median time-to-first-token for Qwen3.8-Max is roughly 2–4 seconds, and Artificial Analysis measures around 45–48 output tokens per second: not slow, but verbose, which is why agentic tasks on this model can burn surprising token counts despite the cheap rate.

Qwen3.8-2.4T-A95B has no list price, because the price is your infrastructure. The BF16 weights run to roughly 4.9 TB and even the official FP8 build is about 2.4 TB, so this is rack-scale hardware: the smallest documented serving configs start around eight B300 or GB300 GPUs, and a full GB300 NVL72 rack is the reference deployment. Aggressive quantization changes the floor — an NVFP4 build fits in roughly 1.3 TB, and Unsloth's 1-bit layered quantization compresses the model to around 397 GB, which is what finally makes a single well-provisioned node feasible — but every one of those paths assumes you operate GPUs at data-center scale. Third-party hosts that serve the open checkpoint will sell you tokens below the Max's per-token price, but you give up the multimodal input, the non-thinking mode, the native tooling, and the 0902 post-training in the process, and no independent benchmark of that downloadable checkpoint itself has been published yet. Alibaba's own model card is candid about the relationship: it describes Qwen3.8-Max as the official version built on Qwen3.8-2.4T-A95B, adding vision input, non-thinking support, a default 1M context, and built-in tools on top of the open core.

Independent numbers versus vendor claims
The sourcing honesty matters more here than in most matchups, because the two sides have very different evidence ages. Qwen3.8-Max has been independently scored: the Code Arena: WebDev 1,691 result for the 0902 build and Artificial Analysis' Intelligence Index of 56 are third-party measurements, and the pricing that makes the Pareto-frontier claim interesting is a published rate card. Qwen3.8-2.4T-A95B does not yet have that — the benchmark table Alibaba published describes the Max-class core, not an independent run of the downloadable checkpoint, so the open side of this comparison is still largely "vendor says the core scores like this." If you are deciding between them on capability grounds, treat the open weights' headline numbers as claims until a third party runs them.
Who should pick which
The decision falls out of the list above. Reach for the hosted Qwen3.8-Max — today, specifically the 0902 build — when you need the September post-training, image or video input, a non-thinking mode, native tools and structured output, or a million-token window without standing up infrastructure. That is the coding-and-agentic frontier position, and at $2/$6 with $0.25 cached input it is aggressively priced for what it is.
Choose Qwen3.8-2.4T-A95B when control outweighs convenience: you need the weights on your own hardware or a provider you already operate, you want a fixed checkpoint you can audit and reproduce, or your sustained volume makes per-token cost the dominant term and you are prepared to run a 2.4T MoE. Accept the trade list — text-only, thinking always on, no native tooling, no 0902 post-training yet — and verify the licence terms for your revenue band, because the custom Qwen3.8-Max license is not Apache 2.0 and adds conditions for large commercial operators.
If your workload is high-volume, single-GPU, or latency-sensitive, neither of these may be the right lane — the generation's 27-billion-parameter sibling, Qwen3.8-27B, is the one sized for that, and it is covered separately in our Qwen3.8-27B vs Qwen3.8-Max comparison.
How to try both without betting your stack
The clean way to make this call is to run both sides on your own prompts, and this is where a routing layer earns its keep rather than bolting one on. Qwen3.8-Max and the pinned Qwen3.8-Max-0902 snapshot both sit behind one OpenAI-compatible endpoint on OrcaRouter, so switching between the rolling flagship and the reproducible checkpoint is a model-id change, not a redeploy. When a team decides to bring the open weights in-house, the routing DSL can front a self-hosted vLLM or SGLang endpoint serving Qwen3.8-2.4T-A95B and place it in the same call graph — bulk traffic to your own deployment, hard prompts and multimodal requests spilled to the hosted Qwen3.8-Max, with automatic failover between them. Trying a new checkpoint that way costs a config change rather than a migration, which is exactly what you want when the two sides of this comparison are still moving at different speeds.
The bottom line
Qwen3.8-Max and Qwen 3.8 are not two models fighting for the same job. They are one 2.4-trillion-parameter core delivered as a rented API and as downloadable weights, and the September 2 refresh made the two deliveries mean different things. Rent the API and you get the 0902 post-training that took the Code Arena: WebDev lead, plus vision, a non-thinking mode, native tooling, and a native million-token window, at $2/$6 with $0.25 cached input. Run the open weights and you get the same architecture frozen at its August 12 state — text-only, reasoning locked on — with no independent scores yet and a data-center hardware bill. If the September coding and agentic gains matter to you, only one of the two names has them today. If reproducible control matters more, only one of the two names is yours to keep.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
