
Qwen3.8-Max Review: Alibaba's 2.4T Flagship at $2/$6 — 0902 Build Tops Code Arena WebDev
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Alibaba Qwen previewed Qwen3.8-Max at the World AI Conference in Shanghai on July 19, 2026, and for two weeks it was the most talked-about model nobody could actually price. There was no rate card, no benchmark table, no model card, no license, and no active-parameter count — just a 2.4-trillion-parameter headline and a claim that the model sat "second only to Claude Fable 5."
On August 3, 2026, that changed. Qwen3.8-Max is generally available: a published rate card at $2 per million input tokens and $6 per million output, an OpenAI- and DashScope-compatible endpoint, a full benchmark table, and open weights that landed on August 12. A week later, on August 14, it went live on DigitalOcean Serverless Inference as Alibaba's "Day 0 launch partner" — the first major Western cloud to serve it. On September 2, Alibaba shipped a named refresh, Qwen3.8-Max-0902, further post-trained on Coding and Cowork, which Qwen says strengthens performance on complex enterprise and scientific work. This review is written against the GA facts, with that day-zero launch and the 0902 build folded in, and it answers the question teams actually have now that the model is buyable: is Qwen3.8-Max ready to take production traffic, or is it still a watchlist model?
The short answer is that it graduated further than most people expected — the pricing in particular is aggressive in a way that reprices the frontier, and the September 2 refresh doubles down on the two things the model was already good at — coding and collaborative work. The independent scorecard that has since landed is genuinely good, but it carries a caveat worth reading before you commit traffic.
TL;DR. Qwen3.8-Max is now GA at $2 / $6 per 1M tokens, flat across the entire 1M-token context with no long-prompt surcharge — undercutting Kimi K3 ($3/$15) and Claude Opus 5 ($5/$25) on output, the cost driver for reasoning work. Alibaba published a strong benchmark table (Terminal-Bench 2.1 86.6, GPQA Diamond 92.6, OSWorld-Verified 86.1) and the coding jump over its predecessor is dramatic: FrontierSWE 73.5 versus 40.7. But the quality table is Alibaba's own, and the first independent referee — the Artificial Analysis Intelligence Index — has now weighed in: Qwen3.8-Max scores 56 at $1.14 per task, on par with Claude Opus 4.8 (56) and one point behind open-weights leader Kimi K3 (57) at a 25% higher cost per task. The active-parameter count is no longer a question mark: 95B activated out of 2.4T total, confirmed on the official model card, which at last lets outsiders reason about serving economics. Since this review first went up, a second managed path opened too: Qwen3.8-Max went live on DigitalOcean Serverless Inference on August 14 as Alibaba's Day 0 partner, with DigitalOcean-published serving numbers — prefill scaling to ~16,000 tokens/sec and ~1,937 output tokens/sec at 256 concurrent requests with zero errors — that are the first hard latency data from a provider other than Alibaba. On September 2, Alibaba refreshed the flagship as Qwen3.8-Max-0902, further post-trained on Coding and Cowork; Qwen says the new snapshot is stronger on complex enterprise and scientific workloads — a vendor claim with no independent numbers yet on enterprise or scientific work. The 0902 build's coding side now has an independent arena result of its own: Qwen3.8-Max-0902 debuted at #1 on the Code Arena: WebDev board at 1,691 points — up 22 from the previous build and ahead of Claude Opus 5 and Kimi K3 — per Alibaba, citing the live third-party arena listing. The agentic-commerce story gained a scoreboard the same week: on CommerceAgentBench, a new benchmark from Alibaba's Accio team built on stateful replicas of real commerce software, Qwen3.8-Max posts the strongest open-weight result — 156 aggregate passes across three harnesses, two ahead of DeepSeek V4 Pro — a vendor-run table, not an independent referee. Verdict: genuinely buyable now, and cheap enough to justify a real evaluation — but route it deliberately, not wholesale.
Key takeaways
• GA since August 3, 2026. The preview-and-credits-campaign era is over; there is a real rate card and a real endpoint.
• $2 / $6 per 1M tokens, one flat tier across 1M context. Most rivals surcharge long prompts. Alibaba does not, which matters enormously for long-document and large-repo work.
• Alibaba's benchmark table is strong but unaudited: Terminal-Bench 2.1 86.6, GPQA Diamond 92.6, PaperBench 93.0, IFBench 82.8, OSWorld-Verified 86.1, OmniDocBench 1.5 92.1.
• The coding generational jump is the real story: FrontierSWE 73.5 (from 40.7) and DeepSWE 1.1 56.6 (from 21.6) — a near-doubling on both.
• The first independent score is in: Qwen3.8-Max scores 56 on the AA Intelligence Index at $1.14/task — a 10-point jump over Qwen3.7-Max (46), on par with Claude Opus 4.8 (56), one point behind Kimi K3 (57). LMArena's general leaderboard still has no public score.
• Active parameters confirmed at 95B — the official model card states 2.4T total and 95B activated, closing the biggest open question in the serving-cost math.
• Open weights are out — Qwen3.8-2.4T-A95B (BF16) and Qwen3.8-2.4T-A95B-FP8 landed on Hugging Face on August 12 under a custom Qwen3.8-Max license, not Apache 2.0. The Qwen3.8-27B on-premise path shipped on August 14 under Apache 2.0.
• A second managed path opened August 14. Qwen3.8-Max is live on DigitalOcean Serverless Inference as Alibaba's "Day 0 launch partner" — NVFP4 on NVIDIA HGX B300, $2/$6 with $0.20 cached input, served at the open checkpoint's native 262K context. DigitalOcean's measured figures (prefill ~16K tokens/sec, zero errors at 256 concurrency) are the first serving-performance data on a managed platform outside Alibaba.
• September 2 refresh: Qwen3.8-Max-0902. Alibaba post-trained the flagship further on Coding and Cowork and is serving it on QwenCloud at the same $2/$6 and 1M context. Qwen says enterprise and scientific performance are stronger — still vendor claims — while the coding side gained an independent arena result the same day: the build debuted at #1 on Code Arena: WebDev at 1,691 (Code Arena bullet below).
• CommerceAgentBench: open-weight leader, vendor-run. Alibaba's Accio team released CommerceAgentBench this week — 107 long-horizon tasks inside stateful replicas of real commerce and business software, graded state-based across the Accio, OpenClaw, and Pi harnesses. Qwen3.8-Max leads open-weight models with 156 aggregate passes, two ahead of DeepSeek V4 Pro; closed-model Claude Opus 5 leads overall at 191, and the table is Alibaba's own, not an independent referee.
• Code Arena: WebDev #1 — the 0902 build's independent arena result. Qwen3.8-Max-0902 debuted at the top of the Code Arena: WebDev leaderboard at 1,691 points — up 22 from the previous build and ahead of Claude Opus 5 and Kimi K3 — and it leads that board's cost-performance Pareto frontier at a blended ~$5/MToken, about a quarter of Claude Opus 5's blended price. Qwen's official QwenCloud account confirmed the placement. Unlike CommerceAgentBench, this board is third-party: a blind human-vote arena, not an Alibaba table, though 'debuting at #1' is Alibaba's announcement of a point-in-time listing.
Accuracy note: the Qwen3.8-Max benchmark table in this review is Alibaba-reported and has not been independently replicated. The Artificial Analysis Intelligence Index score (56 at $1.14/task) is independently measured by AA on Alibaba's official API — the model's first independent referee. Pricing is from Alibaba Model Studio's GA rate card. The open-weights repositories (Qwen3.8-2.4T-A95B and Qwen3.8-2.4T-A95B-FP8), the 95B-active figure, and the license text are first-party: they come from Qwen's own Hugging Face organization, verified August 13, 2026. The DigitalOcean availability, its rate card, the serving-performance measurements, and the GPQA spot-check of 88 on the quantized build come from DigitalOcean's own docs and community tutorial, published alongside the August 14 announcement; the "Day 0 launch partner" characterization is Alibaba's announcement language. Where we cite competitor numbers they come from Artificial Analysis or the vendor named inline. The Qwen3.8-Max-0902 refresh announced September 2 is vendor-stated throughout: its specs, its Coding-and-Cowork post-training description, and the claimed enterprise and scientific gains all come from Qwen's own announcement and its QwenCloud model page, and none of its numbers have been independently reproduced. The Code Arena: WebDev position covered in the benchmark section is the one exception: that board is a third-party, blind-vote arena rather than an Alibaba table — though 'debuting at #1 with 1,691' is Alibaba's announcement of a point-in-time listing, and arena scores keep moving as votes accrue. The CommerceAgentBench result covered below is in the same category: the benchmark was built and published by Accio, Alibaba International's team, so its tables are Alibaba-reported — an open-source benchmark, but not an independent referee in the way Artificial Analysis is. Vendor benchmark tables run optimistic as a rule — treat these as a hypothesis to test on your own workload, not a settled ranking.

What is Qwen3.8-Max?
Qwen3.8-Max is the flagship of Alibaba's Qwen family: a sparse Mixture-of-Experts model with 2.4 trillion total parameters, a 1M-token context window, and native multimodal input. It accepts text, images, and video, and returns text. It supports thinking mode, function calling, built-in tools, and structured outputs, and it is served over an OpenAI-compatible /v1/chat/completions endpoint, which means most existing integrations need a base-URL change rather than a rewrite. The current production snapshot, shipped September 2, is Qwen3.8-Max-0902, further post-trained on Coding and Cowork.
The practical context numbers are slightly more specific than the marketing "1M." Alibaba's cloud metadata lists a 983,616-token window with thinking enabled, a maximum input of roughly 991K tokens, and a maximum output of 131,072 tokens. That output ceiling is unusually generous and matters for tasks like full-file code generation or long report drafting, where lesser ceilings force you to chunk and stitch. The open checkpoint that DigitalOcean now serves is a different configuration: its model card lists 262,144 tokens natively, extensible toward roughly 1,010,000, so third-party serving that runs the open base will typically land at the smaller native number.
The active-parameter count is now public, and the repository name carries it: A95B means 95 billion activated. The official model card for Qwen3.8-2.4T-A95B states "2.4T in total and 95B activated," a fine-grained Mixture-of-Experts with 512 experts, of which 10 routed plus one shared are active per token. That number matters more than the 2.4T headline. For a dense model, total parameters tell you roughly what inference costs; for an MoE model, they tell you almost nothing — only the active count does, because that is what actually runs per token. A 2.4T model activating 30B behaves completely differently, in latency and in serving economics, from one activating 200B. At 95B active, per-token compute lands in the same order as a dense model around the 90B mark — a fraction of what a dense 2.4T would demand — which is the number that finally makes the self-hosting question below answerable.
The September 2 refresh: Qwen3.8-Max-0902
On September 2, 2026, Qwen shipped its first named production refresh. Qwen3.8-Max-0902 keeps the same 2.4T-parameter sparse-MoE architecture, the same 95B active parameters, and the same 1M-token context window, but it has been further post-trained on two capabilities — Coding and Cowork — and it is live on QwenCloud's API as qwen3.8-max-0902 (also listed as qwen3.8-max-2026-09-02). It is the current recommended build rather than a side branch: Alibaba's undated qwen3.8-max endpoint now serves the 0902 snapshot, so a call to the standard model name lands on the refresh, while the dated id is there for teams that want to pin the exact build. From September 3 the snapshot has also been listed as its own routable model id on third-party managed APIs. The rate card did not move: $2 per 1M input, $6 per 1M output, with the same cache rates.
The performance claims are Qwen's. The announcement and the QwenCloud model page describe Coding capability that "breaks new ground" on engineering-scale projects and long-horizon autonomous development, a "significantly enhanced" collaborative agent experience across multi-tool orchestration and end-to-end task delivery, and — as the company puts it — stronger performance on complex enterprise tasks, scientific research, and long-cycle workflows. The first independent check landed the same day. Qwen3.8-Max-0902 debuted at #1 on the Code Arena: WebDev board — a third-party, blind-vote arena for agentic web development, the project formerly known as WebDev Arena — with 1,691 points, up 22 from the previous build and ahead of Claude Opus 5 and Kimi K3, per Alibaba citing the live board — a placement Qwen's official account confirmed the same day, pointing to the QwenCloud API. What that does and does not prove is unpacked in the benchmark section below. The refresh's remaining claims — enterprise reasoning, scientific work, and general long-horizon autonomy — are still vendor-stated, so keep the discipline that applied to the August table: treat those as a hypothesis until an Artificial Analysis pass or a real workload confirms them.
What "Cowork" means in practice is where the refresh gets interesting. The GA-era Qwen3.8-Max already leaned on long-horizon autonomy — the 16-day oh-my-cli build, the 2,000-plus-round virtual business — and the 0902 build is Alibaba doubling down on that axis: agents that orchestrate multiple tools and deliver a task end-to-end rather than single-shot answers. The same week, Alibaba's Accio team published CommerceAgentBench, a benchmark built to measure that axis directly: 107 long-horizon tasks inside stateful replicas of real commerce and business software, where Qwen3.8-Max leads the open-weight field — detailed in the benchmark section below. For teams evaluating the model for enterprise work, that is the claim to test first, because it is also the hardest one to verify from a benchmark table.
Pricing: the most aggressive thing about this launch
Alibaba Model Studio lists Qwen3.8-Max at $2.00 per 1M input tokens and $6.00 per 1M output tokens. Output includes thinking tokens, which matters because reasoning-heavy configurations bill their internal deliberation at the output rate. Cache economics: cached input hits run $0.25 per 1M, explicit cache creation $2.50, and explicit cache reads $0.17.
The structurally interesting part is not the headline number, it is the flat tier. Alibaba charges the same $2/$6 whether your prompt is 5,000 tokens or 900,000. Most competitors apply a long-context surcharge precisely because serving a near-million-token prompt is expensive. Alibaba absorbing that is a deliberate land-grab aimed at exactly the workloads where long context is the point: whole-repository code review, contract and filing analysis, multi-document research synthesis.
Worked example. Say you run a repo-analysis agent: 200,000 input tokens of code context and 8,000 output tokens per call.
• Per call: 0.2M × $2 = $0.40 input, plus 0.008M × $6 = $0.048 output. ≈ $0.45 per call.
• At 1,000 calls/day: ≈ $448/day, or about $13,400/month.
• The same call against Claude Opus 5 at $5/$25: $1.00 + $0.20 = $1.20 — roughly 2.7× more.
• With prompt caching on a stable code context, cached input at $0.25 drops the input side from $0.40 to $0.05, taking the call to ≈ $0.10 — a nearly 4.5× reduction on repeat traffic.
That cache differential is where the real budget lives. If your agent re-reads a mostly-unchanged context on every turn — which describes most coding and support agents — the effective price is far closer to $0.25/$6 than $2/$6. Teams that skip cache configuration routinely overpay by 3–4× on this model.
For comparison at list prices: Qwen3.8-Max $2/$6; its own predecessor Qwen3.7-Max $2.50/$7.50; Kimi K3 $3/$15; GPT-5.6 Terra $2/$12; Claude Opus 5 $5/$25. On input Qwen is merely competitive. On output — the side that dominates spend in reasoning and agentic work — it is the cheapest frontier-claimed model on that list.
Artificial Analysis's independent measurement exposes the flip side of those cheap tokens. On its Intelligence Index, Qwen3.8-Max costs $1.14 per task — more than double the predecessor's $0.53, and 25% higher than Kimi K3's $0.86 even though Kimi K3 lists at $3/$15 per token. The gap is behavioral, not a price hike: the model averaged 64 steps per GDPval-AA task versus 14 for Qwen3.7-Max, and because the harness resends the full conversation at each step, input tokens grew roughly 15× and output ~45%. Cheap tokens do not equal cheap agents — the verbosity the preview testers flagged now has a measured price tag. Because OrcaRouter passes list pricing through at 0% markup behind one OpenAI-compatible endpoint, that $1.14-versus-$0.86 question is something you can measure on identical prompts without a second contract.

The benchmark table, and what each number actually measures
At GA Alibaba published the table it withheld at preview. Reported scores:
• Terminal-Bench 2.1 — 86.6. Measures whether a model can operate a real shell to completion: run commands, read output, recover from errors. This is the closest proxy for "can it function as a coding agent rather than a code suggester."
• GPQA Diamond — 92.6. Graduate-level physics, chemistry and biology questions written to be resistant to retrieval. A high score indicates genuine scientific reasoning rather than recall. 92.6 is at the top of the current field.
• PaperBench — 93.0. Reproducing results from research papers — reading a methodology and re-implementing it. Rewards sustained multi-step comprehension.
• IFBench — 82.8. Instruction following under constraint: format, length, and negative instructions. Notably the lowest score in Alibaba's own table, which is consistent with the preview-era complaint that the model is thorough to the point of ignoring scope limits.
• OSWorld-Verified — 86.1. Computer use: driving a real desktop GUI to finish tasks. Strong here supports the agentic positioning.
• OmniDocBench 1.5 — 92.1. Document parsing — tables, layout, mixed text and figures. Pairs with the 1M flat-rate context to make a real case for document pipelines.
• FrontierSWE — 73.5, up from a predecessor's 40.7. DeepSWE 1.1 — 56.6, up from 21.6.
Those last two deserve emphasis. Near-doublings on hard software-engineering benchmarks in one generation are not normal incremental gains, and they are consistent with Alibaba's decision to market this model at developers rather than chat users. If the FrontierSWE figure survives independent replication, Qwen3.8-Max is a serious coding model and not just a large one.
The caveat from preview has shrunk but not disappeared. The table above was produced by Alibaba, on its own harness, under conditions nobody else can inspect. Vendor tables are chosen, not sampled — labs publish the benchmarks they do well on. But as of August 6 there is now a genuine independent referee: Artificial Analysis has scored Qwen3.8-Max at 56 on the Intelligence Index at $1.14 per task, measured on Alibaba's official API — a 10-point jump over the predecessor's 46, level with Claude Opus 4.8 (56), and one point behind Kimi K3 (57). On AA's agentic GDPval-AA leaderboard the model reaches 1,739 Elo, passing Kimi K3 (1,685) and GPT-5.6 Sol (1,730) with only Claude Opus 5 (1,852) ahead. AA's own caveats deserve equal weight: an earlier run scored 53 before endpoint issues were resolved and the official-API retest produced 56; AA-Omniscience fell 10 points on a hallucination rate that rose from 23% to 40%; and the per-task cost is high precisely because the model grinds through far more steps. LMArena's general leaderboard still has no public score.
The DigitalOcean launch added a third data point, this one on the quantized build rather than the flagship. DigitalOcean's own internal spot-check of the NVFP4 weights it serves scored 88 on GPQA Diamond, against Alibaba's published 92.6 for the unquantized flagship. DigitalOcean itself flags the comparison as "an indicative data point rather than a controlled comparison" — the differences run on three axes (open-weights base rather than the hosted flagship, NVFP4 rather than BF16, and a different evaluation stack) — but it is the first independent check on a real serving deployment of these weights, and a reminder that the open checkpoint is not byte-for-byte the number on Alibaba's table.
CommerceAgentBench: the agentic-commerce scoreboard
Alongside the refresh, Alibaba International's Accio team released CommerceAgentBench, a benchmark built to score exactly the long-horizon work Qwen keeps marketing. An agent starts each task inside a fresh container in one of 14 offline replicas of real commerce and business software — product publishing, freight booking, storefront administration, payment operations, supplier analysis — and grading is state-based: a task passes only when the underlying mock services and produced files actually change, so an agent that narrates a plausible workflow without changing state scores nothing. The suite runs 107 tasks across CLI, browser, file/document, and API/MCP surfaces, executed through three harnesses — Accio, OpenClaw, and Pi — and the harness code (Apache 2.0) and task data (CC BY 4.0) are open for inspection or re-runs.
On the published tables, Qwen3.8-Max records the strongest open-weight result: 56 of 107 tasks on the Accio harness, 47 on OpenClaw, and 53 on Pi — 156 aggregate passes, two ahead of DeepSeek V4 Pro's 154. The entire open-weight edge comes from the Accio harness; the two models tie on OpenClaw and Pi. Open-weight leader is not the overall lead: closed-model Claude Opus 5 clears the field at 191 aggregate passes, 35 ahead. And the table is Alibaba's own — Accio is Alibaba's team, and the repository itself flags the audit limits: the Accio harness is reference-only and not reproducible from a checkout (OpenClaw is the rerunnable path), task-level results lack immutable public checksums, and Qwen's rows relied on a reasoning-content replay protocol to stay comparable. Treat it as a vendor-stated data point on the same agentic axis OSWorld-Verified and AA's GDPval-AA approach from different angles — worth probing against your own workflows, not a settled ranking.
Code Arena WebDev: the 0902 build's independent referee
Code Arena is the piece of independent evidence the refresh was missing. It is a third-party leaderboard for agentic web development — the project formerly known as WebDev Arena — where blind, pairwise human votes on which model's output a user would rather ship are converted to an Elo-style rating. On its WebDev board, Qwen3.8-Max-0902 debuted at the top with 1,691 points, up 22 from the previous build and ahead of Claude Opus 5 and Kimi K3. The placement has a cost-efficiency edge as well: at a blended ~$5 per million tokens — the $2/$6 list rate weighted toward the output-heavy mix that web-app generation actually produces — the 0902 build is the highest-scoring model on the board's Pareto frontier, costing about a quarter of Claude Opus 5's ~$20/M blended price and well under Kimi K3's ~$12/M, which is why those two sit off the frontier despite ranking #2 and #3 on score. Alibaba announced the position on September 2 and Qwen's official account confirmed it, pointing users to QwenCloud; the measurement is not Alibaba's: unlike the benchmark table above or the CommerceAgentBench run just above, this score is produced by a third-party arena from human preference votes.
Two caveats keep it honest. First, arena ratings are point-in-time: 1,691 is where the board sat when Alibaba announced the debut, and it will drift as votes accumulate, so read it as a strong signal on web-development coding rather than a permanent crown. Second, it only covers that one axis. The refresh's enterprise, scientific, and general long-horizon claims still have no independent referee, which is why the vendor table and the AA Index 56 on the base model remain the reference points for everything outside agentic front-end work. Third, the frontier price is a modeling choice, not a new rate card: the ~$5/MToken figure blends the unchanged $2/$6 list at an output-heavy usage assumption, so treat it as a cost-efficiency ranking on the arena's own terms rather than a repricing by Alibaba.

Open weights, Qwen3.8-27B, and the on-premise question
Alibaba's open-weight promise is now kept. On August 12, 2026, Qwen published two repositories on Hugging Face — Qwen3.8-2.4T-A95B in BF16 and Qwen3.8-2.4T-A95B-FP8 — each with 213 weight files. The license is a custom Qwen3.8-Max license, not Apache 2.0; Hugging Face's license field reads "other." The terms permit use, modification, distribution, and fine-tuning, including commercial use, with attribution, plus two scale-based gates: products above 100 million monthly active users or $20 million monthly revenue must display the model name prominently, and Model-as-a-Service or AI-Work-Assistant businesses above $50 million in aggregate trailing-twelve-month revenue need a separate license from Qwen. Most teams are inside those gates; if your business crosses them, read the license text before you build on it. The download counters corroborate the recency — 978 on the BF16 repo and 3,851 on the FP8 repo as of this writing, low for a frontier release and consistent with a repository created August 8 and flipped public only on August 12, not one that had been sitting visible for a week. One more honesty note: the open checkpoint is the base model, text-only with thinking always on, while the hosted Qwen3.8-Max API adds vision input, an optional non-thinking mode, and the full 1M context — downloading the weights gets you the model, not every API feature.
Self-hosting the flagship is still out of reach for most teams, but now it is out of reach with known math. The full 2.4T-parameter weight set is roughly 4.9TB in BF16 and about 2.4TB in FP8, because every expert has to be resident in memory even though only 95B are activated per token. At 4-bit quantization that is on the order of 1.2TB against roughly 141GB per H200, so you are looking at eight or more top-end accelerators before you serve a single token. What the now-public active count adds is the throughput side: with 95B activated per token, compute per token is comparable to a dense model in the ~90B class — a fraction of the cost a dense 2.4T model would impose — so once the weights are resident, the per-token cost is not the nightmare the total-parameter number suggests. That combination of a large memory footprint and modest per-token compute is exactly why Alibaba can price output at $6/1M, and it points self-hosting teams at multi-GPU nodes running the FP8 checkpoints rather than anything single-GPU.
DigitalOcean's serving choice is a working example of that math in production. It runs the model NVFP4-quantized on NVIDIA HGX B300 GPUs specifically so the 2.4T weights fit on a single node, avoiding cross-node expert routing — a live deployment of the 4-bit economics this section has been doing on paper. The trade-off is exactly the one the GPQA spot-check above quantifies: a smaller footprint, at some measurable cost in quality against the unquantized flagship.
Which is what makes Qwen3.8-27B the more consequential release for most readers — and unlike the flagship, its weights landed on schedule. On August 14, 2026, Qwen published Qwen/Qwen3.8-27B on Hugging Face and ModelScope under Apache 2.0: a dense 27B model, natively multimodal, with a 262K-token native context extensible to 1M and a configurable reasoning_effort mode. Unsloth's Daniel Han had estimated it would run in roughly 17GB of VRAM — a single prosumer card — and that is now testable: the official repo ships BF16 safetensors (about 55.6GB across 18 shards), with GGUF quants following in community repos. So the generation's release pattern is now complete: the 2.4T flagship's weights came first on August 12, and the small on-premise-path model landed two days later. We tracked its release in our Qwen3.8-27B release-date post.
Where Qwen3.8-Max fits: four scenarios
1. Long-document and contract analysis. This is the strongest fit. The flat rate across 1M tokens plus OmniDocBench 92.1 means you can push a 400-page filing through in one call without a surcharge and without chunking. Competitors charging long-context premiums make the same job meaningfully more expensive. On the DigitalOcean path, remember that path serves the open base at 262K context — still a large filing, but not the full million.
2. Repo-scale coding agents. Terminal-Bench 86.6 and FrontierSWE 73.5 support this, and the cache pricing makes iterative work on a stable codebase cheap. The thing to test first is latency under your own concurrency, because preview-era reviewers found the model thorough but slow on long agentic runs, and GA serving is a different animal from preview serving. AA's GDPval-AA result — a leading 1,739 Elo at the cost of 64 steps per task — is the independent confirmation of both halves of that trade-off. DigitalOcean's August 12 measurements — aggregate throughput scaling close to linearly to ~1,937 output tokens/sec at 256 concurrent requests with zero errors and no saturation found — are the first managed-platform data point on that question, though they are DigitalOcean's own numbers on its own serving, not a head-to-head against Alibaba's.
3. Document-and-screenshot pipelines. Video and image input plus OSWorld-Verified 86.1 make it credible for workflows that read screens — QA automation, support triage from screenshots, RPA-adjacent tasks. Again, the multimodal input is a hosted-API feature; the DigitalOcean open-base path is text-only.
4. Where we would not put it yet. Latency-critical interactive UX and any regulated workload requiring reproducible evaluation. For the first, you want measured throughput you control. For the second, reproducibility now has a model card and a license to cite, but Alibaba's benchmark table is still un-replicated — and AA's Omniscience regression (hallucination rate up to 40%) is a reminder that the independent score cuts both ways.

How to access Qwen3.8-Max
There are several practical paths. Direct via Alibaba Model Studio / DashScope — or Qwen's own QwenCloud API — gets you the rate card above and the earliest access to new dated snapshots — the September 2 Qwen3.8-Max-0902 build appeared there first before rolling out across other endpoints — at the cost of onboarding a new vendor and managing another key. Qwen Chat and the Qwen App are the fastest way to eyeball behavior before writing code. Downloading the open weights from Hugging Face is a fourth path for teams that want to run their own infrastructure — with the size and license caveats above. Through a router is the option most production teams should consider, because it removes the migration decision from the evaluation decision.
Since August 14, 2026, there is a genuinely new option: Qwen3.8-Max is live on DigitalOcean Serverless Inference, which Alibaba announced as its "Day 0 launch partner" — Alibaba's own framing, though the listing is DigitalOcean's and stands on its own. DigitalOcean serves the open-weights checkpoint (platform model id qwen3.8-max), NVFP4-quantized on NVIDIA HGX B300 GPUs in a technical collaboration with Inferact, as a fully managed serverless API: one endpoint, usage-based pricing, no infrastructure to manage. Its rate card matches Alibaba's list — $2.00 input / $6.00 output per 1M, with cached input at $0.20 — and it serves the open base's native 262K context with thinking always on, rather than the hosted flagship's ~1M-token multimodal mode. That context ceiling matters if your workload leans on the long end: the full million-token mode remains Alibaba-API-only.
What the DigitalOcean path adds beyond geography is the first hard serving data from a provider other than Alibaba. DigitalOcean's own measurements against its production endpoint, taken August 12, show prefill scaling linearly at roughly 16,000 tokens/sec — rule of thumb TTFT ≈ (input ÷ 16,000) + 0.7s, so ~1.1s at ~1,080 tokens and ~15.8s at ~254K — and aggregate throughput reaching ~1,937 output tokens/sec at 256 concurrent requests with zero errors and no saturation point found. Those are DigitalOcean's numbers on its own deployment, not a controlled bake-off, but they are the first latency-and-concurrency data on this model outside Alibaba's cloud, and they directly address the "thorough but slow" worry from the preview era.
Qwen3.8-Max is live on OrcaRouter as qwen/qwen3.8-max, and the September 2 refresh is routable under its own dated id — qwen/qwen3.8-max-0902 — both at the same $2.00 / $6.00 list rate. OrcaRouter applies 0% markup and passes provider pricing straight through, so either route costs the same as going direct. Our own serving telemetry currently shows a p50 time-to-first-token of 1.64 seconds over a 7-day window. That is a first-party latency measurement rather than a vendor claim, and it is worth weighing against the preview-era "slowest model I've used" reports: those came from a single reviewer running hour-long agentic builds on preview infrastructure, and they should be re-tested against GA serving rather than repeated as current fact.
The practical value of the router path is that Qwen3.8-Max sits behind the same OpenAI-compatible endpoint as the models you already run, so an evaluation is a model-string change. You can send 5% of production traffic to it, compare against your incumbent on identical prompts, and keep a fallback — which is exactly the posture a model with an unreplicated vendor table deserves. That same posture is the right way to test the DigitalOcean deployment: run a slice against it with a fallback in place, rather than betting a production path on a two-week-old serving stack.

Limitations and honest risks
• Vendor table still unaudited. The independent score is in (AA Index 56), but Alibaba's own benchmark table remains un-replicated, and AA's measurement flags an Omniscience regression with the hallucination rate up to 40%. Independent scoring helps; it does not make the vendor table auditable.
• The DigitalOcean path is the open base, not the flagship. It serves the checkpoint at 262K context, text-only, thinking always on, NVFP4-quantized — and DigitalOcean's own spot-check scores that build 88 on GPQA Diamond versus Alibaba's published 92.6, a gap DigitalOcean calls indicative rather than controlled. If you need the full million-token multimodal API, that is still Alibaba-hosted.
• Qwen3.8-27B shipped, but vendor-numbered. The 27B's weights landed August 14 under Apache 2.0, closing the on-premise gap — but its benchmark claims are Alibaba-reported and unreplicated, and community quantizations were still rolling out as of this update.
• Custom license, not Apache 2.0. The Qwen3.8-Max license allows commercial use with attribution but carries two scale gates — prominent model-name display above 100M MAU or $20M monthly revenue, and a separate license for MaaS/AI-Work-Assistant businesses above $50M trailing-twelve-month revenue. Read it before you scale past the thresholds.
• Verbosity and instruction drift. IFBench 82.8 is the weakest number in Alibaba's own table, and preview testers repeatedly saw the model exceed scope — adding unrequested features. AA's own numbers now put a figure on it: 64 steps per GDPval-AA task is what pushes cost to $1.14, 25% above Kimi K3. Charming in a demo, expensive when output is billed at $6/1M and destabilising when you need exact formats.
• Content guardrails. Like other Chinese-hosted frontier models, responses on politically sensitive topics are constrained. Relevant if your product serves open-ended queries.
• Thinking tokens bill as output. With reasoning enabled, real costs run above naive estimates. Measure with reasoning configured the way you will actually ship it.

FAQ
Is Qwen3.8-Max available now?
Yes. It reached general availability on August 3, 2026, with a published rate card and an OpenAI- and DashScope-compatible API. The current production snapshot — Qwen3.8-Max-0902, shipped September 2 — is live on QwenCloud's API and is what the standard qwen3.8-max endpoint now serves. This is a change from its July 19 preview status, when access was limited to Qwen Chat, the Qwen App, and Qoder's credits campaign.
What is Qwen3.8-Max-0902?
Qwen3.8-Max-0902 is the September 2, 2026 production refresh of Qwen3.8-Max: the same 2.4T-parameter sparse-MoE flagship and 1M-token context, further post-trained on Coding and Cowork, and live on QwenCloud's API at the same $2/$6. Qwen says the new snapshot is stronger on complex enterprise and scientific tasks — a vendor claim, not yet independently benchmarked — while on the coding axis it already holds an independent result: #1 on the Code Arena: WebDev board at 1,691 points and the top of its cost-performance Pareto frontier at a blended ~$5/MToken, per Alibaba's announcement of the live arena listing, confirmed by Qwen's own QwenCloud account.
How much does Qwen3.8-Max cost?
$2.00 per 1M input tokens and $6.00 per 1M output tokens, flat across the full 1M-token context with no long-prompt surcharge. Cached input hits are $0.25 per 1M, explicit cache creation $2.50, and cache reads $0.17. Thinking tokens bill at the output rate. On Artificial Analysis's Intelligence Index the model's measured cost is $1.14 per task — see the pricing section for why that runs above the token math. DigitalOcean's serverless path lists the same $2/$6 with cached input at $0.20.
How many parameters does Qwen3.8-Max have?
2.4 trillion total, in a sparse Mixture-of-Experts configuration. The active-parameter count — the figure that actually determines serving cost and latency — is now confirmed at 95 billion activated, per the official model card for Qwen3.8-2.4T-A95B (512 experts; 10 routed plus one shared active per token).
Are the Qwen3.8-Max weights open?
Yes — since August 12, 2026. Qwen published Qwen3.8-2.4T-A95B (BF16) and Qwen3.8-2.4T-A95B-FP8 on Hugging Face, each with 213 weight files. The license is a custom Qwen3.8-Max license (Hugging Face lists it as "other"), not Apache 2.0: it permits commercial use with attribution and carries two scale-based gates. The open checkpoint is text-only with thinking always on; the hosted API adds vision, an optional non-thinking mode, and the full 1M context.
Has Qwen3.8-Max been independently benchmarked?
Yes — as of August 6, 2026. Artificial Analysis scores it 56 on its Intelligence Index at $1.14 per task, measured on Alibaba's official API: a 10-point jump over Qwen3.7-Max (46), level with Claude Opus 4.8 (56), and one point behind Kimi K3 (57). The benchmark table above, however, remains Alibaba-reported and unreplicated, and LMArena's general leaderboard still has no public score. The independent-arena picture changed on September 2: the 0902 refresh debuted at #1 on the same platform's Code Arena: WebDev board at 1,691 points — a third-party, blind-vote result, not an Alibaba table — and Code Arena's cost-performance frontier lists it as the board's highest-scoring value at a blended ~$5/MToken. DigitalOcean's spot-check of its quantized build (88 on GPQA Diamond) is the first third-party check on a serving deployment, and it is explicitly indicative rather than controlled. One caution for the "independent" column: CommerceAgentBench, where Qwen3.8-Max leads open-weight models, is not a second independent referee — Accio, which built and published it, is Alibaba's own team.
Is Qwen3.8-Max better than Qwen3.7-Max?
On Alibaba's own numbers, substantially — FrontierSWE 73.5 versus 40.7 and DeepSWE 1.1 56.6 versus 21.6 are near-doublings on hard software-engineering tasks. It is also cheaper than its predecessor's $2.50/$7.50. Independently, AA puts the generational jump at 10 points on its Intelligence Index (56 versus 46). Both vendor claims should still be verified on your workload.
Can I self-host Qwen3.8-Max?
Not practically. The full 2.4T weight set is roughly 4.9TB in BF16 and about 2.4TB in FP8, around 1.2TB at 4-bit — against roughly 141GB per H200, that is eight or more top-end GPUs. With 95B active per token the compute per token is comparatively modest, but the memory footprint is the binding constraint. DigitalOcean's NVFP4 serving on B300s is the working proof that 4-bit can fit the model on a single node. The Qwen3.8-27B — reportedly ~17GB of VRAM — is the realistic on-premise option, and its weights shipped August 14, 2026 under Apache 2.0 — now downloadable rather than a promise.
What is Qwen3.8-27B?
A much smaller model announced alongside the flagship, positioned as the practical on-premise deployment path. It is a dense 27B model, natively multimodal, with a 262K-token native context extensible to 1M, and it ships under Apache 2.0. Its weights landed on Hugging Face and ModelScope on August 14, 2026; Unsloth's Daniel Han reports it runs in roughly 17GB of VRAM — a single high-end consumer card. Its benchmark claims are Alibaba-reported and have not been independently reproduced.
Is Qwen3.8-Max multimodal?
Yes — it accepts text, image, and video input and returns text. It also supports thinking mode, function calling, built-in tools, and structured outputs. The open-weights checkpoint is text-only; multimodal input is a feature of the hosted API, which is what the DigitalOcean path does not carry.
Is Qwen3.8-Max available on DigitalOcean?
Yes — since August 14, 2026. DigitalOcean Serverless Inference serves the open-weights checkpoint (NVFP4 on NVIDIA HGX B300) at $2.00/$6.00 per 1M with $0.20 cached input, a 262K-token context, text-only, thinking always on. Alibaba calls DigitalOcean its "Day 0 launch partner"; the listing itself is DigitalOcean's and independent of that language. DigitalOcean's measured figures: prefill ~16K tokens/sec and ~1,937 output tokens/sec at 256 concurrent requests with zero errors.
Can I use Qwen3.8-Max through OrcaRouter?
Yes. It is live as qwen/qwen3.8-max, and the September 2 snapshot as its own dated id — qwen/qwen3.8-max-0902 — at the same $2.00/$6.00 list pricing, with 0% markup. Routing costs the same as going direct, behind the OpenAI-compatible endpoint you already use. Our 7-day telemetry shows a p50 time-to-first-token of 1.64 seconds.
Should I move production traffic to it today?
Not wholesale. The pricing and the first independent score make a serious evaluation cheap and worth doing now, but with a per-task cost 25% higher than Kimi K3's, the disciplined move is a traffic split against your incumbent on identical prompts, with a fallback in place — not a migration. The DigitalOcean path adds a second managed deployment to test against, which makes the evaluation cheaper still.
Bottom line
Qwen3.8-Max spent two weeks as the most interesting model nobody could buy. At GA it is a genuinely different proposition: $2/$6 flat across a million tokens is aggressive pricing, the cache rates make iterative agent work cheap, and the generational leap on FrontierSWE and DeepSWE — if it replicates — makes this a real coding model rather than a scale flex. The open weights arriving under a real license on August 12 turned "open weights coming" from a promise into a fact, and the day-zero DigitalOcean path announced August 14 — with measured serving numbers and a $0.20 cache read — makes the "try it cheaply" case stronger and gives teams a managed third-party deployment to compare against Alibaba's own API. The September 2 Qwen3.8-Max-0902 refresh keeps that thesis intact: the same $2/$6 and 1M context, with more post-training on the coding and collaborative work the model was already positioned for. The refresh also added an independent, human-preference scoreboard for that coding pitch: Qwen3.8-Max-0902 debuted at #1 on Code Arena's WebDev board at 1,691 points, ahead of Claude Opus 5 and Kimi K3, and at a blended ~$5/MToken it is the highest-scoring model on that board's cost-performance Pareto frontier. The CommerceAgentBench result Alibaba's Accio team published the same week — Qwen3.8-Max leading the open-weight field on a benchmark built from real commerce workflows — gives that collaborative-work thesis a concrete scoreboard of its own, albeit a vendor-run one.
What has changed is that the referee and the weights both landed — and the verdict is mostly good with real asterisks. AA puts Qwen3.8-Max at 56 on the Intelligence Index, a genuine generational jump that lands level with Claude Opus 4.8, though the same scorecard shows Kimi K3 a point ahead at 25% lower cost per task and flags a 10-point Omniscience drop on a rising hallucination rate. The active-parameter question is settled — 95B activated, confirmed on the model card — and the license is published, custom rather than Apache 2.0. What has not changed: Alibaba's own benchmark table is still unreplicated, the per-task cost is still high because the model grinds through so many steps, the open base served on DigitalOcean is a slightly different model from the flagship's headline numbers (262K context, NVFP4, GPQA spot-check 88 versus 92.6), and the Qwen3.8-27B's benchmark claims are vendor-reported even though its weights have shipped. That combination doesn't make Qwen3.8-Max a bad bet — at this price it is close to an obligatory experiment — but it does mean the right posture is evaluate aggressively, route deliberately, keep a fallback. The event to watch now is whether the agentic step-count behind that $1.14-per-task cost is something your workload can absorb, whether the 0902 build's Code Arena: WebDev lead holds as votes accrue, whether its Coding and Cowork gains replicate on real workloads, whether the CommerceAgentBench open-weight lead — two aggregate passes on Alibaba's own harness — survives an independent run, whether the 27B's benchmark claims hold up, and whether the DigitalOcean deployment's measured throughput holds under real traffic — that will decide whether "second only to Fable 5" was positioning or prophecy.

Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
