
Qwen 3.8 vs Claude Fable 5: Alibaba's 2.4T Flagship Finally Has a Price Tag
- qwenNEWQwen: Qwen3.8 Max2026-08-03$2.00 / $6.00 per 1M tokens · 54 tok/s
- deepseekNEWDeepSeek: DeepSeek V4 Flash 07312026-07-3150Intelligence69Coding
- qwenNEWQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 206 tok/s
- orcaNEWOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicNEWAnthropic: Claude Opus 52026-07-2461Intelligence78Coding
- googleNEWGoogle: Gemini 3.6 Flash2026-07-2150Intelligence69Coding
- googleNEWGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1651Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1557Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0951Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0955Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0959Intelligence77Coding
- grokxAI: Grok 4.52026-07-0854Intelligence72Coding
- tencentTencent: Hy32026-07-0641Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3053Intelligence72Coding
- klingKling: Kling 3.0 Turbo2026-06-1757Intelligence52Coding57Math
- z-aiZ.ai: GLM 5.22026-06-1651Intelligence69Coding60Math
- kimiMoonshotAI: Kimi K2.7 Code2026-06-1242Intelligence61Coding61Math
When Alibaba previewed Qwen3.8-Max in Shanghai on July 19, 2026, it made one claim louder than any other: this 2.4-trillion-parameter model is close to frontier-level and behind only Claude Fable 5. At the time that was impossible to evaluate. There was no rate card, no benchmark table, and no license — just a positioning statement.
On August 3, 2026, Qwen3.8-Max went generally available, and the comparison finally has substance on both sides. Alibaba published pricing at $2 / $6 per million tokens and a benchmark table with real numbers on it. Fable 5 still sits at the top of the independent Artificial Analysis Intelligence Index with an audited 95.0% on SWE-bench Verified. So the question sharpens: now that Qwen has shown its numbers, does "second only to Fable 5" hold up?
A note for builders — the fastest way to settle a claim like this is to run both models on your own prompts. OrcaRouter fronts 200+ models behind one OpenAI-compatible endpoint, so you can pit Qwen 3.8 Max against Fable 5 on the same task without wiring up two SDKs.
TL;DR verdict. Qwen3.8-Max at GA is a much stronger challenger than its preview suggested — and dramatically cheaper. At $2 / $6 flat across 1M tokens it undercuts Fable 5's $10 / $50 by roughly 5× on input and 8× on output, and its self-reported numbers are frontier-shaped (GPQA Diamond 92.6, Terminal-Bench 2.1 86.6, FrontierSWE 73.5 against a predecessor's 40.7). But Fable 5 keeps the crown for one reason that has not changed since July: it is the only one of the two with a referee. Fable's ~60 Intelligence Index and 95.0% SWE-bench are third-party audited; every Qwen figure is Alibaba's own, and neither Artificial Analysis nor LMArena has posted a Qwen3.8-Max score. Fable 5 wins on proof; Qwen 3.8 wins decisively on price and, once the weights land, on openness.
Key takeaways
• Qwen3.8-Max is GA as of August 3, 2026 — no longer a preview. Published rate card, OpenAI- and DashScope-compatible API, vendor benchmark table.
• The price gap is enormous: Qwen at $2 / $6 per 1M versus Fable 5 at $10 / $50. That is ~5× on input and ~8× on output, and Qwen's rate is flat across the entire 1M-token context with no long-prompt surcharge.
• Fable 5 remains the only audited side. ~60 on the Artificial Analysis Intelligence Index (#1) and 95.0% SWE-bench Verified, versus Qwen3.8-Max still unscored by every independent evaluator.
• Qwen's self-reported numbers are genuinely strong: GPQA Diamond 92.6, PaperBench 93.0, Terminal-Bench 2.1 86.6, OSWorld-Verified 86.1 — but they are vendor-produced, and labs publish the benchmarks they win.
• Alibaba never published an active-parameter count, so the 2.4T headline tells you very little about real serving cost. Fable 5's architecture is undisclosed too — on transparency, neither side is clean.
• Open weights are promised within the week, plus a Qwen3.8-27B reportedly running in ~17GB of VRAM. Fable 5 is closed and API-only, permanently.
Accuracy note: all Qwen3.8-Max quality figures are Alibaba-reported and have not been independently replicated. Fable 5 figures come from Artificial Analysis and Anthropic and may differ across trackers. Qwen's pricing is the GA rate card from Alibaba Model Studio and is no longer the temporary preview discount that applied in July.
The specs and price, side by side
Here is the hard data, each figure attributed to its source.
• Maker / status — Qwen3.8-Max: Alibaba; GA (August 3, 2026); Claude Fable 5: Anthropic; GA flagship
• Architecture — Qwen3.8-Max: ~2.4T total, sparse MoE (active params still undisclosed); Fable 5: Closed; architecture not disclosed
• Context window — Qwen3.8-Max: 1M tokens (983,616 with thinking; max output 131,072); Fable 5: Long-context flagship
• AA Intelligence Index — Qwen3.8-Max: Still unscored; Fable 5: ~60 — #1 (Artificial Analysis)
• SWE-bench Verified — Qwen3.8-Max: Not reported (Alibaba reports FrontierSWE 73.5 instead); Fable 5: 95.0% (Anthropic / buildfastwithai)
• Vendor-reported strengths — Qwen3.8-Max: GPQA Diamond 92.6, Terminal-Bench 2.1 86.6, PaperBench 93.0, OSWorld-Verified 86.1 (all Alibaba-reported); Fable 5: audited #1 overall
• Pricing (in / out) — Qwen3.8-Max: $2.00 / $6.00 per 1M, flat across 1M context; cached input $0.25; Fable 5: $10 / $50 per 1M
• Open weights — Qwen3.8-Max: Promised within the week (no license yet); Fable 5: No — closed / API-only
• Multimodality — Qwen3.8-Max: Text, image, video in → text out; Fable 5: Multimodal flagship
Two things frame the whole comparison, and both changed on August 3.
First, the price column is now the loudest thing on the page. In July, Qwen's cost advantage was a discount coupon — 90% off credits, temporary, with no per-token rate to plan against. Now it is a rate card, and the gap is structural rather than promotional. At $6 output against Fable's $50, you can run roughly eight Qwen calls for the price of one Fable call. For any workload where you are paying for volume rather than for the single hardest answer, that ratio changes the architecture of what you build.
Second, the benchmark column is no longer empty — but it is also not comparable. This is the subtlety that matters most. Alibaba reports FrontierSWE 73.5; Anthropic reports SWE-bench Verified 95.0%. Those are different benchmarks with different difficulty curves, and setting 73.5 against 95.0 as though they measure the same thing would be straightforwardly misleading. What we can say is narrower and more honest: Fable 5's number was produced by a third party under conditions others can inspect, and Qwen's was produced by Alibaba on a harness nobody else has run. Until Artificial Analysis or LMArena posts a Qwen3.8-Max score, the correct read is unchanged from July — plausibly close, still unproven.

What the price gap actually buys you
Abstract multiples are less useful than a worked example, so here is one. Take a code-review agent that sends 150,000 tokens of repository context and generates 6,000 tokens of review per call.
• Qwen3.8-Max: 0.15M × $2 = $0.30, plus 0.006M × $6 = $0.036. ≈ $0.34 per call.
• Claude Fable 5: 0.15M × $10 = $1.50, plus 0.006M × $50 = $0.30. ≈ $1.80 per call.
• At 2,000 calls/day: Qwen ≈ $672/day; Fable ≈ $3,600/day. Over a month that is roughly $20,000 versus $108,000.
• With prompt caching on a stable repo context, Qwen's cached input at $0.25/1M drops the input side from $0.30 to under $0.04, taking the call to ≈ $0.08 — about 22× cheaper than Fable per call.
That is not a marginal saving; it is the difference between running a reviewer on every pull request and running one on the risky ones. And Qwen's flat-rate context makes the effect stronger as prompts grow: because Alibaba applies no long-prompt surcharge, pushing a 900,000-token context through costs the same per token as a 5,000-token one.
The honest counterweight is that price per token is not price per solved task. If Fable 5 resolves a problem in one pass where a cheaper model needs three attempts plus human correction, the cheaper model may cost more in aggregate — and on the hardest reasoning work, Fable's audited #1 ranking is the best available evidence that it does resolve more in one pass. The cost argument for Qwen is strongest on high-volume, tolerant workloads and weakest on low-volume, high-stakes ones.
Proof, and why it still decides this matchup
The most-cited hands-on evidence for Qwen3.8-Max comes from a preview-era review (thomas-wiegold.com) that put the model through four real builds. The results were genuinely impressive: Qwen one-shotted a Go poker simulation — only the third model ever to clear that prompt in a single pass, alongside Fable 5 and Grok 4.5 — and on a coffee-roaster website prompt it produced the best output the reviewer had ever seen from that test, adding an unrequested shopping cart, wholesale section, and Instagram integration out of sheer thoroughness.
The catch was speed: 30+ minutes on the website and 1 hour 20 minutes on the poker simulation, making it the slowest model that reviewer had used. That finding needs a caveat now that it did not need in July. It was measured on preview infrastructure, by one reviewer, on unusually long agentic runs. GA serving is a different system. For what it is worth, OrcaRouter's own 7-day telemetry currently shows a p50 time-to-first-token of 1.64 seconds for Qwen3.8-Max — a first-party measurement, though TTFT and end-to-end throughput on hour-long builds are different quantities. The honest position is that the preview slowness finding should be re-tested rather than repeated as current fact.
What has not moved is the asymmetry of evidence. Alibaba's GA benchmark table is a real improvement over publishing nothing, but it is still a vendor artifact: labs choose which benchmarks to report, and Alibaba's own weakest reported number — IFBench 82.8, instruction-following — happens to align exactly with the preview complaint that the model over-delivers and ignores scope. Fable 5, meanwhile, is graded by an evaluator with no stake in the result. On the question "which model will handle work I cannot afford to get wrong," that difference is not a technicality. It is the entire basis for choosing.

Openness: the axis where Qwen may win permanently
Fable 5 will never be self-hostable. Qwen3.8-Max might be, and Alibaba now says the weights arrive within the week. But two caveats deserve equal weight.
The first is legal: there is still no license text and no model card. "Open weights coming" is not a grant of commercial rights until there is a document to read, and the license will decide whether the release is usable in a product at all.
The second is physical. A 2.4T model is roughly 1.2TB at 4-bit quantization against about 141GB per H200 — eight or more top-end accelerators before you serve a token. And because Alibaba still has not disclosed the active-parameter count, you cannot even estimate the throughput that outlay would buy. For nearly every team, self-hosting the flagship is not a real option.
Which makes Qwen3.8-27B — announced alongside, also going open-weights, reportedly running in ~17GB of VRAM — the more practically important release. If your reason for preferring Qwen over Fable 5 is data residency or on-premise control rather than raw capability, 27B is the model that actually delivers it. Comparing the 2.4T flagship to Fable 5 on openness is mostly theoretical; comparing the 27B on openness is concrete.
FAQ
Is Qwen 3.8 better than Claude Fable 5?
Not on the evidence that exists. Alibaba's GA benchmark table is strong (GPQA Diamond 92.6, Terminal-Bench 2.1 86.6), but every number is self-reported, and no independent evaluator has scored the model. Fable 5 holds an audited ~60 Intelligence Index (#1) and 95.0% SWE-bench Verified. Fable 5 leads on proof; Qwen 3.8 leads overwhelmingly on price.
Is the "behind only Fable 5" claim true?
It remains Alibaba's own positioning. GA brought a benchmark table but not third-party verification — Artificial Analysis and LMArena are both still unscored. For reference, the predecessor Qwen3.7-Max scored 46 on the AA Intelligence Index against Fable 5's ~60, so the claim requires a very large generational jump to be literally true.
How much cheaper is Qwen 3.8 than Fable 5?
Substantially, and it is now a real rate rather than a discount. Qwen3.8-Max is $2 / $6 per 1M tokens against Fable 5's $10 / $50 — about 5× cheaper on input and 8× on output. Qwen's rate is also flat across the full 1M context, and cached input at $0.25/1M pushes repeat-context workloads cheaper still.
Did Qwen 3.8 Max leave preview?
Yes. It reached general availability on August 3, 2026 with a published rate card and an OpenAI- and DashScope-compatible endpoint. The Qoder credits campaign and 90%-off preview framing that circulated in July no longer describe how the model is sold.
Is Qwen 3.8 still the slowest model, as early reviews said?
That finding came from one reviewer on preview infrastructure running hour-long agentic builds, and it should be re-tested against GA serving rather than treated as current fact. OrcaRouter's 7-day telemetry shows a p50 time-to-first-token of 1.64 seconds, though that measures first-token latency rather than throughput on very long builds.
Can I self-host Qwen 3.8 instead of paying for Fable 5?
Not yet, and probably not the flagship. Weights are promised within the week but there is still no license; and at ~1.2TB in 4-bit you would need eight or more H200-class GPUs. The concurrently announced Qwen3.8-27B, reportedly ~17GB of VRAM, is the realistic self-hosting path. Fable 5 is closed permanently.
Which should I use today?
Fable 5 for work where a wrong answer is expensive and you need audited reliability — hard reasoning, high-stakes production paths. Qwen3.8-Max for high-volume work where the 8× output-price advantage compounds: bulk code review, long-document analysis, anything where you can tolerate verification. Many teams should run both and route by task value.
Bottom line
Qwen 3.8 vs Claude Fable 5 is a materially different comparison than it was two weeks ago. Qwen3.8-Max is no longer a preview with a coupon — it is a generally available model with a rate card that undercuts Fable 5 by 5× on input and 8× on output, flat across a million tokens, with self-reported numbers that look frontier-shaped and a coding jump over its predecessor that would be remarkable if it replicates.
But Fable 5 keeps the crown, and for exactly the reason it kept it in July: it is the side with a referee. An audited #1 Intelligence Index and 95.0% SWE-bench Verified are claims someone else checked. Alibaba's table, however strong, is Alibaba's. Watch two events: the first independent Intelligence Index score for Qwen3.8-Max, and the weights landing with a license attached. If both go Qwen's way, "behind only Fable 5" stops being marketing. Until then the sensible answer is not to pick a winner but to split by stakes — Fable where you cannot be wrong, Qwen where you cannot afford to be expensive — and to test both on your own prompts.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
