Qwen 3.8 vs Claude Fable 5: The 0902 Build Takes the Coding Lead
Guides & Insights

Qwen 3.8 vs Claude Fable 5: The 0902 Build Takes the Coding Lead

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

When Alibaba previewed Qwen3.8-Max in Shanghai on July 19, 2026, it made one claim louder than any other: this 2.4-trillion-parameter model is close to frontier-level and behind only Claude Fable 5. At the time that was impossible to evaluate. There was no rate card, no benchmark table, and no license — just a positioning statement.

On August 3, 2026, Qwen3.8-Max went generally available — a real rate card at $2 / $6 per million tokens, a benchmark table with numbers on it, an OpenAI- and DashScope-compatible endpoint. On September 2 Alibaba shipped the next step without changing the version number: Qwen3.8-Max-0902, a post-training refresh aimed at coding and agentic work that debuted at #1 on the Code Arena: WebDev leaderboard at 1,691 Elo. Claude Fable 5, the model Alibaba spent July positioning itself directly behind, sits at 1,628 on that same board. The question from July has sharpened: is Qwen 3.8 still “second only to Fable 5,” or has the coding axis already flipped?

A note for builders — the fastest way to settle a claim like this is to run both models on your own prompts. OrcaRouter fronts 200+ models behind one OpenAI-compatible endpoint, so you can pit Qwen 3.8 Max against Fable 5 on the same task without wiring up two SDKs.

TL;DR verdict. Qwen3.8-Max-0902 is the strongest build of Alibaba’s 2.4T flagship yet, and the new part of its case is no longer self-reported: it debuted at #1 on the independent Code Arena: WebDev board at 1,691 Elo, above Claude Fable 5 (1,628, #8) on that board. Price is unchanged and still decisive — at $2 / $6 flat across 1M tokens, Qwen undercuts Fable 5’s $10 / $50 by roughly 5× on input and 8× on output. What has not closed is the general-purpose gap: Artificial Analysis scores Qwen3.8-Max at 56 on its Intelligence Index against Claude Fable 5’s 62, and on the general Arena board’s latest published week (Aug 24–30) Fable 5 holds #1 (1,508 Elo) while Qwen3.8-Max sits at #20 (1,479). Two verdicts, then: on coding, Qwen’s 0902 refresh now has a referee and a win; on broad, audited reasoning, Claude Fable 5 — and Anthropic’s newer Claude Fable 5.1, which leads the AA index at 66 — still hold the ground.

Key takeaways

Qwen3.8-Max reached GA on August 3, 2026 and its 0902 refresh followed on September 2 — no version bump, same $2 / $6 rate card, same 1M-token context.

The price gap is enormous: Qwen at $2 / $6 per 1M versus Fable 5 at $10 / $50. That is ~5× on input and ~8× on output, and Qwen's rate is flat across the entire 1M-token context with no long-prompt surcharge.

Both sides now have independent scores — and they point in different directions. Qwen3.8-Max measures 56 on the Artificial Analysis Intelligence Index (Claude Fable 5: 62), and its 0902 build tops Code Arena: WebDev at 1,691 Elo against Fable 5’s 1,628.

Qwen’s own benchmark table is still strong and still vendor-produced. GPQA Diamond 92.6, PaperBench 93.0, Terminal-Bench 2.1 86.6, OSWorld-Verified 86.1 — but the September 2 refresh is the first Qwen result an independent arena has ranked #1.

The architecture gap in the spec sheet is closed. Qwen3.8-2.4T-A95B is ~95B active parameters out of 2.4T total, across 512 experts — disclosed with the August 12 open-weights release. Fable 5’s architecture remains undisclosed.

Open weights are out, not promised. Qwen3.8-2.4T-A95B (BF16 and FP8) landed on Hugging Face and ModelScope on August 12 under a custom Qwen3.8-Max license, with Qwen3.8-27B following on August 14 under Apache 2.0. Fable 5 is closed and API-only, permanently.

Accuracy note: figures from Alibaba’s own benchmark table remain vendor-reported and unreplicated. Independent figures here are time-stamped and third-party: the Artificial Analysis Intelligence Index (Qwen3.8-Max 56, Claude Fable 5 62, Claude Fable 5.1 66), the Code Arena: WebDev Elo listing, and the general Arena weekly board — arena scores drift as votes accrue, and the 0902 build’s #1 is per Alibaba’s announcement of a point-in-time listing. Fable 5’s 95.0% SWE-bench Verified is Anthropic-reported; Qwen’s pricing is the GA rate card from Alibaba Model Studio.

The specs and price, side by side

Here is the hard data, each figure attributed to its source.

• Maker / status — Qwen3.8-Max: Alibaba; GA August 3, 2026; 0902 refresh September 2, 2026. Claude Fable 5: Anthropic; GA flagship

• Architecture — Qwen3.8-Max: 2.4T total / ~95B active, sparse MoE, 512 experts (disclosed Aug 12); Fable 5: undisclosed

• Context window — Qwen3.8-Max: 1M tokens (983,616 with thinking; max output 131,072); Fable 5: Long-context flagship

• AA Intelligence Index — Qwen3.8-Max: 56; Claude Fable 5: 62; Claude Fable 5.1 (Sep 1): 66, leads

• SWE-bench Verified — Qwen3.8-Max: Not reported (Alibaba reports FrontierSWE 73.5 instead); Fable 5: 95.0% (Anthropic / buildfastwithai)

• Independent + vendor-reported strengths — Qwen3.8-Max: Code Arena WebDev #1 at 1,691 (0902, independent); vendor table: GPQA Diamond 92.6, PaperBench 93.0, Terminal-Bench 2.1 86.6. Claude Fable 5: AA 62; SWE-bench Verified 95.0%

• Pricing (in / out) — Qwen3.8-Max: $2.00 / $6.00 per 1M, flat across 1M context; cached input $0.25; Fable 5: $10 / $50 per 1M

• Open weights — Qwen3.8-Max: released Aug 12, 2026 (custom license, not Apache); Qwen3.8-27B under Apache 2.0 Aug 14. Fable 5: no — closed / API-only

• Multimodality — Qwen3.8-Max: Text, image, video in → text out; Fable 5: Multimodal flagship

Two things frame the whole comparison. The first changed on August 3 and still holds; the second changed again on September 2.

First, the price column is now the loudest thing on the page. In July, Qwen's cost advantage was a discount coupon — 90% off credits, temporary, with no per-token rate to plan against. Now it is a rate card, and the gap is structural rather than promotional. At $6 output against Fable's $50, you can run roughly eight Qwen calls for the price of one Fable call. For any workload where you are paying for volume rather than for the single hardest answer, that ratio changes the architecture of what you build.

Second, the benchmark column has stopped being one-sided. For most of August the honest read was narrow: Alibaba reported FrontierSWE 73.5 and a table of self-run numbers, Anthropic reported a third-party-checked 95.0% SWE-bench Verified, and no independent evaluator had scored Qwen at all. The 0902 build ended that. Code Arena: WebDev is a blind pairwise human-vote arena — the project formerly known as WebDev Arena, run by no one at Alibaba — and Qwen3.8-Max-0902 entered it at #1 with 1,691 Elo, above Claude Opus 5 (1,688), Kimi K3 (1,674), and Claude Fable 5 (1,628, #8). Two caveats keep this honest: it is one board, measuring agentic front-end development rather than general capability, and “debuted at #1” is Alibaba’s announcement of a point-in-time listing. But the claim that Qwen has no referee at all is no longer true, and that changes the comparison more than any single number on it.

What the price gap actually buys you

Abstract multiples are less useful than a worked example, so here is one. Take a code-review agent that sends 150,000 tokens of repository context and generates 6,000 tokens of review per call.

Qwen3.8-Max: 0.15M × $2 = $0.30, plus 0.006M × $6 = $0.036. ≈ $0.34 per call.

Claude Fable 5: 0.15M × $10 = $1.50, plus 0.006M × $50 = $0.30. ≈ $1.80 per call.

• At 2,000 calls/day: Qwen ≈ $672/day; Fable ≈ $3,600/day. Over a month that is roughly $20,000 versus $108,000.

• With prompt caching on a stable repo context, Qwen's cached input at $0.25/1M drops the input side from $0.30 to under $0.04, taking the call to ≈ $0.08 — about 22× cheaper than Fable per call.

That is not a marginal saving; it is the difference between running a reviewer on every pull request and running one on the risky ones. And Qwen's flat-rate context makes the effect stronger as prompts grow: because Alibaba applies no long-prompt surcharge, pushing a 900,000-token context through costs the same per token as a 5,000-token one.

The honest counterweight is that price per token is not price per solved task. If Fable 5 resolves a problem in one pass where a cheaper model needs three attempts plus human correction, the cheaper model may cost more in aggregate — and on the hardest reasoning work, Fable's audited #1 ranking is the best available evidence that it does resolve more in one pass. The cost argument for Qwen is strongest on high-volume, tolerant workloads and weakest on low-volume, high-stakes ones.

The 0902 naming: capability jumped, the version number did not

The notable thing about September 2 is what Alibaba did not do: release Qwen 3.9. The improvement shipped as a dated checkpoint on the same model name — Qwen3.8-Max-0902 — and the API served it under the same name and the same price. This is the direction the industry is drifting: when a capability jump no longer guarantees a new version number, “which model should I use” becomes a question about builds and dates, not just family names.

It also means a benchmark score attached to “Qwen3.8-Max” has a shelf life: a score measured against the July or August build may not describe the model the API returns today. Check the build identifier on the endpoint you actually call — the 0902 number is the difference between a model that trailed Claude Fable 5 and one that leads it on Code Arena: WebDev.

Proof, and why it still decides this matchup

The most-cited hands-on evidence for Qwen3.8-Max comes from a preview-era review (thomas-wiegold.com) that put the model through four real builds. The results were genuinely impressive: Qwen one-shotted a Go poker simulation — only the third model ever to clear that prompt in a single pass, alongside Fable 5 and Grok 4.5 — and on a coffee-roaster website prompt it produced the best output the reviewer had ever seen from that test, adding an unrequested shopping cart, wholesale section, and Instagram integration out of sheer thoroughness.

The catch was speed: 30+ minutes on the website and 1 hour 20 minutes on the poker simulation, making it the slowest model that reviewer had used. That finding needs a caveat now that it did not need in July. It was measured on preview infrastructure, by one reviewer, on unusually long agentic runs. GA serving is a different system. For what it is worth, OrcaRouter's own 7-day telemetry currently shows a p50 time-to-first-token of 1.64 seconds for Qwen3.8-Max — a first-party measurement, though TTFT and end-to-end throughput on hour-long builds are different quantities. The honest position is that the preview slowness finding should be re-tested rather than repeated as current fact.

What survives the 0902 update is that the strongest evidence on each side still comes from different places. Qwen3.8-Max-0902’s #1 on Code Arena: WebDev is an independent, blind-vote result, but it measures one board, and Alibaba’s own benchmark table remains a vendor artifact — labs choose which benchmarks to report, and Alibaba’s weakest reported number (IFBench 82.8, instruction-following) still lines up with the preview complaint that the model over-delivers and ignores scope. Claude Fable 5’s evidence is broader and mostly third-party: 62 on the AA Intelligence Index, #1 on the general Arena board, 95.0% SWE-bench Verified — and Anthropic’s own newer Claude Fable 5.1 has led the AA index at 66 since September 1. On “which model will handle work I cannot afford to get wrong,” the general-purpose answer is unchanged. On “which model ships the best front-end at a tenth of the price,” the answer flipped on September 2.

Openness: the axis Qwen has already won

Fable 5 will never be self-hostable. Qwen3.8-Max already is: on August 12 the weights for Qwen3.8-2.4T-A95B — BF16 and an official FP8 variant — landed on Hugging Face and ModelScope, making it the first Max-class Qwen anyone can download. Two caveats deserve equal weight.

The first is legal. The release ships under a custom “Qwen3.8-Max” license, not Apache 2.0: providers running model-as-a-service or AI work-assistant businesses above $50M trailing revenue must negotiate a separate license from Qwen, and commercial products passing 100M monthly active users or $20M monthly revenue must display the model name. For internal use and startups the thresholds rarely bind — but this is not a free-for-all.

The second is physical. A 2.4T-total MoE runs to roughly 4.9TB in BF16 (about half in the official FP8 checkpoint) against roughly 141GB of memory per H200 — so self-hosting the flagship means a rack of accelerators before you serve a token. The disclosed ~95B active parameters at least let you estimate what that rack would buy. For nearly every team, the hosted API at $2 / $6 is the practical path.

Which is why Qwen3.8-27B, whose Apache-2.0 weights followed on August 14 and which reportedly runs in ~17GB of VRAM, remains the practically important self-host release. If your reason for preferring Qwen over Fable 5 is data residency or on-premise control rather than raw capability, the 27B is the model that actually delivers it — and on both sides the openness comparison is no longer theoretical.

FAQ

Is Qwen 3.8 better than Claude Fable 5?

It now depends on the axis, which is a genuine change from August. On coding, Qwen3.8-Max-0902 leads: #1 on Code Arena: WebDev at 1,691 Elo against Claude Fable 5’s 1,628, at $2 / $6 versus Fable 5’s $10 / $50. On broad, audited quality, Fable 5 still leads: 62 on the AA Intelligence Index against Qwen’s 56, 95.0% SWE-bench Verified, and #1 on the general Arena board. For work where a wrong answer is expensive, Fable 5 is the better-documented pick; for high-volume coding and agentic web work, the newest Qwen is both cheaper and arena-ranked ahead.

Is the "behind only Fable 5" claim true?

Alibaba’s positioning was made without independent numbers, and the 0902 build has since changed its shape. Qwen3.8-Max-0902 ranks ahead of Claude Fable 5 on Code Arena: WebDev (1,691 vs 1,628) but behind it on the AA Intelligence Index (56 vs 62) and on the general Arena board (Qwen #20 at Elo 1,479; Fable 5 #1 at 1,508). So “second only to Fable 5” has become “ahead of Fable 5 on the coding board this refresh targeted, behind it on the broad general-purpose boards.”

How much cheaper is Qwen 3.8 than Fable 5?

Substantially, and it is now a real rate rather than a discount. Qwen3.8-Max is $2 / $6 per 1M tokens against Fable 5's $10 / $50 — about 5× cheaper on input and 8× on output. Qwen's rate is also flat across the full 1M context, and cached input at $0.25/1M pushes repeat-context workloads cheaper still.

Did Qwen 3.8 Max leave preview?

Yes. It reached general availability on August 3, 2026 with a published rate card and an OpenAI- and DashScope-compatible endpoint. The Qoder credits campaign and 90%-off preview framing that circulated in July no longer describe how the model is sold.

Is Qwen 3.8 still the slowest model, as early reviews said?

That finding came from one reviewer on preview infrastructure running hour-long agentic builds, and it should be re-tested against GA serving rather than treated as current fact. OrcaRouter's 7-day telemetry shows a p50 time-to-first-token of 1.64 seconds, though that measures first-token latency rather than throughput on very long builds.

Can I self-host Qwen 3.8 instead of paying for Fable 5?

Yes for Qwen, with qualifications. Qwen3.8-2.4T-A95B has been downloadable since August 12 under a custom license (not Apache 2.0), and Qwen3.8-27B under Apache 2.0 since August 14. The practical barrier is hardware: a 2.4T-total MoE needs roughly 4.9TB in BF16 — a rack of accelerators — so most teams will still call the hosted API at $2 / $6 rather than serve the flagship. Fable 5 is closed permanently.

Which should I use today?

For general work where a wrong answer is expensive — hard reasoning, high-stakes production paths — Claude Fable 5 is still the better-documented choice, and Anthropic’s Claude Fable 5.1 extends that lead. For high-volume coding and agentic web development, Qwen3.8-Max-0902 now has the independent arena edge and an output price roughly 8× lower: bulk code review, front-end generation, long-document analysis. Many teams should run both and route by task.

Bottom line

Qwen 3.8 vs Claude Fable 5 is a different comparison than it was a month ago, and different again from a week ago. Qwen3.8-Max is a generally available model with a rate card that undercuts Fable 5 by 5× on input and 8× on output, flat across a million tokens — and since September 2 the same name serves the 0902 build, which an independent arena now ranks #1 for agentic web development, above Fable 5 on that board.

Neither model now owns the whole comparison. Claude Fable 5 — and Anthropic’s Claude Fable 5.1, which has led the AA Intelligence Index at 66 since September 1 — hold the broad, audited, general-purpose ground, and Alibaba’s own benchmark table remains a vendor artifact. But the reason this matchup was lopsided in July — that Qwen had no independent referee at all — expired on September 2: Qwen3.8-Max-0902 is the #1 model on Code Arena: WebDev. The sensible answer is still to split by stakes and task — Fable 5 for general reasoning where being wrong is expensive, Qwen3.8-Max for high-volume coding where the price gap and the arena edge both cut its way — and to test both on your own prompts.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily