
Qwen 3.8 vs GPT-5.6: Now That Both Have Price Tags, Which One Wins?
- qwenNEWQwen: Qwen3.8 Max2026-08-03$2.00 / $6.00 per 1M tokens · 56 tok/s
- deepseekNEWDeepSeek: DeepSeek V4 Flash 07312026-07-3150Intelligence69Coding
- qwenNEWQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 198 tok/s
- orcaNEWOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicNEWAnthropic: Claude Opus 52026-07-2461Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2150Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1651Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1557Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0951Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0955Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0959Intelligence77Coding
- grokxAI: Grok 4.52026-07-0854Intelligence72Coding
- tencentTencent: Hy32026-07-0641Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3053Intelligence72Coding
- klingKling: Kling 3.0 Turbo2026-06-1757Intelligence52Coding57Math
- z-aiZ.ai: GLM 5.22026-06-1651Intelligence69Coding60Math
- kimiMoonshotAI: Kimi K2.7 Code2026-06-1242Intelligence61Coding61Math
When Alibaba previewed Qwen3.8-Max in Shanghai on July 19, 2026, the community reaction was immediate: this 2.4-trillion-parameter model was "supposedly ahead of GPT-5.6 and only slightly behind Fable 5." OpenAI's GPT-5.6 in its flagship Sol configuration sits at #2 on the independent Artificial Analysis Intelligence Index, one point below Fable 5, so that was a large claim to make with no published numbers behind it.
Two things have changed since. Qwen3.8-Max went generally available on August 3, 2026 with a real rate card and a real benchmark table. And on July 31, OpenAI cut prices across the GPT-5.6 family — Terra to $2 / $12, Luna to $0.20 / $1.20, with Sol unchanged at the premium tier. So this matchup is no longer a claim against a leaderboard; it is two priced, documented models that can actually be compared.
A note for builders — the fastest way to settle a claim like this is to run both on your own prompts. OrcaRouter fronts 200+ models behind one OpenAI-compatible endpoint, so you can pit Qwen 3.8 Max against GPT-5.6 on the same task without wiring up two SDKs.
TL;DR verdict. Qwen3.8-Max at GA is priced aggressively — $2 / $6 per 1M, flat across 1M tokens — which puts it at the same input price as GPT-5.6 Terra but half Terra's output cost, and far below Sol. Its self-reported numbers are strong (GPQA Diamond 92.6, Terminal-Bench 2.1 86.6). But GPT-5.6 Sol still wins the two things that decide production use: it is the audited #2 on the Artificial Analysis Intelligence Index (~59) and it is fast — up to 750 tokens/sec on Cerebras hardware. Qwen3.8-Max remains unscored by every independent evaluator. So Qwen may well match GPT-5.6 on one-shot quality and it clearly beats Terra on output price; GPT-5.6 wins on refereed proof and on throughput, which is often the whole ballgame.
Key takeaways
• Qwen3.8-Max is GA since August 3, 2026 with published pricing of $2 / $6 per 1M tokens, flat across the full 1M-token context.
• Against GPT-5.6 Terra ($2 / $12 after OpenAI's July 31 cut), Qwen matches on input and halves the output rate. Against Sol, Qwen is dramatically cheaper.
• GPT-5.6 Sol is audited #2 on the AA Intelligence Index (~59), and Artificial Analysis notes it leads on the hardest STEM problems. Qwen3.8-Max is still unscored by AA and LMArena.
• Speed is the sharpest contrast. GPT-5.6 runs up to 750 tokens/sec on Cerebras. Qwen was the slowest model in preview hands-on testing — a finding that predates GA serving and needs re-testing.
• Qwen's vendor table is credible but unrefereed: GPQA Diamond 92.6, PaperBench 93.0, Terminal-Bench 2.1 86.6, OSWorld-Verified 86.1, FrontierSWE 73.5 (up from a predecessor's 40.7).
• Openness splits them permanently: Qwen promises open weights within the week plus a ~17GB-VRAM Qwen3.8-27B; GPT-5.6 is closed and API-only.
Accuracy note: all Qwen3.8-Max quality figures are Alibaba-reported and unverified by third parties. GPT-5.6 figures come from Artificial Analysis and may differ from OpenAI's own reporting. GPT-5.6 family pricing reflects OpenAI's July 31, 2026 reduction. Qwen pricing is the GA rate card and supersedes July's preview discount.
The specs and price, side by side
Here is the hard data, each figure attributed to its source.
• Maker / status — Qwen3.8-Max: Alibaba; GA (August 3, 2026); GPT-5.6 Sol: OpenAI; closed flagship (variants Sol / Terra / Luna)
• Architecture — Qwen3.8-Max: ~2.4T total, sparse MoE (active params still undisclosed); GPT-5.6 Sol: Closed; architecture not disclosed
• Context window — Qwen3.8-Max: 1M tokens (983,616 with thinking; max output 131,072); GPT-5.6 Sol: Long-context flagship
• AA Intelligence Index — Qwen3.8-Max: Still unscored; GPT-5.6 Sol: ~59 — #2 (Artificial Analysis)
• Audited strengths — Qwen3.8-Max: None third-party; GPT-5.6 Sol: Leads hardest STEM problems (Artificial Analysis)
• Vendor-reported strengths — Qwen3.8-Max: GPQA Diamond 92.6, Terminal-Bench 2.1 86.6, OSWorld-Verified 86.1 (Alibaba-reported); GPT-5.6 Sol: n/a
• Speed — Qwen3.8-Max: p50 TTFT 1.64s (OrcaRouter 7-day telemetry); slowest model in preview hands-on tests; GPT-5.6 Sol: Up to 750 tok/s on Cerebras (Artificial Analysis)
• Pricing (in / out) — Qwen3.8-Max: $2.00 / $6.00 per 1M, flat; cached input $0.25; GPT-5.6: Terra $2 / $12, Luna $0.20 / $1.20, Sol premium (~1/3 of Fable's cost per AA)
• Open weights — Qwen3.8-Max: Promised within the week (no license yet); GPT-5.6: No — closed / API-only
Two things frame this comparison, and both are new since July.
First, the price story got genuinely interesting rather than just cheap. In July, Qwen's advantage was an unbudgetable discount. Now the comparison is concrete and it splits by tier. Against Terra at $2 / $12, Qwen matches the input rate exactly and halves the output rate — a real advantage on reasoning-heavy work where output dominates spend. Against Luna at $0.20 / $1.20, Qwen is 10× more expensive, and Luna is the better choice for high-volume simple tasks. Against Sol, Qwen is far cheaper. So "which is cheaper" now has three different answers depending on which GPT-5.6 you mean — and the honest framing is that OpenAI's July 31 cut narrowed the gap considerably.
Second, the speed row is still where these two diverge hardest, and it still favors OpenAI. Artificial Analysis clocks GPT-5.6 Sol at up to 750 tokens/sec on Cerebras silicon. Nothing in Alibaba's GA materials suggests Qwen3.8-Max is engineered for that kind of throughput, and the preview-era reviewer who called it the slowest model he had used was measuring exactly the long agentic runs where throughput compounds. If your workload is interactive, agentic, or high-volume, that gap can decide the matchup before quality enters the conversation.

Proof: what the GA benchmark table does and does not settle
Alibaba's decision to publish numbers at GA is a real improvement over the preview, where the community line "ahead of GPT-5.6" rested on nothing published at all. The table is strong: GPQA Diamond 92.6 on graduate-level science, PaperBench 93.0 on reproducing research results, Terminal-Bench 2.1 86.6 on operating a real shell, OSWorld-Verified 86.1 on driving a desktop GUI. The generational coding jump is the standout — FrontierSWE 73.5 against a predecessor's 40.7, and DeepSWE 1.1 56.6 against 21.6.
What the table does not do is settle the GPT-5.6 comparison, for two reasons.
The first is provenance. Every figure was produced by Alibaba on its own harness. Vendor tables are chosen, not sampled — a lab publishes the benchmarks where it looks good and omits the rest. OpenAI's #2 Intelligence Index ranking, by contrast, was assigned by an evaluator with no stake in the outcome. Notably, Artificial Analysis specifically credits GPT-5.6 Sol with leading the hardest STEM problems, which is precisely the territory Qwen's GPQA Diamond 92.6 is meant to claim. One of those claims has been checked.
The second is that Alibaba's own weakest number is telling. IFBench 82.8 — instruction following under constraint — is the lowest score in its table, and it lines up exactly with the most consistent preview-era criticism: the model over-delivers, exceeds scope, and adds things nobody asked for. That is charming in a demo and expensive when output bills at $6 per million tokens and you need an exact format. For agentic pipelines where downstream steps parse the output, instruction-following discipline matters more than a GPQA point.
For anchoring: the predecessor Qwen3.7-Max scored 46 on the AA Intelligence Index against GPT-5.6 Sol's ~59. Closing thirteen points in one generation would be an extraordinary leap. Possible — Alibaba's own FrontierSWE near-doubling suggests real progress — but until a referee posts a number, "ahead of GPT-5.6" stays an unproven claim rather than a result.

Cost in practice: a worked example across the tiers
Take a reasoning agent sending 100,000 tokens of context and generating 10,000 tokens of output per call — a shape typical of document analysis or multi-step planning.
• Qwen3.8-Max: 0.1M × $2 = $0.20, plus 0.01M × $6 = $0.06. ≈ $0.26 per call.
• GPT-5.6 Terra: 0.1M × $2 = $0.20, plus 0.01M × $12 = $0.12. ≈ $0.32 per call.
• GPT-5.6 Luna: 0.1M × $0.20 = $0.02, plus 0.01M × $1.20 = $0.012. ≈ $0.03 per call.
• At 5,000 calls/day: Qwen ≈ $1,300, Terra ≈ $1,600, Luna ≈ $160.
Two lessons fall out. Against Terra the Qwen saving is real but modest — roughly 19% on this shape — and nowhere near the order-of-magnitude advantage Qwen enjoys against premium models like Fable 5 or Sol. Anyone who assumed OpenAI was structurally expensive should update: the July 31 cut did most of the work of neutralizing Qwen's pricing story in the mid tier.
Where Qwen pulls ahead sharply is long context and caching. Because the $2/$6 rate is flat across the entire 1M-token window, a 900,000-token prompt costs the same per token as a short one, while many competitors surcharge long inputs. And cached input at $0.25 per 1M cuts the input side by 8× on repeat-context workloads: the call above drops from $0.26 to about $0.09 once the context is cached. If your agent re-reads a mostly-stable context every turn, that is where the durable advantage lives — not in the headline rate.
Openness and the 27B sibling
GPT-5.6 will never be self-hostable. Qwen3.8-Max might be — Alibaba now says weights arrive within the week — but the flagship is not a practical target: roughly 1.2TB at 4-bit against about 141GB per H200 means eight or more top-end accelerators, and with the active-parameter count still undisclosed you cannot even model the throughput that buys. There is also still no license text, which means commercial usability is formally undetermined.
The concurrently announced Qwen3.8-27B is the more consequential release for anyone whose interest in Qwen is about control rather than raw capability. Also going open-weights, it reportedly runs in about 17GB of VRAM — a single high-end consumer card — which makes it a genuine on-premise option in a way the 2.4T flagship never will be. If you are weighing Qwen against GPT-5.6 because you need data residency, 27B is the model to evaluate.
FAQ
Is Qwen 3.8 better than GPT-5.6?
Not on the evidence that exists. Alibaba's GA table is strong (GPQA Diamond 92.6, Terminal-Bench 2.1 86.6) but entirely self-reported, and no independent evaluator has scored the model. GPT-5.6 Sol holds an audited #2 Intelligence Index (~59), leads the hardest STEM problems per Artificial Analysis, and runs up to 750 tokens/sec. GPT-5.6 wins on proof and speed.
Is the "Qwen 3.8 is ahead of GPT-5.6" claim true?
It began as community sentiment and is still unverified. GA brought Alibaba's own benchmark table but no third-party score — Artificial Analysis and LMArena remain unscored. The predecessor Qwen3.7-Max scored 46 against Sol's ~59, so the claim needs a very large one-generation jump to hold literally.
Which is cheaper, Qwen 3.8 or GPT-5.6?
It depends on the tier, and the answer changed on July 31 when OpenAI cut prices. Qwen3.8-Max is $2 / $6 per 1M. GPT-5.6 Terra is $2 / $12 — same input, double the output, so Qwen wins there. Luna is $0.20 / $1.20, making it about 10× cheaper than Qwen. Sol is premium, so Qwen is far cheaper than Sol.
Which is faster?
GPT-5.6, decisively on throughput. Artificial Analysis reports up to 750 tokens/sec on Cerebras hardware. Qwen3.8-Max shows a p50 time-to-first-token of 1.64 seconds in OrcaRouter's 7-day telemetry, but preview-era reviewers found it very slow on long agentic builds — that finding predates GA serving and deserves re-testing.
Did Qwen 3.8 Max leave preview?
Yes, on August 3, 2026. It now has a published rate card and an OpenAI- and DashScope-compatible endpoint, so the July framing of a Qoder credits campaign with a 90% discount no longer describes how it is sold.
Can I self-host Qwen 3.8 to avoid OpenAI pricing?
Not the flagship, realistically. Weights are promised within the week but no license has been published, and at ~1.2TB in 4-bit you would need eight or more H200-class GPUs. The concurrently announced Qwen3.8-27B, reportedly ~17GB of VRAM, is the practical option.
Which should I use today?
GPT-5.6 Sol for anything latency-sensitive, interactive, or reasoning-critical where audited quality matters. Luna for high-volume simple tasks, where it is cheapest by a wide margin. Qwen3.8-Max where long context and prompt caching dominate your bill — its flat 1M-token rate and $0.25 cached input are the strongest parts of its offer.
Bottom line
Qwen 3.8 vs GPT-5.6 is a fairer fight than it was in July, and slightly less one-sided on price than Qwen's fans expected. Qwen3.8-Max arrived at GA with real pricing, a credible self-reported benchmark table, and a flat 1M-token rate plus cheap caching that make it genuinely attractive for long-context work. Those are structural advantages, not promotional ones.
But OpenAI's July 31 price cut blunted the cost argument in the mid tier — Terra now matches Qwen on input — and GPT-5.6 still owns the two decisive properties: an audited #2 ranking from an evaluator with no stake in the result, and throughput up to 750 tokens/sec that Qwen has shown no evidence of approaching. Watch two events: the first independent Intelligence Index score for Qwen3.8-Max, and the weights landing with a license. Until then, route by shape — GPT-5.6 when speed and proof matter, Qwen when context length and cache hits dominate the bill — and verify on your own prompts.

Compared in this article4
Detected from this article · Benchmarks: Artificial Analysis · updated daily
