
Qwen 3.8 vs GLM-5.2: The First Benchmark Where They Actually Overlap
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
These two models are the clearest expression of opposite bets in Chinese open-weight AI. Zhipu's GLM-5.2 is lean and proven: 753B total parameters with only 40B active, an audited Artificial Analysis Intelligence Index of 51, weights downloadable since June 2026, and a published price of $0.95 / $3.00. Alibaba's Qwen3.8-Max is the maximalist bet: ~2.4T parameters, an undisclosed active count, and — until recently — no numbers at all.
General availability on August 3, 2026 gave this matchup something no other article in this series has: a benchmark both models have a score on. Alibaba reports Terminal-Bench 2.1 at 86.6; Artificial Analysis has GLM-5.2 audited at 82.7 on that same test. For the first time, Qwen's claim can be set beside an independently measured competitor number on identical ground — with the important caveat that only one of the two was produced by a referee.
A note for builders — the fastest way to settle this is to run both on your own prompts. OrcaRouter fronts 200+ models behind one OpenAI-compatible endpoint, so you can pit Qwen 3.8 Max against GLM-5.2 on the same task without wiring up two SDKs.
TL;DR verdict. On the one shared benchmark, Qwen claims a 3.9-point lead on Terminal-Bench 2.1 (86.6 vs 82.7) — a real result if it replicates, though Qwen's figure is self-reported and GLM's is audited. Everywhere else GLM-5.2 holds the practical advantages: it is ~2× cheaper at $0.95 / $3.00 against Qwen's $2 / $6, its weights are downloadable today, its 40B active parameters make it genuinely deployable, and it carries an independent index score of 51 where Qwen has none. Qwen answers with a 1M-token flat-rate context, native multimodality GLM lacks, and a higher potential ceiling. Lean-and-proven still beats big-and-unverified for production — but the gap narrowed on August 3.
Key takeaways
• First direct benchmark overlap in the series: Terminal-Bench 2.1 — Qwen3.8-Max 86.6 (Alibaba-reported) against GLM-5.2 82.7 (Artificial Analysis, audited). Qwen leads by 3.9 points, but the two figures are not equally trustworthy.
• GLM-5.2 is roughly half the price: $0.95 / $3.00 per 1M (AA blended ~$0.90) versus Qwen3.8-Max's $2 / $6.
• Active parameters are the real architectural story: GLM-5.2 runs 40B active of 753B total. Alibaba has still not disclosed Qwen's active count, which makes its serving cost unmodellable.
• GLM's weights are downloadable today; Qwen's are promised within the week, with no license or repository yet.
• GLM-5.2 has an audited index of 51. Qwen3.8-Max remains unscored — though its predecessor Qwen3.7-Max scored 46, below GLM.
• Qwen's unique offers: a 1M-token window billed flat with no long-prompt surcharge, plus native text/image/video input and OmniDocBench 1.5 92.1. GLM-5.2 is text-focused.
Accuracy note: all Qwen3.8-Max quality figures are Alibaba-reported and unverified by third parties. GLM-5.2 figures come from Artificial Analysis. Comparing a self-reported score to an audited one on the same benchmark is informative but not conclusive — vendor harnesses and evaluator harnesses differ. Qwen pricing is the GA rate card and supersedes July's preview discount.
The specs and price, side by side
Here is the hard data, each figure attributed to its source.
• Maker / status — Qwen3.8-Max: Alibaba; GA (August 3, 2026); GLM-5.2: Zhipu; released June 2026
• Architecture — Qwen3.8-Max: ~2.4T total, sparse MoE; GLM-5.2: 753B total, sparse MoE (Artificial Analysis)
• Active parameters — Qwen3.8-Max: Still undisclosed by Alibaba; GLM-5.2: 40B active (Artificial Analysis)
• Context window — Qwen3.8-Max: 1M tokens, flat-rate (983,616 with thinking; max output 131,072); GLM-5.2: 1M tokens (Artificial Analysis)
• Terminal-Bench 2.1 — Qwen3.8-Max: 86.6 (Alibaba-reported); GLM-5.2: 82.7 (Artificial Analysis, audited)
• AA Intelligence Index — Qwen3.8-Max: Still unscored; GLM-5.2: 51 (Artificial Analysis)
• Other vendor-reported scores — Qwen3.8-Max: GPQA Diamond 92.6, OSWorld-Verified 86.1, FrontierSWE 73.5 (Alibaba-reported); GLM-5.2: standing is third-party measured
• Modalities — Qwen3.8-Max: Text, image, video in → text out; GLM-5.2: Text-focused
• Pricing (in / out) — Qwen3.8-Max: $2.00 / $6.00 per 1M, flat; cached input $0.25; GLM-5.2: $0.95 / $3.00 per 1M (AA blended ~$0.90)
• Open weights — Qwen3.8-Max: Promised within the week (no license yet); GLM-5.2: Yes — downloadable today
Two things frame this comparison, and the first is the most useful data point in the whole series.
First, the Terminal-Bench overlap is genuinely informative. Terminal-Bench 2.1 measures whether a model can operate a real shell to completion — run commands, read output, recover from errors. It is the closest available proxy for "can this function as a coding agent rather than a code suggester," and it is the benchmark both vendors chose to report. Qwen's 86.6 against GLM's audited 82.7 is a 3.9-point lead in Qwen's favor.
The caveat matters but should not be overstated. Vendor harnesses and evaluator harnesses differ in scaffolding, retry policy, and prompt construction, and those choices can move a Terminal-Bench score by more than four points. So this is not proof that Qwen is the better agentic model. What it does establish is that Qwen's self-reported claim lands in the same competitive band as an audited frontier open-weight model, rather than in implausible territory. That is a meaningfully stronger position than "unscored."
Second, the architecture comparison is where GLM quietly wins. GLM-5.2 activates 40B parameters out of 753B. That single number tells you its serving cost, its latency profile, and that a well-equipped team can actually run it. Alibaba has still not published Qwen's active count — which means that for all the 2.4T headline conveys, nobody outside Alibaba can estimate what Qwen3.8-Max costs to serve or how it would behave self-hosted. In a Mixture-of-Experts comparison, that is not a footnote; it is the most important missing number.

Lean and proven vs big and unverified
GLM-5.2's case is built on a coherent engineering position: do more with less, publish the numbers, ship the weights. A 40B-active model is cheap to serve, fast by construction, and deployable by a team with a serious but not absurd GPU budget. Its audited index of 51 and audited Terminal-Bench of 82.7 mean a buyer is not taking anything on trust. And at $0.95 / $3.00, it is roughly half Qwen's price.
Qwen3.8-Max's case is the opposite wager: build the largest model you can and let capability justify the cost. GA made that case much stronger than it was, because there are finally numbers attached — GPQA Diamond 92.6 on graduate science, OSWorld-Verified 86.1 on GUI operation, FrontierSWE 73.5 against a predecessor's 40.7, and the Terminal-Bench 86.6 discussed above. Those are frontier-shaped figures, and the near-doubling on FrontierSWE is the kind of generational jump that is hard to fake entirely.
But two gaps persist, and they are the same two that mattered in July. There is no independent score — Artificial Analysis and LMArena remain unscored on Qwen3.8-Max, so every quality claim is Alibaba's own. And there is no license, so the open-weight promise cannot yet be evaluated commercially.
The historical anchor cuts against Qwen here more than in most matchups: the predecessor Qwen3.7-Max scored 46 on the AA index, below GLM-5.2's 51. So Alibaba's implicit claim is not merely that Qwen3.8-Max is good, but that it leapt from below GLM to well above it in one generation. The FrontierSWE improvement lends that some credibility. An independent score would settle it.

Cost in practice: where 2× lands
Take an agentic coding pipeline: 180,000 tokens of repository context, 10,000 tokens of output per run.
• GLM-5.2: 0.18M × $0.95 = $0.171, plus 0.01M × $3.00 = $0.03. ≈ $0.20 per run.
• Qwen3.8-Max: 0.18M × $2 = $0.36, plus 0.01M × $6 = $0.06. ≈ $0.42 per run.
• At 3,000 runs/day: GLM ≈ $603/day; Qwen ≈ $1,260/day — roughly $20,000/month apart.
• With Qwen's cached input at $0.25/1M on a stable repo, its run drops to ≈ $0.11, which actually undercuts GLM's uncached cost.
That last line is the most practically useful thing in this comparison. Qwen's headline price is double GLM's, but its cache-hit rate of $0.25 per 1M is aggressive enough that a well-cached pipeline can flip the ranking. If your agent re-reads a mostly-unchanged context on every turn — which describes most coding agents — Qwen may be the cheaper option in practice despite the worse rate card. Teams that do not configure caching will pay double for nothing.
Both models bill thinking tokens as output, so reasoning-heavy configurations cost more than these estimates on either side.
Deployability: the axis GLM owns
Both are Chinese open-weight MoE models with 1M-token context claims, which makes deployability the sharpest practical divide.
GLM-5.2's weights have been downloadable since June, and at 40B active it is a realistic self-hosting target — quantizable, servable on a sane number of GPUs, and already exercised by the community.
Qwen3.8-Max's weights are promised within the week, but the flagship is not a realistic target for anyone: roughly 1.2TB at 4-bit against about 141GB per H200 means eight or more top-end accelerators, and the undisclosed active-parameter count makes the resulting throughput impossible to predict. Add the missing license and self-hosting the flagship is, for now, not a plan.
The concurrently announced Qwen3.8-27B is Alibaba's genuine answer on this axis. Also going open-weights and reportedly running in about 17GB of VRAM — a single high-end card, well below GLM-5.2's footprint — it competes directly on deployability. If your interest in either model is running it yourself, the Qwen 27B versus GLM-5.2 comparison is the one that will actually matter, and it is a much closer fight than 2.4T versus 753B.
FAQ
Is Qwen 3.8 better than GLM-5.2?
On the one benchmark they share, Qwen claims a lead — Terminal-Bench 2.1 86.6 against GLM's audited 82.7. But Qwen's figure is self-reported and GLM's was measured by Artificial Analysis, and harness differences can exceed a four-point gap. GLM-5.2 also holds an audited index of 51 where Qwen has no independent score at all. GLM is the safer choice; Qwen has the higher unverified ceiling.
How do they compare on Terminal-Bench 2.1?
Qwen3.8-Max reports 86.6; Artificial Analysis measured GLM-5.2 at 82.7. That is a 3.9-point nominal lead for Qwen and the first same-benchmark overlap between them. Read it as evidence Qwen is competitive in that band rather than proof it is ahead, since only GLM's number was independently produced.
Which is cheaper?
GLM-5.2 on the rate card — $0.95 / $3.00 per 1M (AA blended ~$0.90) versus Qwen's $2 / $6, roughly half. But Qwen's cached input at $0.25 per 1M is aggressive enough that a heavily cached pipeline can end up cheaper on Qwen than on GLM uncached.
How many active parameters does Qwen 3.8 Max use?
Alibaba has never published the figure, even at general availability. That is the number that determines serving cost and latency in a Mixture-of-Experts model, so its absence is a real gap. GLM-5.2, by contrast, is documented at 40B active out of 753B total.
Which can I actually self-host?
GLM-5.2 today — weights are out and 40B active makes it tractable. Qwen3.8-Max's weights are promised within the week but the flagship needs eight or more H200-class GPUs at ~1.2TB in 4-bit, with no license published yet. The concurrently announced Qwen3.8-27B, reportedly ~17GB of VRAM, is the deployable Qwen option and a closer match to GLM's niche.
Did Qwen 3.8 Max leave preview?
Yes, on August 3, 2026, with a published rate card and an OpenAI- and DashScope-compatible endpoint. July's 90%-off Qoder credits campaign no longer describes how it is sold.
What does Qwen offer that GLM-5.2 does not?
Native multimodality — image and video input, with a self-reported OmniDocBench 92.1 for document parsing — and a 1M-token context billed at a flat rate with no long-prompt surcharge. GLM-5.2 is text-focused, so for document, screenshot, or video pipelines Qwen is not competing with it so much as doing something else.
Bottom line
Qwen 3.8 vs GLM-5.2 is the matchup where general availability delivered the most genuinely new information. For the first time Qwen has a score on a benchmark an independent evaluator has also run, and it lands above GLM: Terminal-Bench 2.1 86.6 against 82.7. That moves Qwen from "unscored" to "plausibly competitive on measured ground," which is real progress even accounting for the self-reported caveat.
But GLM-5.2 keeps the practical crown, and it does so on unglamorous strengths: half the price, an audited index of 51, 40B active parameters that make its economics predictable and its deployment achievable, and weights you can download right now. Qwen's answers — a flat-rate million-token context, native multimodality, and a bigger ceiling — are meaningful but conditional, and its two July gaps are unchanged: no independent score and no license. Watch three things: the first independent Intelligence Index score for Qwen3.8-Max, whether an evaluator reproduces that 86.6 Terminal-Bench figure, and the Qwen3.8-27B release, which is where these two philosophies will really collide. Until then, GLM-5.2 is what you ship and Qwen3.8-Max is what you evaluate — with caching switched on.

Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
