
GLM-5.3 vs GLM-5.2: Same Brain, Doubled Benchmarks
- obsidianNEWQwen3.8 27B Uncensored (Aggressive)2026-08-1552Intelligence68Coding
- qwenNEWQwen: Qwen3.8 27B (free)2026-08-1359 tok/s
- deepseekNEWDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokNEWSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaNEWMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 231 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
- grokxAI: Grok 4.52026-07-0856Intelligence72Coding
- tencentTencent: Hy32026-07-0642Intelligence59Coding
Zhipu AI's own framing for the upgrade from GLM-5.2 to GLM-5.3 is the honest one: the textbook did not change, only the student's training did. GLM-5.3, released August 14, 2026, is built on the exact same 743-billion-parameter MoE base as GLM-5.2, which crowned the open-weights field on the Artificial Analysis Intelligence Index in June. The architecture is unchanged, the footprint is unchanged, the $1.40 / $4.40 per-million pricing is unchanged — and yet Z.ai reports that the retrained model doubles or better its predecessor on several of the coding and security benchmarks both versions publish. Whether that makes GLM-5.3 a must-upgrade or a wait-and-see depends on how much you trust a model that improved this much without changing its own architecture, and on whether its newly emerged cyber capability is a feature you want gated behind a two-week safety window.
The same brain, retrained
Everything that made GLM-5.2 a practical server carries over: the 743B MoE with roughly 40B active parameters, the IndexShare sparse attention that cut FLOPs at 1M context, the MIT license, the 1M context window, and the lean token throughput that measured ~145–168 tokens/sec in independent observation. GLM-5.3 differs in one dimension only — post-training. Z.ai scaled the reinforcement-learning stage with the IndexShare, SAO, and Slime frameworks, added dozens of times more long-horizon task environments, and extended them from code debugging into security vulnerability discovery and systems engineering. The result is a model that behaves differently on the same hardware, which is precisely why the upgrade question is not "do I need new GPUs" but "do I need new behavior."
The benchmark delta, in both directions
Read as a before-and-after on Z.ai's published figures, the jump is the biggest in open-weights history. Terminal-Bench 3.0 — the harder revision, where both were scored — goes from 4.6 to 28.3, a sixfold jump. DeepSWE v1.1 goes from 46.2 to 66.9, which is the difference between "good open model" and "competitive with the closed frontrunners." The Agents' Last Exam CLI subset moves 23.8 to 28.5. CyberGym vulnerability discovery rises 77.2 to 84.5. ExploitBench more than doubles, 24.4 to 54.4, and ExploitGym throughput goes from 29 and 39 completed tasks (two and six hours) to 105 and 130. Z.ai summarizes it as roughly a 50% coding improvement, and the shape of the gains — concentrated in long-horizon coding, terminal automation, and security — is consistent with a post-training-only model: the raw knowledge is the same, but the ability to actually do multi-step work is what got stronger.
• Terminal-Bench 3.0 — GLM-5.2 4.6 vs GLM-5.3 28.3
• DeepSWE v1.1 — GLM-5.2 46.2 vs GLM-5.3 66.9
• Agents' Last Exam (CLI) — GLM-5.2 23.8 vs GLM-5.3 28.5
• CyberGym vulnerability discovery — GLM-5.2 77.2 vs GLM-5.3 84.5
• ExploitBench — GLM-5.2 24.4 vs GLM-5.3 54.4


The capability nobody scheduled
The security gains deserve their own paragraph because Z.ai says they were not the plan. The post-training team added vulnerability-discovery environments expecting modest improvement; the model instead began chaining complete exploits, an emergent behavior that forced Z.ai to hold the weights back roughly two weeks for safety hardening. In real-world testing with security teams, GLM-5.3 contributed to finding 2,436 vulnerabilities across 269 projects, 1,097 of them medium- or high-risk, including flaws that had sat undetected for decades in widely deployed software. GLM-5.2 had security capability in the ordinary sense — it could read code for bugs. GLM-5.3 has it in a different sense: it is the thing the safety window is about. That is a change in kind, not degree, and it is the single strongest reason to treat the two models as different products despite the shared base.

The consistency caveat
The counterweight to every number above is that early third-party testing describes GLM-5.3 as inconsistent in a way GLM-5.2 was not. Independent reviewers report the same model performing like a barely-passing outsourced engineer on some tasks and delivering professional-grade results on others, with the difference tracking how clearly the requirements were specified. That is exactly the failure profile you would predict from a model whose gains came entirely from environment-heavy post-training: it generalizes from the task shapes it was trained on, and poorly-specified prompts fall outside them. None of the independent results are as clean as Z.ai's published tables, and none of Z.ai's numbers are yet independently reproduced. For production, that argues for explicit evaluation on your own workload before the migration — the same caveat GLM-5.2 warranted, now with a wider spread between best and worst case.
Price, access, and the wait question
Pricing is identical, which removes the usual upgrade friction: both models are $1.40 / $4.40 per million tokens, with GLM-5.3 requiring thinking enabled. Access is the real difference. GLM-5.2 is downloadable today under MIT, self-hostable on the sub-terabyte FP8 footprint, and callable through OrcaRouter at list price with no markup — the model you can route to right now, stable and independently scored. GLM-5.3 is usable now through Z.ai's ZCode / AutoClaw and GLM Coding Plan, but the public API and the weights are both gated for the safety window, and its benchmark figures are all vendor-reported pending third-party re-runs. So the honest answer to "should I wait" has two parts: for production traffic that must not regress, GLM-5.2 remains the safe open-weights bet today, and the moment GLM-5.3's API opens and independent runs check out, the move is a route change, not a rewrite — you keep the same one-key setup and swap the lane. For exploratory and security-eval work, GLM-5.3 is worth trying now through Z.ai's tools, with the understanding that its published lead is unverified and its behavior is more variable than the model it replaces.
The honest verdict
GLM-5.3 is the better model on Z.ai's own numbers — dramatically better on long-horizon coding and categorically different on security — and it is free to your workflow because the base, footprint, and price are unchanged. But "better on the vendor's evals" and "better in your pipeline" are separated by the two things GLM-5.2 already has and GLM-5.3 does not yet: independent reproduction and shipping weights. If you already run GLM-5.2, the upgrade path is the cheapest in open-weights history — same hardware, same price, same API surface, and a two-week wait for the weights. The decision is less about whether GLM-5.3 is better than GLM-5.2 and more about whether you are willing to bet production on a model whose improvements, and whose newly emerged security capability, have not yet been confirmed by anyone outside Z.ai.
