Hero title card for the comparison of GLM-5.3 and GLM-5.2, subtitled 'Same brain, doubled benchmarks', with two model cards sharing a single connected 743B MoE base block and a curved upgrade arrow labelled 'post-training only' pointing from GLM-5.2 to GLM-5.3.
Guides & Insights

GLM-5.3 vs GLM-5.2: Same Brain, Doubled Benchmarks

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Zhipu AI's own framing for the upgrade from GLM-5.2 to G​LM-5.3 is the honest one: the textbook did not change, only the student's training did. G​LM-5.3, released August 14, 2026, is built on the exact same 743-billion-parameter MoE base as GLM-5.2, which crowned the open-weights field on the Artificial Analysis Intelligence Index in June. The architecture is unchanged, the footprint is unchanged, the $1.40 / $4.40 per-million pricing is unchanged — and yet Z.ai reports that the retrained model doubles or better its predecessor on several of the coding and security benchmarks both versions publish. Whether that makes G​LM-5.3 a must-upgrade or a wait-and-see depends on how much you trust a model that improved this much without changing its own architecture, and on whether its newly emerged cyber capability is a feature you want gated behind a two-week safety window.

The same brain, retrained

Everything that made GLM-5.2 a practical server carries over: the 743B MoE with roughly 40B active parameters, the IndexShare sparse attention that cut FLOPs at 1M context, the MIT license, the 1M context window, and the lean token throughput that measured ~145–168 tokens/sec in independent observation. G​LM-5.3 differs in one dimension only — post-training. Z.ai scaled the reinforcement-learning stage with the IndexShare, SAO, and Slime frameworks, added dozens of times more long-horizon task environments, and extended them from code debugging into security vulnerability discovery and systems engineering. The result is a model that behaves differently on the same hardware, which is precisely why the upgrade question is not "do I need new GPUs" but "do I need new behavior."

The benchmark delta, in both directions

Read as a before-and-after on Z.ai's published figures, the jump is the biggest in open-weights history. Terminal-Bench 3.0 — the harder revision, where both were scored — goes from 4.6 to 28.3, a sixfold jump. DeepSWE v1.1 goes from 46.2 to 66.9, which is the difference between "good open model" and "competitive with the closed frontrunners." The Agents' Last Exam CLI subset moves 23.8 to 28.5. CyberGym vulnerability discovery rises 77.2 to 84.5. ExploitBench more than doubles, 24.4 to 54.4, and ExploitGym throughput goes from 29 and 39 completed tasks (two and six hours) to 105 and 130. Z.ai summarizes it as roughly a 50% coding improvement, and the shape of the gains — concentrated in long-horizon coding, terminal automation, and security — is consistent with a post-training-only model: the raw knowledge is the same, but the ability to actually do multi-step work is what got stronger.

• Terminal-Bench 3.0 — GLM-5.2 4.6 vs G​LM-5.3 28.3

• DeepSWE v1.1 — GLM-5.2 46.2 vs G​LM-5.3 66.9

• Agents' Last Exam (CLI) — GLM-5.2 23.8 vs G​LM-5.3 28.5

• CyberGym vulnerability discovery — GLM-5.2 77.2 vs G​LM-5.3 84.5

• ExploitBench — GLM-5.2 24.4 vs G​LM-5.3 54.4

The OrcaRouter GLM-5.2 model page showing Z.ai's open-weights flagship with its per-million-token pricing, context window and model ID.Scoreboard contrasting GLM-5.2 (Terminal-Bench 3.0 4.6, DeepSWE v1.1 46.2, Agents' Last Exam CLI 23.8, CyberGym 77.2, ExploitBench 24.4, Price $1.40/$4.40 per M) with GLM-5.3 (Terminal-Bench 3.0 28.3, DeepSWE v1.1 66.9, Agents' Last Exam CLI 28.5, CyberGym 84.5, ExploitBench 54.4, Price $1.40/$4.40 per M).

The capability nobody scheduled

The security gains deserve their own paragraph because Z.ai says they were not the plan. The post-training team added vulnerability-discovery environments expecting modest improvement; the model instead began chaining complete exploits, an emergent behavior that forced Z.ai to hold the weights back roughly two weeks for safety hardening. In real-world testing with security teams, G​LM-5.3 contributed to finding 2,436 vulnerabilities across 269 projects, 1,097 of them medium- or high-risk, including flaws that had sat undetected for decades in widely deployed software. GLM-5.2 had security capability in the ordinary sense — it could read code for bugs. G​LM-5.3 has it in a different sense: it is the thing the safety window is about. That is a change in kind, not degree, and it is the single strongest reason to treat the two models as different products despite the shared base.

Infographic titled 'The upgrade delta' showing a two-column before-and-after comparison of GLM-5.2 vs GLM-5.3 across five benchmarks with upward arrows: Terminal-Bench 3.0 4.6 to 28.3, DeepSWE v1.1 46.2 to 66.9, Agents' Last Exam CLI 23.8 to 28.5, CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4.

The consistency caveat

The counterweight to every number above is that early third-party testing describes G​LM-5.3 as inconsistent in a way GLM-5.2 was not. Independent reviewers report the same model performing like a barely-passing outsourced engineer on some tasks and delivering professional-grade results on others, with the difference tracking how clearly the requirements were specified. That is exactly the failure profile you would predict from a model whose gains came entirely from environment-heavy post-training: it generalizes from the task shapes it was trained on, and poorly-specified prompts fall outside them. None of the independent results are as clean as Z.ai's published tables, and none of Z.ai's numbers are yet independently reproduced. For production, that argues for explicit evaluation on your own workload before the migration — the same caveat GLM-5.2 warranted, now with a wider spread between best and worst case.

Price, access, and the wait question

Pricing is identical, which removes the usual upgrade friction: both models are $1.40 / $4.40 per million tokens, with G​LM-5.3 requiring thinking enabled. Access is the real difference. GLM-5.2 is downloadable today under MIT, self-hostable on the sub-terabyte FP8 footprint, and callable through OrcaRouter at list price with no markup — the model you can route to right now, stable and independently scored. G​LM-5.3 is usable now through Z.ai's ZCode / AutoClaw and G​LM Coding Plan, but the public API and the weights are both gated for the safety window, and its benchmark figures are all vendor-reported pending third-party re-runs. So the honest answer to "should I wait" has two parts: for production traffic that must not regress, GLM-5.2 remains the safe open-weights bet today, and the moment G​LM-5.3's API opens and independent runs check out, the move is a route change, not a rewrite — you keep the same one-key setup and swap the lane. For exploratory and security-eval work, G​LM-5.3 is worth trying now through Z.ai's tools, with the understanding that its published lead is unverified and its behavior is more variable than the model it replaces.

The honest verdict

G​LM-5.3 is the better model on Z.ai's own numbers — dramatically better on long-horizon coding and categorically different on security — and it is free to your workflow because the base, footprint, and price are unchanged. But "better on the vendor's evals" and "better in your pipeline" are separated by the two things GLM-5.2 already has and G​LM-5.3 does not yet: independent reproduction and shipping weights. If you already run GLM-5.2, the upgrade path is the cheapest in open-weights history — same hardware, same price, same API surface, and a two-week wait for the weights. The decision is less about whether G​LM-5.3 is better than GLM-5.2 and more about whether you are willing to bet production on a model whose improvements, and whose newly emerged security capability, have not yet been confirmed by anyone outside Z.ai.

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

Contact us

Join our community

DiscordEmailXGitHubYouTube