
Nex-N2.5 Pro vs GLM-5.2: The Same 82.7 on Terminal-Bench, Two Very Different Provenances
- openaiNEWOpenAI: GPT-6 Astra2026-09-0455Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0247Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0247Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0157Intelligence82Coding
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2646Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1849Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1541Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1242Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1251Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0547Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0347Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3141Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2454Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2140Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2128Intelligence49Coding
Two open-weights agentic models, one suspicious coincidence: Nex-N2.5 Pro and GLM-5.2 both surface at 82.7 on Terminal-Bench 2.1. Nex-AGI prints 82.7 for Nex-N2.5 Pro on its own model card, measured on the alliance's own NexCUA harness, which the public has not seen. Z.ai's best-reported figure for GLM-5.2 on the same suite is also 82.7 — and when an independent harness re-ran the model, it landed at 77.9, about five points under the vendor's own number. A perfect tie between a figure nobody can check and a figure that already shrank once it was checked is not a tie at all. It is the cleanest possible demonstration of why the difference between "open weights" and "weights that will exist eventually" decides this matchup before a single benchmark row is read.
Both models sit in the same commercial neighborhood and could not be more different in kind. Nex-N2.5 Pro, announced September 8, 2026, is a multimodal computer-use model — 397 billion total parameters with roughly 17 billion active, built from Qwen3.5 lineage, architected to read a screen and drive a browser or desktop session with a visual feedback loop. GLM-5.2 is Z.ai's (Zhipu AI's) MIT-licensed flagship for long-horizon reasoning and coding — roughly 743 billion total parameters with about 40 billion active, text-only, GA since June 16, 2026, and the model that has anchored the open-weights tier's independent scores for months. One looks at the world through pixels and acts on it. The other thinks in text and ships code. The license on both says you can build a business on them; the difference is that on one of them, that license is currently a promise.
The same number, two different provenances
Start with the tie, because it teaches the whole matchup. Nex-AGI's card lists Terminal-Bench 2.1 at 82.7 for Nex-N2.5 Pro, alongside OSWorld-2 at 56.4, SWE-Bench Pro at 61.2 and BrowseComp at 89.7 — every one of them produced through the NexCUA harness and none independently reproduced in the first days of the model's life, for the simple reason that the weights are not downloadable. Z.ai's figure of 82.7 for GLM-5.2 on Terminal-Bench 2.1 is the vendor's own best-reported number, but it sits on a model anyone can pull from Hugging Face and re-run this afternoon; the independent harness measurement of 77.9 already exists, and it is roughly five points below Z.ai's claim. When a vendor number shrinks under independent testing, that is not a scandal — it is the normal tax on self-measurement. When a competing vendor number cannot even be tested, the absence of the tax is the problem, not a sign the number is pristine.

Read the surrounding rows the same way. GLM-5.2 holds the highest independent index reading the open tier has produced — the low 50s on Artificial Analysis' current Intelligence Index — which is why it has been the reference open model for months despite not being the flashiest new release. Nex-N2.5 Pro has no independent index entry at all. On the rows where both vendors claim numbers, the verifiable one is the incumbent's; on the rows where only Nex-AGI has published, the numbers are unreproduced. That is not a verdict on which model is smarter. It is a verdict on which model you are allowed to verify.
The spec sheets agree more than you would expect
Strip the benchmarks away and the two models line up closely on the fundamentals:
• Released — Nex-N2.5 Pro: September 8, 2026 (hosted endpoints live, weights pending). GLM-5.2: GA June 16, 2026, weights live.
• License — Nex-N2.5 Pro: Apache-2.0, weights coming soon. GLM-5.2: MIT, downloadable now.
• Scale — Nex-N2.5 Pro: 397B total / ~17B active. GLM-5.2: ~743B total / ~40B active.
• Modality — Nex-N2.5 Pro: text and image in, text out, computer-use vision loop. GLM-5.2: text-only reasoning model.
• Context / output — Nex-N2.5 Pro: 262,144-token context. GLM-5.2: 1M-token context, 128K output ceiling.
• Reasoning control — Nex-N2.5 Pro: none / medium (adaptive, default) / high. GLM-5.2: reasoning-effort controls on the same model string.
• Price per 1M tokens — Nex-N2.5 Pro: free hosted tier, no list price. GLM-5.2: $1.40 input / $4.40 output, $0.26 cached input.
The envelopes differ in the ways the two models' jobs differ: GLM-5.2 brings a full million tokens of text context and a 128K output ceiling to long engineering sessions, while Nex-N2.5 Pro trades context size for a vision channel and a smaller 262K window aimed at screen-sized workloads. Neither spec is wrong; they are specced for different tasks.
Open weights, two different meanings

This is where the comparison stops being about numbers entirely. GLM-5.2 is open the way you ship a product you can build a business on: MIT license, weights on Hugging Face, no restriction on commercial use, no data-sharing clause, no vendor watching over your shoulder. You can download it, fine-tune it, serve it behind your own compliance boundary, or take it to any host you like — and because the weights are out, its price has a floor under it that no single vendor can raise. Nex-N2.5 Pro is promised under Apache-2.0, an equally permissive license, but the banner on its Hugging Face card says the weights are coming soon, and no date is attached. For a team under data-governance rules, or one that wants to escape per-token pricing entirely at scale, that difference is the whole decision: GLM-5.2 is a model you can own today, and Nex-N2.5 Pro is a model you have been promised. The self-hosting math also favors the newcomer when the weights finally land — 17B active parameters on an eight-H100 node is a far more accessible footprint than GLM-5.2's ~40B-active box — but "when the weights land" is doing a lot of work in that sentence.
The families are moving in opposite directions
Both models sit inside fast-moving families, and the trajectories pull in different directions. GLM-5.2's lineage is transparent and iterative: Z.ai shipped GLM-5.3 in August as a post-training refresh of the same 5.2 base and GLM-5.3 Flash in early September, so the 5.2 weights you build on today are the foundation the family keeps improving — your architecture knowledge and your fine-tunes carry forward. Nex-AGI's family is a bet on a brand-new generation: Nex-N2.5 Mini at 35B, Nex-N2.5 Pro at 397B, and a text-only Nex-N2.5 Max at 1.6T whose own base is DeepSeek-V4-Pro-Base, with the Pro weights that would let anyone validate the family's claims still unpublished. A team choosing a platform to build on is choosing between a lineage with a published roadmap of refreshes and a first release whose second act has not arrived.
Which one earns your build
If your requirement is "the model must be ours" — for data governance, for audit, for a product you do not want a single vendor controlling — then GLM-5.2 is not the better choice in this matchup; it is the only choice, because it is the only model in it you can download. It is also the only one with an independent score, and it is priced at Z.ai's list of $1.40 per million input and $4.40 per million output. If your requirement is a multimodal computer-use agent that could eventually undercut the closed leaders, Nex-N2.5 Pro is the more interesting long-term thesis — a 17B-active vision operator from an alliance that clearly knows what it wants to build — but a thesis is not a dependency, and a free hosted endpoint is not a foundation.

The practical way to act on all of this is to build on what exists and measure the promise when it lands. GLM-5.2 is on OrcaRouter at Z.ai's list price — the same $1.40 / $4.40, passed through with no markup — and because it is open-weights it is also self-hostable if you want the measurement to include your own serving stack. It is the reference point to run your actual agentic workload against today, with the routing DSL standing in for the day Nex-N2.5 Pro appears on a provider at list: when that happens, switching the comparison is a routing rule, not a migration. Until the shards drop and an independent harness confirms — or corrects — that 82.7, Nex-N2.5 Pro vs GLM-5.2 is a matchup between a model you can verify and a number you cannot.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
