A generated hero card titled 'NV-Reason-CT vs GLM-5.2' with the subtitle 'One of these is deprecated, and it is not the one you think'. Two facing cards are separated by a pair of blue arrows: the left card is labelled 'NV-Reason-CT' with a body-scan icon and 'released September 2026 - current - chest and abdomen CT only'; the right card is labelled 'GLM-5.2' with a chip icon and 'released June 2026 - superseded by GLM-5.3 - deprecated at Artificial Analysis'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

NV-Reason-CT vs GLM-5.2: One of These Is Deprecated, and It Is Not the One You Think

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The older model in this comparison is the one being retired. GLM-5.2 was released on 16 June 2026; NV-Reason-CT has repository activity dated between 2 and 25 September 2026. You would expect the September one to be the uncertain quantity and the June one to be the settled choice. The opposite is true on the axis that matters for a text pipeline: GLM-5.2 has been superseded by GLM-5.3, which shipped on 18 August 2026, and it now carries a deprecation marker on the independent leaderboard where its figures are published. NV-Reason-CT, whatever else is unresolved about it, is the current release of the only thing it does.

So the first useful sentence in this article is the one a reader is least likely to have: if you are starting a text build today and reaching for GLM-5.2 out of habit, stop and look at GLM-5.3 first. It costs $1.26 per million input tokens and $3.96 per million output against GLM-5.2's $1.40 and $4.40 — it is both newer and cheaper — and its independent intelligence index is 44.8 against 33.7, which is a step up the same scale rather than a sideways one. That is a separate decision from anything to do with CT, and it deserves to be made separately.

The two models, and the two different kinds of stale

There is a real distinction between a model that is old and a model that is superseded, and this pairing makes it concrete.

• What it does — NV-Reason-CT reads a chest or abdominal CT volume and emits structured findings by anatomical region; GLM-5.2 is a text-only language model for reasoning, code and long-document work

• Released — NV-Reason-CT's artefacts date from 2 to 25 September 2026; GLM-5.2 on 16 June 2026, with GLM-5.3 following on 18 August 2026

• Generation status — NV-Reason-CT is version 0.1, an initial release, unannounced; GLM-5.2 is the previous generation of a shipping product line

• Independent standing — no independent measurements of NV-Reason-CT exist; GLM-5.2 is measured at 33.7 on the Artificial Analysis Intelligence Index, with a deprecation marker, against GLM-5.3 at 44.8

• Licence and access — OpenMDW-1.1 weights you download; against an open-weight model served at $1.40 and $4.40 per million tokens with cached reads at $0.26

• Context — 13,824 visual tokens per volume with no chat window; against 1,000,000 tokens in and 128,000 out

• Stated use — research and education only, not a medical device; against general-purpose deployment

The word "deprecated" is doing precise work here and it is worth not over-reading. A deprecation marker on a leaderboard means the evaluator has stopped treating the model as current, usually because a successor has taken its place. It does not mean the weights have been withdrawn, the API turned off, or the model stopped working. Teams with pipelines tuned against GLM-5.2 are not in trouble. What it does mean is that its figures are a snapshot from before GLM-5.3 existed, and that mixing them with current numbers for other models is comparing a June measurement to a September field.

The numbers each side has, and what they are worth

GLM-5.2's record is unusually well-populated and it comes from two sources that disagree in instructive ways.

From Z.ai's own reporting, the model posts 99.12 on τ²-Bench, 62.1 on SWE-bench Pro, 54.7 on Humanity's Last Exam with tools, and a best-reported 82.7 on Terminal-Bench 2.1. On the independent side, Artificial Analysis measures the same model at 77.9 on Terminal-Bench 2.1, 33.7 on its Intelligence Index, 89.5 on GPQA Diamond and 41.1 on Humanity's Last Exam. There are other vendor figures — 99.2 on AIME 2026, 92.5 on HMMT February 2026, 91 on IMOAnswerBench, 74.4 on FrontierSWE, 63.7 on ProgramBench, 76.8 on MCP-Atlas — and they are all Z.ai's.

Two things about that list are worth naming. First, the 4.8-point spread on Terminal-Bench 2.1 between the vendor's best-reported 82.7 and the independent 77.9 is not a contradiction. Vendor best-reported figures are typically the best of several harnesses, and Z.ai's own Terminus-2 number for the same benchmark is 81 — the independent measurement sits inside the vendor's own range. Second, the Humanity's Last Exam split is larger: 41.1 independently against a vendor figure of 40.5 without tools and 54.7 with them, which means the headline 54.7 is a tool-assisted number and the 41.1 is the bare one. Quoting 54.7 next to another model's bare score would be a real error.

NV-Reason-CT's record has no such two-source structure, and that is precisely where it is weaker. Its CT-RATE results — 0.614 F1 and 0.871 AUROC across eighteen labels, with a fixed uniform threshold and a direct yes/no prompt and no classification head or task-specific adaptation — are NVIDIA's own. The models it is compared against in that paper (VoxelFM 0.581, Pillar-0 0.544, ClinFusion-8B 0.442, CT-CLIP 0.398, Merlin 0.358, MedGemma 1.5 0.303) are all imaging models. Nothing has been reproduced independently, and the model has been downloadable for weeks. It also reports a report-derived macro-F1 of 0.592 and a preliminary expert radiologist study linking assisted review to a 50% reduction in average reported interpretation time, which the authors themselves flag as severe small-sample.

So the provenance comparison does not favour the newer model, and this is the point worth sitting with. GLM-5.2 is the older, superseded model, and its numbers are more trustworthy than NV-Reason-CT's, because somebody outside the company ran the harness. Being current is not the same as being verified.

A generated scoreboard titled 'Provenance, not recency, decides what to trust' with two columns. The left column, headed 'NV-Reason-CT', lists five rows: CT-RATE F1 0.614 and AUROC 0.871; provenance, the authors paper only; independent checks, none; model status, version 0.1, unannounced, current; licence, OpenMDW-1.1, research and education only. The right column, headed 'GLM-5.2', lists five rows: AA Intelligence 33.7, GPQA Diamond 89.5; provenance, vendor and Artificial Analysis; independent checks, yes, with a 4.8-point vendor-versus-independent gap on Terminal-Bench 2.1; model status, superseded by GLM-5.3 and deprecated by the evaluator; licence, open weights at $1.40 and $4.40 per million. A footer reads 'The superseded model is the better-verified one. Recency is not the same as evidence.'

Why the deprecation matters more than it sounds

There is a specific failure this comparison is designed to prevent, and it is not about CT at all.

A team building a clinical text pipeline goes looking for an open-weight model to run on their own hardware, because the data is sensitive. They find GLM-5.2, which is MIT-licensed at 743 billion parameters with roughly 40 billion active, fits the bill, and has a well-documented benchmark record. They tune against it. Six months later they discover that the current generation is GLM-5.3, that it is 1M-context with a 128K output ceiling, that it is cheaper at $1.26 and $3.96, and that the decision they thought they were making — which open-weight model to standardise on — was actually a decision about which generation, and they made it by accident.

That is an expensive accident to discover late, and it is avoidable with one lookup. The concrete guidance:

• Check whether the model you are about to standardise on is the current generation — a leaderboard deprecation marker is the fastest signal, and a vendor's own model list is the authoritative one

• Do not mix a June snapshot with September figures — GLM-5.2's index of 33.7 was measured before GLM-5.3 existed, and putting it beside a current model's index compares two different moments in a moving field

• Be careful with the name — a "GLM-5.3 Prime" listing is a serving tier carrying the same weights at roughly twice the list price, not a separate model. Check the weights behind any SKU before paying a premium for it

• Treat a lower headline rate as a question, not an answer — GLM-5.3 being cheaper than its predecessor is unusual and worth verifying against your own workload rather than assuming

None of that makes GLM-5.2 a bad choice. An MIT-licensed 743B model at $1.40 and $4.40 that a team has already tuned against is a perfectly reasonable thing to keep running, and a mid-generation model with a settled track record is sometimes exactly what a regulated workflow wants. The point is that it should be a choice, made deliberately, rather than a default inherited from a list that was current in July.

The trap in comparing a text model to a CT model

The reason this pairing looks plausible at all is that both are open-weight models released in 2026, and a reader scanning a catalogue sees two comparable line items. They are not comparable and the mismatch has a specific shape.

GLM-5.2 holds 1,000,000 tokens. NV-Reason-CT's single CT volume costs 13,824 visual tokens. By that arithmetic GLM-5.2 could hold seventy CT studies and still have room, which invites the conclusion that a sufficiently long-context text model makes the imaging model redundant. The conclusion does not follow, and the reason is the interface rather than the budget.

Those 13,824 tokens are not a summary. They are a 384-millimetre cube resampled to 2 mm isotropic, cut into 8 x 8 x 8 patches on a 24 x 24 x 24 grid, handed to a language backbone with no spatial downsampling and no merging layer — which is why the active pathway is 4.35B parameters of the 4.69B total, and why the compute is spent on keeping resolution rather than compressing it away. GLM-5.2 is text-only. It has no vision path at all, so the question of how it would handle a scan does not arise; the answer is that it cannot be given one in any form. The two model cards describe a six-hundred-fold difference in token budget that has nothing to do with the comparison, because one of them has no way to receive the input.

The reverse mistake is just as available. NV-Reason-CT supports chest and abdomen, takes a NIfTI file rather than a prompt, has no tool use, and its model card restricts it to research and education and states that it is not a medical device and that outputs can be incorrect. It cannot normalise a report corpus, write an evaluation harness, or reason over a guideline document. A 4.7B imaging model is not a substitute for a 743B text model any more than the reverse.

A generated scoreboard titled 'Two open-weight models, two different jobs' listing six paired rows. Row one: what it reads, one-channel NIfTI volumes in Hounsfield units against text only. Row two: context, 13,824 visual tokens per volume against 1,000,000 tokens in and 128,000 out. Row three: parameters, 4.69B total and 4.35B active against 743B total and about 40B active. Row four: generation status, version 0.1 current, unannounced against superseded by GLM-5.3. Row five: price, self-hosted with no rate card against $1.40 and $4.40 per million, or $1.26 and $3.96 for the successor. Row six: stated use, research and education, not a medical device against general purpose. A footer reads 'GLM-5.2 figures per the provider and Artificial Analysis; NV-Reason-CT figures per the authors paper.'

Putting the text half on one key

If the question is which text model to run the pipeline work on, the answer today is probably GLM-5.3 rather than GLM-5.2 — and both are on our catalogue under one key. z-ai/glm-5.2 is served at Z.ai's provider list rate with zero markup, and GLM-5.3 alongside it, inside a single API covering more than 200 models. Keeping both available matters more than usual during a generation change: you can measure the successor on your own corpus before committing, keep the incumbent as a fallback while you do, and switch without a second contract or a code change.

The pass-through rate is worth a sentence for the same reason it matters on any model whose vendor reprices. With no markup layer in between, a vendor rate change lands on our side the same day. Automatic failover covers a provider degrading mid-batch, which for a long extraction job is the failure that actually costs a night, and the routing DSL lets you send individual requests to the model that should handle them rather than pointing everything at one endpoint.

NV-Reason-CT is not on our platform, and no NVIDIA model is. That is worth stating plainly rather than leaving to inference: the CT side of this architecture is downloaded weights on your own hardware, under a licence that permits it and a model card that forbids clinical use. We do not host it and we are not going to write a sentence that suggests otherwise.

The order to make these decisions in

The right sequence for a clinical build is not the one the comparison implies.

Start with the text generation question, and start it by checking which generation is current — GLM-5.3 rather than GLM-5.2, on price and on measured capability, unless you have a reason to stay put. Then measure the CT model's per-study latency on your own hardware, because that number is unpublished and it governs the rest of the budget. Then design the pipeline so the two halves are independently replaceable, since one is on a preview-adjacent lifecycle and the other is at version 0.1.

The one thing not to do is treat the two as a choice. A superseded 743B text model and a current 4.69B CT model have no task in common. The comparison is worth making only for what it exposes: that the newer artefact has the weaker evidence behind it, that the older one is quietly out of generation, and that both facts are invisible to anyone reading a catalogue row that lists them side by side as if they were alternatives.

A screenshot of the OrcaRouter model page for GLM 5.2, showing the identifier z-ai/glm-5.2 with a Featured badge and tags for Tools, JSON and Reasoning, the description of it as Z.ai's flagship model for long-horizon tasks with a 1M-token context window and up to 128K output tokens, a side panel listing a 1M-token context window and 128K maximum output with text input and text output, and a stat strip reading $1.40 per million input tokens, $4.40 per million output tokens, and a p50 time to first token under five seconds.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily