A generated title card for the article 'Mistral Large 4 vs GLM-5.5' showing the title in bold geometric sans-serif, the subtitle 'A real model against a name' beneath it, and a row of three flat line icons above — a solid model card beside an empty dashed-outline card — on a white background with blue-and-cyan gradient accents and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Mistral Large 4 vs GLM-5.5: A Real Model Against a Name

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

There is no GLM-5.5 to compare against Mistral Large 4. As of 6 October 2026 the name has no model card, no API identifier, no pricing row, no checkpoint and no licence file, and Z.ai's own site uses it nowhere. The model at the end of that number line is GLM-5.3, which Z.ai released on 18 August 2026 — the same Z.ai open-weight flagship Mistral benchmarks Mistral Large 4 against by name in its own launch post. So the comparison that can actually be run is a preview released today against a two-month-old open-weight model with a live price, four and a half billion tokens a week of real traffic, and a licence file you can read. That turns out to be the more useful matchup anyway.

It is also a cleaner test of Mistral's central claim than the GLM comparison the launch post performs. Mistral says Mistral Large 4 is the strongest open-weight model built outside China. GLM-5.3 is the strongest open-weight model built inside it. The two claims are checkable against each other on the same index, and one of them currently rests on weights that have not been released.

What exists, and what does not

The check takes about ten minutes, and each of these is a place a released Z.ai flagship would appear:

• Z.ai's own model-api page lists exactly three cards — GLM-5.3, GLM-5.3-FlashX and GLM-5.2. There is no fourth entry under any 5.4 or 5.5 name.

• Artificial Analysis has a live GLM-5.3 page; artificialanalysis.ai/models/glm-5-5 and /glm-5-4 both 404.

• The zai-org Hugging Face organisation's newest repositories are the GLM-5.3 series, created 25 August 2026 — GLM-5.3, GLM-5.3-BF16 and GLM-5.3-Flash among them. Nothing later exists.

• Our own catalogue returns model-not-found for z-ai/glm-5.5, z-ai/glm-5.4 and z-ai/glm-6.0, while z-ai/glm-5.3 resolves with a live price.

• Z.ai's 31 August 2026 earnings call named GLM-6.0 as the next-generation base, skipping 5.4 and 5.5 entirely.

The name circulates because Z.ai shipped 5.3 in August and the next flagship reads naturally as the next number. That is a naming convention, not a product. Nothing in this article should be read as a spec sheet for GLM-5.5, because producing one would mean inventing every figure in it.

A screenshot of Mistral's own documentation, the Models Overview page at docs.mistral.ai, showing the GENERALIST MODELS card grid with 'Mistral Large 4' at version v26.10 described as 'A state-of-the-art, open-weight, general-purpose multimodal model' and no licence badge, beside Mistral Large 3 at v25.12 with an APACHE 2.0 badge, and Mistral Small 4, Mistral Medium 3.5, Z.ai GLM 5.3 and the Ministral 3 sizes. Z.ai GLM 5.3 appears as the first card in the grid.

The real comparison: Mistral Large 4 preview vs GLM-5.3

Mistral's figures come from its own launch post and model card; GLM-5.3's come from its live catalogue entry on OrcaRouter and its Artificial Analysis page, both read on 6 October.

• Architecture — 1.05T total / 49B active sparse MoE with a 1.6B vision encoder vs a text-only MoE with published open weights

• Modalities — text and image in, text out vs text in, text out

• Context window — 1M stated by Mistral, 524,288 on the configuration Artificial Analysis is testing, vs 1M served

• Output ceiling — not published vs 128K tokens

• Price per million, input / output — $1.36 / $4.18 list, currently $0.68 / $2.09 promoted, vs $1.40 / $4.40 list, currently $1.26 / $3.96 served

• Cached input per million — $0.14, currently $0.07, vs $0.26 list, $0.234 served

• Intelligence Index v4.3 — 38.4 vs 44.8

• Cost per index task — $1.13 vs $2.01

• Terminal-Bench 4.0 — 26.8% vs 41.9%

• AutomationBench — 59.9% vs 62.2%

• Weights — promised end of October, licence unnamed vs available now under Z.ai's own terms

Read plainly: GLM-5.3 is the better model on the independent index by 6.4 points and it is already downloadable. Mistral Large 4 is 44% cheaper per completed task and is, on the Terminal-Bench row, the weaker coder by fifteen points. The pricing symmetry is coincidental and close enough to be confusing — $1.36 / $4.18 against $1.40 / $4.40 for the list rates, though OrcaRouter serves GLM-5.3 at $1.26 / $3.96, which undercuts Mistral's list rate outright. Mistral's promotional $0.68 / $2.09 is what makes the preview cheaper today, and it is a launch price.

Where Mistral has the better of it

Two rows matter, and neither is the index.

The first is multimodal. GLM-5.3 is a text-in, text-out model, and its catalogue card is explicit about it. Mistral Large 4 ships a 1.6-billion-parameter vision encoder and Mistral claims state-of-the-art visual grounding among open models — citing a 42% against GPT-6 Astra's 41% on Dense 200, and demos of gigapixel satellite imagery, engineering drawings and dense document retrieval. On the same harness GLM-5.3's multimodal score does not appear because there is nothing to score. For any pipeline reading charts, drawings, filings or screenshots, the comparison collapses to one candidate.

The second is security work. Mistral claims 82% on the Artificial Analysis Cyber Index test that asks a model to reproduce a real vulnerability and then patch it — described as the highest of any model — plus 93% on Cybench's 40 competition exercises, and a top-five Cyber Index rank overall that leads open-weight models built outside China. Z.ai has published cybersecurity claims for GLM-5.3 too, and GLM-5.3's own description credits it with matching Mythos 5 on selected cybersecurity capabilities. Both sets of figures are vendor-cited and both should be labelled that way. The difference is that Mistral's are placed on a named independent index and tied to a specific test, which makes them easier to check and easier to falsify.

A generated two-column scoreboard titled 'Mistral Large 4 vs GLM 5.3 — the scoreboard'. Left column 'Mistral Large 4 Preview': Intelligence Index v4.3 38.4; cost per index task $1.13; Terminal-Bench 4.0 26.8%; AutomationBench 59.9%; multimodal vision encoder; weights promised end of October. Right column 'GLM 5.3': Intelligence Index v4.3 44.8; cost per index task $2.01; Terminal-Bench 4.0 41.9%; AutomationBench 62.2%; multimodal text only; weights open and available now.

Everything else favours the incumbent, and the cost-of-a-preview argument is the one to be careful about. Mistral's launch post states that the preview is served on the same European infrastructure it was trained on — 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own datacenters — and that the weights follow at the end of October. What that buys you today is a callable model with a stated end to its preview period; what it does not buy you is the thing Mistral's whole pitch rests on.

What would make GLM-5.5 a model

If Z.ai does ship a 5.4 or a 5.5, the artifacts would appear in a fixed order and each is checkable without asking Z.ai anything:

• A model card on Z.ai's own model-api page, which today lists three entries and no more

• An API identifier that resolves, the way z-ai/glm-5.3 resolves and z-ai/glm-5.5 does not

• A pricing row with a per-million-token rate on Z.ai's documentation

• A Hugging Face repository from the zai-org organisation with a licence file — the current series tops out at 25 August 2026

• A page on a neutral benchmark house, where glm-5.5 currently 404s

The forward-looking evidence points away from the name rather than toward it: the company's own earnings call skipped 5.4 and 5.5 and named GLM-6.0 as the next generation. Anyone writing a GLM-5.5 spec sheet this month is writing fiction with a plausible title.

Choosing between the two that exist

If you want the better model today, the answer is GLM-5.3 and it is not close on the independent index. It is 6.4 points ahead on Intelligence Index v4.3, fifteen points ahead on Terminal-Bench 4.0, ahead on workflow automation, an actual open-weight release with a checkpoint you can download, and it moves more than five billion tokens a week across our catalogue — volume that tells you its failure modes are known. Its context is a served million tokens and its output ceiling a published 128K, against a preview whose context figure differs by a factor of two depending on whose page you read and whose output ceiling has not been stated at all.

If you need vision, if you need security work, or if a preview with a stated end date is acceptable risk for a model that may become your own to run, Mistral Large 4 is the interesting side of the trade. On OrcaRouter, GLM-5.3 sits in the catalogue at the provider's rate with 0% markup in the same OpenAI-compatible endpoint as 200-plus other models, with automatic failover across providers and a routing DSL that can send classification and extraction to whichever model is cheapest that week. Mistral Large 4 is not in our catalogue; the preview is reached through Mistral's own API and several third-party platforms.

A screenshot of the OrcaRouter model page for z-ai/glm-5.3, showing the model identity, a 1M-token context window, a 128K maximum output, text-only input with text output, the 2026-08-18 release date, a p-50 time-to-first-token figure of 3.99 seconds, pricing of $1.26 per million input and $3.96 per million output, and a code sample against the api.orcarouter.ai base URL. The page describes the model as Z.ai's latest flagship for complex software engineering and long-horizon agentic tasks, text-in and text-out.

Be explicit about what you are buying. GLM-5.3 is a product. Mistral Large 4 is a product with a promise attached, and the promise is the part that has not shipped. The gap between them is not the 6.4 index points — those will move — it is that one of these models exists in a form you can hold, and the other is a naming convention with a delivery date.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily