
GLM 5.5 or GLM 5.3-Vision? An Anonymous Model Matching Z.ai's Fingerprint Just Surfaced
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 150 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 127 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 250 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 229 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
An unidentified model whose serving profile — speed, time-to-first-token, cache hit rate — matches Z.ai's GLM stack has started drawing attention from capacity analysts, and the working guess is that it could be named GLM 5.3-Vision or GLM 5.5. On August 20, compute-and-capacity analyst teortaxesTex posted that an anonymous model he is tracking is "clearly not DeepSeek, at least," that "there's nothing ruling out GLM so far," and that "speed, ttft, cache hit rate match GLM too." His read: expect it to be called GLM 5.3-Vision or 5.5, scoring roughly 61 on the Artificial Analysis Intelligence Index — one point above where the shipped GLM-5.3 sits today. Nothing here is confirmed. There is no model card, no Z.ai announcement, and no API, and this page is treating the sighting as what it is: an early, unreplicated leak signal, worth tracking precisely because GLM-5.5 has been the expected Z.ai flagship all month and has not appeared. Five days after that post, a second and this time fully checkable data point arrived: on August 25, vLLM — the inference runtime behind most large-scale model serving — opened a Do-Not-Merge pull request enabling masked-MHA kernel support for the GLM-5 head dimensions. It names no variant and proves no launch; it is serving-stack groundwork, the kind of background activity a new model leaves behind.
What the signal actually says
The source is a single X post from teortaxesTex, dated August 20 — the same account whose "roughly a week" timing signal on GLM-5.3's launch held to the day back in August. The framing is explicit about the evidence level: "strong evidence (haven't replicated)." The analyst rules DeepSeek out, finds nothing that contradicts GLM, and cites three serving-level measurements — throughput, time-to-first-token, and cache hit rate — as matching Z.ai's stack. Then the two candidate names, and a projected score: "Currently I expect it to be called GLM 5.3-Vision or 5.5, and score roughly 61 on AA."
Take each piece separately. The identity claims (not DeepSeek, matches GLM) are fingerprint inference from how the model is served, not a model card. The score is an expectation, not a measurement. The names are guesses. What makes the signal worth more than noise is that it is directionally consistent with everything else we know about Z.ai's roadmap this month: GLM-5.3 shipped text-only, the community's single loudest request was vision, and GLM-5.5 has been reported by multiple outlets (via JPMorgan, Reuters, CGTN, and a Goldman note on August 18) as arriving "within August" — the month's finale is now.
Why the serving fingerprint matters
Time-to-first-token and cache hit rate are infrastructure-level signatures. They are set by how a provider builds its inference stack — the scheduler, the cache layout, the batching policy — not by which model name a prompt calls. When an analyst says an anonymous model's TTFT and cache hit rate match GLM, they are saying the serving stack behind it behaves like Z.ai's, which is the closest thing to a fingerprint you can get without weights or a model card. It is not proof. Another vendor could mirror the same stack, or Z.ai could serve a model for a partner. But it is a specific, checkable claim, and the "haven't replicated" caveat tells you the analyst is not presenting it as settled.
The DeepSeek exclusion matters too. DeepSeek V4 Pro GA'd on August 13 and its Flash build has been the other high-velocity Chinese open-weight release this month. An anonymous model that rules out DeepSeek but points at GLM narrows the field to Z.ai's pipeline — and Z.ai's pipeline has two plausible occupants, which is exactly the fork the signal names.
A second signal: the serving stack is building for GLM-5 attention
On August 25, a second data point landed — this one fully checkable. vLLM, the inference runtime behind most large-scale model serving, received pull request #53785, titled "[Do Not Merge][Attention] Enable masked MHA for GLM-5 head dimensions." Verified against the repository, it enables the masked-MHA prefill path for the GLM-5 attention geometry: 64 heads, a KV rank of 512, QK head dimension 192+64, and V head dimension 256 — a 256/256 QK/V layout, per the PR description. It depends on a FlashAttention change for head-dimension-256 mask support, temporarily pins FlashAttention to a fork commit (which is what the Do Not Merge marker is for), and adds benchmarked routing thresholds for the new path: FlashMLA pure-prefill limits of 8K / 20K / 48K / 112K tokens at TP 1/2/4/8, and a rounded 64K limit for FlashInfer at TP8. The author notes that no other open PR covered GLM-5 masked-MHA support.
Read it as what it is: serving-stack groundwork, not an announcement. The PR does not name GLM 5.5 or GLM 5.3-Vision — it says "GLM-5" — and enabling a masked-MHA prefill path for GLM-5 head dimensions is infrastructure work on a family the vLLM ecosystem already runs over the sparse-MLA path (GLM-5.3's own lineage is served there today). It is also not isolated: sibling PRs this month covered FlashInfer MLA dimension checks for GLM-5, a Triton sparse-MLA fallback for GLM-5-class models, and FlashAttention's reported dimension support. Two details keep it on a leak tracker's radar. First, the geometry it targets — 64 heads, KV rank 512, 256/256 QK/V — is distinct from DeepSeek's 192/128, which is the same "not DeepSeek, matches GLM" separation the anonymous-model fingerprint draws, now visible at the attention-kernel level. Second, the timing: the PR opened August 25, inside the same August window in which GLM-5.5 has been expected all month. None of that proves the anonymous model is a next-gen GLM — this could equally be engineering on the model that already shipped — but it is the kind of background activity the serving ecosystem does ahead of a model, and it is verifiable in a way no X post is.
Two names, one fork: GLM 5.3-Vision vs GLM 5.5
The two candidate identities are genuinely different products, and the leak does not resolve which one is on the board.
• GLM 5.3-Vision would be a vision-capable variant of the model that shipped on August 14. GLM-5.3 is text-input only — Artificial Analysis lists it as text in, text out, no image support — and that gap is the community's loudest complaint. When Z.ai's Tang Jie ran a global feedback campaign, the comments came back essentially unanimous for vision, and Z.ai already has the building blocks (the GLM-5V-Turbo multimodal model and the CogVLM encoder) it has not yet merged into the flagship line. A 5.3-Vision variant would be the fastest way to close that gap.
• GLM 5.5 is the flagship itself. Reported but unconfirmed specs point at a base above one trillion parameters (up from GLM-5.3's 743B), a carried-over 1M-token context, open weights, and a heavier agentic/coding focus. Every analyst note this month treats 5.5 as the big one — the vehicle Z.ai co-founder Tang Jie has hinted will close the gap to Fable-class closed models. It is expected within August and has not shipped as of this writing.
The same fingerprint can't tell you which of the two this is — a vision-tuned 5.3 and a fresh flagship could both ride the same serving stack. That ambiguity is the honest state of the signal, and the two names in the analyst's post reflect it rather than resolving it.
img src="2.png" alt="A two-column comparison scoreboard titled 'GLM 5.5 vs GLM-5.3 — the scoreboard'. Left column GLM 5.5 (leaked): 'AA Index ~61 (projected, unverified)', 'Status: unreleased', 'Size >1T params (rumored)', 'Vision: expected', 'Context: 1M (expected)', 'Price: TBD'. Right column GLM-5.3 (shipped): 'AA Index 60', 'Status: API live Aug 19', 'Size 743B MoE / 40B active', 'Vision: text-only', 'Context: 1M', 'Price: $1.40 / $4.40'. Footer reads 'GLM 5.5 figures are unverified projections; GLM-5.3 per Artificial Analysis.'" />

What a ~61 on the AA Intelligence Index would mean
The projected score is the sharpest part of the leak. The Artificial Analysis Intelligence Index (v4.1.1) currently puts GLM-5.3 at 60 — tied with Kimi K3 for the top open-weights score, and 7 points clear of GLM-5.2's 53. One point up from that, at 61, would be GPT-5.6 Sol's territory, with Claude Fable 5 at 62 and Claude Opus 5 at 63. In plain terms: a GLM-family model scoring 61 would be the single highest open-weights score on the index, out on its own above Kimi K3 and GLM-5.3, and effectively at the closed-frontier line.
It is worth stressing what 61 is not. It is teortaxesTex's expectation of what an independent evaluation will return, not a published Artificial Analysis score and not a vendor number. A projection can be wrong in either direction — a vision variant might trade a point or two on the text-heavy index, and a >1T flagship might land higher or lower depending on how its post-training carries. The value of the number is the claim it encodes: whoever this anonymous model turns out to be, it is expected to beat the current open-weights leader rather than merely match it.

What is confirmed, and what is not
Keep the ledger clean, because a leak piece lives or dies on that.
• Confirmed: GLM-5.3 is real, shipped August 14, API live August 18/19, scored 60 on the AA Intelligence Index (measured independently by Artificial Analysis), priced at $1.40 / $4.40 per 1M tokens, text-input only, with open weights confirmed for Friday, August 28. Those facts do not depend on the leak.
• Confirmed: GLM-5.5 is still unreleased. As of this writing there is no model card, no pricing, and no availability; the August window and the >1T parameter figure are analyst-reported, not Z.ai commitments.
• Unconfirmed: that the anonymous model exists as a separate product, that it is GLM, that its name will be GLM 5.3-Vision or 5.5, and that it scores 61. All four rest on one analyst's unreplicated post.
• Also unconfirmed: whether this anonymous model is even distinct from GLM-5.3 itself. A live GLM-5.3 endpoint under an unlabeled alias would produce the same serving fingerprint. Until a name, a model card, or a weight drop resolves it, the "new model" reading is the strong prior but not a fact.
What to watch next
• Friday, August 28 — the GLM-5.3 weights. If the anonymous model is actually a renamed 5.3 or a 5.3-Vision, the weight drop and the model card will settle it fast.
• A second corroborating sighting. One analyst's fingerprint is a hypothesis; a second independent read on the same anonymous entry, or an Artificial Analysis listing that names it, upgrades it to evidence.
• Any Z.ai statement. GLM-5.5 has been reported for the August window, and the month is nearly over. An official naming, a pricing page, or an API changelog entry is the event that turns this leak into a launch story.
• Whether the AA score materializes at 61. If the anonymous model is real and GLM-family, an independent index score near 61 would be the first open-weights number to clear 60.
• Whether the vLLM attention work gets a model name attached. PR #53785 targets "GLM-5" head dimensions without naming 5.3-Vision or 5.5; a merge, a supported-models entry, or a model card that pairs that geometry with a specific name would be the hardest signal yet.
How to act on a leak like this
A leak is a reason to prepare, not to re-platform. If the next GLM-family model lands at or near the top of the open-weights leaderboard, the practical question is whether to route real traffic to it — and an unproven model with vendor claims and one analyst's projection is exactly the case where you want a safety net rather than a bet. The GLM-5.3 that shipped last week is already live on OrcaRouter — the model page shows it at $1.26 / $3.96 per 1M tokens, a 10% launch discount off the $1.40 / $4.40 vendor list price, passed through with zero markup — and the routing setup is what a 5.5 or 5.3-Vision would inherit the day it ships. Point a slice of traffic at the new model through the router with automatic failover to a proven model — GLM-5.3, DeepSeek V4 Pro, GPT-5.6 Sol — and you get a quality signal on your own workload instead of a slide deck. If the leak's projection is right, you found out first; if it is wrong, the failover absorbs it and nothing ships broken. Same key, no second contract, and no need to choose between the two candidate names until Z.ai does.
The next GLM is coming — the only open question is which name it wears. GLM 5.3-Vision would be Z.ai closing its most complained-about gap on the model it just shipped; GLM 5.5 would be the trillion-parameter flagship the whole month has been pointing at. The anonymous entry teortaxesTex fingerprinted on August 20 is the first concrete object that looks like either one, and its projected 61 on the AA Intelligence Index would put it ahead of every open-weights model currently on the board. Five days after the fingerprint, the serving stack moved too: vLLM's GLM-5 attention work is exactly the kind of groundwork a new model leaves behind, even though it names no variant. None of that is confirmed, and this page will update the moment the name, the card, or the weights resolve it. Until then, the low-risk move is to have the routing ready — not to bet the stack on a fingerprint.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
