
LongCat-2.5-Preview vs GLM-5.2: Two 1M-Token Windows, One of Them Measured
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 610 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 189 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1306 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 225 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
LongCat-2.5-Preview and GLM-5.2 carry the same headline number: a context window of one million tokens. That single coincidence is doing all the work in most of the comparisons being written this week, and it is not enough to compare two models on. GLM-5.2 has been on the record since 16 June 2026 with a public benchmark profile, a published rate card and a long-context recall score from an independent evaluator. LongCat-2.5-Preview appeared on Meituan's LongCat API Platform on 25 September 2026 with a discounted rate card, a parameter count reported by Chinese trade coverage, and no published benchmark of any kind. One window is measured. The other is asserted. Everything below keeps those two categories apart, because a spec sheet and a test result are not the same kind of fact.
This is the honest version of the matchup, and it is more useful than the flattering one.
What each side has actually published
Meituan's changelog carries a dated entry — "Version: 2026-09-25, LongCat-2.5-Preview Now Available" — listing three claimed capabilities: image understanding, coding, and compatibility with "Claude Code and other mainstream dev environments", naming Hermes, OpenClaw, OpenCode and Kilo Code. The vendor's pricing page states a pay-as-you-go rate of $0.30 per million uncached input tokens, $0.006 per million cached input tokens and $1.20 per million output tokens, flagged on the page itself as a limited-time discount. The context window is one million tokens and the maximum output is 131,072. Roughly 1.6 trillion total parameters with about 48 billion active per token, on a mixture-of-experts backbone, comes from Meituan's own site metadata and from Chinese trade coverage of the listing — not from the API documentation. No technical report, no model card and no benchmark table accompanied any of it. There is no LongCat-2.5-Preview repository on HuggingFace and no repository by that name in Meituan's GitHub organisation; the model-catalogue metadata that ships with OpenCode records open weights as false.

Z.ai's GLM-5.2 sits in the opposite position. It is a June release with a rate card we serve at list price: $1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens on our catalogue, with a 1,000,000-token context window and 128,000 tokens of maximum output. Its published scores come from two different places, and the distinction matters: Z.ai's own numbers for AIME 2026, HMMT February 2026, GPQA-Diamond and the FrontierSWE suite are vendor-reported; the Coding Index and Intelligence Index figures our catalogue carries are sourced to Artificial Analysis, which runs its own evaluations.
The one dimension where they are genuinely equal
Both models claim exactly 1,000,000 tokens of context. Only one of them has a published score that tests whether a model can still use what is in there. Artificial Analysis records a Long-Context Recall figure of 78.33 for GLM-5.2. LongCat-2.5-Preview has no equivalent number, from anyone. That asymmetry is the whole matchup in one line: a one-million-token window is a statement about the architecture's capacity, not about whether the model retrieves correctly at token 900,000. Capacity is cheap to print and expensive to verify, which is why every vendor prints it and almost none of them publish the recall test.
So if long-context work is what you are choosing between these two for, you are choosing between one option with a published recall score and one without — and the one without is also the one with the unmeasured benchmark profile. That is not an argument against it. It is an argument against treating the two as interchangeable on the strength of a matching integer.

What the price gap looks like on a real job
The rate cards are not close, and the direction is not the one people expect from a Chinese open-weights challenger against a Western frontier lab. LongCat-2.5-Preview is roughly 4.7 times cheaper on uncached input and 3.7 times cheaper on output than GLM-5.2 — a ratio that is arithmetic on the two published cards, not a measurement. On a 500,000-token document-processing pass that writes 20,000 tokens back, that is about 17 cents against about 79 cents. Both models have to read the same document. Only one of them has a published score for whether it will still be paying attention at the end of it.
There is a second caveat that belongs in the same breath: Meituan labels its own rate card a limited-time discount. The $0.30 figure is the promotional number, and the vendor does not say what follows it. A cost model built on a promotional price has an expiry date you cannot read. That is a reason to keep the migration between providers cheap rather than a reason to avoid the model.
On OrcaRouter, the routes we serve sit behind one key at provider list price with zero markup, so a vendor's price change reaches your bill the same day instead of at the next renewal, and automatic failover moves a degraded request to a healthy provider without your application knowing. GLM-5.2 is one of our routes. LongCat-2.5-Preview is not — we do not serve it, and nothing here should be read as claiming otherwise. Z.ai has also moved on since June: GLM-5.3 is live on our catalogue as of 18 August 2026, so a comparison framed as "new model versus current model" is already framing GLM-5.2 as the incumbent rather than the frontier.

What would settle this, and what will not
• A Long-Context Recall score for LongCat-2.5-Preview, from Meituan or from an independent evaluator. Until that exists, its 1M window is a claim with no test attached, and GLM-5.2's 78.33 is the only number in the comparison that speaks to the same question.
• A vendor rate card without the limited-time flag. The discounted price tells you what the model costs during a promotion, not what it costs.
• A benchmark table of any kind. Not a leaderboard position — just a published evaluation run, so that the comparison stops being a spec sheet against a results page.
• An image-input confirmation. The changelog presents image understanding as a headline addition, but the example response in Meituan's own "Retrieve Model" documentation still shows input_modalities as ["text"] with a text->text modality string. That sample may simply be stale. It is nonetheless the vendor's published contract, so treat image input as claimed rather than available. GLM-5.2, by contrast, is text-in and text-out on our catalogue with no image modality at all — so on this dimension neither model gives you a verified answer, and one of them gives you a claim.
• Your own numbers. The gap between what is published and what you need is closed by running both models on your tasks, and the free window on LongCat-2.5-Preview makes that cheap right now. A reasoning trace that arrives in an interleaved reasoning_content field is the harness detail to check first, because tooling that silently drops it will lose most of the model's usefulness without erroring.
Where this leaves the comparison
If you need a decision today and it involves long context, GLM-5.2 is the one you can actually evaluate: it has published recall, an independent coding score of 68.8 and an Intelligence Index of 33.7 from Artificial Analysis, and a price you can budget against. Those figures describe a model that is competent rather than frontier — Z.ai's own successor, GLM-5.3, outscores it — but they describe something. LongCat-2.5-Preview is a cheaper, larger-context, entirely unmeasured proposition whose most attractive quality, its price, is explicitly temporary. The correct posture toward it is not scepticism. It is a two-week evaluation with your own scoring harness attached, run before the promotional rate disappears.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
