
LongCat-2.5-Preview vs Kimi K3: The Long-Context Recall Question Nobody Can Answer Yet
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 610 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 189 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1306 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 225 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
Kimi K3 scores 88.67 on Long-Context Recall. That is the highest figure of any model on our catalogue for the specific question of whether a model still uses information buried deep inside a long input. LongCat-2.5-Preview has no Long-Context Recall score at all — from Artificial Analysis, from Meituan, from anyone. And LongCat-2.5-Preview is the one advertising a one-million-token window as its headline feature. That mismatch, a model whose central claim is long context and for which no long-context measurement exists, is the entire substance of this comparison. Everything else is secondary.
It is worth being precise about what the missing number does and does not mean, because it is easy to slide from "unmeasured" into "probably bad" and neither direction is supported.
What the two models each put on the record
MoonshotAI's Kimi K3 shipped 15 July 2026. Our catalogue carries it with a 1,048,576-token context window, text and image input, and a rate card of $3.00 per million input tokens, $0.30 per million cached input and $15.00 per million output. Its independent profile from Artificial Analysis runs: Coding Index 76.2, ninth; Intelligence Index 43.6, nineteenth; GPQA Diamond 93.5%; Humanity's Last Exam 46.9%; SciCode 59.5%; Terminal-Bench v2.1 85.0; τ²-Bench banking 46.0. And Long-Context Recall 88.67, the best we carry.

Meituan's LongCat-2.5-Preview appeared on the LongCat API Platform on 25 September 2026. Its changelog entry names three claimed capabilities — image understanding, coding, and compatibility with "Claude Code and other mainstream dev environments", listing Hermes, OpenClaw, OpenCode and Kilo Code. The pricing page gives $0.30 per million uncached input, $0.006 per million cached input and $1.20 per million output, flagged as a limited-time discount. Context window 1,000,000; maximum output 131,072. The ~1.6T total / ~48B active parameter figure comes from Meituan's site metadata and Chinese trade coverage, not from documentation. No benchmark table, no model card, no technical report, no repository on HuggingFace or in Meituan's GitHub organisation, and catalogue metadata recording open weights as false.
Why the missing number matters more here than in any other matchup
Long-Context Recall is not one score among many for these two models. It is the score that tests the thing they both sell. A one-million-token window is a statement about architecture capacity; recall measures whether the model can still retrieve and use a fact placed at token 900,000 instead of token 900. Vendors publish the first because it is a specification. Independent evaluators publish the second because it is a test, and most models do not do as well on it as their window size implies.
So the honest position is narrower than it sounds. Kimi K3's 88.67 does not mean Kimi K3 is the better long-context model than LongCat-2.5-Preview. It means Kimi K3 is the one for which the question has been asked and answered. LongCat-2.5-Preview could be better. It could be substantially worse. The point is that nobody outside Meituan currently knows, and a 1M window printed on a pricing page is not evidence in either direction.

The cost picture, with the same arithmetic warning
On published rates, LongCat-2.5-Preview is 10 times cheaper on uncached input and 12.5 times cheaper on output than Kimi K3. A 500,000-token read that writes 20,000 tokens back comes to roughly $1.80 on Kimi K3 and roughly 17 cents on LongCat-2.5-Preview. Both are arithmetic on card prices rather than measurements, and the LongCat figure carries an extra caveat the Kimi figure does not: Meituan labels its own rate as a limited-time discount, so the number that survives the promotion is unknown.
Set that against the coverage question. If you are processing a 500,000-token input and your model's recall at that depth is unmeasured, the cost difference is not a saving — it is the price of an unknown failure rate. A 17-cent request that silently misses the relevant section is more expensive than a $1.80 request that finds it, because you pay for the re-run and the debugging either way. This is the specific shape of the risk, and it is why the missing benchmark is a cost problem and not just a marketing one.
What to do while the number does not exist
Generate your own. This is the rare case where the free window is genuinely well suited to the gap: LongCat-2.5-Preview costs nothing to call through OpenCode for a period of unstated length, with a zero-retention policy that OpenCode documents — model training "Not used", data retention "0 days" — which means you can point it at a private repository without a data-processing conversation first. Build a needle-in-a-haystack test at the depth you actually use, run it on both models, and you will have the number that neither vendor has published. Two things to check before you trust the result: that your harness surfaces the interleaved reasoning_content field the model returns, and that image input works at all, since Meituan's changelog claims it while its own "Retrieve Model" example still shows input_modalities as ["text"] with a text->text string.

OrcaRouter's part in this one
We serve Kimi K3 at MoonshotAI's list price with zero markup. LongCat-2.5-Preview is not one of our routes — we do not serve it, and nothing in this piece should be read as claiming we do.
For this particular comparison the useful part of a router is not the price. It is that the model you are evaluating and the model you are relying on can sit behind the same key. Kimi K3's 88.67 recall makes it the model you would route a long-context pass through today; LongCat-2.5-Preview is the model you want to test against it, because if its recall holds up at a twelfth of the output price the economics of long-document work change. Model fusion is the mechanism that makes that concrete rather than theoretical: a long-context retrieval pass and a final synthesis step can be two different models on one request, so the cheap unmeasured model and the expensive measured one do not have to be an either-or. Automatic failover then covers the case where one of them degrades, which matters more than usual when one of your two candidates has no serving history you can inspect.
The verdict, stated with its own uncertainty attached
If you need long-context work to be reliable this week, Kimi K3 is the defensible choice: 88.67 recall is the best number available for the exact capability in question, and its broader profile — Coding 76.2, Intelligence 43.6, GPQA Diamond 93.5% — is solid rather than spectacular, all of it independently sourced. If you are willing to spend an afternoon generating your own evaluation, LongCat-2.5-Preview is the more interesting bet: a million-token window, a 131,072-token output ceiling, zero-retention access and a rate card an order of magnitude below Kimi K3's, all of it resting on vendor claims that no one has yet tested in public.
What is not defensible is treating them as equivalent because both list a million tokens. One of them has a number attached to that window, and the other has a price. Those are different kinds of evidence, and the gap between them is the only thing this comparison is actually about.
