
LongCat-2.5-Preview vs Grok 4.6: One Has a Benchmark Card, One Has a Successor
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 610 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 189 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1306 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 225 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
Comparing LongCat-2.5-Preview with Grok 4.6 is awkward for a reason that has nothing to do with either model's quality: the question has a shelf life. Grok 4.6 shipped on 12 August 2026 and Grok 4.7 shipped on 21 September 2026 — six days before this is being written — at the same $2.00 / $6.00 per million rates and the same 500,000-token context window. LongCat-2.5-Preview, meanwhile, landed on Meituan's LongCat API Platform on 25 September 2026 with no announced successor anywhere in sight and no published benchmark at all. So the real choice is not "which model is better". It is whether you want a measured capability with a measured successor already standing behind it, or an unmeasured one that is at least the current thing.
That framing is less exciting than a scoreboard and considerably more useful.
Start with what Grok 4.6 actually has on file
Grok 4.6 is a well-documented model. On our catalogue it carries a 500,000-token context window, text, image and file input, and a rate card of $2.00 per million input tokens, $6.00 per million output tokens and $0.50 per million cached input tokens, stepping to $4.00 and $12.00 above 200,000 input tokens. Its independent profile from Artificial Analysis includes a Coding Index of 76.8 — fifth among the models that index tracks — an Intelligence Index of 44.3 at eighteenth, GPQA Diamond at 94.9%, Humanity's Last Exam at 42.9%, SciCode at 56.5%, Terminal-Bench v2.1 at 88.4 and τ²-Bench for banking at 50.7. Its Long-Context Recall is 80.3.

LongCat-2.5-Preview has none of that. Meituan's changelog entry for 25 September 2026 is a dated availability notice listing three claimed capabilities — image understanding, coding, and compatibility with "Claude Code and other mainstream dev environments" including Hermes, OpenClaw, OpenCode and Kilo Code — and nothing that resembles a result. The vendor's pricing page gives $0.30 per million uncached input, $0.006 per million cached input and $1.20 per million output, explicitly flagged as a limited-time discount. The context window is 1,000,000 tokens; maximum output is 131,072. The roughly 1.6-trillion-total / 48-billion-active parameter count comes from Meituan's site metadata and Chinese trade coverage rather than documentation. There is no repository for it on HuggingFace or in Meituan's GitHub organisation, and the catalogue metadata OpenCode ships records open weights as false.
The dimension a scoreboard would miss entirely
These two models differ on something no benchmark measures, and for a lot of teams it decides the question on its own: what happens to the prompts you send.
OpenCode's own documentation states that LongCat 2.5 Preview Free is served by a provider that follows a zero-retention policy and does not use your data for training; the Zen privacy table puts numbers on it — model training "Not used", data retention "0 days". The same privacy table lists thirty days of retention for Grok 4.6 and Grok 4.7. Both of those statements are OpenCode's, published on OpenCode's pages, and they describe the free listing and the comparison listing specifically. But if you are pointing an agent at a private repository, that is the difference between a conversation you do not have to have with legal and one you do — and it has nothing to do with which model scores better on SciCode.

What "measured" does and does not buy you
Grok 4.6's numbers are better than LongCat-2.5-Preview's in the only sense that matters here: they exist. That is not the same as saying Grok 4.6 is the better model. Its Coding Index of 76.8 is below Grok 4.7's generation, and the model it was measured against is no longer the frontier of its own family. Its published successor has an Intelligence Index of 46.4 against Grok 4.6's 44.3, and a Long-Context Recall of 76.67 against Grok 4.6's 80.33 — so the successor is not better on every axis, which is exactly the kind of detail a generation number hides.
There is also a live-serving caveat worth stating plainly, because it cuts against the flattering reading for both models. On our own seven-day playground sample, Grok 4.6 recorded no measured throughput and no token traffic, which is what an absence of data looks like in that column rather than a performance verdict. LongCat-2.5-Preview has no serving sample of ours at all, because we do not route it. So any latency comparison between these two is one-sided in a way that prose usually papers over: you cannot benchmark a model you are not sending traffic to.
Where the routing layer actually helps here
Three of the facts above are the kind that a router is built for rather than merely compatible with. Grok 4.6's rate card steps up above 200,000 input tokens, so the same logical task can be billed at two different rates depending on a number your application already knows before it sends anything. Grok 4.6 has a published successor at identical rates, which means the switch from one to the other should be a one-line change and never a re-integration. And LongCat-2.5-Preview's most attractive property — its price — is explicitly labelled temporary by the vendor.
On OrcaRouter we serve Grok 4.6 at xAI's list price with zero markup, so a vendor's rate change reaches your bill the same day rather than at the next renewal, and automatic failover moves a degraded request to a healthy provider without your code knowing which one answered. A routing DSL covers the tiered-pricing case directly, and model fusion covers the case where the model you want for the long read is not the model you want for the final synthesis. LongCat-2.5-Preview is not one of our routes — we do not serve it, and nothing here claims otherwise. The point of naming that is not to route you elsewhere; it is that a model you are evaluating rather than running belongs behind the same key as the models you run, so that evaluating it costs a config entry instead of a project.

How to decide
• Choose Grok 4.6 if you need an answer you can defend. It has an independent card, a documented input profile, and a price that does not expire. Its main risk is not quality but relevance: Grok 4.7 exists at the same rates, and the cost of staying on 4.6 out of inertia is the difference between an Intelligence Index of 44.3 and 46.4 that you did not have to pay for.
• Choose LongCat-2.5-Preview if the deciding factor is retention or reach. Zero-retention access through OpenCode, a one-million-token window, a 131,072-token output ceiling and a rate card roughly a seventh of Grok 4.6's are real advantages — the first of them verified by a third party's privacy policy, the rest asserted by the vendor alone.
• Do not choose between them on the strength of a benchmark table, because only one exists. Grok 4.6's 76.8 coding score and 80.3 long-context recall describe Grok 4.6. Nothing describes LongCat-2.5-Preview yet. Anyone printing a LongCat-2.5-Preview score beside a Grok 4.6 score this week is either quoting a different LongCat model or inventing the number.
• Verify the input side before you build on it. Meituan's changelog leads with image understanding, but the example response in its own "Retrieve Model" documentation still shows input_modalities as ["text"] with a text->text modality string — vendor-claimed against a published contract that says otherwise. Grok 4.6 takes text, images and files with no such ambiguity. And check that your harness surfaces an interleaved reasoning_content field before you conclude anything about LongCat-2.5-Preview's reasoning quality; a client that drops that field will fail quietly.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
