A generated hero title card reading 'LongCat-2.5-Preview vs Grok 4.6' with the overline 'One has a benchmark card, the other has a successor', and two panels: 'LongCat-2.5-Preview: 1M context, zero retention via OpenCode' and 'Grok 4.6: AA Coding 76.8, Grok 4.7 already shipped'. OrcaRouter logo bottom-right.
Guides & Insights

LongCat-2.5-Preview vs Grok 4.6: One Has a Benchmark Card, One Has a Successor

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Comparing LongCat-2.5-Preview with Grok 4.6 is awkward for a reason that has nothing to do with either model's quality: the question has a shelf life. Grok 4.6 shipped on 12 August 2026 and Grok 4.7 shipped on 21 September 2026 — six days before this is being written — at the same $2.00 / $6.00 per million rates and the same 500,000-token context window. LongCat-2.5-Preview, meanwhile, landed on Meituan's LongCat API Platform on 25 September 2026 with no announced successor anywhere in sight and no published benchmark at all. So the real choice is not "which model is better". It is whether you want a measured capability with a measured successor already standing behind it, or an unmeasured one that is at least the current thing.

That framing is less exciting than a scoreboard and considerably more useful.

Start with what Grok 4.6 actually has on file

Grok 4.6 is a well-documented model. On our catalogue it carries a 500,000-token context window, text, image and file input, and a rate card of $2.00 per million input tokens, $6.00 per million output tokens and $0.50 per million cached input tokens, stepping to $4.00 and $12.00 above 200,000 input tokens. Its independent profile from Artificial Analysis includes a Coding Index of 76.8 — fifth among the models that index tracks — an Intelligence Index of 44.3 at eighteenth, GPQA Diamond at 94.9%, Humanity's Last Exam at 42.9%, SciCode at 56.5%, Terminal-Bench v2.1 at 88.4 and τ²-Bench for banking at 50.7. Its Long-Context Recall is 80.3.

A screenshot of Meituan's LongCat API Platform change-log page, headlined 'Version: 2026-09-25 - LongCat-2.5-Preview Now Available', listing the model's three stated features - multimodal image understanding, coding capability, and compatibility with Claude Code, Hermes, OpenClaw, OpenCode and Kilo Code - with the platform's Guide, API, Tools and Pricing navigation and its column of earlier dated releases visible.

LongCat-2.5-Preview has none of that. Meituan's changelog entry for 25 September 2026 is a dated availability notice listing three claimed capabilities — image understanding, coding, and compatibility with "Claude Code and other mainstream dev environments" including Hermes, OpenClaw, OpenCode and Kilo Code — and nothing that resembles a result. The vendor's pricing page gives $0.30 per million uncached input, $0.006 per million cached input and $1.20 per million output, explicitly flagged as a limited-time discount. The context window is 1,000,000 tokens; maximum output is 131,072. The roughly 1.6-trillion-total / 48-billion-active parameter count comes from Meituan's site metadata and Chinese trade coverage rather than documentation. There is no repository for it on HuggingFace or in Meituan's GitHub organisation, and the catalogue metadata OpenCode ships records open weights as false.

The dimension a scoreboard would miss entirely

These two models differ on something no benchmark measures, and for a lot of teams it decides the question on its own: what happens to the prompts you send.

OpenCode's own documentation states that LongCat 2.5 Preview Free is served by a provider that follows a zero-retention policy and does not use your data for training; the Zen privacy table puts numbers on it — model training "Not used", data retention "0 days". The same privacy table lists thirty days of retention for Grok 4.6 and Grok 4.7. Both of those statements are OpenCode's, published on OpenCode's pages, and they describe the free listing and the comparison listing specifically. But if you are pointing an agent at a private repository, that is the difference between a conversation you do not have to have with legal and one you do — and it has nothing to do with which model scores better on SciCode.

A screenshot of OrcaRouter's own model page for xAI Grok 4.6 at /models/grok/grok-4.6, showing the grok/grok-4.6 identifier, a 500K-token context window, text + image + file input and text output, Vision, Tools, JSON and Reasoning capability chips, a release date of 2026-08-12, $2.00 and $6.00 rate tiles, and a performance strip whose first-token and token-traffic tiles both read 'Collecting' rather than a figure. The page copy describes Grok 4.6 as the current flagship of the Grok line and cites a headline GPQA Diamond score of 94.9.

What "measured" does and does not buy you

Grok 4.6's numbers are better than LongCat-2.5-Preview's in the only sense that matters here: they exist. That is not the same as saying Grok 4.6 is the better model. Its Coding Index of 76.8 is below Grok 4.7's generation, and the model it was measured against is no longer the frontier of its own family. Its published successor has an Intelligence Index of 46.4 against Grok 4.6's 44.3, and a Long-Context Recall of 76.67 against Grok 4.6's 80.33 — so the successor is not better on every axis, which is exactly the kind of detail a generation number hides.

There is also a live-serving caveat worth stating plainly, because it cuts against the flattering reading for both models. On our own seven-day playground sample, Grok 4.6 recorded no measured throughput and no token traffic, which is what an absence of data looks like in that column rather than a performance verdict. LongCat-2.5-Preview has no serving sample of ours at all, because we do not route it. So any latency comparison between these two is one-sided in a way that prose usually papers over: you cannot benchmark a model you are not sending traffic to.

Where the routing layer actually helps here

Three of the facts above are the kind that a router is built for rather than merely compatible with. Grok 4.6's rate card steps up above 200,000 input tokens, so the same logical task can be billed at two different rates depending on a number your application already knows before it sends anything. Grok 4.6 has a published successor at identical rates, which means the switch from one to the other should be a one-line change and never a re-integration. And LongCat-2.5-Preview's most attractive property — its price — is explicitly labelled temporary by the vendor.

On OrcaRouter we serve Grok 4.6 at xAI's list price with zero markup, so a vendor's rate change reaches your bill the same day rather than at the next renewal, and automatic failover moves a degraded request to a healthy provider without your code knowing which one answered. A routing DSL covers the tiered-pricing case directly, and model fusion covers the case where the model you want for the long read is not the model you want for the final synthesis. LongCat-2.5-Preview is not one of our routes — we do not serve it, and nothing here claims otherwise. The point of naming that is not to route you elsewhere; it is that a model you are evaluating rather than running belongs behind the same key as the models you run, so that evaluating it costs a config entry instead of a project.

A generated two-column scoreboard titled 'LongCat-2.5-Preview vs Grok 4.6' with the subhead 'One model has a measured card and a measured successor; the other has a rate card'. Left column 'LongCat-2.5-Preview' (Meituan, listed 2026-09-25, no benchmark published): 1,000,000-token context, 131,072 max output, text with image input claimed not confirmed, $0.30 uncached / $0.006 cached input, $1.20 output flagged limited-time, coding index not published, long-context recall not published, data retention 0 days on the free listing per OpenCode. Right column 'Grok 4.6' (xAI, released 2026-08-12, successor shipped 2026-09-21): 500,000-token context, max output not published, text, image and file to text, $2.00 input up to 200K then $4.00, $6.00 output up to 200K then $12.00, coding index 76.8 at rank 5 by Artificial Analysis, long-context recall 80.33, data retention 30 days on the comparison listing per OpenCode. Footnote credits the left column to Meituan's changelog and pricing page and the right column to OrcaRouter's catalogue with index figures sourced to Artificial Analysis, and notes that LongCat-2.5-Preview is not routed by OrcaRouter. OrcaRouter logo bottom-right.

How to decide

• Choose Grok 4.6 if you need an answer you can defend. It has an independent card, a documented input profile, and a price that does not expire. Its main risk is not quality but relevance: Grok 4.7 exists at the same rates, and the cost of staying on 4.6 out of inertia is the difference between an Intelligence Index of 44.3 and 46.4 that you did not have to pay for.

• Choose LongCat-2.5-Preview if the deciding factor is retention or reach. Zero-retention access through OpenCode, a one-million-token window, a 131,072-token output ceiling and a rate card roughly a seventh of Grok 4.6's are real advantages — the first of them verified by a third party's privacy policy, the rest asserted by the vendor alone.

• Do not choose between them on the strength of a benchmark table, because only one exists. Grok 4.6's 76.8 coding score and 80.3 long-context recall describe Grok 4.6. Nothing describes LongCat-2.5-Preview yet. Anyone printing a LongCat-2.5-Preview score beside a Grok 4.6 score this week is either quoting a different LongCat model or inventing the number.

• Verify the input side before you build on it. Meituan's changelog leads with image understanding, but the example response in its own "Retrieve Model" documentation still shows input_modalities as ["text"] with a text->text modality string — vendor-claimed against a published contract that says otherwise. Grok 4.6 takes text, images and files with no such ambiguity. And check that your harness surfaces an interleaved reasoning_content field before you conclude anything about LongCat-2.5-Preview's reasoning quality; a client that drops that field will fail quietly.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily