A generated hero title card reading 'LongCat-2.5-Preview vs MiniMax M3' with the overline 'The cheap seat is already taken', and two panels: 'LongCat-2.5-Preview: $0.30 / $1.20 per 1M, 131,072 max output' and 'MiniMax M3: $0.30 / $1.20 per 1M, 512,000 max output'. OrcaRouter logo bottom-right.
Guides & Insights

LongCat-2.5-Preview vs MiniMax M3: The Cheap Seat Is Already Taken

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The premise of most LongCat-2.5-Preview coverage is that Meituan has undercut the field. Against MiniMax M3, that premise evaporates on contact. MiniMax M3 lists at $0.30 per million input tokens and $1.20 per million output tokens on our catalogue. LongCat-2.5-Preview's published rate card is $0.30 per million uncached input tokens and $1.20 per million output tokens. Those are not similar figures; they are the same figures. The one place they diverge is cached input, where Meituan quotes $0.006 per million against the incumbent's $0.06 — a tenfold gap on the cheapest line of the bill. So the interesting question is not whether LongCat-2.5-Preview is cheap. It is what you are buying and giving up when a new model arrives at exactly the incumbent's price.

The short answer: a much larger output ceiling, a much thinner evidence base, and a cache discount.

Same headline price, very different published ceilings

MiniMax M3 has been on the record since 31 May 2026. Our catalogue carries it with a 1,048,576-token context window and a maximum output of 512,000 tokens — more than most models allow, and four times LongCat-2.5-Preview's 131,072. It accepts text, images and video as input and returns text, with tool calling, JSON mode and reasoning exposed.

A screenshot of Meituan's LongCat API Platform change-log page, headlined 'Version: 2026-09-25 - LongCat-2.5-Preview Now Available', listing the model's three stated features - multimodal image understanding, coding capability, and compatibility with Claude Code, Hermes, OpenClaw, OpenCode and Kilo Code - with the platform's Guide, API, Tools and Pricing navigation and its column of earlier dated releases visible.

LongCat-2.5-Preview, listed on Meituan's LongCat API Platform on 25 September 2026, matches on context at 1,000,000 tokens and undershoots badly on output at 131,072. Its input profile is the live question: the changelog presents image understanding as its headline addition, but the example response in Meituan's own "Retrieve Model" documentation still shows input_modalities/ as ["text"]/ with a text->text/ modality string. MiniMax M3 also accepts video, which LongCat-2.5-Preview does not claim at all. If your pipeline feeds screen recordings or frame sequences, this comparison ends here.

The cache asymmetry runs the other way and is worth stating in Meituan's favour, because it is the one dimension where the new model genuinely differentiates. $0.006 per million cached input against $0.06 is a real advantage for agent workloads that resend the same long prefix on every turn — which is most of them. It is also the line item most likely to change without notice, since Meituan flags the whole card as a limited-time discount.

The evidence gap, measured honestly in both directions

MiniMax M3's benchmark position is not strong, and saying so is part of doing this comparison properly. Artificial Analysis records a Coding Index of 58.6 for it — forty-sixth of the models tracked — and an Intelligence Index of 29.2 at sixtieth. GPQA Diamond 92.9%, Humanity's Last Exam 39%, SciCode 47.1%, Long-Context Recall 83, τ²-Bench banking 15.3, Terminal-Bench v2.1 65.2. MiniMax's own published figures cover different ground: BrowseComp 83.5, GDPval rubrics 76.7, MCP Atlas 74.2, BankerToolBench 76.1. Those are vendor-reported numbers and belong in a different column from the independent ones, but they exist, which is the operative word.

A screenshot of OrcaRouter's own model page for MiniMax M3 at /models/minimax/minimax-m3, showing the minimax/minimax-m3 identifier, a 1M-token context window, 512K maximum output, text + image + video input and text output, Vision, Tools, JSON and Reasoning capability chips, a release date of 2026-05-31, a p50 first-token figure of 3.35 s, $0.30 and $1.20 rate tiles, and the OpenAI-compatible code samples.

LongCat-2.5-Preview has no column at all. No published benchmark table, no model card, no technical report, no repository on HuggingFace or in Meituan's GitHub organisation, and catalogue metadata recording open weights as false. The approximately 1.6-trillion-total / 48-billion-active parameter count is reported from Meituan's site metadata and Chinese trade coverage rather than from documentation. So the accurate statement is: M3's scores are mediocre and independently sourced; LongCat-2.5-Preview's are absent. A model that scores in the forties on the coding index is a known quantity you can plan around. A model with no score is a different kind of risk, and the price being identical removes the compensation that usually justifies taking it.

Where the same price changes what you should do

When a new model undercuts the field, the rational move is to evaluate it — the upside is a real discount. When a new model arrives at exactly the same price as an incumbent that has been shipping since May, the upside is not price. It is capability, and capability is the one thing LongCat-2.5-Preview has not published. That reframes the evaluation entirely: you are not testing whether it is cheaper, you are testing whether it is better, on your tasks, from a standing start with no evidence to anchor on.

There is a second-order detail that makes that test cheaper than it looks. LongCat-2.5-Preview is free and unlimited through OpenCode for a window of unstated length, with a zero-retention policy OpenCode documents explicitly — model training "Not used", data retention "0 days", against thirty days for the comparison listing. So the cost of generating your own evaluation is close to zero, and the retention policy means you can run it against private code. The one harness detail to settle first is that the model returns its reasoning trace in an interleaved reasoning_content/ field; clients that parse chat completions without expecting it will drop the thinking silently and make the model look worse than it is.

A generated two-column scoreboard titled 'LongCat-2.5-Preview vs MiniMax M3' with the subhead 'Identical headline rates, four times the output ceiling, ten times the cache discount'. Left column 'LongCat-2.5-Preview' (Meituan, listed 2026-09-25, no benchmark published): 1,000,000-token context, 131,072 max output, text with image input claimed not confirmed, $0.30 uncached / $0.006 cached input, $1.20 output flagged limited-time, coding index not published, long-context recall not published, no repo and catalogue open_weights false. Right column 'MiniMax M3' (MiniMax, released 2026-05-31, independently scored): 1,048,576-token context, 512,000 max output described as four times as much, text, image and video to text, $0.30 uncached / $0.06 cached input, $1.20 output, coding index 58.6 at rank 46 and intelligence index 29.2 at rank 60 by Artificial Analysis, long-context recall 83. Footnote credits the left column to Meituan's changelog and pricing page and the right column to OrcaRouter's catalogue with index figures sourced to Artificial Analysis, and notes that LongCat-2.5-Preview is not routed by OrcaRouter. OrcaRouter logo bottom-right.

The routing argument, specific to a same-price comparison

We serve MiniMax M3 at MiniMax's list price with zero markup. LongCat-2.5-Preview is not one of our routes — we do not serve it, and nothing here should be read as claiming otherwise.

A same-price matchup is exactly the case where routing earns its keep rather than being a convenience. If two models cost the same, the decision between them is per-request and per-task: M3 for the 512,000-token output ceiling and video input, LongCat-2.5-Preview for the cached-prefix discount on long agent loops once you have satisfied yourself about its quality. Model fusion is the mechanism that makes that literal — a retrieval pass and a synthesis step can be different models on a single request, so you are not forced to pick a winner before you have evidence. And because we pass provider list prices through with no markup, a vendor's rate change lands on your bill the same day instead of at the next contract renewal, which matters more than usual when one of the two cards is explicitly promotional. Automatic failover then covers the case where one provider degrades, and it matters most for the model with no serving history to inspect.

Frequently asked

• Is LongCat-2.5-Preview cheaper than MiniMax M3? No. On published cards the input and output rates are identical at $0.30 and $1.20 per million. The only price advantage is on cached input, $0.006 against $0.06, and Meituan labels its whole card a limited-time discount without saying what follows.

• Which has the bigger usable output? MiniMax M3 by a wide margin: 512,000 tokens against 131,072. For long-form generation, structured extraction over large documents, or anything that writes more than a chapter, that ceiling is the deciding spec.

• Which one can I verify? MiniMax M3, and the verification does not flatter it — Coding 58.6 at forty-sixth and an Intelligence Index of 29.2 at sixtieth from Artificial Analysis. LongCat-2.5-Preview has no published evaluation from anyone. If a LongCat-2.5-Preview benchmark appears this week, treat it as unverifiable until you can trace it to a named evaluator.

• Which should I put in production? Neither without a test of your own. Between the two, M3 is the lower-variance choice because its weaknesses are documented and its input modalities are confirmed; LongCat-2.5-Preview is the higher-variance one, with a larger context claim, a much smaller output ceiling, a cache discount, and zero public evidence. Run it free while the window is open, keep both behind one key, and let your own numbers make the call.