
LongCat-2.5-Preview vs MiniMax M3: The Cheap Seat Is Already Taken
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 610 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 189 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1306 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 225 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
The premise of most LongCat-2.5-Preview coverage is that Meituan has undercut the field. Against MiniMax M3, that premise evaporates on contact. MiniMax M3 lists at $0.30 per million input tokens and $1.20 per million output tokens on our catalogue. LongCat-2.5-Preview's published rate card is $0.30 per million uncached input tokens and $1.20 per million output tokens. Those are not similar figures; they are the same figures. The one place they diverge is cached input, where Meituan quotes $0.006 per million against the incumbent's $0.06 — a tenfold gap on the cheapest line of the bill. So the interesting question is not whether LongCat-2.5-Preview is cheap. It is what you are buying and giving up when a new model arrives at exactly the incumbent's price.
The short answer: a much larger output ceiling, a much thinner evidence base, and a cache discount.
Same headline price, very different published ceilings
MiniMax M3 has been on the record since 31 May 2026. Our catalogue carries it with a 1,048,576-token context window and a maximum output of 512,000 tokens — more than most models allow, and four times LongCat-2.5-Preview's 131,072. It accepts text, images and video as input and returns text, with tool calling, JSON mode and reasoning exposed.

LongCat-2.5-Preview, listed on Meituan's LongCat API Platform on 25 September 2026, matches on context at 1,000,000 tokens and undershoots badly on output at 131,072. Its input profile is the live question: the changelog presents image understanding as its headline addition, but the example response in Meituan's own "Retrieve Model" documentation still shows input_modalities/ as ["text"]/ with a text->text/ modality string. MiniMax M3 also accepts video, which LongCat-2.5-Preview does not claim at all. If your pipeline feeds screen recordings or frame sequences, this comparison ends here.
The cache asymmetry runs the other way and is worth stating in Meituan's favour, because it is the one dimension where the new model genuinely differentiates. $0.006 per million cached input against $0.06 is a real advantage for agent workloads that resend the same long prefix on every turn — which is most of them. It is also the line item most likely to change without notice, since Meituan flags the whole card as a limited-time discount.
The evidence gap, measured honestly in both directions
MiniMax M3's benchmark position is not strong, and saying so is part of doing this comparison properly. Artificial Analysis records a Coding Index of 58.6 for it — forty-sixth of the models tracked — and an Intelligence Index of 29.2 at sixtieth. GPQA Diamond 92.9%, Humanity's Last Exam 39%, SciCode 47.1%, Long-Context Recall 83, τ²-Bench banking 15.3, Terminal-Bench v2.1 65.2. MiniMax's own published figures cover different ground: BrowseComp 83.5, GDPval rubrics 76.7, MCP Atlas 74.2, BankerToolBench 76.1. Those are vendor-reported numbers and belong in a different column from the independent ones, but they exist, which is the operative word.

LongCat-2.5-Preview has no column at all. No published benchmark table, no model card, no technical report, no repository on HuggingFace or in Meituan's GitHub organisation, and catalogue metadata recording open weights as false. The approximately 1.6-trillion-total / 48-billion-active parameter count is reported from Meituan's site metadata and Chinese trade coverage rather than from documentation. So the accurate statement is: M3's scores are mediocre and independently sourced; LongCat-2.5-Preview's are absent. A model that scores in the forties on the coding index is a known quantity you can plan around. A model with no score is a different kind of risk, and the price being identical removes the compensation that usually justifies taking it.
Where the same price changes what you should do
When a new model undercuts the field, the rational move is to evaluate it — the upside is a real discount. When a new model arrives at exactly the same price as an incumbent that has been shipping since May, the upside is not price. It is capability, and capability is the one thing LongCat-2.5-Preview has not published. That reframes the evaluation entirely: you are not testing whether it is cheaper, you are testing whether it is better, on your tasks, from a standing start with no evidence to anchor on.
There is a second-order detail that makes that test cheaper than it looks. LongCat-2.5-Preview is free and unlimited through OpenCode for a window of unstated length, with a zero-retention policy OpenCode documents explicitly — model training "Not used", data retention "0 days", against thirty days for the comparison listing. So the cost of generating your own evaluation is close to zero, and the retention policy means you can run it against private code. The one harness detail to settle first is that the model returns its reasoning trace in an interleaved reasoning_content/ field; clients that parse chat completions without expecting it will drop the thinking silently and make the model look worse than it is.

The routing argument, specific to a same-price comparison
We serve MiniMax M3 at MiniMax's list price with zero markup. LongCat-2.5-Preview is not one of our routes — we do not serve it, and nothing here should be read as claiming otherwise.
A same-price matchup is exactly the case where routing earns its keep rather than being a convenience. If two models cost the same, the decision between them is per-request and per-task: M3 for the 512,000-token output ceiling and video input, LongCat-2.5-Preview for the cached-prefix discount on long agent loops once you have satisfied yourself about its quality. Model fusion is the mechanism that makes that literal — a retrieval pass and a synthesis step can be different models on a single request, so you are not forced to pick a winner before you have evidence. And because we pass provider list prices through with no markup, a vendor's rate change lands on your bill the same day instead of at the next contract renewal, which matters more than usual when one of the two cards is explicitly promotional. Automatic failover then covers the case where one provider degrades, and it matters most for the model with no serving history to inspect.
Frequently asked
• Is LongCat-2.5-Preview cheaper than MiniMax M3? No. On published cards the input and output rates are identical at $0.30 and $1.20 per million. The only price advantage is on cached input, $0.006 against $0.06, and Meituan labels its whole card a limited-time discount without saying what follows.
• Which has the bigger usable output? MiniMax M3 by a wide margin: 512,000 tokens against 131,072. For long-form generation, structured extraction over large documents, or anything that writes more than a chapter, that ceiling is the deciding spec.
• Which one can I verify? MiniMax M3, and the verification does not flatter it — Coding 58.6 at forty-sixth and an Intelligence Index of 29.2 at sixtieth from Artificial Analysis. LongCat-2.5-Preview has no published evaluation from anyone. If a LongCat-2.5-Preview benchmark appears this week, treat it as unverifiable until you can trace it to a named evaluator.
• Which should I put in production? Neither without a test of your own. Between the two, M3 is the lower-variance choice because its weaknesses are documented and its input modalities are confirmed; LongCat-2.5-Preview is the higher-variance one, with a larger context claim, a much smaller output ceiling, a cache discount, and zero public evidence. Run it free while the window is open, keep both behind one key, and let your own numbers make the call.
