
Claude Sonnet 5.5 vs MiniMax M3: Where the 15× Cheaper Model Actually Ties
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 931 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 192 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1177 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 70 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 107 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
There is one row on the independent board where MiniMax M3 and Claude Sonnet 5.5 are the same model, and it is not the one anyone would guess. On long-context recall, MiniMax M3 scores 83.0% and Claude Sonnet 5.5 scores 82.7%. MiniMax M3 costs $0.30 per million input tokens and $1.20 per million output. Claude Sonnet 5.5 costs $2.00 and $10.00. Across the whole Intelligence Index, M3's measured cost is $0.51 per task against $7.60 for Claude Sonnet 5.5 — fifteen times cheaper, and on the one evaluation that measures whether a model can still find something in a very long document, it is not behind at all.
Then you open the agentic evaluations and the picture inverts so hard it stops being a comparison. On Terminal-Bench 4.0, which asks a model to drive a shell and actually finish a software task, MiniMax M3 scores 2.0% and Claude Sonnet 5.5 scores 63.6%. A thirty-fold gap on the one benchmark that most closely resembles production work is not a price question, and no per-token saving survives it. The useful question this matchup raises is not "which is better" but which of these two failure modes your workload can absorb: MiniMax M3 is the open-weight, million-token, native-video model that matches a frontier model on retrieval and collapses on execution, and Claude Sonnet 5.5 is the closed model that shipped on September 28, 2026 to do the opposite.
What each one is, stated plainly
MiniMax M3 is the Chinese lab's flagship open-weight foundation model, announced June 1, 2026 and downloadable under MiniMax's own community licence. It is a mixture of experts at roughly 428 billion total parameters with about 23 billion active per token, and it is natively multimodal in the strong sense: text, image and video go in and text comes out, with the multimodal pipeline rebuilt from the first pretraining step rather than bolted on. The window is 1,048,576 tokens — with a guaranteed usable minimum of 512,000 in our own catalogue entry — and the maximum output is 512,000 tokens, which is our figure rather than the vendor's, because MiniMax does not publish one. The architecture is MiniMax Sparse Attention, the subject of arXiv 2606.13392, and the vendor's claim for it is more than 9× faster prefilling and more than 15× faster decoding than the previous generation at a million-token window. Those are vendor numbers and we have not seen them independently reproduced.

Claude Sonnet 5.5 is Anthropic's second 5.5-generation model, one day old at the time of writing, and it is closed. Anthropic describes it as the best combination of speed and intelligence in the lineup, and the specs are 1,000,000 tokens of context, 128,000 tokens of synchronous output with 300,000 on the Message Batches API, text, image and file input, adaptive thinking with selectable effort, a June 2026 knowledge cutoff, and zero data retention available from launch. Its price did not move from Claude Sonnet 5: $2.00 input, $10.00 output, $0.20 for cache reads, $2.50 for cache writes. Anthropic says it runs 30% faster and costs up to 30% less per task than Claude Sonnet 5 — vendor figures, unreproduced, and to be read as claims about token counts rather than rates, since the rates are identical.
The benchmark shape is the whole story
Both models are measured on the same index, revision v4.3.2, and the per-evaluation results are what makes this pairing interesting rather than lopsided.
• Long-context recall — MiniMax M3 83.0% vs Claude Sonnet 5.5 82.7%, a rounding-error difference on a genuine tie
• SciCode — MiniMax M3 47.1% vs Claude Sonnet 5.5 61.0%, a real but survivable gap
• Humanity's Last Exam — MiniMax M3 39.0% vs Claude Sonnet 5.5 55.0%
• Terminal-Bench 4.0 — MiniMax M3 2.0% vs Claude Sonnet 5.5 63.6%, where the comparison stops being useful
• Intelligence Index — MiniMax M3 29 vs Claude Sonnet 5.5 56
• Cost per Index task — MiniMax M3 $0.51 vs Claude Sonnet 5.5 $7.60
• Output tokens generated across the Index — MiniMax M3 120M vs Claude Sonnet 5.5 410M
• Output speed — MiniMax M3 115.3 tokens/sec vs Claude Sonnet 5.5 138.7 tokens/sec
Read those rows in order and a coherent model emerges. MiniMax M3 is competent at static reasoning and retrieval and weak at sustained execution. Its Terminal-Bench result is not a modest shortfall; a 2.0% score means the model essentially does not complete the tasks, and the same pattern shows in its automation-benchmark score of 0.21 against 0.71, and in an omniscience score of 1.35 against 32.3. Where MiniMax M3 does hold its ground is exactly where its architecture claims an advantage: a genuinely long input, read and answered. That is MSA doing what MiniMax sold it for.

The cost line adds a second, less obvious point. MiniMax M3 is not only cheaper per token — 1/6.7 on input and 1/8.3 on output — it also spends fewer tokens reaching an answer, roughly 120 million against 410 million across the same board. Two effects multiply into the fifteen-fold per-task gap. Some of that gap is genuine efficiency and some of it is simply that MiniMax M3 gives up sooner on tasks it is failing; a cheap task that returns the wrong answer is not a saving, and this is the cheapest lesson in the article.
The dimension the price list cannot show
MiniMax M3 has a property Claude Sonnet 5.5 does not and cannot acquire: you can take the weights. That matters for exactly one class of buyer, and for that buyer it outweighs everything above. If a request cannot leave a jurisdiction, cannot cross a network boundary, or has to run where no API is reachable at all, a model with no downloadable weights is not an option at any price, and a 428-billion-parameter open-weight model is. The same property is a liability in the other direction: self-hosting 428B parameters at 23B active is a capital decision, not an API key, and MiniMax's own sparse-attention claims exist because the alternative at a million tokens is unaffordable infrastructure.
For everyone else the honest framing is that this is a build-versus-buy line with a price tag attached. Claude Sonnet 5.5 is the better model for anything agentic. MiniMax M3 is the better model for reading something enormous and answering a question about it, and it is dramatically the better model for processing video, which Claude Sonnet 5.5 does not accept at all — its input surface is text, images and files.
• Weights — MiniMax M3 open under MiniMax's community licence vs Claude Sonnet 5.5 closed
• Input modality — MiniMax M3 text, image and video vs Claude Sonnet 5.5 text, image and file
• Maximum output — MiniMax M3 512,000 tokens per our catalogue vs Claude Sonnet 5.5 128,000 synchronous, 300,000 on Batches
• Cache pricing — MiniMax M3 $0.06 per million read vs Claude Sonnet 5.5 $0.20 read and $2.50 write
• Effort control — MiniMax M3 thinking enabled, adaptive or disabled vs Claude Sonnet 5.5 adaptive thinking where disabling now requires the between_tools form
• Data retention — MiniMax M3 governed by your own deployment if self-hosted vs Claude Sonnet 5.5 zero data retention available from Anthropic at launch
Running both without running two contracts
Mix an open-weight Chinese model and a closed Anthropic model in one pipeline and the operational cost is usually not the tokens; it is the second contract, the second key, the second set of rate limits, and the second place a failure can happen. OrcaRouter collapses that into one OpenAI-compatible endpoint covering 200-plus models, with provider list price passed through unchanged — no markup added to either side — automatic failover between upstream providers, and a routing DSL for composing several models into a single call. MiniMax M3 is routable here as minimax/minimax-m3 at MiniMax's $0.30 and $1.20 with $0.06 cache reads, and it is one of the busiest models in the catalogue by traffic, which is a useful signal in itself: the routing layer can tell you how much real load a model is carrying, not just how it scored.
Claude Sonnet 5.5 is not in our catalogue yet, and the piece will not pretend otherwise. The route to it today is Anthropic's own API. Claude Sonnet 5 is routable here at the same $2.00 and $10.00 and has been since June 30, and it is the natural staging target and failover partner while its successor is a day old.
That combination — the long-context swap and the agentic path behind one credential — is what a router is for. MiniMax M3 carries 63.2 million tokens of traffic a week through this catalogue against Claude Sonnet 5's 7.9 million, so the cheap model is not a theoretical substitute here; it is already carrying the volume. Routing both from one key means the split is a configuration decision you can move traffic across, rather than a second contract you have to sign before you can test the cheaper side.

What to do with the tie
If your workload is document-scale question answering — a very long input, a short factual answer, latency that is tolerable — the long-context tie is real and MiniMax M3's fifteen-fold cost advantage is real with it. That is a straightforward swap and worth testing this week.
If your workload is agentic, the Terminal-Bench line settles it before price enters the conversation. Claude Sonnet 5.5 at $7.60 per Index task against MiniMax M3 at $0.51 is a 15× premium for a model that scores sixty points higher on the evaluation most like your production traffic, and the fifteen-times-cheaper option that fails the task is the expensive one. The measurement to run on your own data is tokens per completed task, not tokens per request, because that is the only figure in this article that MiniMax M3's price advantage does not survive by default.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
