
Mistral Large 4 vs MiniMax M3: What Five Months of Progress Costs
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 151 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 116 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 249 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
MiniMax M3 is the cheapest model in this class and it has been since 31 May 2026. At $0.30 per million input tokens, $1.20 per million output and $0.06 per million cached, it undercuts Mistral Large 4 by roughly half on input and twice on output — and unlike the newer model, you can buy it today without a preview label attached. Mistral Large 4, a 1.05-trillion-parameter open-weight model in public preview since 6 October 2026, costs $0.68 and $2.09 and promises weights by the end of the month. Two open-weight flagships, five months apart, and the five months are visible in exactly one place: M3's independent Intelligence Index score of 29.2 against a Mistral Large 4 score that does not exist yet.
That is the honest shape of this comparison, and it is more interesting than a price table suggests. M3 was built around a specific architectural bet — MiniMax Sparse Attention, sustained over a million tokens of context, with video as a first-class input — and it wins on paper in two categories Mistral Large 4 does not contest: raw output length and native video. Everywhere else, the newer model is the one a buyer would rather have, at a price that is still low enough that the difference rarely decides the purchase.
Two architectures, two bets
MiniMax M3 is MiniMax's flagship open-weight foundation model and the first to combine frontier coding and agentic performance, a million-token context window and native multimodality in one release. It accepts text, images and video and returns text, and it is built on MiniMax Sparse Attention, an architecture designed to sustain up to 1,048,576 tokens of context with a guaranteed floor of 512K. The pretraining pipeline was rebuilt to scale past 100 trillion tokens with multimodal data from the first step rather than bolted on later. Its maximum output is 512,000 tokens.
Mistral Large 4 makes a different bet. It is a granular mixture-of-experts model, 49 billion active parameters out of 1.05 trillion total, with a 1.6-billion-parameter vision encoder, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs inside Mistral's own European datacentres on data spanning more than 160 languages. It takes text and images — no video — across about a million tokens of context, and Mistral positions it around sovereign deployment: European infrastructure, open weights, and the ability to run it under your own policies, with cybersecurity as the flagship use case.
• Context window — MiniMax M3 1,048,576 tokens vs Mistral Large 4 about 1,000,000
• Max output — MiniMax M3 512,000 tokens vs Mistral Large 4 not yet documented
• Input modality — MiniMax M3 text, image and video vs Mistral Large 4 text and image
• Input price — MiniMax M3 $0.30 per 1M vs Mistral Large 4 $0.68 per 1M
• Output price — MiniMax M3 $1.20 per 1M vs Mistral Large 4 $2.09 per 1M
• Cached input — MiniMax M3 $0.06 per 1M vs Mistral Large 4 $0.07 per 1M
• Parameters — Mistral Large 4 1.05T total / 49B active vs MiniMax M3 undisclosed
• Released — MiniMax M3 31 May 2026 vs Mistral Large 4 6 October 2026, preview
• Weights — both open licence, Mistral Large 4's promised by end of October 2026

The scoreboard is not close, and it is not fair either
On the Artificial Analysis Intelligence Index v4.3.2, MiniMax M3 scores 29.2 and sits 62nd of 147 models. That is a mid-table result and it is not a criticism of the model — the Index was revised after M3's release and includes evaluations that did not exist when MiniMax was training. Its coding sub-score tells a better story: 58.6 on the AA Coding index, 46th of 138, and 65.2% on Terminal-Bench 2.1. It holds 92.9% on GPQA Diamond, 83% on long-context recall, and 47.1% on SciCode.
Mistral Large 4 has no Index score on the same board; the page is live under the title "Mistral Large 4 Preview" and the number is absent. What Mistral does publish is that ML4 leads DeepSeek V4 Pro 0813 and Qwen3.8 Max on the combined Coding Agent Index at 49.8%, and that on AutomationBench — 657 business workflows across Gmail, Sheets, Slack and Salesforce — it scores 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro. MiniMax is not named in either comparison, which is itself a signal about which model Mistral considers the real competitor.
So the fair reading is this: M3 is a capable, cheap, well-documented model with a mid-table general score and a respectable coding profile, five months old. ML4 is a preview model with a higher claimed ceiling, no published general score, and a set of agentic figures that Artificial Analysis evaluated privately and will publish on the Coding Agent Index when that harness ships. If you need a number today, M3 gives you one. If you can wait a few weeks, ML4 will give you one too, and Mistral is betting it will be higher.
Where M3 has no competition from ML4 at all
Two rows on that list are not a trade-off, they are an absence.
Video input is the first. M3 accepts video natively, alongside text and images. Mistral Large 4 takes text and images and nothing else — Mistral's own deep-dive is about grounding over gigapixel satellite imagery and engineering drawings, which is still image work. If your pipeline ingests screen recordings, surveillance footage or video documentation, ML4 cannot be the model, whatever its price. That is not a gap Mistral will close with a reprice.
Output length is the second. M3 emits up to 512,000 tokens in a single response. Mistral has published no output ceiling for ML4. Generating a long structured artefact in one pass — a full report, a large migration script, a complete spec — is the workload where a 512K ceiling is worth more than an Intelligence Index point, and until Mistral documents ML4's limit, M3 is the only one of the two you can plan that workload around.
M3's vendor-reported agentic numbers reinforce the same profile. MiniMax puts it at 83.5 on BrowseComp for autonomous web research, 70 on OSWorld-verified for computer use, 74.2 on MCP Atlas for tool use, 76.7 on GDPval rubrics and 59 on SWE Bench Pro. Those come from MiniMax's own evaluations and have not been independently reproduced, and they sit oddly against its 29.2 Index score — which is precisely why the vendor-versus-independent distinction matters. A model can lead a vendor's chosen benchmarks and land mid-table on a neutral ten-evaluation average, and both things can be true.
Is the 2x worth it
The price gap here is genuinely small in absolute terms, which changes the calculus compared with a mismatch against a closed flagship. Take a hundred-million-token month at a 4:1 input/output split: 80 million input tokens and 20 million output. M3 bills $24 plus $24, a total of $48. Mistral Large 4 bills $54.40 plus $41.80, a total of $96.20. The premium is about $48 per month — real money at scale, but the kind of gap that a single avoided failure pays for.
Caching narrows it further and in M3's favour only marginally: $0.06 against $0.07 per million cached tokens is a rounding error at this price point. Both are cheap enough that the decision should be made on capability and fit rather than on the invoice, which is the opposite of how this comparison reads at first glance.
What the premium buys is span. ML4 was trained five months later on newer data, in a mixture-of-experts layout at a trillion parameters with 49 billion active, and Mistral is running reinforcement learning on it at a stated 33 billion tokens per day with a run that has not saturated. M3 is finished; ML4 is improving on a published cadence. When a vendor says the model you are buying today is the weakest version you will ever pay for, and the premium over the incumbent is under a dollar per million tokens, the newer model is usually the better default. The exception is the two workloads above, where the older model is the only one that fits.
Routing round the gap

This is a pairing where running both is not a hedge but the actual answer, because the two models are strong in disjoint places. Video in, long artefact out, or a computer-use loop: M3. Document and image analysis, agentic workflows, or anything that needs to run on European infrastructure under your own policies: ML4. There is no routing rule that makes one of them redundant.
Both sit behind one OpenAI-compatible endpoint on OrcaRouter with 0% markup and provider list prices passed through unmodified, which keeps a five-month price spread honest — if MiniMax cuts M3's rate or Mistral ends the preview rate, the number your route is evaluated against changes the same day. The routing DSL is what turns the disjoint strengths into one call surface: a rule that inspects the request's modality sends video-bearing payloads to M3 and everything else to ML4, with the confidence that a preview model's availability blips are absorbed by automatic failover rather than propagated to your users. Where a single output matters more than a single price, model fusion can run both and let a panel pick, which is a reasonable way to buy Mistral's claimed ceiling while keeping M3's documented behaviour as a floor.

Verdict
MiniMax M3 is the right answer for two workloads and a defensible answer for a third. Video input, multi-hundred-thousand-token single-pass generation and computer-use agents are its territory, its price is the lowest of any flagship in this class, and its published mid-table scores mean you know exactly what you are getting. Five months old is not old for a model that has not been superseded in its own niche.
Mistral Large 4 is the better default for everything else. A hundred-million-token month costs about $48 more, the parameter count and European training lineage point at a higher ceiling, the cybersecurity claims are specific enough to be checkable, and the promised open weights plus on-premises deployment settle requirements M3's API-only enterprise story does not. The price of that bet is that you are committing to a model whose Independent Intelligence Index score has not been published and whose behaviour Mistral says will change substantially within weeks.
The two dates that resolve this: the end of October 2026, when ML4's weights either land or the open-weights promise starts to look soft, and the publication of the Coding Agent Index, where Mistral's privately-evaluated 49.8% meets M3's public 58.6 coding score on the same revision. Until then, M3 is the model you can verify and ML4 is the model you are betting on — and at a $48 monthly premium on a hundred million tokens, that is not a large bet.
