
Mistral Large 4 vs Claude Opus 5: The Gap Is Smaller Than the Price Gap
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 151 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 116 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 249 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The only head-to-head number anyone has published for Mistral Large 4 against Claude Opus 5 comes from a blind human evaluation, and it is closer than the benchmark tables suggest. Surge AI hid the model identities and asked professional annotators to rate coding output on a 1–5 scale; Claude Opus 5 finished first at 4.22 and Mistral Large 4 Preview finished second at 3.74, ahead of Kimi K3, GLM-5.3 and GLM-5.2. On the independent aggregate the two are not neighbours at all — 50.8 against 38.4 on Artificial Analysis's Intelligence Index, read on 6 October. What makes the pairing worth an afternoon is that the model scoring 12 points lower costs roughly a fifth as much to run a task on.
Both figures need their provenance attached, because they are not the same kind of number. The Surge AI evaluation is a third-party human study that Mistral commissioned and published — a stronger form of evidence than a self-run benchmark, and still a study whose publication Mistral controlled. The index scores are Artificial Analysis's own runs, comparable to each other, and therefore the right pair to reason from. Neither one tells you which model to build on; together they bracket the question.
Two products, not two generations of one
Claude Opus 5 is a closed-weight flagship from the vendor, released 24 July 2026, served with a 1M-token context window and up to 128K output tokens, at $5 per million input and $25 per million output with cached reads at $0.50. It is the vendor's most capable model in the Opus line and is positioned squarely at long-horizon agentic work — staying coherent across multi-step sessions with heavy tool use.
Mistral Large 4 is a preview, announced 6 October 2026, that Mistral describes as a 1.05-trillion-parameter sparse mixture-of-experts with 49 billion active parameters and a 1.6-billion-parameter vision encoder. Its API id is mistral-large-4-0, it is natively multimodal, and Mistral's docs give it a 1M-token context window while Artificial Analysis reads 524,288 on the configuration it is testing. The weights are promised for the end of October and the licence has not been named.
So this is not a newer model against an older one. It is a proprietary frontier model against a to-be-open-weight challenger, and the two are being priced accordingly.

The same index, both models
Read on 6 October off their Artificial Analysis pages, on Intelligence Index v4.3:
• Intelligence Index — Mistral Large 4 Preview 38.4 vs Claude Opus 5 50.8
• Cost per index task — $1.13 for Mistral Large 4 Preview vs $5.86 for Claude Opus 5
• Terminal-Bench 4.0 — 26.8% vs 49.0%, the widest single gap between them
• Humanity's Last Exam — 35.0% vs 54.9%
• MMMU-Pro, multimodal reasoning — 76.4% vs 84.7%
• AutomationBench, business workflows — 59.9% vs 56.6%, the one row Mistral Large 4 leads
• Context window — 1M stated by Mistral and 524,288 on the tested configuration, vs 1M served
• Output ceiling — not published for the preview vs 128K tokens
• Price per million, input / output — $1.36 / $4.18 list, currently $0.68 / $2.09 on promotion, vs $5.00 / $25.00

The pattern is legible. Opus 5 is ahead on the two hardest reasoning-heavy rows — terminal work and open-ended knowledge — and ahead on multimodal understanding despite Mistral Large 4 being the model sold on its vision. Mistral Large 4 leads the workflow-automation row, which is the row closest to what a business unit actually asks a model to do, and it does so at a fifth of the price per task.
Where Mistral Large 4 genuinely wins
The most interesting claim in Mistral's launch post is not a benchmark score. It is that on one of the Artificial Analysis Cyber Index tests — reproduce a real vulnerability in open-source software, then patch it — Mistral Large 4 scores 82%, said to be the highest of any model, and that several leading closed models score near zero on the same test because they refuse the task. Mistral names Claude Opus 5.5 and GPT-6 Astra as examples of that refusal.
That is a vendor interpretation of a vendor-cited figure, and it should be read that way. It is also a real operational difference if it holds: a model that will not reproduce a vulnerability cannot be used to prove one is real, which is the first step of most incident response. Mistral pairs the 82% with a 93% score on Cybench's 40 competition exercises and a counter-claim that its refusal rate on cyber prompts drawn from JailbreakBench, StrongREJECT and AgentHarm is higher than other open-weight models' — the argument being that you want the capability and the discipline in the same model rather than trading one for the other.
Beyond cyber, the case for Mistral Large 4 rests on three things Opus 5 does not offer. It is trained and served on European infrastructure under European law, which matters to buyers with data-residency obligations. It will, Mistral promises, be downloadable — the point of the whole exercise, since self-deployment is what removes the mid-incident access risk Mistral is at pains to describe. And it is cheap enough per task that using it as a first pass and escalating the hard residue to a frontier model is economically sensible rather than a rounding error.
Where Claude Opus 5 still earns its price
A 12-point index gap is not small, and the row-by-row picture says where it lives. Terminal-Bench 4.0 is nearly double — 49.0% against 26.8% — and terminal work is the closest public proxy for agentic coding, which is most of what teams are buying frontier models for this year. Humanity's Last Exam separates them by 20 points. On long-horizon autonomy, Opus 5's adaptive reasoning design, its 128K output ceiling, and a served 1M context are the configuration you want if your workload is a multi-hour agent session rather than a single hard question.
The output ceiling is the quieter difference. Mistral has not published a maximum output length for the preview, and Opus 5's 128K is a real constraint-lifter for code generation and long structured documents. If your pipeline emits files rather than answers, check that number before you switch, because it is the kind of gap that only shows up in production.
Cost per task is the number to hold on to
Per-token prices flatter cheap models and mislead about expensive ones. The index's own per-task cost — total spend to complete one evaluation task, cache and all — is the number that survives contact with a budget: $1.13 against $5.86.
That is not a licence to move everything to the cheaper model, because a task the cheap model fails is a task you pay for twice. It is a licence to stop treating a single frontier model as the only option. On OrcaRouter, Claude Opus 5 sits in the catalogue at the vendor's list price with 0% markup, in the same OpenAI-compatible endpoint as 200-plus other models, which is what makes a two-model pipeline an afternoon's work rather than a second vendor contract. Mistral Large 4 is not in the catalogue today — the preview is reached through Mistral's own API and several third-party platforms. When its weights land and a provider hosts it, the same pass-through rule applies: whatever the vendor charges is what you are charged, and a vendor price cut shows up on your bill the same day.

For an unproven preview the routing feature that matters most is failover. Sending a slice of production traffic to Mistral Large 4 while Opus 5 holds the rest means a regression turns into a reroute instead of an outage, and you learn the model's real failure modes on your own data rather than on a leaderboard.
What resolves this in three weeks
Three facts are still open, and each one moves the answer. If the weights arrive under a permissive licence, Mistral Large 4 becomes the only model in this pair you can self-host, and the calculus for regulated buyers tilts hard. If the promotional $0.68 / $2.09 pricing reverts to the $1.36 / $4.18 on the rate card, the per-task gap narrows but stays decisive. And if Artificial Analysis moves the privately-evaluated coding figures onto its public index unchanged, the coding story becomes third-party-verified rather than vendor-published-on-a-third-party's-behalf.
Until then: Opus 5 is the safe production default and worth its multiple for agentic and terminal-heavy work. Mistral Large 4 is the interesting bet — cheap enough per task to run beside it, capable enough on automation and cyber to earn traffic today, and holding a promise of self-deployment that no closed model in this comparison can match.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
