
Mistral Large 4 vs GPT-6 Sol: A Preview Price Against a Shipped Default
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 144 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 79 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 293 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
On paper this is not a close fight. Mistral Large 4, Mistral's 1.05-trillion-parameter open-weight flagship, is priced at $0.68 per million input tokens and $2.09 per million output tokens. GPT-6 Sol, the cost-efficient high-end tier, is priced at $2.00 and $10.00 — about three times as much on input and nearly five times as much on output, and it reprices to $4.00 and $15.00 once a request crosses 272,000 tokens. A 100-million-token month at a 4:1 input/output split costs roughly $140 on Mistral Large 4 and roughly $480 on GPT-6 Sol. That is the whole case, and it is a strong one, until you read the second line of Mistral's own documentation: Mistral Large 4 is in public preview, and it is the only one of these two models that is still moving.
That distinction is the entire comparison. GPT-6 Sol shipped on 22 September 2026 and is stable; its behaviour, its rate card and its independent scores are settled. Mistral Large 4 arrived on 6 October 2026 with a preview API, a promised weight release "by the end of the month", and a reinforcement-learning run that Mistral says is still in flight. The version you call today is not the version you will have in November. So the interesting question is not "which is better" — on measured intelligence they are not in the same bracket yet — but "which one belongs in your stack, and under what conditions".
What each vendor actually shipped
Mistral Large 4 is a granular mixture-of-experts model with 49 billion active parameters out of 1.05 trillion total, plus a 1.6-billion-parameter vision encoder. It takes text and images, serves a context window of about one million tokens, and is trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacentres. Mistral describes it as its largest and most capable model to date and says the weights — released under an open licence — are coming before the end of October 2026, pending a red-teaming period run with cybersecurity partners and state authorities. Until then the only way in is the preview API.
GPT-6 Sol is the middle tier of OpenAI's GPT-6 series, below the flagship GPT-6 Astra and above the fast GPT-6 Luna. It accepts text, images and files, carries a context window just over one million tokens with up to 128,000 output tokens, and exposes configurable reasoning effort, function calling, structured outputs and seeded sampling. It is a closed, moderated API — the opposite of Mistral's pitch, which is that an organisation can eventually run ML4 on its own hardware under its own policies.
• Total parameters — Mistral Large 4 1.05T (49B active) vs GPT-6 Sol undisclosed
• Context window — roughly 1M tokens each, Mistral Large 4 advertised at 1M, GPT-6 Sol at ~1.05M
• Max output — GPT-6 Sol 128K tokens vs Mistral Large 4 not yet documented
• Input modality — Mistral Large 4 text and image vs GPT-6 Sol text, image and file
• Input price — Mistral Large 4 $0.68 per 1M vs GPT-6 Sol $2.00, rising to $4.00 past 272K tokens
• Output price — Mistral Large 4 $2.09 per 1M vs GPT-6 Sol $10.00, rising to $15.00 past the same threshold
• Cached input — Mistral Large 4 $0.07 per 1M vs GPT-6 Sol $0.20, rising to $0.40
• Weights — Mistral Large 4 open licence, promised by end of October 2026 vs GPT-6 Sol closed
• Release status — Mistral Large 4 public preview, 6 October 2026 vs GPT-6 Sol generally available, 22 September 2026

Where GPT-6 Sol is simply ahead
Mistral's launch post is unusually direct about this, which saves everyone the guesswork. On the Artificial Analysis Intelligence Index — the v4.3.2 revision, ten evaluations from AA-Briefcase to CritPt — GPT-6 Sol scores 47.6 and sits around 14th of the 147 models on that board. Mistral Large 4 has no published Index score at all. What exists is a partial row on Artificial Analysis, which lists the model as "Mistral Large 4 Preview" and reports its position as not publicly available. On the individual evaluations that are visible, GPT-6 Sol leads clearly: 47.9 against a Humanity's Last Exam bar it clears while ML4 does not appear, 83.7 against 80.3 on AA-LCR long-context recall, 43.9 against 28.3 on Terminal-Bench 4.0.
Those Terminal-Bench figures deserve a caveat that cuts the other way. Mistral's 28.3% and the 59.4% on SWE-Atlas-QnA come from numbers Artificial Analysis evaluated privately ahead of a harness launch, and Mistral is explicit that they will appear on the Artificial Analysis Coding Agent Index when that harness ships. They are not vendor-invented, but they are also not reproducible by you today. Treat them as promising and unverified.
Where Mistral does publish numbers, several of them are worth reading properly because they are the kind of claim a vendor makes when it is confident. Mistral reports that on one Artificial Analysis Cyber Index test — reproduce a real vulnerability in open-source software, then patch it — ML4 scores 82%, the highest of any model, and that it solves 93% of the 40 challenges in Cybench. It also reports that leading closed models including Claude Opus 5.5 and GPT-6 Astra score near zero on the same reproducibility test because they refuse the task. Those are vendor-reported and have not been independently reproduced, but the underlying mechanism is not exotic: a model that refuses more actually completes less, and Mistral has been selling that argument since its open-weights turn.
On general capability the honest read is that GPT-6 Sol is the better model today and Mistral Large 4 is the cheaper one. If your workload is a straight question-answer or coding task with no preference for self-hosting, GPT-6 Sol's extra five dollars per million output tokens buys you a model whose behaviour will not change under you.
The case for the preview, which is not just price
Three things make Mistral Large 4 more than a discount option, and none of them is the sticker.
The first is sovereignty in the literal sense. ML4 was trained in European datacentres on European hardware and will be servable in a European region Mistral operates end to end. For a regulated buyer, that is not a marketing line; it is a procurement requirement that GPT-6 Sol cannot currently satisfy at any price. Mistral raised €3 billion in a Series D specifically to fund this, the largest equity round a European technology company has raised, and this model is the first milestone against it.
The second is the weights. When they land, an organisation that has the GPUs can run ML4 on-premises with no provider in the loop, which changes both the refusal behaviour and the data-residency story. GPT-6 Sol will still be an API call when that happens. If your roadmap has an on-premise phase, ML4 is the only one of the two that has one.
The third is cadence. Mistral says the RL run behind this preview has not saturated and that it expects "large and rapid improvements in the weeks and months to come". When a vendor ships a preview and states that the curve is still climbing, the price you are quoted today is attached to the worst version of the model you will ever be billed for. GPT-6 Sol's price is attached to a fixed capability.
Running both without betting on either

The awkward part of a preview is that you cannot pin it. If your product depends on ML4's current behaviour, an update next week is a production incident. If it depends on GPT-6 Sol, you are paying a fixed premium for stability. The sensible shape is usually both, behind one endpoint, so the decision is a routing rule and not an architecture.
That is what OrcaRouter is for. Both models sit behind a single OpenAI-compatible API with 0% markup, which matters more than usual here: Mistral's preview pricing is a launch rate and will move — upward when the preview ends, possibly downward when the weights land and competition on the open model bites — and because provider list prices pass straight through, a vendor price cut is live on our side the same day rather than whenever someone updates a hard-coded table. The routing DSL lets you send long-context document work to whichever model is cheapest for that shape, and keep GPT-6 Sol as the fallback for anything that has to be stable. Automatic failover covers the case a preview API invites most: the endpoint being busy or briefly unavailable, which is normal for any model in its first month.

Which one to pick
Choose GPT-6 Sol if the workload is production, the behaviour matters more than the unit price, and you have no requirement to run the model yourself. Its 47.6 on the Intelligence Index, its 128K output ceiling and its $0.20 cache reads make it the safer default, and the long-context reprice at 272,000 tokens is easy to plan around because the threshold is published.
Choose Mistral Large 4 if you are building something where one of three things is true: you need a European-deployed model, you need open weights eventually, or you are in the security-research business where a refusal is a failure — 82% on the reproduce-and-patch test against near-zero for the closed alternatives is a real capability gap, vendor-reported or not. And accept that you are buying a model mid-flight. Budget for the API surface to change, and do not hard-code the preview price into a forecast for next quarter.
What to watch before committing either way: the actual weight release at the end of October, which is the moment the open-weights claim stops being a promise; the first published Artificial Analysis Intelligence Index score for ML4, which will settle whether the price gap is a bargain or a discount for less; and the Coding Agent Index, which is where Mistral's private Terminal-Bench numbers either get corroborated or quietly do not.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
