
Mistral Large 4 Is a 1T-Parameter Preview: The Weights Have Not Dropped
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 151 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 116 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 249 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Mistral Large 4 arrived on 6 October 2026 as a public preview — a one-trillion-parameter mixture-of-experts model with 49 billion active parameters, a 1.6-billion-parameter vision encoder, and a claim from its maker that it is the strongest open-weight model built outside China. What has not arrived is the part the words "open-weight" describe. There are no checkpoints to download today, no licence file, and no repository. Mistral says the weights land "end of this month," after a red-teaming window during which vetted partners and national authorities get an unlocked build with reduced moderation and expanded cyber capability. So the accurate description of Mistral Large 4 this week is a preview API you can call and a model you cannot yet own.
That split is the story, not a footnote to it. Every number Mistral published alongside the announcement — the coding scores, the cyber scores, the financial and legal evaluations — was produced on a build that is still being changed, and every headline comparing it to a closed frontier model is comparing an artifact that will not be the artifact you download. The useful thing to do with a launch like this is separate what you can verify today from what you are being asked to take on faith, and mistral.ai's own post is unusually generous about which is which.
What is actually live on 6 October
The preview API is up on Mistral Studio under the model id mistral-large-4-0. Mistral's own model card describes it as "a state-of-the-art, open-weight, general-purpose multimodal model" built on a granular MoE architecture, and breaks the parameter count down as 49B active against 1.05T total, plus the vision encoder. The company trained it from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, and the preview is served on that same infrastructure — a sovereignty argument Mistral is making as loudly as the performance one.
The specification, as the vendor states it:
• Total vs active parameters — 1.05 trillion total, 49 billion active per token, sparse MoE
• Modalities — text and image in, text out; natively multimodal rather than a bolted-on adapter
• Context window — Mistral's docs state 1M tokens; Artificial Analysis reads 524,288 on the configuration it is testing, so treat the 1M as the vendor ceiling rather than the served window
• Price — $1.36 per million input tokens and $4.18 per million output, with cached input at $0.14, per Mistral's published card
• Launch offer — the same card shows a struck-through $0.68 / $2.09 promotion with $0.07 cached input, i.e. half price
• Weights — not released; "end of this month," per the announcement
• Licence — unnamed in the announcement and on the docs page, which tags the entry only as "Open"
Two of those lines deserve to be read twice. The context window differs by a factor of two depending on whether you read the vendor or the third-party lab, which is not unusual for a preview whose serving configuration is still moving but is worth knowing before you size a workload on it. And the price you will pay depends on whether the promotional rate survives the launch window; $1.36 / $4.18 is the number Mistral's own rate card lists, and $0.68 / $2.09 is the number it is currently charging. Assume the former when you model your bill.

The benchmark numbers, and who ran them
Mistral's post is careful about provenance and the caution is worth carrying forward. Three of the headline agentic-coding figures — 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.0, combining into a 49.8% Coding Agent Index score — are, in Mistral's own words, "the numbers evaluated privately by Artificial Analysis ahead of the harness' public launch." They are not on Artificial Analysis's public index yet. They will be, when the harness ships. Until then they are vendor-published figures that a third party produced but has not yet published, which is a stronger provenance than most launch claims and still not the same thing as a reproducible public score.

What is public and independent today, read off Artificial Analysis's model page on 6 October: Mistral Large 4 Preview scores 38.4 on Intelligence Index v4.3, taking about 507 seconds per task, with Terminal-Bench 4.0 at 26.8% and Long-Context Reasoning at 81.3%. The same page lists it as an open-weights model that is not, in fact, open-weights — the flag reads isOpenWeights: false with the categorisation "proprietary," because a preview served from the vendor's own cluster is what exists. That flag is the cleanest single illustration of the gap this article is about.

The most quotable result is the one Mistral did not run itself. Surge AI, a third-party annotation firm, ran a blind human evaluation on coding quality with model identities hidden, and professional annotators rated Mistral Large 4 Preview at 3.74 on a 1–5 scale — second of five, ahead of Kimi K3 at 3.59, GLM-5.3 at 3.60 and GLM-5.2 at 3.40, and behind only Claude Opus 5 at 4.22. A blind human eval with a named lab on the other end is a different kind of evidence from a self-run benchmark, and it is the single strongest independent data point in the announcement.
Then there is cyber, where Mistral makes its sharpest claim. On Artificial Analysis's Cyber Index, the company says Mistral Large 4 ranks in the global top five and leads open-weight models developed outside China. On one specific test — reproduce a real vulnerability in open-source software, then patch it — it scores 82%, described as the highest of any model, with 93% on Cybench's 40 competition exercises. Mistral adds a pointed observation: several leading closed models, Claude Opus 5.5 and GPT-6 Astra named among them, score near zero on that same test because they refuse the task. Without the third-party page in front of you, the 82% is a vendor figure; the refusal-rate argument is a vendor interpretation. Both are plausible and neither is independently confirmed today.
Why the preview framing changes the price maths
A preview is not merely a stage label. It means the model can be re-pointed, re-priced, or retired with less notice than a GA release, and it means the serving configuration you benchmark against may not be the configuration you deploy against. For a routing decision, that is a real cost: a pipeline tuned on this week's build carries a re-validation bill you did not budget for.
On OrcaRouter, where 200-plus models sit behind one OpenAI-compatible endpoint at the provider's list price with 0% markup passed straight through, a price change on a vendor's own card is a price change on ours the same day — which matters here precisely because this launch ships two prices, one struck through. Automatic failover is the other half of the answer for an unproven preview: it lets you put a fraction of production traffic on Mistral Large 4 while a first-party model holds the rest, so a preview that regresses under your real workload degrades into a reroute rather than an incident. What careful readers should not assume is that Mistral Large 4 is routable on OrcaRouter today. It is not in the catalogue — the preview is reached through Mistral's own API and several third-party platforms, and its weights, when they come, will be a self-hosted proposition.
The open-weight claim is currently a promise
Mistral's framing is that Mistral Large 4 "pushes the frontier of open-weight performance" and that open weights are the mechanism by which enterprises keep control of a model. That is a defensible argument for the model Mistral intends to ship. For the model it shipped this week, the mechanism is absent: no Hugging Face repository, no licence text, no downloadable checkpoint. The company's own Hugging Face organisation, checked on 6 October, has nothing newer than July, and its newest entries are unrelated to the Large line.
This is not a criticism of the plan; staggered weight release after a red-teaming window is a coherent policy, and Mistral is explicit about it. It is a caution about the adjective. "Open-weight" is doing marketing work in a paragraph where the weights are four weeks away, and the licence that governs them has not been named. A permissive licence and Apache-style terms are one outcome; a research-only or revenue-capped licence is another, and the announcement does not choose between them.
What to check before the end of October
Five things would turn this preview into a settled fact, and each is checkable without asking Mistral anything:
• A Hugging Face repository under the Mistral organisation carrying the checkpoint and a licence file — the release Mistral has committed to this month
• The licence name itself, which determines whether "open weight" means Apache-2.0-style freedom or a usage-capped grant
• Whether the $0.68 / $2.09 promotional rate survives, or reverts to the $1.36 / $4.18 on the rate card
• Whether Artificial Analysis moves the privately-evaluated coding figures onto its public index when the harness ships, and whether they hold at the same values
• The served context window, currently listed as 1M by the vendor and read as 524,288 by the independent lab
Until those resolve, the right posture toward Mistral Large 4 is the one its own launch invites: call it, measure it on your workload, and keep the production path on a model whose behaviour is not scheduled to change. The weights are the event. The preview is the trailer.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
