A generated title card for the article 'Mistral Large 4 vs DeepSeek V4 Pro' showing the title in bold geometric sans-serif, the subtitle 'Twin architectures, opposite bets' beneath it, and a row of three flat line icons above — two matching network nodes, one with a download arrow and one greyed out — on a white background with blue-and-cyan gradient accents and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Mistral Large 4 vs DeepSeek V4 Pro: Twin Architectures, Opposite Bets

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The two models in this comparison are near-identical on paper and could not be more different in practice. Mistral Large 4 is a 1.05-trillion-parameter sparse mixture-of-experts with 49 billion active parameters, announced 6 October 2026. DeepSeek V4 Pro is a 1.6-trillion-parameter sparse mixture-of-experts with 49 billion active parameters, released 24 April 2026. Same activation budget, same architectural family, same claim of frontier performance at open-weight economics — and exactly one of them has weights you can download today.

The pairing is not one this article invented. Mistral's own launch post benchmarks against DeepSeek V4 Pro by name in three separate places: agentic coding, long-horizon knowledge work, and business-workflow automation. Asking the same question of the other side — what does DeepSeek V4 Pro already do that Mistral Large 4 is promising — turns out to be the fastest way to see what the preview actually adds.

The specification, side by side

Read on 6 October, Mistral's figures from its own announcement and model card, DeepSeek's from its live catalogue entry:

• Total vs active parameters — 1.05T total / 49B active vs 1.6T total / 49B active

• Modalities — text and image in, text out vs text only

• Context window — 1M stated by Mistral, 524,288 on the configuration Artificial Analysis is testing, vs 1M served

• Output ceiling — not published for the preview vs 384K tokens

• Price per million, input / output — $1.36 / $4.18 list, currently $0.68 / $2.09 on promotion, vs $0.66 / $1.98

• Cached input per million — $0.14, currently $0.07, vs $0.022

• Weights — promised end of October, licence unnamed vs available now

• Training infrastructure — 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters vs DeepSeek's own cluster

A screenshot of Mistral's own documentation, the Models Overview page at docs.mistral.ai, showing the GENERALIST MODELS card grid in which 'Mistral Large 4' appears at version v26.10 described as 'A state-of-the-art, open-weight, general-purpose multimodal model', with no licence badge, next to a Mistral Large 3 card at v25.12 that carries an APACHE 2.0 badge. Mistral Small 4, Mistral Medium 3.5, Z.ai GLM 5.3 and the Ministral 3 sizes are also visible.

The symmetry of the 49B activation figure is the headline and also the trap. Active parameters govern what a token costs to generate; total parameters govern what the model knows. DeepSeek's 1.6T against Mistral's 1.05T means DeepSeek is carrying roughly 50% more capacity per token computed — a real advantage in knowledge-heavy work and no advantage at all in tasks where the bottleneck is instruction-following or tool use.

What the independent index says

Both models have pages on Artificial Analysis, which makes this one of the few comparisons in the series where the two sides are measured on the same harness. Intelligence Index v4.3, read 6 October:

• Intelligence Index — 38.4 for Mistral Large 4 Preview vs 36.0 for DeepSeek V4 Pro 0813

• Cost per index task — $1.13 vs $0.67

• Terminal-Bench 4.0 — 26.8% vs 14.1%

• Humanity's Last Exam — 35.0% vs 41.0%

• AutomationBench, business workflows — 59.9% vs 56.7%

• τ²-Bench, tool-use agentics — not published for the preview vs 96.2%

• SciCode — 54.2% vs 51.0%

This is a much tighter race than the price tags imply, and the disagreements are informative rather than noise. Mistral Large 4 wins the agentic-coding row by nearly thirteen points and the workflow-automation row narrowly, while DeepSeek V4 Pro wins open-ended knowledge by six points and costs 40% less per task. Note also that the index spreads costs across cache reads, cache writes, reasoning tokens and answer tokens — Mistral's cache is discounted 90% while DeepSeek's is discounted 97%, which is why two models with similar headline rates land at $1.13 and $0.67 per completed task.

A generated two-column scoreboard titled 'Mistral Large 4 vs DeepSeek V4 Pro — the scoreboard'. Left column 'Mistral Large 4 Preview': total and active params 1.05T / 49B; Intelligence Index v4.3 38.4; Terminal-Bench 4.0 26.8%; AutomationBench 59.9%; modalities text and image in; weights promised end of October. Right column 'DeepSeek V4 Pro': total and active params 1.6T / 49B; Intelligence Index v4.3 36.0; Terminal-Bench 4.0 14.1%; AutomationBench 56.7%; modalities text only; weights open and available now.

Where DeepSeek V4 Pro has already won

The decisive difference is not a benchmark. It is that DeepSeek V4 Pro is open-weight today and Mistral Large 4 is a preview served from its maker's own cluster. Mistral says the weights drop at the end of October — a commitment, not an artifact. Its Hugging Face organisation, checked on 6 October, holds nothing newer than July. Until a repository appears with a licence file, the self-deployment argument Mistral is making belongs to DeepSeek.

The second difference is availability breadth. DeepSeek V4 Pro has been routable since April and is, by a wide margin, the most-used model on our own catalogue — 1.13 billion tokens across the last seven days at the time of writing, against the tens of millions a typical frontier model moves. That volume matters to a buyer in a way a leaderboard does not: it means the failure modes are known, the throughput characteristics are known, and the providers carrying it have been stress-tested by other people's production traffic.

DeepSeek also wins on the boring mechanics. A 384K output ceiling against an unpublished one, 1M context served rather than stated, and a text-only scope that makes long-context throughput predictable. If your workload is document analysis, long-form generation, or knowledge-heavy extraction over a million tokens, DeepSeek V4 Pro is cheaper, more open, and more proven than the Mistral preview on every axis.

Where Mistral Large 4 is the better bet

Mistral chose the rows it benchmarks carefully, and the coding gap is not cosmetic. 26.8% against 14.1% on Terminal-Bench 4.0 is roughly double, and it lines up with Mistral's claim of 61.7% on DeepSWE v1.1 and 59.4% on SWE-Atlas-QnA — figures it says Artificial Analysis evaluated privately ahead of the harness's public launch, which is why they are not on the public index yet. When those land, or do not, will settle the coding question.

Multimodality is the second clear split. Mistral Large 4 takes images natively, with a dedicated 1.6B vision encoder, and Mistral claims state-of-the-art visual grounding among open models — naming a 42% against GPT-6 Astra's 41% on Dense 200. DeepSeek V4 Pro is text-only and its catalogue card says so plainly. Any pipeline involving screenshots, engineering drawings, scanned filings or chart reading has one candidate here, not two.

And then there is cyber, where Mistral's claim is sharper than anything DeepSeek has published: 82% on the Artificial Analysis Cyber Index test that asks a model to reproduce a real vulnerability and patch it, said to be the highest of any model, and 93% on Cybench's competition exercises. Those are vendor-cited figures and should be labelled as such — but they sit on an independent index, and no comparable DeepSeek result has been published. For security work, the honest position is that Mistral Large 4 is claiming a lead that has not been shown to be false.

The last differentiator is jurisdictional. Mistral trained the model on 3,800 Grace Blackwell GPUs in its own European datacenters and serves the preview there, with a European deployment operated end-to-end under European law. For a European regulated buyer whose alternative is a Chinese open-weight model or a US closed one, that is a category of its own.

The pricing, read properly

Neither model's sticker price is the whole story. DeepSeek V4 Pro's catalogue entry carries two timed windows — 01:00–04:00 and 06:00–10:00 UTC — during which the listed rate is multiplied, so a workload that lands in those hours costs double the headline. Mistral Large 4's $0.68 / $2.09 is a launch promotion against a $1.36 / $4.18 rate card; the promotion is the price today and the rate card is the price to model your bill on.

Where OrcaRouter clears this up is the pass-through: 0% markup on provider list prices, which means a vendor's price cut is your price cut the same day and a vendor's promotional rate is passed along rather than absorbed. DeepSeek V4 Pro is routable today behind one OpenAI-compatible endpoint at its provider rate, in the same credential as 200-plus other models, with automatic failover across providers when one degrades. Mistral Large 4 is not in our catalogue — the preview is reached through Mistral's own API and several third-party platforms — so if you want to run both sides of this comparison at once today, that means one route on us and one direct.

A screenshot of the OrcaRouter model page for deepseek/deepseek-v4-pro, showing the model identity, a 1M-token context window, a 384K maximum output, text-only input with text output, the 2026-04-24 release date, a p-50 time-to-first-token figure of 2.23 seconds, pricing of $0.66 per million input and $1.98 per million output, and a code sample using the api.orcarouter.ai base URL. The page describes the model as a 1.6T-total, 49B-active flagship MoE with a 1,048,576-token context and 384,000-token maximum output that is text-only.

Which one to pick

If you need weights you can hold, 384K outputs, a million tokens of served context, the lowest cost per completed task in the pair and a model whose behaviour other people have already broken in production, DeepSeek V4 Pro is not a compromise — it is the finished product versus a trailer. It is the right default for document, knowledge and long-form work, and its open weights mean the ceiling on what you can do with it is your own infrastructure rather than a vendor's roadmap.

If your workload is agentic coding, if it involves images, if it involves security work, or if European jurisdiction is a requirement rather than a preference, Mistral Large 4 is the model worth the preview risk — and the risk is bounded, because Mistral has committed to the weights and to a red-teaming window with a stated end. Put a slice of traffic on it now, keep DeepSeek V4 Pro on the rest, and let the two settle the architecture question on your data. The 49B activation budget is going to look the same from either side; what differs is everything around it.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily