
Mistral Large 4 vs Gemini 3.1 Pro: Eight Months, Two Directions
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 151 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 116 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 249 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Gemini 3.1 Pro has been on the market since 19 February 2026 and still ranks near the top of every broad capability table, including the ones where it is being compared to models released this autumn. Mistral Large 4 is three weeks old at most — a preview announced 6 October 2026, with its weights promised for the end of October and no licence named. On Artificial Analysis's Intelligence Index, read on 6 October, the eight-month-old Gemini 3.1 Pro scores 29.7 against the newcomer's 38.4. That is the honest starting point for this comparison, and it is the opposite of what the release dates suggest.
The explanation is not mysterious. Gemini 3.1 Pro is a preview-line flagship Google has never promoted to a stable release, and Artificial Analysis's page for it is labelled "Gemini 3.1 Pro Preview" — the same page where a lower index sits next to a top-tier GPQA Diamond. It is a model built for breadth: a very large context window, a wide modality set, and reasoning that is strong on knowledge questions. Mistral Large 4 is built for a different map of the world entirely, and it is priced accordingly.
What each one is
Gemini 3.1 Pro is Google's frontier multimodal reasoning model, API id gemini-3.1-pro-preview, with a 1,048,576-token context window and a 65,536-token output ceiling. It accepts text, images, video, audio and files. It is served from OrcaRouter's catalogue at Google's own rates. Google's documentation lists no non-preview Gemini 3.1 Pro; the preview name is the product.
Mistral Large 4 is a 1.05-trillion-parameter sparse mixture-of-experts with 49 billion active parameters and a dedicated 1.6-billion-parameter vision encoder, offered as "Mistral Large 4 Preview" under the API id mistral-large-4-0. Mistral's docs state a 1M-token context window; Artificial Analysis reads 524,288 on the configuration it is testing. Maximum output is not published. Mistral describes it as its largest and most capable model to date and says the weights arrive at the end of October.
The numbers, with their provenance attached
Both models have independent pages, so the comparison below is same-harness rather than vendor-versus-vendor:
• Intelligence Index v4.3 — Mistral Large 4 Preview 38.4 vs Gemini 3.1 Pro 29.7
• Cost per index task — $1.13 vs $1.30
• Long-context reasoning — 81.3% vs 82.0%
• GPQA Diamond — not published for the preview vs 94.1%
• Humanity's Last Exam — 35.0% vs 47.0%
• Terminal-Bench 4.0 — 26.8% vs 4.0%
• SciCode — 54.2% vs 58.7%
• AutomationBench, business workflows — 59.9% vs 35.4%
• MMMU-Pro, multimodal reasoning — 76.4% vs 82.4%
• Context window — 1M stated / 524,288 tested vs 1,048,576 served
• Output ceiling — not published vs 65,536 tokens
• Price per million, input / output — $1.36 / $4.18 list, $0.68 / $2.09 promoted, vs $2.00 / $12.00 under 200K tokens and $4.00 / $18.00 above
• Weights — promised, licence unnamed, vs proprietary

The single most striking row is Terminal-Bench 4.0. Gemini 3.1 Pro scores 4.0% — not a typo, and not a hallucination in the table: it reflects a model from February being run on a terminal-agent harness that postdates it. Mistral Large 4's 26.8% is six times higher on the same test. Anyone who has ruled Gemini out on the basis of an old agentic-coding number should rule it back in on the basis of its GPQA Diamond, and anyone who has ruled Mistral out on index rank should look at the terminal row first.
Where Gemini 3.1 Pro still holds the better hand
Knowledge and breadth. A 94.1% GPQA Diamond is the highest figure anywhere in this article, on either side, and it is a graduate-level science benchmark — not a number a model talks its way to. Humanity's Last Exam at 47.0% against 35.0%, and a SciCode score that edges the newcomer, tell the same story. If your workload is answering hard questions accurately from a corpus of knowledge, Gemini 3.1 Pro remains the stronger model eight months after release.
Modality breadth is the second advantage. Gemini 3.1 Pro takes video and audio as well as text and images. Mistral Large 4 takes text and images. If your pipeline ingests recorded meetings, screen recordings or sensor audio, only one model here can see the whole input.
The output ceiling is the third. 65,536 tokens is a published, guaranteed ceiling; Mistral has not published one for the preview at all. Nothing in a launch post is more likely to bite you in production than an unstated output limit.
And Gemini's context window is genuinely served at 1,048,576 tokens, whereas Mistral's 1M is the vendor figure against an independently-read 524,288. For a workload that lives at the far end of the window, the difference between a stated and a measured number is a rewrite.
Where Mistral Large 4 changes the decision
Agentic work, and specifically anything that touches a terminal, is where the two models stop being comparable. Mistral's own launch post bundles 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4.0 into a Coding Agent Index of 49.8%, and states plainly that it placed ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max on that combined measure. Those three figures are, in Mistral's words, numbers Artificial Analysis evaluated privately before its harness went public; they are not on the public index yet and should be labelled vendor-published until they are.
Independent evidence points the same way. On AutomationBench — 657 business workflows spanning Gmail, Sheets, Slack and Salesforce — Mistral Large 4 scores 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro, and on the same harness Gemini 3.1 Pro does not appear at all because the benchmark postdates it. A blind human study run by Surge AI rated Mistral Large 4 Preview second of five models on coding quality at 3.74, ahead of GLM-5.3 and Kimi K3 and behind only Claude Opus 5.
The cyber claim is the sharpest one in Mistral's post and worth stating exactly: 82% on the Artificial Analysis Cyber Index test that asks a model to reproduce a real vulnerability and then patch it, described as the highest of any model, with 93% on Cybench's 40 competition exercises, and Mistral's assertion that it ranks in the global top five on the Cyber Index overall while leading open-weight models developed outside China. Those are vendor-cited figures on an independent index. Their practical edge is that Mistral built the model to do the work rather than refuse it.
The cheapest line in the table belongs to Mistral too, but only narrowly: $1.13 per task against $1.30, a 13% edge that the promotional rate produces and the rate card would erase. Gemini 3.1 Pro is a 2026 model on 2026 pricing and it is not expensive to run an index task on.
The tiebreaker nobody benchmarks
Two things are in play that the tables cannot show. The first is licensing. Gemini 3.1 Pro is proprietary and will remain so; Mistral Large 4 is a preview whose entire pitch is that it becomes a downloadable artifact you can run under your own policy, on your own hardware, under European law — Mistral trained it on 3,800 Grace Blackwell GPUs in its own datacenters and serves the preview there. That argument is worth nothing until the repository appears and the licence has a name. It is worth a great deal the day it does.

The second is operational risk. A preview that is still being refined will change under you. On OrcaRouter both sides of this comparison are reachable through one OpenAI-compatible endpoint at provider list prices with 0% markup — Gemini 3.1 Pro is live in the catalogue at Google's rates, while Mistral Large 4 must be reached through Mistral's own API and several third-party platforms, since it is not a route we carry. The routing DSL is the useful piece here: send classification and extraction to whichever model is cheaper that week, escalate the hard residue, and let automatic failover absorb the preview's bad days rather than your on-call rotation absorbing them.

The verdict
Ask what the workload is, not which model is better. For knowledge work, audio and video input, a guaranteed output ceiling and a served million-token window, Gemini 3.1 Pro is still the stronger model and its age is not a flaw — it is eight months of other people finding its edges. For agentic coding, terminal work and security operations, Mistral Large 4 is ahead by a margin wide enough that the index rank is misleading, and it is the only one of the two that may end up self-hostable.
The tiebreaker to watch is the end of October. If the weights arrive under a permissive licence, Mistral Large 4 stops competing with Gemini 3.1 Pro on price and starts competing with it on control, and that is a different fight. If they arrive under a restrictive one, this article's answer does not change much — but the reason does.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
