
Mercury 2.5 Preview vs Gemini 3.1 Pro: Two Previews, Six Months Apart
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0259Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0258Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0166Intelligence82Coding
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2658Intelligence72Coding
- DeepSeekNEWDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
Both models in this matchup carry the word "preview" in their names, and that is where the resemblance ends. Mercury 2.5 Preview and Gemini 3.1 Pro are both reasoning models you can call from an API today, but they are previews at opposite ends of their lives. Gemini 3.1 Pro, the flagship reasoning model, has been in preview since February 19, 2026 — six months of independent evals, public leaderboard placements, and production traffic behind it, with an Artificial Analysis Intelligence Index of 57 at launch, multimodal input across text, image, audio, and video, a 1M-token context, and a $2.00 / $12.00 price per million tokens. Mercury 2.5 Preview, Inception's diffusion LLM released August 31, 2026, is two days old: a text-only model with a 256K context, a claimed 1,107 tokens per second, a list price of $0.20 / $0.75 (about $0.04 / $0.15 under a discount through September 8), and not a single independent benchmark. Calling them both "previews" hides the actual question, which is how much a six-month evidence head start is worth against a fundamentally faster way of generating text.
What each model actually is
Gemini 3.1 Pro is a closed, multimodal, general-purpose reasoning flagship. It reads text, images, audio, and video, handles a 1M-token context window with a 64K output ceiling, and adjusts its reasoning depth dynamically — the vendor describes a four-tier reasoning intensity that scales with task complexity. Its launch numbers — ARC-AGI-2 at 77.1%, GPQA Diamond at 94.3%, SWE-bench Verified at 80.6%, Terminal-Bench 2.0 at 68.5% — are vendor-reported, but they sit on top of six months of independent scrutiny, including a re-measurement by Artificial Analysis on a stricter index scale where the model reads around 46.5. That re-measurement is exactly what a mature model earns: it has been tested hard enough for the index to get harder around it. Mercury 2.5 Preview is a diffusion model — it generates and refines tokens in parallel instead of one at a time — with tunable reasoning, parallel tool calls, and schema-aligned JSON, built for the latency-compounding workloads Inception names: search agents, voice pipelines, coding subagents. Every quality claim attached to it, including a 10-plus-point jump over Mercury 2, is the vendor's own, and no third party has measured it.

The spec contrast
• Price — Mercury 2.5 Preview: $0.20 / $0.75 per 1M in/out, ~$0.04 / $0.15 intro through Sep 8. Gemini 3.1 Pro: $2.00 / $12.00 per 1M, stepping to $4.00 / $18.00 above 200K input.
• Context / output — Mercury 2.5 Preview: 256K context. Gemini 3.1 Pro: 1M context, 64K max output.
• Inputs — Mercury 2.5 Preview: text only. Gemini 3.1 Pro: text, image, audio, video, file.
• Architecture — Mercury 2.5 Preview: diffusion LLM, parallel token generation. Gemini 3.1 Pro: autoregressive, dynamic reasoning depth.
• Speed — Mercury 2.5 Preview: 1,107 tokens/sec claimed. Gemini 3.1 Pro: no comparable headline claim.
• Evidence — Mercury 2.5 Preview: vendor-reported only. Gemini 3.1 Pro: AA Index 57 at launch (46.5 newer scale) plus a long independent eval record.

The multimodality gap is the decisive difference
For most applications this comparison is resolved before pricing or benchmarks enter the picture, because Mercury 2.5 Preview cannot see. Gemini 3.1 Pro reads screenshots, diagrams, audio, and video — it is the model you point at a PDF, a chart, a meeting recording, or a UI when you want an answer about something that is not already text. Mercury 2.5 Preview is text-only by design; a diffusion language model that generates text in parallel has no image or audio input path at all. For a document-heavy application, a vision agent, or any multimodal product, Gemini 3.1 Pro is the only candidate here, full stop. Mercury 2.5 Preview's text-only constraint is not a small print caveat; it is the boundary that decides which workloads the model is even eligible for.
Six months of testing versus two days
On evidence, this is not a contest, and it is important to say plainly why. Gemini 3.1 Pro has been picked apart by independent labs for six months; its score has been placed, re-measured, and argued about, which is what "trustworthy" looks like for a model. Mercury 2.5 Preview has one source of information — its maker — and that source claims a 10-plus-point intelligence jump over Mercury 2 and parity with the cost-optimized tier of the frontier. Both statements may be true. Neither has been confirmed by anyone outside Inception, and the model is too new for that to be anything other than expected. The asymmetry is not that one is better; it is that one is knowable and the other is not yet.
Where Mercury's speed actually wins
The honest case for Mercury 2.5 Preview is narrow, text-only, and real. Parallel token generation changes the latency math in a way no autoregressive model — including Gemini 3.1 Pro — can follow, because an autoregressive model is serial by construction. For a voice pipeline, the irony is that the input and output are audio but the thinking between them is text: a fast text reasoning layer is exactly what makes a voice agent feel responsive. For a search agent making repeated tool calls, and for a coding subagent firing many small completions per task, per-call latency compounds into perceived speed. At the intro price — $0.04 in, $0.15 out — Mercury 2.5 Preview is roughly 80 times cheaper than Gemini 3.1 Pro's standard output rate, which makes it not just faster but a different cost class for high-volume text reasoning. The catch that has to travel with every one of those sentences: the speed is claimed, the quality parity is claimed, and no independent measurement exists.
Pricing: Gemini's long-context premium versus Mercury's teaser
Gemini 3.1 Pro's rate card has a detail worth knowing before you compare the headline numbers: requests that pass 200K input tokens step from $2.00 / $12.00 to $4.00 / $18.00 per million. On a 1M-context model that long-context premium is not a corner case — it is the price of actually using the window the model is famous for. Mercury 2.5 Preview has no such step, and its 256K context is unlikely to reach long-context territory often. But Mercury's low price is a teaser with a September 8 expiry, while Gemini's is a stable published rate. When comparing, compare Mercury's list price of $0.20 / $0.75 against Gemini's standard tier, not the discounted figure — the discount is the incentive to test, not the price to plan around.
The routing decision
Both models are routable today, and that changes how you should treat each one. Gemini 3.1 Pro is on OrcaRouter as google/gemini-3.1-pro-preview, at the provider's list price passed through at 0% markup, so the standard and long-context tiers cost exactly what the provider publishes, alongside 200+ other models on one API key. A preview badge is precisely the kind of thing to route around rather than bet on — the model page's error-rate panel is there for a reason, and automatic failover means a degraded preview response routes to a second provider without a code change. Mercury 2.5 Preview is not on OrcaRouter yet; Inception serves it through its own API and several third-party platforms. A two-day-old model you cannot yet put behind failover is a test candidate, not a production dependency. The day a provider lists it, the same pass-through applies and the intro discount lands here same-day — which is when this matchup becomes a routing rule instead of a decision.
The honest verdict is that these two previews answer different questions. Gemini 3.1 Pro answers "which model should I build my product on today" — it is multimodal, long-context, independently tested, and stable, at a price that reflects all of that. Mercury 2.5 Preview answers "how fast can text reasoning physically get" — and its speed architecture is the most interesting thing in the model, waiting for an independent lab to confirm whether the quality that comes with it is real. If your workload is multimodal or long-context, the choice is already made. If it is text-only, latency-compounding, and price-sensitive, Mercury 2.5 Preview is the bet to watch — and September 8 is the date to watch it against.

