A hero title card for 'Mercury 2.5 Preview vs Gemini 3.1 Pro' with the subtitle 'Two previews, six months apart'. Left panel labelled 'Mercury 2.5 Preview' shows a lightning bolt, a '256K context' pill and a 'text only' tag and a '$0.04 intro' pill; right panel labelled 'Gemini 3.1 Pro' shows a multimodal glyph set (image, audio, video), a '1M context' pill and an 'AA Index 57 at launch' badge and a '$2.00 / $12.00' pill. A thin 'vs' badge sits between them. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Mercury 2.5 Preview vs Gemini 3.1 Pro: Two Previews, Six Months Apart

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Both models in this matchup carry the word "preview" in their names, and that is where the resemblance ends. Mercury 2.5 Preview and Gemini 3.1 Pro are both reasoning models you can call from an API today, but they are previews at opposite ends of their lives. Gemini 3.1 Pro, the flagship reasoning model, has been in preview since February 19, 2026 — six months of independent evals, public leaderboard placements, and production traffic behind it, with an Artificial Analysis Intelligence Index of 57 at launch, multimodal input across text, image, audio, and video, a 1M-token context, and a $2.00 / $12.00 price per million tokens. Mercury 2.5 Preview, Inception's diffusion LLM released August 31, 2026, is two days old: a text-only model with a 256K context, a claimed 1,107 tokens per second, a list price of $0.20 / $0.75 (about $0.04 / $0.15 under a discount through September 8), and not a single independent benchmark. Calling them both "previews" hides the actual question, which is how much a six-month evidence head start is worth against a fundamentally faster way of generating text.

What each model actually is

Gemini 3.1 Pro is a closed, multimodal, general-purpose reasoning flagship. It reads text, images, audio, and video, handles a 1M-token context window with a 64K output ceiling, and adjusts its reasoning depth dynamically — the vendor describes a four-tier reasoning intensity that scales with task complexity. Its launch numbers — ARC-AGI-2 at 77.1%, GPQA Diamond at 94.3%, SWE-bench Verified at 80.6%, Terminal-Bench 2.0 at 68.5% — are vendor-reported, but they sit on top of six months of independent scrutiny, including a re-measurement by Artificial Analysis on a stricter index scale where the model reads around 46.5. That re-measurement is exactly what a mature model earns: it has been tested hard enough for the index to get harder around it. Mercury 2.5 Preview is a diffusion model — it generates and refines tokens in parallel instead of one at a time — with tunable reasoning, parallel tool calls, and schema-aligned JSON, built for the latency-compounding workloads Inception names: search agents, voice pipelines, coding subagents. Every quality claim attached to it, including a 10-plus-point jump over Mercury 2, is the vendor's own, and no third party has measured it.

A screenshot of the Artificial Analysis model page for Inception's Mercury 2, the diffusion model Mercury 2.5 Preview builds on, showing its independently measured output speed of 684 tokens per second, its $0.25 / $0.75 per-1M pricing, its 128K context window, and its Intelligence Index score of 22.

The spec contrast

Price — Mercury 2.5 Preview: $0.20 / $0.75 per 1M in/out, ~$0.04 / $0.15 intro through Sep 8. Gemini 3.1 Pro: $2.00 / $12.00 per 1M, stepping to $4.00 / $18.00 above 200K input.

Context / output — Mercury 2.5 Preview: 256K context. Gemini 3.1 Pro: 1M context, 64K max output.

Inputs — Mercury 2.5 Preview: text only. Gemini 3.1 Pro: text, image, audio, video, file.

Architecture — Mercury 2.5 Preview: diffusion LLM, parallel token generation. Gemini 3.1 Pro: autoregressive, dynamic reasoning depth.

Speed — Mercury 2.5 Preview: 1,107 tokens/sec claimed. Gemini 3.1 Pro: no comparable headline claim.

Evidence — Mercury 2.5 Preview: vendor-reported only. Gemini 3.1 Pro: AA Index 57 at launch (46.5 newer scale) plus a long independent eval record.

A two-column scoreboard titled 'Mercury 2.5 Preview vs Gemini 3.1 Pro — the scoreboard'. Left column 'Mercury 2.5 Preview': 'Price: $0.20 / $0.75 (intro $0.04 / $0.15)', 'Inputs: text only', 'Context: 256K', 'Speed: 1,107 tok/s (claimed)', 'Independent score: none yet', 'Preview since: Aug 31, 2026'. Right column 'Gemini 3.1 Pro': 'Price: $2.00 / $12.00', 'Inputs: text, image, audio, video', 'Context: 1M', 'Speed: autoregressive', 'Independent score: AA Index 57 at launch', 'Preview since: Feb 19, 2026'. A footer reads 'Gemini index per Artificial Analysis; Mercury figures vendor-reported.' The OrcaRouter logo is composited in the bottom-right corner.

The multimodality gap is the decisive difference

For most applications this comparison is resolved before pricing or benchmarks enter the picture, because Mercury 2.5 Preview cannot see. Gemini 3.1 Pro reads screenshots, diagrams, audio, and video — it is the model you point at a PDF, a chart, a meeting recording, or a UI when you want an answer about something that is not already text. Mercury 2.5 Preview is text-only by design; a diffusion language model that generates text in parallel has no image or audio input path at all. For a document-heavy application, a vision agent, or any multimodal product, Gemini 3.1 Pro is the only candidate here, full stop. Mercury 2.5 Preview's text-only constraint is not a small print caveat; it is the boundary that decides which workloads the model is even eligible for.

Six months of testing versus two days

On evidence, this is not a contest, and it is important to say plainly why. Gemini 3.1 Pro has been picked apart by independent labs for six months; its score has been placed, re-measured, and argued about, which is what "trustworthy" looks like for a model. Mercury 2.5 Preview has one source of information — its maker — and that source claims a 10-plus-point intelligence jump over Mercury 2 and parity with the cost-optimized tier of the frontier. Both statements may be true. Neither has been confirmed by anyone outside Inception, and the model is too new for that to be anything other than expected. The asymmetry is not that one is better; it is that one is knowable and the other is not yet.

Where Mercury's speed actually wins

The honest case for Mercury 2.5 Preview is narrow, text-only, and real. Parallel token generation changes the latency math in a way no autoregressive model — including Gemini 3.1 Pro — can follow, because an autoregressive model is serial by construction. For a voice pipeline, the irony is that the input and output are audio but the thinking between them is text: a fast text reasoning layer is exactly what makes a voice agent feel responsive. For a search agent making repeated tool calls, and for a coding subagent firing many small completions per task, per-call latency compounds into perceived speed. At the intro price — $0.04 in, $0.15 out — Mercury 2.5 Preview is roughly 80 times cheaper than Gemini 3.1 Pro's standard output rate, which makes it not just faster but a different cost class for high-volume text reasoning. The catch that has to travel with every one of those sentences: the speed is claimed, the quality parity is claimed, and no independent measurement exists.

Pricing: Gemini's long-context premium versus Mercury's teaser

Gemini 3.1 Pro's rate card has a detail worth knowing before you compare the headline numbers: requests that pass 200K input tokens step from $2.00 / $12.00 to $4.00 / $18.00 per million. On a 1M-context model that long-context premium is not a corner case — it is the price of actually using the window the model is famous for. Mercury 2.5 Preview has no such step, and its 256K context is unlikely to reach long-context territory often. But Mercury's low price is a teaser with a September 8 expiry, while Gemini's is a stable published rate. When comparing, compare Mercury's list price of $0.20 / $0.75 against Gemini's standard tier, not the discounted figure — the discount is the incentive to test, not the price to plan around.

The routing decision

Both models are routable today, and that changes how you should treat each one. Gemini 3.1 Pro is on OrcaRouter as google/gemini-3.1-pro-preview, at the provider's list price passed through at 0% markup, so the standard and long-context tiers cost exactly what the provider publishes, alongside 200+ other models on one API key. A preview badge is precisely the kind of thing to route around rather than bet on — the model page's error-rate panel is there for a reason, and automatic failover means a degraded preview response routes to a second provider without a code change. Mercury 2.5 Preview is not on OrcaRouter yet; Inception serves it through its own API and several third-party platforms. A two-day-old model you cannot yet put behind failover is a test candidate, not a production dependency. The day a provider lists it, the same pass-through applies and the intro discount lands here same-day — which is when this matchup becomes a routing rule instead of a decision.

The honest verdict is that these two previews answer different questions. Gemini 3.1 Pro answers "which model should I build my product on today" — it is multimodal, long-context, independently tested, and stable, at a price that reflects all of that. Mercury 2.5 Preview answers "how fast can text reasoning physically get" — and its speed architecture is the most interesting thing in the model, waiting for an independent lab to confirm whether the quality that comes with it is real. If your workload is multimodal or long-context, the choice is already made. If it is text-only, latency-compounding, and price-sensitive, Mercury 2.5 Preview is the bet to watch — and September 8 is the date to watch it against.

A screenshot of the OrcaRouter model page for Gemini 3.1 Pro Preview (google/gemini-3.1-pro-preview) showing the model header, the Home > Models > Google breadcrumb, the February 19, 2026 preview start date, and the description of Google's frontier reasoning model.
© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube