
Mercury 2.5 vs Mercury 2.5 Preview: One Diffusion Model, Two Launch Dates
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Here is a head-to-head in which both contenders share a model card. Inception's Mercury 2.5 Preview, released August 31, 2026, and Mercury 2.5, released September 8, 2026, are the same diffusion language model eight days apart: a preview channel Inception opened to developers, and the "enterprise-ready" production version it launched a week later. Inception has not disclosed a different checkpoint, a different speed figure, or a different price for Mercury 2.5 — and the numbers that did move between the two announcements look like a tidy-up, not a new model. The context figure was corrected from 256K on the preview page to 260K at launch, and the intelligence gain over Mercury 2 was reframed from "more than 10 points" (preview materials) to "40 percent" (launch materials). The engine, the 1,107-tokens-per-second claim, and the price sheet are identical. So this is not the usual spec-sheet fight, and padding one out would be dishonest. The comparison that matters between these two names is about packaging and timing: what the production launch actually changed, what the Preview label was really telling you, and which name your API calls should be using now.
Why a model is being compared with itself
Vendors usually preview a model and then quietly fold it into the main name; the preview listing disappears and nobody writes the matchup. Inception instead ran the two names side by side for a week, and the preview identity is only now being retired — which is exactly why this comparison has content. The Preview was the fast path Inception used to get its diffusion bet in front of developers. It appeared August 31 with the spec sheet that is now Mercury 2.5's spec sheet: a diffusion LLM (dLLM) that generates and refines blocks of tokens in parallel instead of emitting one token at a time; a claimed 1,107 tokens per second on standard GPUs; a sub-300ms time-to-first-token claim; a 260K-token context window with 65,536-token output; tunable reasoning levels; native tool use, parallel tool calls, and structured output. Every one of those figures is Inception's own, and no independent lab had measured the model when the Preview went out.
On September 8 Inception launched Mercury 2.5 as "the next tier of intelligence for diffusion LLMs" and its fastest reasoning model in production — enterprise-ready, aimed at voice agents, coding subagents, enterprise search, and support workloads. Same specs, same positioning, same price sheet. The launch also previewed two siblings that clarify where Inception is headed: Mercury Voice, a dLLM tuned to a sub-170ms first token for voice agents, and Mercury Router, which routes prompts across models by quality, speed, and cost. Mercury 2.5 is the flagship of a family built for latency-sensitive, agentic workloads rather than for the frontier-quality tier.
What the September 8 launch actually changed
Compare the two names line by line and the delta is not the engine — it is everything wrapped around the engine:
• Status — Mercury 2.5 Preview: a closed preview that could "change behavior or availability." Mercury 2.5: a production model launched with an enterprise-ready commitment, a sales channel, and a dated production listing.
• Identity — Inception's models page now lists Mercury 2.5, Mercury Voice, and Mercury Router, with no Preview in sight, and its support footnote names Mercury 1, Mercury 2, and Mercury Edit 2 as the models that "remain supported for existing customers." The Preview appears in neither list. As far as the vendor is concerned, the preview label has been absorbed into Mercury 2.5.
• Endpoint — Inception's developer platform serves the model as mercury-2.5, and the production listing that went live September 8 is where new traffic is served. On the catalog that first carried the Preview in late August, the Preview name now resolves to no live serving endpoint while the Mercury 2.5 listing added September 8 answers. Anyone who coded against the Preview by name should re-point at Mercury 2.5 and confirm what their provider is actually billing — the id you integrated in week one is the one being retired.
• Price — identical list and identical discount. Both are $0.20 per million input and $0.75 per million output, with cached input at $0.02, and both launched at roughly 80% off: $0.04 / $0.15, cached $0.004. The Preview's discount was explicitly time-boxed to September 8; Mercury 2.5 launched with the same 80% offer, and the discounted rate is what its production listing still shows at the time of writing. That is the trap, unpacked below.
• Onboarding — the production launch added a free-token allowance for new developer accounts — the launch materials cite 100 million tokens — and opened an enterprise contact path that a preview never had.
• Evidence — unchanged. Neither name has an independent benchmark, and the vendor claims are the same claims wearing slightly different wording: a "more than 10 points" intelligence jump over Mercury 2 in the preview materials becomes a "40 percent" increase in the launch materials, alongside the same stated parity with the cost-optimized frontier tier (GPT-5.6 Luna in low-effort mode, Gemini 3.5 Flash-Lite, Claude Haiku 4.5) and the same vendor-reported 79% on GPQA Diamond and 77% on IFBench. A claim restated is not a claim verified.

The pricing trap at the heart of the pairing
Because the two names carry the same list price and the same 80% teaser, it is easy to assume they always cost the same. They do not have to. The Preview's discount had a hard stop on September 8; Mercury 2.5's launch discount is running on a fresh clock. For a week it was entirely possible to be paying $0.20 / $0.75 for a call to the Preview while the same engine answered at $0.04 / $0.15 under its production name — or, more likely, to be on a preview id that had quietly stopped serving while the bill kept arriving. A launch discount that expires by endpoint is exactly the kind of fine print that turns a "same model, same price" story into a real cost difference.
The durable number is the list price, and the durable advice is to budget around it: $0.20 / $0.75 per million, cached input $0.02 — cheap for a reasoning model with a 260K context, comfortably under the flagship tier, but not the $0.04 / $0.15 that the launch banners advertise. Treat the discounted rate as a test-drive price with a lifespan, verify what your provider bills for the exact model id in your code rather than for the name on the marketing page, and if you onboarded on the Preview, migrate to Mercury 2.5 before an expiry or a retired endpoint decides for you.
This is also the class of problem a routing layer exists to remove. A model router that passes providers' list prices through at 0% markup shows the price a vendor actually bills, updated the same day a launch discount lands or expires, and automatic failover keeps calls flowing when an endpoint is retired underneath you. Mercury 2.5 is not on OrcaRouter as of this writing — Inception is serving it through its own API and several third-party platforms — so for now the vendor's models page is the source of truth on the rate. The moment a provider lists it, those pass-through mechanics are what you get on one key, and a price cut or a discount expiry shows up on our side the same day it is announced.
The evidence problem both names share
The honest scorecard for this pairing is short, because the two columns are the same model with the same absence of proof. Neither Mercury 2.5 nor Mercury 2.5 Preview has an independent index score, a reproduced run, or an arena placement as of September 9, 2026. Inception's own claims — the intelligence jump over Mercury 2, the parity with the Haiku 4.5 tier, the 79% GPQA Diamond, the 1,107 tokens per second — are the company's word, and they are identical for both names because the model is identical.
The one piece of independent evidence in the family points at how to read the speed claim. Artificial Analysis measured Mercury 2, the predecessor, at 684 tokens per second on the production endpoint, scoring it 22 on the Intelligence Index — genuinely fast, and proof the diffusion approach is not a mirage, but well short of the roughly 1,000 tokens per second Inception advertises for that model on Blackwell hardware. Expect a similar gap between Mercury 2.5's 1,107 figure and what a shared, multi-tenant endpoint actually delivers under real load. The same caution applies to quality: a brand-new diffusion model measured only by its maker should not be taken at its word on matching Claude Haiku 4.5 until a third party runs it.

The trackers that have started to cover the release reflect exactly that state of play. A BenchLM record for Mercury 2.5, posted September 8, gives the model no public score and no public rank; it carries Inception's 79% GPQA Diamond and 77% IFBench figures explicitly labeled as provider-reported, and notes that independent runtime speed has not been measured. The first Artificial Analysis index entry, an LMArena placement, or a reproduced GPQA run will turn Mercury 2.5 from a claim into a comparable object — and it will apply to both names at once, because there is only one model to measure.

Which name should you call
• New integration, no legacy: call Mercury 2.5 and ignore the Preview. The production name carries the same engine, the same discount for now, the enterprise terms, and the guarantee that it is the id the vendor is actually serving. The Preview adds instability risk and nothing else.
• Already integrated the Preview: check, today, what your provider serves under that name and what it bills for it. Re-point your client at the Mercury 2.5 id and confirm the rate. Keeping the Preview name in production code is now a bet that an explicitly time-boxed preview channel outlives its own retirement.
• Enterprise or production-critical traffic: Mercury 2.5 through Inception's enterprise path, not a preview label with no commitment behind it. Voice and routing needs point at Mercury Voice and Mercury Router respectively, which are still previews themselves.
• Anyone evaluating the diffusion tier: both names are the same evaluation target, so run your eval against Mercury 2.5 once and treat the result as applying to the model line. Budget at list price, and weight the vendor claims as hypotheses until an independent score lands.
What to watch
Three things will settle this pairing over the next few weeks. First, the first independent benchmark — an index placement or a reproduced GPQA or AIME run — which will test the parity claim that is the entire commercial bet. Second, whether the Preview identity disappears from the model ecosystem entirely, as Inception's own models page already implies. Third, what happens to the 80% launch discount now that the September 8 marker the Preview was pinned to has passed: an extension makes Mercury 2.5 structurally cheap, while a snap-back to $0.20 / $0.75 makes it merely competitive. Until then the correct answer to "Mercury 2.5 or Mercury 2.5 Preview?" is that it is one model, and you should be calling the one that is still a product.
