Hero title card reading "GPT-6 vs Gemini 3.1 Pro", with a chip reading "Gemini 3.1 Pro: in preview since 2026-02-19", two price lines reading "GPT-6 Sol $2 / $10" and "Gemini 3.1 Pro $2 / $12", and a footer reading "Index figures per Artificial Analysis v4.3.2; GPT-6 vendor-reported items labelled in text." The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

GPT-6 vs Gemini 3.1 Pro: Google's Pro Tier Hasn't Shipped a New Model Since February

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Gemini 3.1 Pro has been on Goo​gle's price list for seven and a half months and has never left preview. It was announced on 19 February 2026, it is still served under an endpoint named Gemini 3.1 Pro Preview, and as of today there is no general-availability commitment and no shutdown date — the two dates that would tell you where the model sits in its own vendor's plan. GPT-6 Sol, by contrast, shipped on 22 September 2026, and on 7 October 2026 it became the model behind the paid tiers of ChatGPT's global rollout to more than 1.2 billion weekly users.

So the headline comparison is not between two flagship tiers. It is between a shipping product and an open-ended preview, and that difference turns out to be worth more than the benchmark gap does. Put them side by side on Artificial Analysis's Intelligence Index v4.3.2 and Goo​gle's seven-month-old preview scores 29.7 against GPT-6 Sol's 47.6 while charging 1.25× more per completed index task — $1.30 against $1.04. Both of those numbers are real, and both need a caveat attached before you spend money on them.

What makes this worth reading rather than dismissing is that the preview does win things. It beats GPT-6 Sol on GPQA Diamond, on τ²-Bench agentic tool use, and on one long-context measurement — and it takes audio and video as input, which GPT-6 Sol does not. A seven-month-old preview that still takes two of four fights is not a model to write off. It is a model whose lifecycle, not its capability, is the problem.

The numbers, and the revision caveat that has to travel with them

Artificial Analysis re-runs its suite and republishes scores under the current index revision, which means the same model can carry two different numbers depending on when you read it. For Gemini 3.1 Pro that variance is unusually wide: earlier coverage of this model reported 57 on an older revision, while the page as it stands today reports 29.7 on v4.3.2. Both figures are genuine and both sit on the same website.

That matters because only figures from the same revision are subtractable. The 29.7 and the 47.6 are the pair to use here, since both come from v4.3.2 — the same run that scores Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7. The same-run reading, one line per dimension:

• Intelligence Index, v4.3.2 — GPT-6 Sol 47.6 vs Gemini 3.1 Pro 29.7
• Cost per completed index task — Gemini 3.1 Pro $1.30 vs GPT-6 Sol $1.04; the older model is 25% more expensive per finished answer
• Input price — $2.00 per 1M on both, level at the headline
• Output price — GPT-6 Sol $10.00 vs Gemini 3.1 Pro $12.00; Sol is 17% cheaper
• Long-context clause — GPT-6 Sol reprices the whole request at $4.00/$15.00 above 272,000 input tokens; Gemini 3.1 Pro steps to $4.00/$18.00 above 200,000
• Context window — GPT-6 Sol 1,050,000 tokens vs Gemini 3.1 Pro 1,048,576, effectively level
• Maximum output — GPT-6 Sol 128,000 tokens vs Gemini 3.1 Pro 65,536; Sol is nearly double
• Input modalities — GPT-6 Sol text, image, file vs Gemini 3.1 Pro text, image, audio, video, PDF

A two-column comparison scoreboard for GPT-6 Sol and Gemini 3.1 Pro, showing GPT-6 Sol at an Intelligence Index of 47.6, $1.04 per index task and a 128,000-token output ceiling, against Gemini 3.1 Pro at 29.7, $1.30 and 65,536 tokens, with input and output prices of $2.00/$10.00 against $2.00/$12.00 and a preview-since-2026-02-19 chip. A footer reads "Index figures per Artificial Analysis v4.3.2; Gemini 3.1 Pro is a preview product."

Two lines there deserve expanding because they invert the obvious reading. The cost-per-task line says the cheaper-looking rate card belongs to the model that is more expensive per finished job — Gemini 3.1 Pro's $12 output rate plus its reasoning volume does more work than its $2 headline input suggests. And the long-context step is worth re-reading on both sides: Gemini 3.1 Pro's tier begins at 200,000 tokens, lower than GPT-6 Sol's 272,000, but reprices output to $18.00 where Sol reprices to $15.00. Neither is a flat rate card, and the crossover points do not line up.

Where the preview still wins

Goo​gle's model is not a relic. Pull the evaluations both cards carry:

• GPQA Diamond — Gemini 3.1 Pro 94.1% vs GPT-6 Sol, no comparable figure published at this revision
• τ²-Bench agentic tool use — Gemini 3.1 Pro 95.6% vs GPT-6 Sol, not published on the same board
• Long-context recall — Gemini 3.1 Pro 82.0% vs GPT-6 Sol 83.7%, a 1.7-point spread
• SciCode — Gemini 3.1 Pro 58.7% vs GPT-6 Sol 57.6%, effectively level
• Terminal-Bench 4.0 — Gemini 3.1 Pro 4.0% vs GPT-6 Sol 43.9%; the one category where the preview is not competitive

A screenshot of the Artificial Analysis leaderboard: the top of the table under the Model, Context Window, Creator, Intelligence Index, Cost per Task, Tokens/s, First Chunk and Response headings - Claude Opus 5.5 (max with fallback) at 58, Claude Sonnet 5.5 at 56, Claude Fable 5.1 at 53, GPT-6 Astra (max) at 53 and $3.26, Gemini 4 Argon (high) at 53 and GPT-6.1 Sol (max) at 52 and $0.72 - with the Gemini 3.1 Pro Preview row set below it, showing an index of 30 at $1.30 per index task, 67M tokens generated and 22M output tokens at 23.40 tokens per second.

A 94.1% GPQA Diamond and a 95.6% agentic tool-use score are frontier numbers, not legacy ones, and they are the reason the index gap of 17.9 points overstates the practical distance. The index is a composite over ten evaluations, and Gemini 3.1 Pro's losses are concentrated where the suite is weighted heavily toward software engineering — Terminal-Bench 4.0 at 4.0% is a catastrophic outlier next to its 58.7% SciCode and its 73.8% Terminal-Bench 2.1. A model that scores near the top of one agentic benchmark and near the bottom of its successor is telling you about its training date, not its ceiling.

Audio and video input is the other concrete advantage, and it is a gate rather than a gradient. If your pipeline takes a meeting recording or a screen capture as input, Gemini 3.1 Pro can enter it and GPT-6 Sol cannot. No amount of index leadership substitutes for that.

What "still in preview" costs you, concretely

Preview status is not a cosmetic label, and the costs are the kind that only show up after you have built on something.

A preview endpoint carries no general-availability commitment, which means the vendor has not promised the model will still be there on a schedule you can plan around. It also carries no shutdown date, which sounds like the good version of that — nothing is being retired — but reads the other way too: Goo​gle has not published a lifecycle for a model that has been in this state since February while shipping Flash-tier models through the same period. Gemini 3.8 Flash, Gemini 3.5 Flash-Lite, the 3.6 and 3.8 Flash line: the Pro tier stayed where it was, and the successor arrived instead as Gemini 4 Argon, which Goo​gle announced on 30 September 2026 in a phased rollout to a named cohort of security partners with no general-availability date attached either.

That is the pattern this comparison is really about. If you build a durable pipeline on Gemini 3.1 Pro Preview, you are building on a model whose own vendor has not said when it will graduate, replace or retire it. GPT-6 Sol's position is the opposite: it is a generally available API product with a published rate card, and since 7 October 2026 it is also the default in the app most of your colleagues use. Vendor-reported and as yet unreproduced, Ope​nAI says that in ChatGPT GPT-6 composes answers using text, graphics, charts and interactive controls through a capability it calls Intelligent UI, answers that need web search begin 44% sooner than GPT-5.6 Instant, and the Extra High effort level begins responding as fast as GPT-5.6 at Medium. None of that changes an API integration. All of it is why GPT-6 is now the ambient default and the preview is not.

There is also a plain version of the argument. You are paying 25% more per finished task, at a lower capability score, for a model whose lifecycle is undeclared. The counter-argument is real — you are getting GPQA Diamond at 94.1% and audio input — but it is a specific argument for a specific workload, not a general reason to prefer the older model.

Testing a preview without betting on it

The mistake to avoid with a model in Gemini 3.1 Pro's position is to treat "should we use it" as a single binary decision. It is not. Both Gemini 3.1 Pro Preview and GPT-6 Sol sit on OrcaRouter's catalogue behind one Ope​nAI-compatible endpoint and one key, at 0% markup over provider list price — so the preview is something you can put in a real request path, measure against your own workloads, and keep or drop without a second vendor contract or a code change.

Automatic failover is what makes that safe for a model on a preview lifecycle. Route the audio and video traffic to Gemini 3.1 Pro where it is the only one of the two that can take the input; route the software-engineering and long-output traffic to GPT-6 Sol; and configure the fallback so that if the preview endpoint is deprecated, rate-limited or quietly repriced, the request lands on a generally available model instead of failing. That is the structure a seven-month-old preview earns — not avoidance, but a path that does not depend on it.

If you want to go further, the routing DSL composes several models into one logical call, so a request can be tried against Gemini 3.1 Pro and GPT-6 Sol and the answer selected by your own criteria rather than one vendor's. Model fusion does the same thing in the other direction — a panel of models answering one question — which is a reasonable way to find out where a preview's specific strengths are worth paying for before committing a production path to it.

A screenshot of the OrcaRouter model page for google/gemini-3.1-pro-preview, showing the Google vendor label, a 2026-02-19 release date, a 1M-token context window, a 65K-token maximum output, text, image, audio, video and file input, and pricing of $2.00 per million input tokens and $12.00 per million output tokens.

The verdict, and the number to watch

For new work, GPT-6 Sol is the safer pick on four of the six dimensions that decide a deployment, and it is cheaper per finished task while scoring 17.9 index points higher at the same revision. For workloads that need audio or video input, or where a 94.1% GPQA Diamond is the thing you are buying, Gemini 3.1 Pro Preview still wins and the preview label is a risk to manage rather than a reason to walk away.

The number to watch is not the index. It is the date on Gemini 3.1 Pro's endpoint name. When the preview suffix comes off — or when Argon reaches general availability with a real API identifier — this comparison changes shape entirely, and the seven-month gap between the Pro tier's last release and today becomes the thing the article is about.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily