
GPT-6 vs Gemini 3.1 Pro: Google's Pro Tier Hasn't Shipped a New Model Since February
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 65 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Gemini 3.1 Pro has been on Google's price list for seven and a half months and has never left preview. It was announced on 19 February 2026, it is still served under an endpoint named Gemini 3.1 Pro Preview, and as of today there is no general-availability commitment and no shutdown date — the two dates that would tell you where the model sits in its own vendor's plan. GPT-6 Sol, by contrast, shipped on 22 September 2026, and on 7 October 2026 it became the model behind the paid tiers of ChatGPT's global rollout to more than 1.2 billion weekly users.
So the headline comparison is not between two flagship tiers. It is between a shipping product and an open-ended preview, and that difference turns out to be worth more than the benchmark gap does. Put them side by side on Artificial Analysis's Intelligence Index v4.3.2 and Google's seven-month-old preview scores 29.7 against GPT-6 Sol's 47.6 while charging 1.25× more per completed index task — $1.30 against $1.04. Both of those numbers are real, and both need a caveat attached before you spend money on them.
What makes this worth reading rather than dismissing is that the preview does win things. It beats GPT-6 Sol on GPQA Diamond, on τ²-Bench agentic tool use, and on one long-context measurement — and it takes audio and video as input, which GPT-6 Sol does not. A seven-month-old preview that still takes two of four fights is not a model to write off. It is a model whose lifecycle, not its capability, is the problem.
The numbers, and the revision caveat that has to travel with them
Artificial Analysis re-runs its suite and republishes scores under the current index revision, which means the same model can carry two different numbers depending on when you read it. For Gemini 3.1 Pro that variance is unusually wide: earlier coverage of this model reported 57 on an older revision, while the page as it stands today reports 29.7 on v4.3.2. Both figures are genuine and both sit on the same website.
That matters because only figures from the same revision are subtractable. The 29.7 and the 47.6 are the pair to use here, since both come from v4.3.2 — the same run that scores Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7. The same-run reading, one line per dimension:
• Intelligence Index, v4.3.2 — GPT-6 Sol 47.6 vs Gemini 3.1 Pro 29.7
• Cost per completed index task — Gemini 3.1 Pro $1.30 vs GPT-6 Sol $1.04; the older model is 25% more expensive per finished answer
• Input price — $2.00 per 1M on both, level at the headline
• Output price — GPT-6 Sol $10.00 vs Gemini 3.1 Pro $12.00; Sol is 17% cheaper
• Long-context clause — GPT-6 Sol reprices the whole request at $4.00/$15.00 above 272,000 input tokens; Gemini 3.1 Pro steps to $4.00/$18.00 above 200,000
• Context window — GPT-6 Sol 1,050,000 tokens vs Gemini 3.1 Pro 1,048,576, effectively level
• Maximum output — GPT-6 Sol 128,000 tokens vs Gemini 3.1 Pro 65,536; Sol is nearly double
• Input modalities — GPT-6 Sol text, image, file vs Gemini 3.1 Pro text, image, audio, video, PDF

Two lines there deserve expanding because they invert the obvious reading. The cost-per-task line says the cheaper-looking rate card belongs to the model that is more expensive per finished job — Gemini 3.1 Pro's $12 output rate plus its reasoning volume does more work than its $2 headline input suggests. And the long-context step is worth re-reading on both sides: Gemini 3.1 Pro's tier begins at 200,000 tokens, lower than GPT-6 Sol's 272,000, but reprices output to $18.00 where Sol reprices to $15.00. Neither is a flat rate card, and the crossover points do not line up.
Where the preview still wins
Google's model is not a relic. Pull the evaluations both cards carry:
• GPQA Diamond — Gemini 3.1 Pro 94.1% vs GPT-6 Sol, no comparable figure published at this revision
• τ²-Bench agentic tool use — Gemini 3.1 Pro 95.6% vs GPT-6 Sol, not published on the same board
• Long-context recall — Gemini 3.1 Pro 82.0% vs GPT-6 Sol 83.7%, a 1.7-point spread
• SciCode — Gemini 3.1 Pro 58.7% vs GPT-6 Sol 57.6%, effectively level
• Terminal-Bench 4.0 — Gemini 3.1 Pro 4.0% vs GPT-6 Sol 43.9%; the one category where the preview is not competitive

A 94.1% GPQA Diamond and a 95.6% agentic tool-use score are frontier numbers, not legacy ones, and they are the reason the index gap of 17.9 points overstates the practical distance. The index is a composite over ten evaluations, and Gemini 3.1 Pro's losses are concentrated where the suite is weighted heavily toward software engineering — Terminal-Bench 4.0 at 4.0% is a catastrophic outlier next to its 58.7% SciCode and its 73.8% Terminal-Bench 2.1. A model that scores near the top of one agentic benchmark and near the bottom of its successor is telling you about its training date, not its ceiling.
Audio and video input is the other concrete advantage, and it is a gate rather than a gradient. If your pipeline takes a meeting recording or a screen capture as input, Gemini 3.1 Pro can enter it and GPT-6 Sol cannot. No amount of index leadership substitutes for that.
What "still in preview" costs you, concretely
Preview status is not a cosmetic label, and the costs are the kind that only show up after you have built on something.
A preview endpoint carries no general-availability commitment, which means the vendor has not promised the model will still be there on a schedule you can plan around. It also carries no shutdown date, which sounds like the good version of that — nothing is being retired — but reads the other way too: Google has not published a lifecycle for a model that has been in this state since February while shipping Flash-tier models through the same period. Gemini 3.8 Flash, Gemini 3.5 Flash-Lite, the 3.6 and 3.8 Flash line: the Pro tier stayed where it was, and the successor arrived instead as Gemini 4 Argon, which Google announced on 30 September 2026 in a phased rollout to a named cohort of security partners with no general-availability date attached either.
That is the pattern this comparison is really about. If you build a durable pipeline on Gemini 3.1 Pro Preview, you are building on a model whose own vendor has not said when it will graduate, replace or retire it. GPT-6 Sol's position is the opposite: it is a generally available API product with a published rate card, and since 7 October 2026 it is also the default in the app most of your colleagues use. Vendor-reported and as yet unreproduced, OpenAI says that in ChatGPT GPT-6 composes answers using text, graphics, charts and interactive controls through a capability it calls Intelligent UI, answers that need web search begin 44% sooner than GPT-5.6 Instant, and the Extra High effort level begins responding as fast as GPT-5.6 at Medium. None of that changes an API integration. All of it is why GPT-6 is now the ambient default and the preview is not.
There is also a plain version of the argument. You are paying 25% more per finished task, at a lower capability score, for a model whose lifecycle is undeclared. The counter-argument is real — you are getting GPQA Diamond at 94.1% and audio input — but it is a specific argument for a specific workload, not a general reason to prefer the older model.
Testing a preview without betting on it
The mistake to avoid with a model in Gemini 3.1 Pro's position is to treat "should we use it" as a single binary decision. It is not. Both Gemini 3.1 Pro Preview and GPT-6 Sol sit on OrcaRouter's catalogue behind one OpenAI-compatible endpoint and one key, at 0% markup over provider list price — so the preview is something you can put in a real request path, measure against your own workloads, and keep or drop without a second vendor contract or a code change.
Automatic failover is what makes that safe for a model on a preview lifecycle. Route the audio and video traffic to Gemini 3.1 Pro where it is the only one of the two that can take the input; route the software-engineering and long-output traffic to GPT-6 Sol; and configure the fallback so that if the preview endpoint is deprecated, rate-limited or quietly repriced, the request lands on a generally available model instead of failing. That is the structure a seven-month-old preview earns — not avoidance, but a path that does not depend on it.
If you want to go further, the routing DSL composes several models into one logical call, so a request can be tried against Gemini 3.1 Pro and GPT-6 Sol and the answer selected by your own criteria rather than one vendor's. Model fusion does the same thing in the other direction — a panel of models answering one question — which is a reasonable way to find out where a preview's specific strengths are worth paying for before committing a production path to it.

The verdict, and the number to watch
For new work, GPT-6 Sol is the safer pick on four of the six dimensions that decide a deployment, and it is cheaper per finished task while scoring 17.9 index points higher at the same revision. For workloads that need audio or video input, or where a 94.1% GPQA Diamond is the thing you are buying, Gemini 3.1 Pro Preview still wins and the preview label is a risk to manage rather than a reason to walk away.
The number to watch is not the index. It is the date on Gemini 3.1 Pro's endpoint name. When the preview suffix comes off — or when Argon reaches general availability with a real API identifier — this comparison changes shape entirely, and the seven-month gap between the Pro tier's last release and today becomes the thing the article is about.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
