
GPT-6 Sol Pro vs Gemini 3.1 Pro Preview: Seven Months of Preview, and What That Costs You
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
Gemini 3.1 Pro Preview has been a preview for seven months. The vendor published it on February 19, 2026 at $2.00 per million input tokens and $12.00 per million output tokens, and as of this week the model card still describes it as "offered in preview form," warns that it "may have limitations in stability or availability compared to stable releases," and says it "could experience changes or deprecation without notice." There is no general-availability date and no shutdown date on the deprecation table. GPT-6 Sol Pro — the configuration of GPT-6 Sol, the model that reached general availability on September 22, 2026 at $2.00 input and $10.00 output — is the opposite arrangement: a production API with published pricing, a documented effort dial, and a vendor that has committed to the number.
The interesting comparison between these two models is not the benchmark table. It is what each vendor has promised you about next quarter.
Two price lines, one of which is a preview
Both are vendor lists. Google's and OpenAI's own numbers, passed through unchanged by the platforms that host them.
• Input — GPT-6 Sol $2.00 per million tokens vs Gemini 3.1 Pro Preview $2.00 per million tokens
• Output — GPT-6 Sol $10.00 per million tokens vs Gemini 3.1 Pro Preview $12.00 per million tokens
• Long context — Gemini 3.1 Pro Preview reprices input at $4.00 and output at $18.00 per million tokens above a 200,000-token threshold; GPT-6 Sol applies its own long-context reprice above a higher input threshold
• Cache read — GPT-6 Sol $0.20 per million tokens vs Gemini 3.1 Pro Preview $0.20 per million tokens
• Context window — GPT-6 Sol 1,050,000 tokens vs Gemini 3.1 Pro Preview 1,000,000 tokens
• Maximum output — GPT-6 Sol 128,000 tokens vs Gemini 3.1 Pro Preview 65,000 tokens
• Input modalities — GPT-6 Sol text and image vs Gemini 3.1 Pro Preview text, image, audio, video and file
• Availability status — GPT-6 Sol generally available since September 22, 2026 vs Gemini 3.1 Pro Preview in preview since February 19, 2026 with no GA date announced
The headline rates are close enough that they are not the decision. Input is identical. Output differs by 20%. The real differences are the ones in the last two lines: Gemini takes four more input modalities and twice the output ceiling in the other direction — 65,000 tokens against 128,000 — and has been shipped as a preview for the better part of a year.

The preview status is a real operational property
Google's own catalogue text tells builders to plan for preview deprecation. That sentence is doing a lot of work, and it is worth reading as an engineering constraint rather than a disclaimer.
A preview model is one you can build on and cannot plan around. Three consequences follow, and none of them appear in a price table.
• No stability commitment. The card states plainly that the model may have limitations in stability compared to stable releases, and that OrcaRouter does not provide specific latency guarantees for this preview model. Our own telemetry on the model page shows p50 and p95 time to first token both at 10.00 seconds over the trailing week, which is the top of the display range rather than a measured median.
• No deprecation date. A model with a shutdown date can be scheduled out of a roadmap. A model with neither a GA date nor a shutdown date cannot be scheduled at all, in either direction — you do not know when it will become safe to depend on, and you do not know when it will stop working.
• No roadmap position. Google has shipped newer Flash models in the same family that now outscore Gemini 3.1 Pro on independent composites, including Gemini 3.8 Flash on September 2, 2026. The 3.5 Pro generation was delayed and reportedly set aside. There is no Gemini 3.8 Pro. The model carrying the Pro name is no longer its vendor's top scorer, and it is still a preview.
That is the honest framing for anyone choosing between these two. GPT-6 Sol is six days old and has a flat, published, permanent price. Gemini 3.1 Pro Preview is seven months old and has a status line instead of a roadmap.
What the independent record shows, and why the version label is the whole argument
Artificial Analysis publishes a direct comparison page for this exact pair on the current index, which makes it one of the few matchups in this batch where the neutral numbers are genuinely like-for-like. GPT-6 Sol at max effort scores 48 at $1.06 per index task. Gemini 3.1 Pro Preview scores 30 at $0.67 per index task.
Eighteen points is a large gap, and it is not the gap the model had in February. At launch, Artificial Analysis placed Gemini 3.1 Pro Preview first on its Intelligence Index at 57. The model has not changed since. The index has — it was re-scored on September 19, 2026, and the same unchanged model now reads 30. Both numbers are Artificial Analysis; neither is wrong; putting them next to each other without the version is the single most common error in the coverage of this model, and it is why so many comparison pages describe a February leader as a September equal.
The per-task cost tells a second story worth doing the arithmetic on. Gemini is cheaper per task — $0.67 against $1.06 — but it scores lower. Cost per index point comes out at $0.022 for GPT-6 Sol and $0.022 for Gemini. On that measure the two are a dead heat, which is the cleanest way to say what is actually happening here: Gemini is not a cheaper way to buy the same intelligence, it is a cheaper way to buy less of it.
Two other independent numbers belong beside those, both labelled. On OSWorld-Verified, the computer-use benchmark, Google's own evaluation page reports 76.2% for Gemini 3.1 Pro. That figure appears on Google's materials and in third-party summaries of them, and it does not appear on the independently verified OSWorld board, which lists other Gemini models — 3.6 Flash at 83%, 3.5 Flash at 78.4% — and no 3.1 Pro row at all. On the harder OSWorld 2.0 board the model scores 30.6%. And on GDPval-AA the model sits at roughly 1,316 Elo, behind Claude Sonnet 4.6 and Claude Opus 4.6.
This is not an accusation of bad faith; self-reported numbers are normal and both vendors publish them. It is a statement about what a seven-month-old preview actually gives you: enough time for a vendor-reported figure to circulate widely, and not enough independent replication for the figure to be checked. GPT-6 Sol's neutral record, by contrast, is short, unambiguous, and six days old.
A fifteen-month knowledge gap nobody mentions
Gemini 3.1 Pro Preview has a knowledge cutoff of January 2025. GPT-6 Sol's is April 20, 2026 — fifteen months later.
That gap is absent from every comparison page on this matchup, and it matters more than the output-price difference for a specific class of work: anything that involves recent APIs, recent library versions, recent events, or recent organisational facts. A model with a January 2025 cutoff will answer confidently about a world that has moved on by a year and a quarter, and the failure mode is not an error message — it is a plausible wrong answer. For retrieval-augmented work the cutoff is nearly irrelevant because the context carries the facts. For anything that leans on parametric knowledge, it is the first constraint to check and the last one anyone checks.

The modality gap is the strongest argument for Gemini, and it is real
It would be easy to write this matchup as a straightforward win for the newer model. It is not, and the reason is the input line.
Gemini 3.1 Pro Preview accepts text, image, audio, video and file input natively on the same endpoint. GPT-6 Sol accepts text and image. For a workload built around video understanding, meeting transcription, or document processing that spans formats, that is not a feature difference — it is the difference between one API call and a pipeline of three, with the intermediate steps billed separately and the failure modes multiplied.
The honest caveat attached to that: a preview model with a 77.6% error rate displayed on its hosted model page is a preview model whose multimodal advantage has to be weighed against whether it will answer. That figure is as displayed on the page and it is high enough to warrant verification before anyone builds on it — which is exactly the point. Seven months in, the honest advice on this model is still "measure it yourself," and that is a sentence nobody writes about a GA release.
Where a routing layer changes the arithmetic
Gemini 3.1 Pro Preview is on OrcaRouter's catalogue at $2.00 input, $12.00 output and a $0.20 cache read per million tokens, with a 1,000,000-token context window, a 65,000-token output cap, the /v1/chat/completions and /v1beta/models/{model}:generateContent endpoints, and p50 time to first token of 10.00 seconds as measured on our own model page over the trailing week. Provider list pricing passes through with no markup on top.
GPT-6 Sol is not one of our routes, so nothing here is a claim about its price through us.

A preview model is the textbook case for routing rather than committing. Both the reason to try Gemini 3.1 Pro Preview and the reason not to build a production path on it are the same property — it can change without notice — and a single endpoint with automatic failover is how you take the first without accepting the second. The request shape does not change when the model behind it does, so the fallback path is a configuration entry rather than a migration. That matters more here than in most matchups, because the model most likely to replace this one in a fallback chain is a newer Flash from the same vendor, on the same key, at a fraction of the output rate.
The decision rule
• Video, audio or mixed-format input on one call — Gemini 3.1 Pro Preview. Nothing in the GPT-6 Sol column accepts a video file, and no price difference compensates for an ingestion pipeline you have to build.
• Anything with an availability commitment attached — GPT-6 Sol. Six days old, generally available, published price described as permanent, and a documented effort dial. That is what schedulable looks like.
• Long generations — GPT-6 Sol, on the output ceiling rather than the rate. 128,000 tokens against 65,000 is the constraint that bites first on synthesis work, and it bites before the 20% output-price difference does.
• Evaluating Gemini 3.1 Pro Preview — do it behind a failover path, and measure the error rate yourself before you size the workload. Seven months of preview with a displayed 77.6% error rate on the hosted page is not a settled instrument.
• If your plan depends on this model being cheap next quarter — neither. Gemini's preview can be repriced or retired without notice, and GPT-6 Sol's price is permanent only in the sense that OpenAI has said so this week. A flat rate is a commitment, not a guarantee; it is simply a stronger commitment than a preview card makes.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
