
GPT-6 Sol vs Gemini 3.1 Pro: The Price Gap Is Smaller Than Either Launch Post Admits
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Both models cost $2.00 per million input tokens. GPT-6 Sol charges $10.00 per million output; Gemini 3.1 Pro charges $12.00. On the headline rates, GPT-6 Sol and the six-month-old Gemini 3.1 Pro are within 20% of each other, which makes this the least dramatic price comparison in the current lineup — and the most misleading, because neither model actually bills at its headline rate for long-context work, and the two tier boundaries sit at different places and reprice in different directions.
That is the comparison worth doing properly. GPT-6 Sol shipped on September 22, 2026 at $2/$10, a 50% cut from GPT-5.6 Sol, with a long-context tier above 272,000 input tokens that roughly doubles input and raises output by half. Gemini 3.1 Pro has been available since February 19, 2026 at $2/$12, with a tier boundary at 200,000 tokens that takes it to $4/$18 and a batch rate of $1/$6. Same sticker, different rate card geometry, and the arithmetic below is what decides which one is cheaper for your traffic.
Two rate cards, drawn to scale
• Base input — GPT-6 Sol $2.00 per 1M vs Gemini 3.1 Pro $2.00 per 1M
• Base output — GPT-6 Sol $10.00 per 1M vs Gemini 3.1 Pro $12.00 per 1M
• Long-context threshold — GPT-6 Sol above 272K input vs Gemini 3.1 Pro above 200K input
• Long-context rate — GPT-6 Sol about $4.00 / $15.00 vs Gemini 3.1 Pro $4.00 / $18.00
• Repricing scope — GPT-6 Sol reprices the whole request vs Gemini 3.1 Pro reprices the whole request
• Cache — GPT-6 Sol roughly 90% off cached input vs Gemini 3.1 Pro cache pricing not restated at launch
• Batch — GPT-6 Sol batch rate not published vs Gemini 3.1 Pro $1.00 / $6.00
• Context window — GPT-6 Sol 872K per Artificial Analysis, 1.05M claimed vs Gemini 3.1 Pro 1M input
• Maximum output — GPT-6 Sol 128K vs Gemini 3.1 Pro roughly 64K to 66K
• AA Intelligence Index (v4.3.2) — GPT-6 Sol 48 vs Gemini 3.1 Pro 30

Two structural points fall out of that. First, the long-context boundaries are close to each other but not identical — 272K against 200K — and both apply the higher rate to the entire request rather than to the excess. A call that carries 210K tokens costs the premium rate on Gemini 3.1 Pro and the base rate on GPT-6 Sol. That 72,000-token band is where the two models diverge most on price.
Second, output room is not comparable. GPT-6 Sol allows 128K tokens of output; Gemini 3.1 Pro stops around 64K to 66K. If your workload generates long artifacts — full file rewrites, long-form drafts, extended chain-of-thought traces you actually keep — that ceiling matters more than the $2 difference on output pricing.
A worked example, because the headline rates hide the answer
Take a realistic monthly workload: 10 million input tokens and 2 million output tokens, all short-context.
• GPT-6 Sol — 10M input at $2.00 is $20.00, 2M output at $10.00 is $20.00. Total: $40.00
• Gemini 3.1 Pro — 10M input at $2.00 is $20.00, 2M output at $12.00 is $24.00. Total: $44.00
GPT-6 Sol is 9% cheaper. That is the entire gap, and it is small enough that a single engineering decision elsewhere in the stack can erase it.
Now make the same workload long-context — an agent that carries 300K tokens of working state on every call. Both models cross their thresholds.
• GPT-6 Sol — input at roughly $4.00 gives $40.00, output at roughly $15.00 gives $30.00. Total: $70.00
• Gemini 3.1 Pro — input at $4.00 gives $40.00, output at $18.00 gives $36.00. Total: $76.00
Both bills nearly double and GPT-6 Sol stays ahead, but the gap is still under 10%. Long context does not break this matchup; it just makes both options expensive.
The interesting case is caching and batching, where the two vendors made different bets. If 80% of your input is a stable prefix, GPT-6 Sol's roughly 90% cache discount changes the arithmetic substantially — 8M cached tokens at about $0.20 per million is $1.60, plus 2M fresh at $2.00 for $4.00, plus $20.00 of output, for a total near $25.60. Gemini 3.1 Pro's batch rate does the same kind of work from the other direction: at $1.00 input and $6.00 output the same 10M/2M workload lands at $10.00 plus $12.00, or $22.00.
So the honest summary is: GPT-6 Sol wins on interactive traffic and on cached-prefix workloads, Gemini 3.1 Pro wins on anything you can batch, and the difference between the two best cases is a few dollars per ten million tokens. This is not a price decision. It is a workload-shape decision.
The capability gap is the actual decision
Fifty-one Index points separate these models on the same board: GPT-6 Sol at 48, Gemini 3.1 Pro at 30. That is not close, and it is the reason the price parity is a curiosity rather than a competitive threat.
Artificial Analysis's Intelligence Index is a composite, and composites can flatter. But a gap that size holds up across the component benchmarks each vendor publishes:
• AA Intelligence Index — GPT-6 Sol 48, ranked 18th of 212 vs Gemini 3.1 Pro 30, ranked 81st of 212
• Output speed — GPT-6 Sol 104.4 tokens/sec vs Gemini 3.1 Pro 115.0 tokens/sec
• Time to first token — GPT-6 Sol 107.18s vs Gemini 3.1 Pro 32.96s, against a 3.87s board median
• GPQA Diamond — Gemini 3.1 Pro 94.3% (Google-reported)
• ARC-AGI-2 — Gemini 3.1 Pro 77.1% (Google-reported)
• HLE — Gemini 3.1 Pro 44.4% (Google-reported)
• SWE-bench Verified — Gemini 3.1 Pro 80.6% (Google-reported)
• Terminal-Bench 2.0 — Gemini 3.1 Pro 68.5% (Google-reported)
• LiveCodeBench Pro — Gemini 3.1 Pro 2887 Elo (Google-reported)
Read those Google figures carefully, because they are not weak. A 94.3% on GPQA Diamond and 77.1% on ARC-AGI-2 are strong results, and they were published in February. The reason the Index still puts Gemini 3.1 Pro thirty points down is that the Index weights agentic and long-horizon tasks heavily, and that is where the six months between these two releases shows. Gemini 3.1 Pro is a strong reasoning model. GPT-6 Sol is a strong agent.
The latency line is where Gemini 3.1 Pro wins outright and it is not a small win. A 33-second first token is slow in absolute terms — the board median is 3.87 seconds — but it is a third of GPT-6 Sol's 107 seconds. For anything a person waits on, Gemini 3.1 Pro is the only one of these two that is usable at all. GPT-6 Sol's 107-second first token is the number that will kill the most pilots, and it is worth noting that its throughput once it starts is excellent: 104.4 tokens per second is not meaningfully behind Gemini's 115.0. The model is not slow. It is slow to begin.
The thing Gemini 3.1 Pro has not escaped
Gemini 3.1 Pro has been in preview since February. Six months is a long time to carry a preview label, and Google has shipped 3.5, 3.6 and 3.8 Flash models in the meantime without promoting 3.1 Pro to general availability. On OrcaRouter the routable slug is still gemini-3.1-pro-preview, and a separate gemini-3.1-pro-preview-customtools variant exists alongside it.
That matters for procurement in a way benchmarks do not. Preview endpoints can change under you, deprecate on short notice, and carry different rate limits than a GA model. If you are building on Gemini 3.1 Pro today you are building on a preview, and you should price the migration risk accordingly — particularly when Google's own newer Flash models are already past it on several axes.
GPT-6 Sol does not have that problem, but it has a different one: availability at launch was narrower than the announcement implied. It is in the API, in ChatGPT Work, and in Codex on paid plans — and Enterprise administrators have to enable it before their users can see it. It is not in ChatGPT Chat. A flagship that skips the consumer surface is a flagship being aimed at developers.
Running the cheap one and the good one together
This is the rare comparison where the answer is obviously "both," and the pricing supports it. Gemini 3.1 Pro is the model you put in front of a human: fast first token, batch rates that make bulk jobs cheap, strong published reasoning scores. GPT-6 Sol is the model you put behind a queue: slow to start, excellent throughput, fifty-one Index points better, and priced within 20% of the model it is replacing on interactive work.
Gemini 3.1 Pro is on OrcaRouter at Google's list price, under the pass-through model that means a Google price change is live on our side the same day rather than waiting on a reseller to renegotiate. GPT-6 Sol is not in our catalogue as of this writing — it is reachable through OpenAI's own API — so a route spanning both means one key for Gemini and a direct OpenAI integration beside it. Routing between a fast preview model and a slower frontier model is exactly the case automatic failover was built for: when the preview endpoint degrades or gets deprecated, the route moves, and you find out from a log line rather than from your users.

The decision rule
Choose Gemini 3.1 Pro when a human is waiting, when the work is batchable, or when output stays short. A 33-second first token is not good, but it is three times better than GPT-6 Sol's, and the $1/$6 batch rate is the cheapest way to move a large corpus through a strong reasoning model. Accept that you are on a preview endpoint and plan for the migration.
Choose GPT-6 Sol when the work is agentic, long-horizon, and unattended, and when your context stays under 272K tokens. It scores 48 against 30, it produces 128K of output against roughly 64K, and its cached-prefix pricing at a 90% discount makes repetitive agentic loops unusually cheap. Its 107-second first token is a feature of the way it reasons, not a bug, and it only becomes a problem when you put it in front of someone.
Do not choose on price. The two are within 9% of each other on a short-context workload and within 10% on a long-context one. The nine-dollar monthly difference on a ten-million-token book is not a reason to pick either model, and the fifty-one-point Index gap is.

Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
