
GPT-6.1 Sol vs Gemini 3.1 Pro: Seven Months of Preview Against Eight Days of Shipping
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 161 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 79 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 300 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Gemini 3.1 Pro has been on the vendor's price list for seven and a half months and has never left preview. Announced February 19, 2026, documented today at an endpoint ending in -preview, with no general-availability commitment and — more tellingly — no shutdown date either. GPT-6.1 Sol is eight days old, shipped September 29, 2026 at $2.00 per million input and $10.00 per million output tokens. Put them side by side on the Artificial Analysis Intelligence Index and the seven-month-old preview is not merely competitive, it is within 22 points of the newer model on the newest index revision — 29.7 against 51.8 on v4.3.2 — while costing 1.8× more per index task, $1.30 against $0.72. That combination is the reason this comparison is interesting and not a formality: on this board the older model is both behind on capability and ahead on cost, which is the wrong way round for a preview that was supposed to graduate. But the index revision caveat matters here more than anywhere else in this set, because a 29.7 and a 53 both get quoted on the same site and they are not the same measurement.
Both are large multimodal models with 1M-class context windows. Gemini 3.1 Pro takes text, images, video, audio and PDFs with a 65,536-token output ceiling and a 200,000-token price tier above which the rate doubles. GPT-6.1 Sol takes text, images and files with a 1,050,000-token window, a 128,000-token output ceiling, five reasoning effort levels and a long-context tier above 272,000 prompt tokens. One is a shipping product with a support commitment. The other is a preview with an open-ended lifecycle, and that difference turns out to be worth more than the index gap.
The index number needs its revision attached
Artificial Analysis re-runs its suite and republishes scores under the current index revision, which means the same model can carry two different numbers depending on when you read. For Gemini 3.1 Pro that gap is unusually wide: earlier house coverage of this model reported 57 on an older revision, while the page as it stands today reports 29.7 on v4.3.2, the ten-evaluation version that also carries GPT-6.1 Sol's 51.8. Both figures are real and both are on the same website.
What this means for a comparison is narrow and important. Only numbers from the same revision are subtractable. A 29.7 and a 51.8 from v4.3.2 produce a 22-point gap; a 57 and a 51.8 produce a five-point gap in the other direction. The 22-point figure is the one to use here, because it is the one where both models were measured by the same suite on the same scale — and because the cost-per-task figures, $1.30 against $0.72, come from that same run.
• Intelligence Index, v4.3.2 — GPT-6.1 Sol 51.8 vs Gemini 3.1 Pro 29.7
• Cost per index task — GPT-6.1 Sol $0.72 vs Gemini 3.1 Pro $1.30
• Cost to run the full suite — GPT-6.1 Sol $1,082 vs Gemini 3.1 Pro $1,935
• Context window — GPT-6.1 Sol 1,050,000 tokens vs Gemini 3.1 Pro 1,000,000 tokens
• Maximum output — GPT-6.1 Sol 128,000 tokens vs Gemini 3.1 Pro 65,536 tokens
• Input modalities — text, image and file on both, plus video, audio and PDF on Gemini 3.1 Pro
Two of those lines favour the preview and deserve to be said plainly. Gemini 3.1 Pro accepts video and audio, which GPT-6.1 Sol does not; if your pipeline ingests a screen recording or a call, that is a capability gap no score closes. And its 65,536-token output ceiling is half of GPT-6.1 Sol's, which matters for long-form generation but rarely binds in agentic work, where outputs are short and turns are many.
Why the newer number is the one that should move your decision
The price question is more interesting than the index question.
Gemini 3.1 Pro's rate card is $2.00 in and $12.00 out per million tokens, with cached input at $0.20 and cache writes at $0.375, and a tier boundary at 200,000 prompt tokens above which the rates become $4.00 and $18.00. GPT-6.1 Sol is $2.00 in and $10.00 out, with cached input at $0.10, cache writes at $2.50, and a boundary at 272,000 tokens above which its rates become $4.00 and $15.00. Headline input price is identical. Output price is 20% lower on GPT-6.1 Sol, and the long-context step is lower on both axes. Where the two cards separate is caching: $0.10 against $0.20 per million cached input tokens, with the preview model writing cache cheaper, at $0.375 against $2.50.

That last pair inverts the casual reading. If your workload is cache-heavy — a long fixed system prompt re-sent every turn, or a document re-read across a session — Gemini 3.1 Pro's cache-write line is dramatically cheaper and its cache-read line is double. Which model wins depends on how many times you re-read what you wrote, and that is a number your own logs carry and no benchmark publishes. What is not ambiguous is the index cost: on identical evaluation work, the eight-day-old model finished 33% cheaper, and the seven-month-old preview finished 79% more expensive than the newer model overall.
The preview badge is the operational fact
This is the part of the comparison that a score table cannot carry, and for a production deployment it outranks everything above.
Gemini 3.1 Pro has been in preview since February with no general-availability announcement and no shutdown date. Both absences matter and they cut in opposite directions. No GA means no lifecycle commitment: preview endpoints are not covered by the same deprecation-notice terms as generally available ones, and can be revised under a caller without a version bump, which is the standard reason to pin an exact model identifier in production rather than an alias. No shutdown date means you cannot plan a migration window either — the model has already outlasted the quarter in which everyone expected it to graduate, and this family has shut a Pro model down before.
GPT-6.1 Sol's position is the mirror image. It is eight days old, which means its independent evaluation record is one configuration deep and its failure modes are undocumented; there is no practitioner corpus to consult when it misbehaves. Its release itself was the news, and it ships under a support structure the preview does not have.
Those are different risks, not different amounts of risk. Choose the preview and you are betting that the vendor keeps serving an endpoint with no contractual obligation to. Choose the eight-day-old model and you are betting that one benchmark run and eight days of availability are enough evidence for your workload. The index will not help with either.
The one thing that would make this easy
A general-availability announcement for Gemini 3.1 Pro, with a model identifier out of the -preview namespace and a published lifecycle commitment, would resolve most of this. It would not close the 22-point index gap or the 1.8× per-task cost gap, but it would remove the depreciation risk that currently makes the comparison lopsided, and it would let a team adopt the model without pinning a preview identifier and hoping.
The absence of that announcement, seven and a half months in, is itself information. A preview that was meant to validate updates before going GA "soon", in the vendor's February wording, has had two full quarters to do so. Either the validation is ongoing, or the preview is the product — and a caller cannot tell which from the outside. What a caller can do is price the uncertainty and route accordingly.
Both on one key, so the answer can be per-workload

Gemini 3.1 Pro is on OrcaRouter at its own listed rates — $2.00 / $12.00, stepping to $4.00 / $18.00 above 200,000 prompt tokens, with cache reads at $0.20. GPT-6.1 Sol is on OrcaRouter at $2.00 / $10.00, stepping to $4.00 / $15.00 above 272,000 prompt tokens, with cache reads at $0.10. Provider list price on both, nothing added, one key and one bill.
That makes the awkward parts of this comparison cheap to resolve. Video and audio ingest stays on the preview because nothing else in the pair accepts it; the high-volume text and coding work goes to the newer model, which is cheaper per task and 33% cheaper on identical evaluation work. If the preview endpoint is revised or withdrawn, automatic failover moves those calls to another route rather than failing them — which is the specific hedge a preview's lifecycle risk calls for. And because both sit behind one OpenAI-compatible endpoint, the price comparison that actually matters can be run on your own traffic in an afternoon instead of argued from a rate card.

The defensible reading, then, is not which model is better. It is that these two are priced as near-equivalents while being seven months apart in lifecycle maturity and 22 points apart on the current independent board, and that the preview model's real advantages — video and audio input, cheap cache writes — are narrow and specific. If your workload needs those modalities, Gemini 3.1 Pro is the only option of the two and the index gap is the price of admission. If it does not, you are being offered a preview endpoint with an open-ended future at a per-task cost 80% above a model that shipped last week, and the burden of proof runs the other way.
