
GPT-6 Sol vs GPT-5.6 Sol: One Index Point, Half the Price, and a Regression Nobody Announced
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 610 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 189 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1306 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 225 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
Yes, upgrade — but read the second paragraph before you do. GPT-6 Sol, released September 22, 2026, scores 48 on Artificial Analysis's Intelligence Index. GPT-5.6 Sol, released July 9, 2026, scores 47. One point. The model that replaced it costs half as much: $2.00 per million input and $10.00 per million output against $4.00 and $20.00. On the headline numbers this is the easiest upgrade decision in the current lineup, and the price cut is real, permanent, and the largest single reduction the vendor has made to a flagship.
The complication is that Artificial Analysis's own assessment of GPT-6 Sol found a mix of gains and regressions rather than a clean step up. There are regressions on GDPval-AA v2.1. There is also a long-context repricing tier above 272,000 input tokens that roughly doubles the input rate and raises output by half — and 272K is exactly the band where GPT-5.6 Sol has a published strength. A migration that looks like a straight halving of your bill can, on long-context traffic, come out closer to a wash with a capability trade attached.
What a one-point Index gain actually means
Both models were measured on the same board, version 4.3.2, which makes this the cleanest comparison in the whole GPT-6 Sol series — no harness differences, no vendor-versus-independent mismatch, one lab measuring two models from one vendor.
• AA Intelligence Index — GPT-6 Sol 48, ranked 18th of 212 vs GPT-5.6 Sol 47, ranked 19th of 212
• Cost per Index task — GPT-6 Sol roughly half of GPT-5.6 Sol
• Output speed — GPT-6 Sol 104.4 tokens/sec vs GPT-5.6 Sol 72.6 tokens/sec
• Time to first token — GPT-6 Sol 107.18s vs GPT-5.6 Sol faster, both against a 3.87s board median
• Input price — GPT-6 Sol $2.00 per 1M vs GPT-5.6 Sol $4.00 per 1M
• Output price — GPT-6 Sol $10.00 per 1M vs GPT-5.6 Sol $20.00 per 1M
• Cached input — GPT-6 Sol about 90% off vs GPT-5.6 Sol $0.40 per 1M
• Batch rate — GPT-6 Sol not published vs GPT-5.6 Sol $2.50 / $15.00
• Long-context tier — GPT-6 Sol reprices above 272K input vs GPT-5.6 Sol no equivalent step
• Context window — GPT-6 Sol 872K per AA, 1.05M claimed vs GPT-5.6 Sol roughly 1.05M to 1.1M
• Maximum output — 128K on both
• Released — GPT-6 Sol 2026-09-22 vs GPT-5.6 Sol 2026-07-09

Look at the two speed lines together. GPT-6 Sol generates output about 44% faster than its predecessor, and it is dramatically slower to produce a first token — 107.18 seconds against a board median of 3.87. GPT-5.6 Sol is not a fast model by board standards either, but it is not in that territory. If you migrated to GPT-5.6 Sol for an interactive product, GPT-6 Sol is not a drop-in replacement, and that is a functional regression independent of any benchmark.
One line is quietly favourable: GPT-6 Sol's standard rate of $2/$10 undercuts GPT-5.6 Sol's batch rate of $2.50/$15. Even the discounted tier of the old model is more expensive than the list price of the new one. That is the strongest evidence that the cut is structural rather than promotional.
The regression is the interesting part
Artificial Analysis found GPT-6 Sol roughly halves cost per task while producing a mix of gains and regressions. GDPval-AA v2.1 is the named regression. Separately, OpenAI's GPT-5.6 Luna also regressed on the Coding Agent Index — which suggests a pattern in this generation rather than an isolated miss on one model.
This is worth taking seriously precisely because it is the independent finding. OpenAI's launch materials are built around cost per task and around benchmark rows it selected; Artificial Analysis has no product to sell and reports what its harness produced. A model that is cheaper and roughly equal on the composite, but measurably worse on specific sub-tasks, is not a strictly better model. It is a different point on the frontier.
The practical version of that: if your workload is heavily GDPval-AA-shaped — economically valuable knowledge work, the kind of task that benchmark was built to price — then the one-point composite gain does not cover you, and the regression is the number that matters. If your workload is general, the halving of cost per task is the number that matters and the composite says you gave up nothing measurable.
Where GPT-5.6 Sol still wins
GPT-5.6 Sol has a published long-context retrieval result that GPT-6 Sol does not match with anything equivalent: MRCR v2, 8-needle, 91.5% at the 256K-to-512K range. That is a retrieval-accuracy figure in the band where GPT-6 Sol's rate card changes.
Put the two facts side by side. Above 272,000 input tokens, GPT-6 Sol's entire request reprices to roughly $4 input and $15 output. That is roughly GPT-5.6 Sol's old rate on input and three-quarters of it on output — so the 50% saving you migrated for collapses to about 0% on input and 25% on output, exactly in the context range where GPT-5.6 Sol has a documented strength and GPT-6 Sol has no published counterpart.
GPT-5.6 Sol's other published results remain respectable and are the reason some teams will not move:
• Agents' Last Exam — GPT-5.6 Sol 53.6
• Terminal-Bench 2.1 — GPT-5.6 Sol 88.8 / 88.0
• SWE-Bench Pro — GPT-5.6 Sol 64.6
• OSWorld 2.0 — GPT-5.6 Sol 62.6
• BrowseComp — GPT-5.6 Sol 90.4
• GDPval-AA v2 — GPT-5.6 Sol 1747.8 Elo
• HLE — GPT-5.6 Sol 49.5
• MRCR v2, 8-needle — GPT-5.6 Sol 91.5% at 256K to 512K
These are OpenAI-reported. Note that BrowseComp at 90.4 and the MRCR figure are the two that are hardest to replace: web-research agents and long-context retrieval are workloads where a two-month-old model with a published number is a safer bet than a new model with no counterpart measurement.
What GPT-6 Sol brings that is not on the rate card
The pricing is the story, but three other things changed.
First, the context ceiling. GPT-6 Sol is listed at 872,000 tokens by Artificial Analysis against a 1.05M claim elsewhere in OpenAI's materials, with 922K maximum input. GPT-5.6 Sol sits around 1.05M to 1.1M depending on source. On paper that is a small step backward in usable window — with the caveat that the 272K tier boundary, not the ceiling, is what governs cost.
Second, availability. GPT-6 Sol is in the API and in ChatGPT Work and Codex on Plus, Pro, Business, Enterprise and Edu — and Enterprise administrators have to enable it before users see it. It is not in ChatGPT Chat. GPT-5.6 Sol was a broader rollout. If your team's access path runs through an enterprise admin, the migration is not just a code change.
Third, the rollout of a new flagship generally means the old one becomes the cheaper fallback rather than being retired. GPT-5.6 Sol at $4/$20 is still an excellent model at 47 on the Index, and for long-context retrieval work it has a published number GPT-6 Sol does not.
A migration that costs less than it looks
The reason this particular upgrade is easy is that the two models are the same vendor, the same API shape, and adjacent in the ranking — which means the migration is a model-string change rather than a re-architecture, and the fallback path is already in your code. GPT-5.6 Sol is on OrcaRouter at OpenAI's list price, under the pass-through model that puts an OpenAI price change live on our side the same day rather than after a reseller renegotiates.
GPT-6 Sol is not in our catalogue as of this writing — it is reachable through OpenAI's own API. So the honest shape of a migration is a route with GPT-5.6 Sol behind it on the same key: send new traffic to GPT-6 Sol, keep the previous model one config change away, and let automatic failover cover the window where a new flagship's availability is still settling. For a model that is two days old, with an Enterprise enablement toggle in front of it and a long-context tier that changes the economics above 272K tokens, having the previous generation still wired in is not caution — it is the difference between a bad afternoon and a bad week.

The migration checklist
Measure your actual context distribution before anything else. If the 95th percentile of your requests sits below 272,000 input tokens, migrate — the saving is a genuine 50% on both sides of the bill, cost per task halves, and the composite index says you gave up nothing. If a meaningful share of traffic sits above that line, model the tier carefully: at roughly $4/$15 the advantage over GPT-5.6 Sol's $4/$20 is 25% on output and nothing on input, and you are paying for it with the loss of a documented long-context retrieval result.
Check whether anyone waits on the first token. 107 seconds is roughly twenty-eight times the board median and it is the failure mode that will surface first in production. Anything with a human on the other end should stay on GPT-5.6 Sol or route elsewhere.
Then check the enablement path. Enterprise plans need an administrator to turn GPT-6 Sol on, and it is absent from ChatGPT Chat entirely. Confirm your users can actually reach it before you announce a migration internally.
Finally, watch GDPval-AA v2.1 on your own evals for the first month. The independent lab found a regression there, and a one-point composite gain does not tell you whether your specific task set moved up or down. Run your evals on both models before you cut over, not after.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
