
GPT-6.1 Sol vs GPT-6 Sol: What One Week of Extra Work Bought, and What It Didn't
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 397 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1141 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 106 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
If you have GPT-6 Sol in production today and you are trying to decide whether GPT-6.1 Sol is worth a migration, the answer depends entirely on which of your costs is bigger — and the two models are built so that the answer is almost never "both." GPT-6 Sol shipped on September 22, 2026 at $2.00 per million input tokens, $0.20 cached, and $10.00 output. GPT-6.1 Sol shipped on September 29, 2026 at $2.00 input, $0.10 cached, and $10.00 output, with a 6.4-percentage-point vendor-reported improvement on complex software-engineering tasks at a lower reasoning effort. The input and output prices did not move. The cache rate halved. That single changed line is the whole financial case, and everything else is a capability argument that, at this point in the release's life, comes from exactly one source.
Three things changed. One of them is a bill.
Strip the marketing and the delta between these two deployments is short enough to hold in your head.
• Cached input — $0.10 per million tokens on GPT-6.1 Sol vs $0.20 on GPT-6 Sol. Measured rather than claimed, and the only change that shows up as arithmetic.
• Capability, per the vendor — OpenAI reports a 6.4-point lead on DeepSWE v1.1 at a lower reasoning effort and cost, more than double the score on Terminal-Bench Science 0.1 at maximum effort, seven points on the OSWorld 2.0 offline set at maximum effort, 4.8 points on AutomationBench at medium effort, and a factual-error rate falling from 11.4% to 7.7% on a deliberately error-inducing prompt set. All of it unreproduced.
• Reasoning ladder — GPT-6.1 Sol supports low, medium, high, xhigh and max, and does not support none; GPT-6 Sol supports all six. This is the delta that breaks code, and it is the one nobody puts in a launch post.
Everything else is identical or near enough: the same 1,050,000-token context window, the same 128,000-token output ceiling, the same $2.50 cache-write rate, the same 272K-token repricing threshold above which an entire request bills at 2× input and cache and 1.5× output, the same Responses-API requirement for tool calling, and the same absence of fine-tuning support. The knowledge cutoff moved from April 20 to April 30, 2026 — ten days, which is the kind of figure that tells you how much of a refit this was rather than a rebuild.
Why the cache line is worth more than it looks

A halved cache read is a strange thing to lead a product refresh with, until you look at what agent traffic actually looks like. An agent loop resends a stable prefix — system prompt, tool schemas, retrieved context, conversation history — on every turn, and on a long session the same prefix is billed dozens or hundreds of times. That is the traffic where a cache rate stops being a rounding error and becomes the dominant line on the invoice.
Run the numbers on a single long agent session. Assume a 120,000-token prefix reused across 40 turns, so 4.8 million cached input tokens per session, plus a modest 150,000 fresh output tokens. On GPT-6 Sol the cached reads bill $0.96 and the output $1.50 — roughly $2.46 per session. On GPT-6.1 Sol the same session bills $0.48 for cached reads and the same $1.50 for output: about $1.98. Per session the difference is 48 cents, which is not a decision. Multiply it across a hundred thousand sessions a month and it is $48,000 — which is.
Two caveats keep that honest. Cached reads only bill at the cache rate when the prefix actually hits, so your realized saving is your hit rate times the difference, not the difference itself. And the prefix has to be under 272,000 tokens for the standard rates to apply at all: cross that line and the entire request reprices at 2× input and cache and 1.5× output, which can swamp a 10-cent saving on the prefix. Long-context sessions are the ones where the cache math is least likely to look like the brochure.
The benchmark gap is real, and it comes from one place
OpenAI's launch post is unusually specific about its evaluations, which is a point in its favour, and every one of those numbers is the vendor grading its own model, which is the point against. The pattern across them is consistent: the comparison that travels is always against GPT-6 Astra, and the Astra comparison is always a cost comparison. On DeepSWE v1.1 the claim is matching Astra at roughly one-fifth the cost. On GDP.pdf, approaching Astra's state-of-the-art at roughly one-fifth the cost per task. On OSWorld 2.0, within 2.1 points of Astra at roughly one-seventh the cost per task. On factuality, within 1.9 points of Astra at less than one-fifth the cost per task. That is not an accident of drafting; it is the product thesis, and it is a thesis about price, not about capability leadership.
The one place OpenAI declines its own framing is worth crediting: on Terminal-Bench Science 0.1 it states plainly that GPT-6 Astra still holds the top score among the models tested at 68.1% and should be used for the hardest scientific research. A launch post that tells you when to buy the more expensive sibling is a launch post with some discipline in it.
Against GPT-6 Sol specifically, the honest state of the comparison is this — there is no independent score for either model at this configuration. Artificial Analysis has a full evaluation of GPT-6 Sol at maximum reasoning effort: an Intelligence Index of 48, $1.06 per index task, 77 million output tokens generated against a board median of 88 million, and a Coding Agent Index of 57 at $2.99 per task. It has no GPT-6.1 Sol entry at all as of September 30; the model's slug 404s and the string does not appear in the live leaderboard. So the only measured, third-party comparison available today is between GPT-6 Sol and other labs' models — not between GPT-6 Sol and its own successor. Anyone showing you a 6.1-versus-6.0 chart right now is showing you OpenAI's slide.
The migration is a string change, except where it isn't
OpenAI's own migration guidance for the GPT-6 family is worth reading before you flip the identifier, because two items in it break working code. The first is the reasoning ladder: if any request sends reasoning_effort: "none", GPT-6.1 Sol rejects it, and the documented remedy is to start at low and compare on representative tasks rather than assume the lowest setting is equivalent. Workflows that used none as a latency baseline lose that baseline. The second is tool calling: GPT-6.1 Sol supports Chat Completions but not tool calling through it — tools require the Responses API. If your stack calls functions on /v1/chat/completions, you are not swapping a string, you are porting an endpoint. The vendor's own instruction to developers already on GPT-6 Sol is to review that guidance before switching, which is a fair signal that this is not the drop-in refresh the one-week gap implies.
The remaining parameters behave: reasoning effort, structured outputs, streaming, prompt caching and the tool set all carry across, and the same 272K repricing rule applies to both models identically. Set model to gpt-6.1-sol, keep your effort setting where it was supported, and remove temperature, top_p and top_logprobs whenever effort is not none — that last one applies to both models and is a common source of 400s after any reasoning-model migration.
Migrating without betting a production path on it

The realistic migration is not a cutover, it is a shadow run: send a slice of production traffic to the new identifier, keep the old one as the path that answers, and compare on your own tasks. That is a routing problem, and it is the one place where the platform you call through changes the shape of the work. OrcaRouter serves GPT-6 Sol today at OpenAI's own list price with zero markup — the $2.00 / $0.20 / $10.00 card, the 272K repricing rule included — so the baseline arm of the comparison is live on the same key as everything else, and the 6.1 arm can be added to the route set as soon as it is callable, without a second contract or a second SDK. Until then the honest position is that GPT-6 Sol is the model available to call, and it is worth saying plainly rather than implying otherwise: openai/gpt-6.1-sol returns "model not found" on our public catalogue endpoint today. Automatic failover is what makes the shadow run safe when it does land — if the new route errors, the request completes on the old one and you find out from the logs rather than from your users.
Who should move now: teams whose cost is dominated by cached input on long agent sessions, and teams whose workflows already use the Responses API and never send none. For them this is close to a free upgrade — the vendor reports the capability gains, the cache rate is halved, and the migration is one string. Who should wait: teams with none in a latency-critical path, teams whose tool calling runs through Chat Completions, and anyone who needs a number from outside OpenAI before committing. For that last group the wait has no published end date, and the sensible thing to do with the time is build the eval set the comparison will need — because when the independent score arrives, it will be for someone else's traffic.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
