A generated two-column comparison scoreboard titled 'GPT-6.1 Sol vs GPT-5.6 Sol - the scoreboard'. Left column 'GPT-6.1 Sol': rows reading 'Intelligence Index: 51.8', 'Cost per index task: $0.72', 'Input / output: $2.00 / $10.00', 'Cached input: $0.10', 'Output tokens on index run: 67M', 'Effort floor: low'. Right column 'GPT-5.6 Sol': rows reading 'Intelligence Index: 47.0', 'Cost per index task: $1.99', 'Input / output: $4.00 / $20.00', 'Cached input: $0.40', 'Output tokens on index run: 90M', 'Effort floor: none'. A footer reads 'Independent figures per Artificial Analysis v4.3.2 at max effort; vendor list prices.'
Guides & Insights

GPT-6.1 Sol vs GPT-5.6 Sol: Four and a Half Index Points, Half the Bill, Two Regressions

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Open​AI released GPT-6.1 Sol on 29 September 2026, eleven weeks after GPT-5.6 Sol, which shipped on 9 July 2026 as the flagship of the previous generation. Both are still on sale, both are still callable, and the difference between them on the neutral board is 4.8 points — 51.8 against 47.0 on Artificial Analysis's Intelligence Index, both measured at maximum reasoning effort on the same ten-evaluation suite. The price difference is much larger than the score difference: $0.72 per completed index task against $1.99, and $2.00/$10.00 per million tokens against $4.00/$20.00. That is the short version. The long version is that the upgrade is not uniform, and two of the evaluations moved backwards.

What each model is, and what replaced what

GPT-5.6 Sol was the top of the 5.6 line — a closed, API-only model with a 1,050,000-token context window, a 128,000-token output ceiling, text, image and file input, and a reasoning ladder running from none through max. It was described by Open​AI as the tier built for the hardest work: long-horizon agentic workflows, multi-file software engineering, deep reasoning, and it was priced accordingly at $4.00 and $20.00 per million tokens with cached input at $0.40 and cache writes at $5.00. Above 272,000 input tokens it repriced to $8.00 and $30.00, and it carried a Batch tier at $2.50 and $15.00.

GPT-6.1 Sol is the mid-tier of the generation that came after: below GPT-6 Astra, above GPT-6 Luna, and now the model Open​AI points cost-sensitive developers at. It keeps the 1,050,000-token window and the 128,000-token output cap, extends the knowledge cutoff from April 20 to April 30, 2026, and cuts the effort ladder from six rungs to five — low, medium (default), high, xhigh, max. The none value that GPT-5.6 Sol accepted now returns an error, and Open​AI's migration guidance is explicit that the substitution is low, not a silent upgrade.

This is worth pausing on because it is the one change in the release that costs money rather than saving it. A pipeline that reached for none to get a cheap, fast, non-reasoning call does not have that configuration any more, on either model family. The cheapest rung on GPT-6.1 Sol is a reasoning rung, and reasoning tokens are billed as output tokens whether or not you can see them.

The ladder, rung by rung

The independent board publishes a score and a run cost for every effort setting on both models, which turns the head-to-head into something more useful than a single comparison. Lined up at equivalent settings:

• Effort ladder, GPT-6.1 Sol — max 51.8 at $0.72 per index task; xhigh 51.0 at $0.39; high 50.2 at $0.32; medium (the default) 47.8 at $0.21 • Output speed — 59.7 tokens per second for GPT-6.1 Sol vs 87.8 for GPT-5.6 Sol
• Time to first chunk — 332 seconds for GPT-6.1 Sol vs a faster figure for GPT-5.6 Sol, against a board median near 4 seconds
• Output tokens on the index run — 67.2 million for GPT-6.1 Sol vs 90.0 million for GPT-5.6 Sol
• Input / output price — $2.00 / $10.00 vs $4.00 / $20.00
• Cached input — $0.10 vs $0.40; cache writes $2.50 vs $5.00
• Batch rate — not published for GPT-6.1 Sol vs $2.50 / $15.00 for GPT-5.6 Sol
• Long-context step — above 272,000 input tokens GPT-6.1 Sol reprices to $4.00 / $15.00 for the whole request vs GPT-5.6 Sol's $8.00 / $30.00
• Effort floor — low vs none

Two things fall out of that list. The first is the price cut is structural rather than promotional: GPT-6.1 Sol's standard rate of $2.00 and $10.00 undercuts GPT-5.6 Sol's batch rate of $2.50 and $15.00, so even the old model's discounted tier is more expensive than the new model's list price. The second is that the upgrade is not a straight line on latency. GPT-6.1 Sol generates output about a third slower and took 332 seconds to first chunk on the board's long-prompt tests — the same order of magnitude as GPT-6 Sol's 107 seconds and roughly eighty times the board median. If you built an interactive product on GPT-5.6 Sol, this is a functional regression independent of any benchmark, and the fix is architectural, not a parameter.

A generated hero card for a GPT-6.1 Sol versus GPT-5.6 Sol comparison, headed 'Four and a half index points, half the bill', with two badges reading 'GPT-6.1 Sol: $0.72 per index task' and 'GPT-5.6 Sol: $1.99 per index task', a footer reading 'Independent figures per Artificial Analysis v4.3.2 at max effort; vendor list prices.', and the OrcaRouter logo in the bottom-right corner.

Two evaluations moved the wrong way

The composite gain is 4.8 points. The components are not uniformly positive, and the board's per-evaluation scores say so:

• SciCode — GPT-6.1 Sol 0.542 vs GPT-5.6 Sol 0.571. A 2.9-point regression on scientific coding.
• GDPval-AA v2.1 — GPT-6.1 Sol Elo 1575.1 vs GPT-5.6 Sol Elo 1611.3. A 36-point Elo regression on economically valuable knowledge work.
• Humanity's Last Exam — GPT-6.1 Sol 0.529 vs GPT-5.6 Sol 0.495. A 3.4-point gain.
• Terminal-Bench 4.0 — GPT-6.1 Sol 0.561 vs GPT-5.6 Sol 0.399. A 16-point gain, the largest single move in the pairing.
• AA-Omniscience — GPT-6.1 Sol 41.5 vs GPT-5.6 Sol 22.0, which is about knowledge reliability and hallucination rather than raw recall.
• AA-Briefcase v1.1, AutomationBench-AA, AA-LCR v1.1, GDP.pdf, CritPt — the remaining five, which fill out the composite and are where the rest of the 4.8-point gain is earned.

A model that is meaningfully better on terminal work, substantially better on factual reliability, and measurably worse on scientific coding and on professional-work Elo is not a strictly better model. It is a different point on the frontier, and it was placed there deliberately: Open​AI's own launch framing for GPT-6.1 Sol is cost per completed task at near-flagship quality, not maximum score. If your workload is SciCode-shaped or GDPval-shaped, the composite number is not the one that applies to you, and the regression is.

GPT-5.6 Sol still has the published long-context result

The one place the older model has a number the newer one does not is retrieval accuracy in the context band where GPT-6.1 Sol's rate card changes. GPT-5.6 Sol carries a published MRCR v2 eight-needle result of 91.5% in the 256,000-to-512,000-token range. GPT-6.1 Sol has no published counterpart, and above 272,000 input tokens its entire request reprices at 2× input and cache and 1.5× output.

Put those two facts together and the segment where the new model's economics are weakest is exactly the segment where the old model has the only documented strength. The savings do not disappear — $4.00 and $15.00 on the repriced tier against $8.00 and $30.00 on GPT-5.6 Sol's own long-context tier is still a halving — but the half-price story is a general-workload story, and a long-context retrieval pipeline is not that. If that is your workload, GPT-5.6 Sol at $4.00 and $20.00 with a published 91.5% is a defensible place to stay.

GPT-5.6 Sol's other published results remain intact and are the reason some teams will not move: BrowseComp 90.4 for web-research agents, Terminal-Bench 2.1 at 88.8 and 88.0, SWE-Bench Pro 64.6, Agents' Last Exam 53.6, OSWorld 2.0 at 62.6. Those are Open​AI-reported, on Open​AI's harness, and they were reported in July. What has changed since is that the board now scores both models on one suite at one set of settings, and the older model loses on the composite and on cost.

The migration itself is a string change, with one exception

Same vendor, same API surface, adjacent in the ranking, same window and output cap — so for most stacks this is a model identifier and a price that falls by half. The exceptions are specific and worth naming before you flip the switch.

If your code sets reasoning.effort to none or minimal, it will fail rather than degrade, and low is the documented substitute. If your traffic runs above 272,000 input tokens, check the repriced tier before you model the savings. If your product depends on time to first token, measure before you ship, because the board's 332-second figure is not a rounding error. And if your evals have a SciCode-shaped or GDPval-shaped slice in them, run that slice on both models rather than trusting the composite.

Both models are on OrcaRouter at Open​AI's own list prices — GPT-6.1 Sol at $2.00 and $10.00 with cached input at $0.10, GPT-5.6 Sol at $4.00 and $20.00 with cached input at $0.40 — under a pass-through model that puts an Open​AI price change on our side the same day it lands rather than after a reseller renegotiates. Because the price list is passed through unchanged rather than marked up, the arithmetic in this article is the arithmetic you get: the same $0.72-per-task model, the same halving, no platform premium layered on either side.

A screenshot of the OrcaRouter model page for openai/gpt-5.6-sol, showing the listing dated 2026-07-09, the model described as the flagship of OpenAI's GPT-5.6 series for deep multi-step reasoning and long-horizon agentic workflows, a 1.05M-token context with 128K maximum output, text, image and file input, and the list price of $4.00 input and $20.00 output per million tokens.

The practical use of having both behind one endpoint is that the comparison stays live. Send new traffic to GPT-6.1 Sol, keep GPT-5.6 Sol one configuration change away as the fallback and as the long-context path, and let automatic failover cover the window in which a two-month-old flagship's availability and Enterprise enablement are still settling. A migration that halves the bill and moves two sub-benchmarks backwards is not a decision to make once and forget; it is one to make per route, and a routing layer is what makes that cheap.

A screenshot of the OrcaRouter model page for openai/gpt-6.1-sol, showing the listing dated 2026-09-29, a 1M-token context with 128K maximum output, text, image and file input, a reasoning effort ladder running from low to max, and the pass-through list price of $2.00 input and $10.00 output per million tokens with cached input at $0.10.

Where each one belongs

• General agentic and terminal work — GPT-6.1 Sol. Sixteen points on Terminal-Bench 4.0 and half the price is not a close call.
• Factual reliability, retrieved-answer products — GPT-6.1 Sol, on an AA-Omniscience gap of nearly twenty points.
• Scientific coding — check your own evals first. The board shows a 2.9-point regression, and the composite will not surface it.
• Economically valuable knowledge work — same caution. The GDPval-AA Elo regression is 36 points.
• Long-context retrieval above 272,000 tokens — GPT-5.6 Sol, until GPT-6.1 Sol has a published number in that band.
• Interactive, latency-sensitive surfaces — measure before migrating. Faster output, much slower first token.
• Cheapest non-reasoning call — neither. That configuration no longer exists on this family.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily