Hero card titled Claude Sonnet 5.5 vs GPT-6 Sol, subtitled the $2 tier has two flagships in it now, with pills reading same price: $2 / $10, Index 56 vs 47, and 1M vs 1.05M context, and a footer citing Artificial Analysis v4.3.2 and vendor pricing.
Guides & Insights

Claude Sonnet 5.5 vs GPT-6 Sol: The $2 Tier Has Two Flagships In It Now

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Sonnet 5.5 and GPT-6 Sol are priced identically — $2 per million input tokens, $10 per million output tokens — and they arrived six days apart in late September 2026. That coincidence is not the interesting part. The interesting part is that when Artificial Analysis measured both on its v4.3.2 Intelligence Index, the mid-tier Anthro​pic model came out ahead of the model Open​AI has been selling as its flagship: 56 for Claude Sonnet 5.5 against 47 for GPT-6 Sol. On the same revision, Claude Sonnet 5 — the model Sonnet 5.5 replaces — sits at 38.

So this is no longer a comparison of a cheap tier against an expensive one. It is a comparison of two models at the same price, from two labs, with an 18-point spread between the Anthro​pic one and its own predecessor, and a nine-point spread in Anthro​pic's favour against Open​AI's top-end release. Everything below is an attempt to work out which of those gaps survives contact with a real workload, and which of them is a benchmark artefact.

What each model actually is

Claude Sonnet 5.5 is the second model in Anthropic's 5.5 generation, released on September 28, 2026, five days after the Claude Opus 5.5 launch that promised it. It carries the id claude-sonnet-5-5 across the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry, keeps the tier's 1M-token context window and 128K output ceiling, and keeps Claude Sonnet 5's prices exactly: $2 in, $10 out, $0.20 per million for cache reads, $2.50 for a five-minute cache write. It is available with zero data retention from day one, and Anthropic's retirement commitment runs to September 28, 2027.

Two-column scoreboard for Claude Sonnet 5.5 and GPT-6 Sol across six rows. Left column Claude Sonnet 5.5: input $2.00 per million, output $10.00 per million, 1M-token context, AA Intelligence Index 56, cost per Index task $7.60, released September 28 2026. Right column GPT-6 Sol: input $2.00 per million, output $10.00 per million up to 272K then $4 / $15, 1,050,000-token context, AA Intelligence Index 47, cost per task not published, released September 22 2026. A footer line reads Index and cost per task per Artificial Analysis v4.3.2; prices per the vendors.

GPT-6 Sol is the successor to the model confusingly called GPT-5.6 Sol — the one OrcaRouter still lists as reachable at $4 in and $20 out, and which Artificial Analysis has since marked deprecated with a pointer to GPT-6 Sol. The new model landed on September 22, 2026 at $2 / $10 with a 1,050,000-token context window, a 128K output ceiling, and a tiered rate: the $2 / $10 headline applies up to 272,000 input tokens, and everything above that bills at $4 / $15.

That last detail is the first place the "identical pricing" framing breaks.

The price tie is only true for short prompts

Both rate cards quote $2 and $10, and almost every launch-day comparison stopped there. Line the two up on the terms that decide a bill:

• Headline rate — Claude Sonnet 5.5 $2 / $10 vs GPT-6 Sol $2 / $10

• Long-context rate — Sonnet 5.5 bills the full 1M window at the standard rate; GPT-6 Sol doubles to $4 / $15 above 272,000 input tokens

• Context window — 1,000,000 tokens vs 1,050,000 tokens

• Maximum output — 128K tokens on the synchronous API for both; Sonnet 5.5 reaches 300K on the Message Batches API behind the output-300k-2026-03-24 header

• Batch discount — 50% off input and output for Sonnet 5.5; check OpenAI's current terms for Sol rather than assuming parity

• Cache reads — $0.20 per million on both

An 800,000-token repository sweep costs twice as much per token on GPT-6 Sol as on Claude Sonnet 5.5, and that is precisely the workload both models are marketed for. If your prompts are short, the rate cards really are a tie and price should not decide anything. If they are long, it already has.

Anthropic's own scoreboard, read carefully

Anthropic published a four-column comparison against Claude Sonnet 5, Claude Opus 5.5 and GPT-6 Sol. The vendor-reported figures, which no third party has yet reproduced:

Artificial Analysis model page for Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback), released September 2026, showing an Intelligence Index of 56, a price of $2.00 per million input tokens and $10.00 per million output tokens, a cost of $7.60 per Intelligence Index task, a verbosity of 410M output tokens against a median of 88M, and a 1M-token context window.

• Terminal-Bench 4.0 — Sonnet 5.5 70.6% vs Sonnet 5 10.3%

• FrontierCode 1.1, Main — Sonnet 5.5 46.2% at Max effort, Sonnet 5 42.4%, Opus 5.5 54.4%, GPT-6 Sol 49.3%

• CursorBench 4.0 — Sonnet 5.5 55.5% vs Sonnet 5 34.1% vs Opus 5.5 57.8%

• GDPval-AA v2.1 — Sonnet 5.5 1844 vs Sonnet 5 1449 vs Opus 5.5 1846 vs GPT-6 Sol 1487

• AA-Briefcase v1.1 — Sonnet 5.5 1811 vs Sonnet 5 1359 vs Opus 5.5 1822 vs GPT-6 Sol 1483

• Humanity's Last Exam, with tools — Sonnet 5.5 64.5% vs Sonnet 5 54.9% vs Opus 5.5 67.7%

• OSWorld 2.1, partial — Sonnet 5.5 80.1% vs Sonnet 5 57.0% vs Opus 5.5 81.8%

• Chartography, no tools — Sonnet 5.5 61.6% vs Sonnet 5 15.6% vs Opus 5.5 64.4% vs GPT-6 Sol 53.6%

Two of those rows carry footnotes and both matter. The GDPval-AA, AA-Briefcase and Chartography figures for GPT-6 Sol are flagged by Anthropic as measured after an image-understanding fix on OpenAI's side, and the FrontierCode number for Sonnet 5.5 is its Max-effort result, which Anthropic notes fell below its own Xhigh result on the same evaluation. A vendor comparing itself to a rival on a benchmark it runs itself, with footnotes qualifying both sides, is not neutral evidence. It is the best evidence available on launch day, and it is not the same thing as measurement.

What the independent numbers say, and what they do not

Artificial Analysis has published a first independent run: Index 56 on v4.3.2, cost of $7.60 per Index task, and a verbosity figure of 410M output tokens against a median of 88M. That verbosity number deserves more attention than it is getting. It means the model tends to spend a great many tokens on the way to an answer, and it is the mechanism behind the one efficiency claim that does not come from Anthropic's own table.

Anthropic says Claude Sonnet 5.5 generates output 30% faster than Claude Sonnet 5 and costs "up to 30% less per task" at the same per-token rate. Those two claims together describe a model that does the same job in fewer tokens, not a cheaper rate card. Whether that holds is exactly what a 410M-token verbosity reading is for, and one independent run is not enough to settle it — but it is enough to say the marketing sentence and the measured token count are pointing in opposite directions, and the measured one is the one that shows up on an invoice.

Note also what the independent index does not yet contain. There is no published OSWorld, Terminal-Bench or CursorBench replication from anyone other than Anthropic. And the 56 is a composite: it says nothing about which of the ten evaluations inside it drove the gap to Sol's 47. A nine-point composite lead can hide a fifteen-point loss on the one eval your workload depends on.

Where they diverge in ways a rate card cannot show

Price and index score are the two easy axes. The ones that decide a migration are the behavioural ones, and here the two models are not close.

OrcaRouter model page for openai/gpt-5.6-sol, attributed to OpenAI and dated 2026-07-09, showing a 1M-token context window, 128K maximum output, text, image and file input, an input price of $4.00 and output price of $20.00 per million tokens, a p50 time to first token of 10.00 seconds, and OpenAI-compatible base URL https://api.orcarouter.ai/v1

Claude Sonnet 5.5 runs adaptive thinking on by default with high as the default effort, and it removes three things the previous generation allowed: sending thinking: {"type": "disabled"} now returns a 400 and must be replaced with between_tools; forced tool use — tool_choice set to "any" or a named tool — returns a 400 outright; and the older computer_20251124 computer-use tool is rejected on the Claude API and Google Cloud in favour of a toolset. Thinking blocks are also now bound to the model that produced them, so a conversation can move from Claude Sonnet 5 onto Claude Sonnet 5.5 and keep its reasoning, but cannot move back out again with it.

If your agent loop sets tool_choice: "any" to force a schema-valid call, Claude Sonnet 5.5 will reject the request until you switch to auto with strict tool use or structured outputs. That is a real migration task, and it is the kind of thing that turns a one-line model-ID swap into an afternoon.

GPT-6 Sol's cost of a switch is different in kind, because the model it replaces is still in the catalogue and still being called. GPT-5.6 Sol is not withdrawn; it is deprecated at Artificial Analysis, priced at double the new model, and still receiving traffic. OrcaRouter's own measurements show it moving roughly 95 million tokens a week, which is a reasonable proxy for how many integrations have not migrated yet. Moving to GPT-6 Sol is a pricing decision you can take later; you do not have to take it now.

Running the comparison on one key

The practical problem with a two-vendor comparison is that it usually costs two accounts, two SDK paths and two sets of rate limits before you learn anything. OrcaRouter exists to collapse that: one OpenAI-compatible endpoint, 200-plus models, provider list price passed through with no markup, and automatic failover across providers.

Both of the models that matter here are routable on it today. GPT-6 Sol is in the catalogue at $2 / $10 with the long-context tier intact, and Claude Sonnet 5 — the model most Claude API traffic still runs on, at a measured 150 output tokens per second and a p50 first-token time of roughly five seconds — has been on the platform since June 30 at Anthropic's own $2 / $10.

Claude Sonnet 5.5 itself is not on our catalogue yet. While that is true, the honest route to it is the vendor's own API and the three clouds Anthropic lists, and the honest way to evaluate it is to point it at a slice of traffic you can afford to lose. That is what the routing layer is for: promote a percentage, watch the failure rate, and keep the previous model as the fallback rather than as a rollback plan.

Which one goes in production

Pick by shape of workload rather than by composite score.

Claude Sonnet 5.5 is the better buy for long-context work, for anything already written against the Anthropic tool-use and computer-use surface, and for teams whose prompts routinely run past a quarter of a million tokens — the flat-rate window is worth more there than any index delta. Accept that the migration has a checklist attached: forced tool use, the thinking toggle, the advisor tool pairings, and a computer-use tool change.

GPT-6 Sol is the safer continuity choice for an OpenAI-shaped stack, and the cheaper one if your prompts are short and your integrations are already built. Its advantage is not capability — on the composite it trails — but that its predecessor is still serving and the migration has no deadline.

On the numbers available today, Claude Sonnet 5.5 wins the head-to-head on the independent index and on long-context economics, and loses on nothing that any published measurement is yet able to detect beyond the confidence of the vendor's own table. That is a narrower claim than the launch material makes, and it is the one worth acting on.

What would change the answer is a second independent run on the agentic evaluations, and a published tokens-per-task figure for both models on the same tasks. Until both exist, the composite index is a nine-point lead for the cheaper-to-run model, and a nine-point composite lead is not the same as a nine-point lead on your workload.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily