Generated hero title card for Claude Sonnet 5.5 vs Claude Opus 5.5, subtitled 'The benchmarks converge, the bill does not', showing $2.00 and $10.00 per million tokens on the left against $4.00 and $20.00 on the right, with the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

Claude Sonnet 5.5 vs Claude Opus 5.5: The Benchmarks Converge, the Bill Does Not

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Sonnet 5.5 reached general availability on September 28, 2026 at $2.00 per million input tokens and $10.00 per million output tokens. Claude Opus 5.5, Anthro​pic's flagship, shipped six days earlier on September 22 at $4.00 and $20.00 — exactly double on both lines. The obvious reading is that Opus 5.5 is the ceiling and Sonnet 5.5 is the compromise you make when the budget is real. Anthro​pic's own launch table does not support that reading. On the numbers the company published, Claude Sonnet 5.5 posts 1844 against Opus 5.5's 1846 on GDPval-AA v2.1, 1811 against 1822 on AA-Briefcase v1.1, and 70.6% against 66.4% on Terminal-Bench 4.0 — it wins the one benchmark that measures work actually done inside a terminal. Two lines below those figures, Anthro​pic's own caption states that Opus 5.5 "remains clearly stronger at complex, open-ended work." Both statements are true, and figuring out how to hold them at once is most of what this comparison is for.

What the launch table says, and what it declines to say

The pattern in Anthropic's benchmark block is a model that has closed most of the gap rather than one that has taken the lead. GDPval-AA v2.1, an economically weighted task set, is a two-point tie. AA-Briefcase v1.1, a document-and-deliverable evaluation, is an eleven-point tie. CursorBench 4.0 goes the other way: 55.5% for Sonnet 5.5 against 57.8% for Opus 5.5. OSWorld 2.1 lands at 80.1% against 81.8%, and Humanity's Last Exam at 64.5% with tools against 67.7%. Only Terminal-Bench 4.0 breaks the pattern, and it breaks hard — 70.6% against 66.4%, a four-point Sonnet advantage on agentic command-line work.

Read the row labels rather than the numbers and something else appears. Several of those Opus 5.5 entries carry an effort annotation — 66.4% at Xhigh — and a footnote states that the Max configuration scores lower than Xhigh on FrontierCode, not higher. Higher effort is not monotonic. A budget-conscious team that assumes "max effort" is a free upgrade is buying tokens for a worse answer on at least one of Anthropic's own tests. And on FrontierCode 1.1 the same footnote places Sonnet 5.5 at 46.2% in the Max configuration against Opus 5.5's 54.4%, which is the clearest single number for where the flagship still pulls away: long-horizon code work under a task definition that rewards persistence.

There is one more caveat worth stating plainly rather than burying. Anthropic notes that the GDPval-AA and AA-Briefcase runs it publishes were executed on a pre-release deployment that carried a structured-outputs bug. That affects both models in the same table, so the comparison is not invalidated, but it does mean the absolute figures should be treated as a snapshot rather than a settled result. Vendor benchmark tables are marketing artifacts with a methodology appendix, and this one comes with its appendix attached.

Where the independent measurement lands

Artificial Analysis scores both models under a single harness, in the same configuration it labels "Adaptive Reasoning, Max Effort, Default Fallback," on Intelligence Index v4.3. Claude Opus 5.5 sits at 58. Claude Sonnet 5.5 sits at 56. That is a two-point gap on a scale where the same evaluator puts Opus 5 at 51, Sonnet 5 at 38, DeepSeek V4 Pro at 36 and Gemini 3.1 Pro Preview at 30. A new Sonnet arriving two points under the flagship and five points over the previous Opus generation is the substantive claim of this release, and it is a third-party claim rather than a vendor one.

Then the same page reverses the economics of it. Artificial Analysis reports that a Sonnet 5.5 evaluation run costs $7.60 per task against $5.98 for Opus 5.5, even though Sonnet 5.5 is half the price per token. The cause is right there in the page: Sonnet 5.5 generated 410 million tokens across the index suite against a median of 88 million for the models it tracks — "very verbose," in the evaluator's words. Opus 5.5 generated 260 million over the same suite. At $10 per million output tokens against $20, the verbosity factor of roughly 1.6 still leaves Sonnet 5.5 paying about 27% more per completed task.

This is the part of the comparison that a per-token price sheet cannot show you, and it is the reason the two figures coexist without contradiction. Per token, Sonnet 5.5 is cheaper by half. Per finished task, on an evaluation suite with open-ended prompts and no output-length discipline, it can be more expensive. Which of those describes your workload is the actual decision.

Effort levels, output ceilings and the shape of the integration

For a buyer, the mechanically important properties are near-identical between these two models.

• Context — 1,000,000 tokens on both, from the same provider, so prompt-engineering assumptions transfer unchanged

• Maximum output — 128,000 tokens on both; bulk generation is not a differentiator here

• Modality — text, image and file input on both; neither accepts audio or video

• Reasoning control — adaptive thinking on both, with depth set by an effort parameter rather than a thinking-token budget. On the Claude Platform the default is high; inside the Claude apps it is medium

• Cache write — $5.00 per million for Claude Opus 5.5 against $2.50 for Claude Sonnet 5.5, the one line where the cheaper model is also the cheaper model to feed with a rebuilt prefix

• Cache read — $0.20 per million on both, an exact tie on the line that dominates long agent runs

Because the surfaces match, migration between them is a model-string change and a re-measurement of effort, not an integration project. That is unusual across vendors and it is the reason this particular pairing is worth comparing at all: the two candidates differ in price and in depth, not in the API you have to write.

One operational difference deserves a line of its own. Anthropic states that Claude Sonnet 5.5 is the first Sonnet to launch with cyber safeguards and fallbacks on the same footing as its top-tier models, meaning higher-risk cybersecurity requests are visibly routed to the previous Sonnet rather than answered by the model you named. If your workload touches that domain, the model string is not a guarantee of which model responded — on either tier. Anthropic also notes that if you run Sonnet with thinking disabled you must switch to a between-tools behaviour, which is a code change rather than a configuration flag. Neither of those is a reason to avoid the model; both are reasons to know before you ship.

On OrcaRouter, Claude Opus 5.5 is routable today on the same credential as every other model in the catalogue, at the provider's list price with 0% markup, with automatic failover available underneath it. Claude Sonnet 5.5 is not yet a route on our platform — the model is one day old, and until it is listed the honest statement is that you would reach it through Anthropic's own API. Anyone currently running Opus 5.5 through us can price the switch the day it lands without changing anything else about their integration.

Which one to put where

The split that follows from the numbers: Claude Sonnet 5.5 for terminal-bound agent work, where it is the stronger of the two on Anthropic's own Terminal-Bench 4.0 figure and half the price; Claude Sonnet 5.5 for anything throughput-shaped, since the cost difference only narrows when the task forces long outputs; Claude Opus 5.5 for FrontierCode-class work, long-horizon refactors across large codebases, and any workload where the 54.4% to 46.2% margin reflects the shape of your actual problem rather than a leaderboard row.

If your tasks are open-ended and you cannot predict output length, treat the per-token price as the wrong number entirely. Measure tokens per completed task on your own traffic for a week before you commit either way; the artificial-analysis figure of $7.60 against $5.98 is a warning that the cheap model can lose on the metric you actually pay.

The honest verdict: this is no longer a capability-versus-cost trade. It is a question of whether your work is bounded. For bounded, high-volume, tool-driven tasks the cheaper model is also the better one on the evidence available, and the flagship is not worth double. For open-ended, long-horizon work Anthropic's own sentence — that Opus 5.5 remains clearly stronger — is the one to believe, and no amount of benchmark convergence changes it yet.

A two-column scoreboard for Claude Sonnet 5.5 and Claude Opus 5.5 across eight rows. Left column Claude Sonnet 5.5: input $2.00, output $10.00, cache read $0.20, cache write $2.50, 1M context with 128K max output, AA Intelligence Index v4.3 score 56 at $7.60 per task, Terminal-Bench 4.0 (vendor) 70.6%, GDPval-AA v2.1 (vendor) 1844. Right column Claude Opus 5.5: input $4.00, output $20.00, cache read $0.20, cache write $5.00, 1M context with 128K max output, AA Intelligence Index v4.3 score 58 at $5.98 per task, Terminal-Bench 4.0 (vendor) 66.4%, GDPval-AA v2.1 (vendor) 1846. The card heading reads "Claude Sonnet 5.5 against Claude Opus 5.5: eight rows, two of them on a different instrument", and a footer line reads "Per-token prices per the OrcaRouter catalogue; index and per-task figures per Artificial Analysis, Intelligence Index v4.3. Terminal-Bench and GDPval-AA rows are Anthropic's own launch table, run on a pre-release deployment, and are not the same instrument as the index."Screenshot of Anthropic's newsroom page for Claude Sonnet 5.5, showing the site navigation bar, the publication date September 28, 2026, the model heading "Claude Sonnet 5.5", and an image-led marketing hero beneath it that carries no machine-readable specification figures.Screenshot of the OrcaRouter model page for Anthropic Claude Opus 5.5 (anthropic/claude-opus-5.5), showing the model header, a 1M-token context window with 128K maximum output, text, image and file input, the Vision, Tools, JSON and Reasoning capability tags, a 4.57-second p50 time to first token, and the $4.00 input and $20.00 output prices per million tokens.