A title card for Claude Sonnet 5.5 headed 'same price, less work', with the subtitle '$2 / $10 per million tokens, unchanged from Sonnet 5' and three stat cards reading 56 (AA Intelligence Index), #3 / 216 (rank on that index) and $7.60 (cost per task at max effort).
Guides & Insights

Claude Sonnet 5.5: The 30% Cost Claim, and the Six Changes That Break Sonnet 5 Code

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Sonnet 5.5 shipped on September 28, 2026 at exactly the same rates as Claude Sonnet 5 — $2 per million input tokens, $10 per million output tokens, $0.20 per million cached reads — while Anthro​pic's announcement leads with "runs 30%+ faster and costs up to 30% less for most work." Both statements are true, and the distance between them is the whole story. The savings are not in the rate card, because the rate card did not move. They come from running the model at a lower effort setting than its predecessor needed, which means the 30% figure is a claim about an operating point you have to choose, not a price you inherit. Independent measurement makes the fork sharp: on Artificial Analysis, Claude Sonnet 5.5 at max effort costs $7.60 per task on the Intelligence Index against $5.09 for Claude Sonnet 5 at max — roughly 49% more, not 30% less. That is not a contradiction of Anthro​pic's number. It is the other end of the same curve, and it is where most migrations will land by default if nobody re-runs their effort sweep.

Two claims wearing one number

Anthropic's post makes the per-task claim in three places and each one leans on effort. The strongest version: on Terminal-Bench 4.0, at Medium effort — the default in the Claude apps — Sonnet 5.5 "far exceeds Sonnet 5's best score for less than a tenth of the cost per task." On AA-Briefcase v1.1, at Medium effort, it beats Sonnet 5's best for about one ninth of the cost. On CursorBench 4.0, at Low effort, it exceeds Sonnet 5's best for less than a tenth.

Read those as a set and the mechanism is plain. Sonnet 5.5's floor sits above Sonnet 5's ceiling on several evaluations. You do not get the 30% by paying less per token; you get it by no longer needing the expensive setting. Anthropic says as much in the migration documentation, one page over from the marketing: "Effort levels are recalibrated. An effort level doesn't produce the same amount of thinking as it did on Claude Sonnet 5. Re-run your effort sweep rather than carrying a setting over."

That sentence is the most operationally important line Anthropic published this week, and it is the one that did not make it into most of the launch coverage.

What Anthropic published — vendor-reported, unreproduced

Everything in this section is Anthropic's own measurement, from the Sonnet 5.5 announcement and system card. It is a vendor claim, not an independent result, and the footnote trail matters as much as the headline numbers.

• Terminal-Bench 4.0 (agentic coding) — Sonnet 5.5 70.6% vs Sonnet 5 10.3%. That is a sevenfold jump on the single most-quoted row, and it is the number most coverage led with.

• FrontierCode 1.1 (Main) — Sonnet 5.5 52.1% at Xhigh and 46.2% at Max; Sonnet 5 42.4%; Claude Opus 5.5 54.4%; GPT-6 Sol 49.3%. Note the inversion: Sonnet 5.5 scores lower at Max than at Xhigh.

• CursorBench 4.0 — 55.5% for Sonnet 5.5 vs 34.1% for Sonnet 5 and 57.8% for Opus 5.5. CursorBench 4.0 does not report GPT-6 Sol publicly.

• GDPval-AA v2.1 (knowledge work) — Sonnet 5.5 1844 Elo, Sonnet 5 1449, Opus 5.5 1846, GPT-6 Sol 1487.

• AA-Briefcase v1.1 — 1811 / 1359 / 1822 / 1483 on the same four models.

• Humanity's Last Exam (with tools) — 64.5% / 54.9% / 67.7%.

• OSWorld 2.1 (computer use, partial credit) — 80.1% / 57.0% / 81.8%.

• Chartography (visual chart recognition, no tools) — 61.6% / 15.6% / 64.4% / 53.6%.

Three footnotes deserve to travel with those rows rather than under them. First, Terminal-Bench 4.0's Opus 5.5 figure of 66.4% is reported at Xhigh effort, Opus 5.5's highest-scoring configuration. Second, the FrontierCode Max inversion is explained: at Max, Sonnet 5.5 more often ran Claude Code's code-review skill, which splits review across many subagents, and in two cases that produced a timeout or edits beyond the task's scope — FrontierCode penalises out-of-scope changes even when they are good. Third, and least comfortable for the vendor, Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment of Sonnet 5.5 that Anthropic says had a bug degrading responses to structured-output requests. Anthropic expects the effect to be small and to understate Sonnet 5.5, and says the bug is fixed. It also notes that OpenAI recently fixed an image-understanding bug in GPT-6 Sol, so the GPT-6 Sol figures in those rows may not yet reflect the current model.

Artificial Analysis model page for Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback), released September 2026, showing an Intelligence Index of 56 at rank #3 of 216, $2.00 input and $10.00 output per 1M tokens, and a $7.60 average cost per index task with 410M tokens generated.

What the independent run says about the same model

Artificial Analysis has Sonnet 5.5 measured, and its page names the exact configuration: "Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)". On the same revision of the index, 216 models in class:

• Claude Sonnet 5.5 — Intelligence Index 56, ranked #3 of 216. $7.60 cost per index task. It generated 410M output tokens during the index run, against a median of 88M.

• Claude Sonnet 5 — Index 38, ranked #58 of 216, on its own page labelled "(Adaptive Reasoning, Max Effort)". $5.09 cost per task. 370M output tokens.

• Claude Opus 5.5 — Index 58, ranked #1 of 216. $5.98 per task. 95.3 output tokens per second.

• GPT-6 Sol — Index 48, ranked #20 of 216, at $2/$10 list. $1.06 per task, 89.8 output tokens per second.

The Intelligence Index gap — 56 against 38, eighteen points — is the strongest independent confirmation in this piece, and it is the number a reader should actually carry away. The cost figure is the one that needs a caveat, and it is worth stating exactly: both points are max-effort configurations, and max effort on a model that reasons for longer is a deliberately expensive place to stand. Sonnet 5.5 burned 410M tokens on the index run where Sonnet 5 burned 370M, and both are far above the 88M median. A model that thinks longer at the top setting costs more at the top setting. What the number falsifies is not "cheaper per task" but any reading of it as a property of the model rather than of the setting.

Artificial Analysis has not published an output-speed figure for Sonnet 5.5 yet — its page reads N/A for Output tokens per second, against 78.7 for Sonnet 5 and 95.3 for Opus 5.5. So the "30%+ faster" half of Anthropic's claim is vendor-reported only at this point. Treat it as directional until a third party measures it.

A two-column comparison scoreboard titled 'Claude Sonnet 5.5 vs Claude Sonnet 5 - the cost-per-task fork'. The left column, Claude Sonnet 5.5, reads AA Intelligence Index 56, rank #3 of 216, cost per task at max effort $7.60, 410M output tokens on the index run, $2 / $10 input and output price, default effort high. The right column, Claude Sonnet 5, reads 38, #58 of 216, $5.09, 370M, $2 / $10, high.

Where the saving actually comes from, and how to get it

If you accept the mechanism, the migration instructions write themselves, and Anthropic has published them:

• Start at high by default, not at whatever your Sonnet 5 config said. Anthropic's guidance is high unless your workload is agentic or latency-sensitive.

• Agentic coding and multistep tool use — start at medium for well-specified tasks and move up to high for harder or longer ones.

• Chat and other latency-sensitive work — start at medium or low. This is the band where the "tenth of the cost" claims live.

There is a coherent explanation for why the lower setting works, and it is not marketing. Anthropic's early testers reported that Sonnet 5.5 batches tool calls together more than Sonnet 5 did, producing fewer steps per task. CodeRabbit's note in the announcement is unusually specific for a launch quote: Sonnet 5's "tendency to reach for web search too often and its high token use are both gone in this new model." Sonnet 5.5 also kept the same tokenizer as Sonnet 5, so identical text costs identical tokens — the reduction is fewer turns and fewer redundant calls, not a cheaper vocabulary. That also means the token counts in your logs stay comparable across the upgrade, which makes before-and-after measurement straightforward.

The six changes that break or alter Sonnet 5 code

This is the part of the release that costs engineering time, and it is where the migration guide is worth more than the benchmark chart. Five of the six return hard errors; the sixth fails silently.

• thinking: {"type": "disabled"} now returns a 400. The replacement is thinking: {"type": "between_tools"}, which turns off up-front thinking and is the lowest thinking setting on this model. It works at low, medium and high effort, needs no beta header, and is available on every platform serving Sonnet 5.5. At xhigh or max effort, between_tools itself returns a 400 — use adaptive thinking (omit the field, or send {"type": "adaptive"}) if you need those levels. With between_tools you also cannot change effort mid-conversation.

• Forced tool use is not supported. tool_choice set to {"type": "any"} or {"type": "tool", "name": "..."} returns a 400 with the message "tool_choice: type 'tool' and 'any' are not supported for this model." The same check applies to the token-counting endpoint. The documented replacement for schema-valid input is tool_choice: {"type": "auto"} plus strict: true on the tool, or moving the schema to structured outputs.

• Thinking blocks are bound to the model and the conversation. Sonnet 5.5 reads thinking blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier — not from Opus 5, Opus 5.5, or any Fable or Mythos model. No other model reads Sonnet 5.5's blocks. Move a conversation from Sonnet 5 to 5.5 and the reasoning survives; move it from Sonnet 5.5 to anything else and the later turns run without it. Separately, the API now checks whether the system prompt, the tool list, or an earlier message changed since a thinking block was produced, and on accounts created on or after August 31, 2026 it returns a 400 when they have. Keep conversations append-only, or send the thinking-binding-controls-2026-08-01 beta header with prefix_mismatch_behavior: "drop_block".

• computer_20251124 is rejected on the Claude API and Google Cloud. Computer use there requires the computer_toolset_20260801 toolset. Amazon Bedrock still accepts the older tool — a platform divergence worth flagging if you run the same agent on both.

• Some advisor pairings are gone. A Sonnet 5.5 executor accepts Mythos 5.1, Fable 5.1, Mythos 5, Fable 5, Opus 5.5, Opus 5 or Sonnet 5.5 itself as advisor. Opus 4.8, Opus 4.7 and Sonnet 5 advisors, which work with a Sonnet 5 executor, return a 400 with a Sonnet 5.5 executor. Every advisor Sonnet 5.5 accepts returns advice encrypted, so your client cannot read the advice text.

• Text between tool calls now arrives inside thinking blocks — and at the default display: "omitted" it is empty. This is the silent one. Notes longer than a sentence or two between tool calls come back as progress-update thinking blocks rather than as text. An application that streams that narration to its users goes quiet between tool calls, with no error and nothing in the logs. The fix is either to set thinking.display so the text is returned, or to switch to between_tools, in which case the text comes back as ordinary text.

Two more things a migration touches, neither an error. Monitoring for refusals: Sonnet 5.5 returns stop_reason: "refusal" with a stop_details object naming one of five policy areas — cyber, bio, frontier_llm, reasoning_extraction, general_harms — and Anthropic's server-side fallback will retry cyber and frontier_llm declines on Sonnet 5 but not the other three. And the security posture genuinely changed: Sonnet 5.5 is the first Sonnet model to ship with cyber safeguards and fallbacks like the Opus class, which means higher-risk cybersecurity requests visibly fall back to Sonnet 5, and the first Sonnet with classifiers against reasoning extraction. If your product's value proposition is a security tool, test that boundary before you ship the model swap.

Pricing, and what is genuinely on our route today

The rate card is unchanged from Sonnet 5 and its full published shape is short enough to state plainly:

• Input — $2.00 per million tokens. Output — $10.00 per million tokens. Both identical to Sonnet 5, confirmed on Anthropic's model page for Sonnet 5.5.

• Cache read — $0.20 per million, a 90% discount on input; 5-minute cache write $2.50 per million; 1-hour cache write $4.00 per million. The minimum cacheable prompt is 512 tokens on Sonnet 5.5, down from 1,024 on Sonnet 5.

• Batch — 50% off both directions, so $1.00 / $5.00 per million. On the Message Batches API, output can go to 300K tokens with the output-300k-2026-03-24 beta header.

• Against Opus 5.5 — Sonnet 5.5 is exactly half: $2/$10 versus $4/$20, with cache writes at half and cache reads identical at $0.20. That is the migration that pays, and Anthropic says it outright: the two models complement each other best when Sonnet runs at lower effort, and Opus "remains clearly stronger at complex, open-ended work requiring sustained judgment."

• Context and output — 1M-token context, 128K maximum output, adaptive thinking on by default, default effort high, reliable knowledge cutoff June 2026, zero data retention available, retirement not sooner than September 28, 2027.

One availability note that we would rather state than imply: Claude Sonnet 5.5 is not on our model catalogue as of September 29, 2026. Its predecessor anthropic/claude-sonnet-5 and its larger sibling anthropic/claude-opus-5.5 both are, and its rates on our side are the provider's own — $2.00 in, $10.00 out, $0.20 cache read, $2.50 cache write for Sonnet 5, with no markup added, so a vendor price change is live here the same day it lands rather than at the next billing cycle. That last property is what makes a rate-card reading useful rather than decorative.

OrcaRouter model page for anthropic/claude-sonnet-5, attributed to Anthropic and dated 2026-06-30, showing a 1M-token context window, up to 128K output tokens, and the unchanged rates of $2.00 per million input tokens and $10.00 per million output tokens. The page carries a link reading 'the Claude Sonnet 5 API'.

What you can do with this today

The honest sequencing for a team that already runs Claude Sonnet 5 in production is three steps, and none of them is "swap the model string and watch the graph."

First, re-run the effort sweep. If you carried a high or max setting over from Sonnet 5 because that is what your quality bar required, you are now paying for reasoning you no longer need to buy. The independent cost-per-task figures say the top of the ladder got more expensive, not less, and the vendor's own cost claims all describe the middle of it. Second, fix the two error paths before the deploy — disabled thinking and forced tool use — because both fail loudly and immediately, which is the good news. Third, watch the narration path, because it fails silently.

If the model is not yet on the route you use, the low-risk way to evaluate it is to run it alongside, not instead: keep Sonnet 5 pinned as the fallback on the same key so that a refusal, a timeout or a bad eval sends the request to the model you already trust, and the automatic failover closes the loop without a second integration. That is the whole argument for evaluating an unproven model behind one endpoint rather than in front of your users — and it applies twice over here, because Sonnet 5.5's refusal categories are new and its effort calibration is explicitly not the same as its predecessor's. When you are ready to move the default, the roster on one key runs to more than 200 models, which makes the comparison a config change rather than a second contract.

What to watch

Three things will settle the open questions in this piece, and none of them has happened yet.

An independent output-speed measurement is the first. Anthropic claims 30%+ faster than Sonnet 5; Artificial Analysis has not published a figure for Sonnet 5.5, and the claim is doing real work in the cost argument, since a faster model finishes the same task in less wall-clock and often in fewer tokens. Claude Haiku 5.5 is the second — Anthropic says it arrives "in the coming weeks" and it will sit under Sonnet 5.5, which changes the calculus for the high-volume chat band where medium and low effort already make Sonnet 5.5 competitive. The third is the correction pass: Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release build with a structured-output bug, Anthropic expects those scores to understate the model, and neither number has been re-run. If they move up, the knowledge-work gap to Opus 5.5 narrows further and the "half the price" argument gets stronger.

Until then, the defensible summary is narrower than the announcement and more useful than the benchmark chart. Claude Sonnet 5.5 is an eighteen-point intelligence gain over Sonnet 5 on the independent index, at the same published rates, with a real and measurable saving available to anyone who moves down the effort ladder — and a real and measurable penalty for anyone who does not.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily