
Claude Sonnet 5.5 Rolls Out on Claude and the API — With Five Changes That Break Existing Code
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 348 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 106 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 987 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 49 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 106 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Claude Sonnet 5.5 is live. The vendor's announcement is dated September 28, 2026, the Claude Platform model page lists it as "Latest", and the identifier is claude-sonnet-5-5 across the Claude API, Amazon Bedrock, the cloud consoles, Microsoft Foundry and Claude Platform on AWS. The rollout to Claude itself and to the API is the part that is happening this week. The part that will cost you a day is elsewhere in the same documentation: five breaking changes affect code already running on Claude Sonnet 5, and a sixth changes the response shape without failing a single request.
That is the shape of this release. The rate card did not move: Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, exactly what Claude Sonnet 5 costs, with $0.20 cache reads and a $2.50 five-minute cache write. The capability moved a long way. On Artificial Analysis's current revision the new model sits 18 points above the one it replaces, which is the widest gap between two consecutive Sonnet models on that board. The migration cost is correspondingly the widest. None of that is a complaint — a model that gets better without getting dearer is the good case — but the two facts arrive together, and the release notes are the half that decides whether your week is fine.
What actually shipped on September 28

The specification is a straight continuation of the tier rather than a new shape:
• Model ID — claude-sonnet-5-5 on the Claude API, Claude Platform on AWS and Microsoft Foundry; anthropic.claude-sonnet-5-5 on Amazon Bedrock
• Price — $2 / MTok input, $10 / MTok output, $2.50 five-minute cache write, $4 one-hour cache write, $0.20 cache read
• Batch — 50% off both directions, so $1 / $5 per MTok
• Context and output — 1M tokens in, 128K tokens out synchronously, 300K out on the Message Batches API behind the output-300k-2026-03-24 beta header
• Thinking — adaptive, with a default effort of high; knowledge and training cutoffs both June 2026
• Lifecycle — zero data retention from launch, and a published retirement floor of not sooner than September 28, 2027
Two details are easy to miss and matter. First, there is no long-context surcharge: Anthropic bills the full 1M-token window at the standard rate, so a 900,000-token request costs the same per token as a 9,000-token one. Second, the $2 / $10 rate on the previous model was originally introductory pricing with a scheduled increase to $3 / $15 on September 1, 2026; Anthropic's pricing page now states that the increase will not occur, which makes the Sonnet tier's current price a standing commitment rather than a promotion. Both of those are vendor statements, not independent findings.
The migration list, in the order it will bite
Anthropic's Sonnet 5.5 overview is unusually direct about this: five breaking changes affect code already running on Claude Sonnet 5. They are not stylistic. Each one returns an error or silently changes behaviour.
• Forced tool use is refused. A tool_choice of "any", or of a named tool, now returns an error rather than constraining the model. If your agent loop pins a tool to guarantee a structured call, that loop stops working on the first request. The replacement is to let the model choose and validate the result, or to move the guarantee into structured outputs.
• thinking: {"type": "disabled"} is refused. Turning thinking off outright is no longer a supported setting. The replacement is between_tools, which turns off up-front thinking and works at high effort or below. It is the lowest thinking setting on the model, not an off switch.
• Thinking blocks are bound to the model and the conversation. Reasoning that a previous turn produced carries forward from Claude Sonnet 5 into Claude Sonnet 5.5, but not out to another family, and not across accounts. Applications that replay a transcript across model versions will find the thinking blocks rejected where they were previously accepted.
• computer_20251124 is not accepted on the Claude API or Google Cloud. Code that declared the earlier computer-use tool version has to move to the current toolset.
• The advisor tool rejects its own family. Claude Opus 4.8, Claude Opus 4.7 and Claude Sonnet 5 are refused as advisors, so any pipeline that used a smaller Claude model as a checker against a larger one needs re-pairing.
Then there is the change that fails nothing. Text the model emits between tool calls now arrives inside thinking blocks rather than as ordinary output. An application that streams text to a user goes quiet between tool calls — no error, no exception, just a UI that appears to stall until you set a display value that returns the text, or switch to between_tools. This is the failure mode most likely to reach production, because nothing in a test suite throws.
One more, for completeness: setting temperature, top_p or top_k to any non-default value returns a 400. If your gateway or client library sets a temperature out of habit rather than intent, that habit is now an outage.
The 30% claim, against the independent number

Anthropic's marketing line for Claude Sonnet 5.5 is that it "costs up to 30% less per task than its predecessor" at identical per-token rates, and that it "generates outputs 30%+ faster" than Claude Sonnet 5. Both are claims about tokens rather than prices, and both are vendor-reported. The mechanism behind the first one is visible in the vendor's own benchmark table: at Medium effort on Terminal-Bench, Anthropic says Claude Sonnet 5.5 beats Claude Sonnet 5's best score for "less than a tenth of the cost per task", and makes similar claims for CursorBench at Low effort and AA-Briefcase at Medium.
The independent picture agrees on direction and disagrees on size. On Artificial Analysis's v4.3.2 Intelligence Index — ten evaluations, run in one configuration across every model — Claude Sonnet 5.5 scores 56, which the page places at rank 3 of 216. Claude Sonnet 5 scores 38 on the same board. That is an 18-point move between two consecutive models in the same tier at the same price, and it is the largest gap Anthropic has opened inside the Sonnet line this year.
The cost figure is where the vendor claim gets harder to hold. Artificial Analysis measures Claude Sonnet 5.5 at $7.60 per completed Intelligence Index task. That number is not 30% below anything in its class, because the model is extremely verbose: it emitted 410 million output tokens to complete the Index evaluations, against a median of 88 million across comparable models. Efficiency per token and efficiency per task are different quantities, and Claude Sonnet 5.5 is unusually good at the first and merely average at the second.
All of Anthropic's own benchmark figures for this release are vendor-reported and unreproduced by third parties, and two of them carry footnotes worth repeating: the GDPval-AA v2.1 and AA-Briefcase v1.1 numbers were run by Artificial Analysis on a pre-release deployment that had a bug possibly understating the scores, and Anthropic's note on its FrontierCode 1.1 figure explains why Claude Sonnet 5.5 scores lower at Max effort than at Xhigh on that evaluation. A table where a model beats itself at a lower effort setting is a table telling you its benchmark configuration is not settled.
How to test it without betting a production path on it
There is a specific reason to run Claude Sonnet 5.5 through a router rather than against the vendor endpoint directly in a migration: you do not have to choose between the old model and the new one while you find out. On OrcaRouter, Claude Sonnet 5 has been routable since June 30 at Anthropic's own $2 / $10 with no markup added, one of 200-plus models behind a single OpenAI-compatible endpoint, and it is still the model most Claude API traffic runs on. Claude Sonnet 5.5 is not in our catalogue yet, so the route to it today is Anthropic's own API.
The pattern that works in the meantime is to send a slice of traffic to the new model and let failover carry the rest — with Claude Sonnet 5 as the fallback, so a 400 from one of the five breaking changes degrades into a slower answer rather than a failed request. That is also the cheapest way to measure the thing the vendor claim does not tell you: your own tokens per task on both models at the same effort setting. The 30% figure is an average over Anthropic's own evaluations; your traffic is not that average.
What lands next
Anthropic's announcement says Claude Haiku 5.5 joins the family "in the coming weeks", which would complete the 5.5 generation in the small tier. Until then, the two decisions in front of a team running Sonnet in production are separable: the migration is a one-day job you can schedule, and the model swap is a measurement you can defer. Doing them in the other order — swapping first and discovering the forced-tool-use error in production — is the version of this rollout that costs a weekend.

