
Claude Sonnet 5.5 vs Claude Sonnet 5: Identical Prices, Different Bills
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 984 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Two models from the same vendor, same tier, same name, and — on every line of the published rate card — the same price. Claude Sonnet 5 launched on June 30, 2026 at $2.00 per million input tokens and $10.00 per million output tokens. Claude Sonnet 5.5 reached general availability on September 28, 2026 at $2.00 and $10.00, with cache reads at $0.20 and five-minute cache writes at $2.50 on both. Nothing about the price changed. What changed is everything a price sheet does not measure, and the result is that one of these two models is substantially cheaper to operate than the other despite costing exactly the same per token. Anthropic's own headline for the release says as much without quite saying it: up to 30% lower cost for most work, on an unchanged rate card. The only way both sentences are true is if the newer model finishes the same job with fewer tokens, which is what the benchmarks show and what the independent per-task figures confirm.
Is Claude Sonnet 5.5 more expensive than Claude Sonnet 5?
No — and the question is worth taking seriously anyway, because a great many migrations have gone wrong by treating a flat rate card as a flat bill.
Per token, the two models are identical: $2.00 in, $10.00 out, $0.20 to read a cached token, $2.50 to write one. There is no tiered band and no long-context surcharge. Whatever the newer model is, it is not a price increase dressed up as a generation.
Per completed task, the picture flips depending on whose workload you are describing. Artificial Analysis reports a cost of $5.09 per task for Claude Sonnet 5 and $7.60 per task for Claude Sonnet 5.5 on the same Intelligence Index v4.3 suite — the newer model 49% more expensive to finish a task with, at an identical rate. The same report notes Sonnet 5.5 emitted 410 million tokens across the suite against Sonnet 5's 370 million, both against a median of 88 million for the tracked field. On the evaluator's open-ended prompts the newer model simply writes more, and at equal prices more tokens is a larger bill.
Anthropic's 30%-cheaper claim is a vendor claim about its own chosen workloads. The third-party $7.60 against $5.09 is a measurement on a different set. Neither is wrong; the disagreement is the whole story, and the deciding factor is whether your tasks bound their own output.
What actually moved between the two generations
The rate card is static; the model is not. Laid out as a single list, the deltas are large and mostly one-directional.
• Terminal-Bench 4.0 — 70.6% for Claude Sonnet 5.5 against 10.3% for Claude Sonnet 5, the largest single-generation jump Anthropic has published on this test
• CursorBench 4.0 — 55.5% against 34.1%
• Chartography, no tools — 61.6% against 15.6%
• GDPval-AA v2.1 — 1844 against 1449
• AA-Briefcase v1.1 — 1811 against 1359
• Humanity's Last Exam, with tools — 64.5% against 54.9%
• OSWorld 2.1, partial completion — 80.1% against 57.0%
• Artificial Analysis Intelligence Index v4.3 — 56 against 38, an eighteen-point gap on a third-party scale measured under the same "Adaptive Reasoning, Max Effort" configuration
The shape of that list is not uniform progress. Chartography moving from 15.6% to 61.6% and Terminal-Bench from 10.3% to 70.6% look like capabilities that were absent and are now present, not capabilities that improved. Humanity's Last Exam moving ten points and OSWorld twenty-three points at an unchanged price is the pattern of a generation that closed specific gaps rather than one that got generally smarter. The eighteen-point index gap is the arithmetic average of that: real, large, and concentrated in tool use, code execution and visual work.
One footnote matters for anyone reading the vendor table closely. Anthropic states that on FrontierCode the Max effort configuration scores lower than Xhigh, and it flags that the GDPval-AA and AA-Briefcase runs were executed on a pre-release deployment carrying a structured-outputs bug. The numbers above are still the best available, but treat them as a snapshot from one deployment rather than a permanent property of the model.
Why the price is equal and the bill is not
Three mechanisms, in descending order of how much they move a real invoice.
The first is output length discipline. Sonnet 5.5 is a more capable agent, which means it does more before answering, and on a suite with no length constraint it writes roughly 11% more than Sonnet 5 while costing the same per token. On a workload that caps the output — extraction into a fixed schema, classification, bounded summaries — that 11% never materialises and the 30% claim becomes credible, because the win is arriving at the answer in fewer turns rather than in fewer tokens per turn.
The second is cache behaviour, and it is the reason an unchanged cache-read price does not mean an unchanged cache cost. Cache reads and writes are priced identically on both models, so any change in cached-token volume flows straight through to the bill. A model that takes fewer turns to finish an agentic task reads its prefix fewer times. That is a saving with no price move behind it at all, and it is invisible in every comparison table that lists only rate-card lines.
The third is the one most teams get wrong on migration day: the effort default. Claude Sonnet 5.5 configures reasoning through an effort parameter with low, medium, high, xhigh and max settings, defaulting to high on the Claude Platform and medium inside the Claude apps. If you migrated from a Sonnet 5 deployment where you had pinned a reasoning budget, the new default may be a different depth than you were paying for. Measure before assuming either a saving or an increase.
There is also a behavioural consequence worth flagging before rollout rather than after. Anthropic states that Claude Sonnet 5.5 is the first Sonnet to launch with cyber safeguards and fallbacks on the same footing as its top-tier models, so higher-risk cybersecurity requests visibly fall back to the previous Sonnet. Your logs will show a model you did not request. Relatedly, if you run Sonnet with thinking disabled you must switch to a between-tools behaviour — a code change, not a flag.
On OrcaRouter, anthropic/claude-sonnet-5 is routable today on one credential at the provider's list price with 0% markup, alongside Claude Opus 5.5, Claude Opus 5, DeepSeek V4 Pro and Gemini 3.1 Pro Preview. Claude Sonnet 5.5 itself is not yet a route on our platform, so the practical form of this comparison right now is that a team already running Sonnet 5 can measure the delta the day it is listed without touching an API key, and can fail over between tiers in the meantime by changing a string.
Questions people ask before switching
Can I run both at once and compare on my own traffic? Not through one OrcaRouter credential yet, since Claude Sonnet 5.5 is not a listed route. The workaround is to run the newer model on Anthropic's own API and the older one through us, then compare tokens per completed task rather than cost per token, because that is the figure the two models actually differ on.
Does the unchanged price mean Anthropic is not charging for the new capability? It means the list price is being held flat while the model improves, which is the same pattern as the previous two Sonnet generations. Whether that persists is a commercial question, not a technical one, and the published retirement window — no sooner than September 2027 — is the only commitment attached.
Which number should I trust, Anthropic's 30% or the third-party 49%? Neither as a general claim. Anthropic's figure describes workloads where the model reaches an answer in fewer turns; the Artificial Analysis figure describes an unconstrained index suite where verbosity is free to grow. Choose the one that resembles your prompt set, and if you cannot tell which that is, that ambiguity is itself the answer — measure it.
Bottom line
Claude Sonnet 5.5 is not a price increase, but it is not automatically a saving either. Anthropic held every line of the rate card flat and moved the model underneath it, which means the entire economic case rests on tokens per completed task — a number that is 30% better on the workloads the vendor chose and 49% worse on the suite an independent evaluator chose. If your tasks bound their own output, switch and take the win. If they do not, the flat price is not a reason to skip the measurement, because you can pay half again as much per task for a model that is unambiguously better at the work.



