
Claude Sonnet 5.5 vs Claude Opus 5: Cheaper on Every Line, and Ahead on the Index
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 984 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Claude Opus 5 shipped on July 24, 2026 at $5.00 per million input tokens and $25.00 per million output tokens. Claude Sonnet 5.5 reached general availability on September 28, 2026 at $2.00 and $10.00. That is a 60% cut on input and a 60% cut on output for a model two generations newer — and, on the harness Artificial Analysis runs across both, a model that lands at 56 on Intelligence Index v4.3 against Claude Opus 5's 51. This is the rare comparison where the newer, cheaper model is also the higher-scoring one on a third-party index, and where the remaining case for the incumbent is not benchmark rank but a specific class of work. Both models take 1M tokens of context, both emit up to 128K tokens, both read text, images and files, and both configure reasoning through an effort parameter rather than a thinking-token budget. Because the surfaces are the same, choosing between them is purely a question of what a finished task costs and which tasks fail.
Two releases, one interface
Start with how little changes, because it is more than in most generation jumps.
• Context window — 1,000,000 tokens on both
• Maximum output — 128,000 tokens on both
• Input modality — text, image and file on both; no audio, no video on either
• Reasoning — adaptive thinking with configurable effort on both, so depth is a request parameter rather than a separate model string
• Capabilities — vision, tools, JSON and reasoning on both, per the OrcaRouter catalogue entries for anthropic/claude-opus-5 and the Sonnet line
The practical consequence is that a prompt written for Claude Opus 5 does not need to be rewritten to run on Claude Sonnet 5.5, and an agent loop tuned around a 128K output ceiling does not need a new tier. Migration is a model-string swap plus a re-check of which effort setting you had pinned — nothing more. For teams that lived through the multi-modal, multi-parameter transitions of the last two years, that is the headline, not the benchmark.
The one thing that does change is the price, and it changes at every line rather than just the two on the marketing slide.
• Input — $2.00 per million against $5.00, a 60% cut
• Output — $10.00 per million against $25.00, a 60% cut
• Cache read — $0.20 per million against $0.50, a 60% cut
• Cache write — $2.50 per million against $10.00 for the five-minute window, a 75% cut
If you run an agent that re-reads a long prefix on every turn, the cache-read line is the one that pays for the migration. At $0.20 against $0.50, every cached token in a long session lands at 40% of its old cost, and cache reads are the bulk of the token count in most agentic loops. The saving compounds in a way the per-token headline does not capture.
Where the five points come from
Artificial Analysis scores Claude Sonnet 5.5 at 56 and Claude Opus 5 at 51 on Intelligence Index v4.3, both measured under the "Adaptive Reasoning, Max Effort" configuration. Five points is not a rounding error on that scale — for context, the same evaluator places Sonnet 5 at 38 and Gemini 3.1 Pro Preview at 30, so the gap between Claude Opus 5 and Claude Sonnet 5.5 is a third of the gap between Claude Opus 5 and the previous Sonnet generation.
Anthropic's own launch numbers for Claude Sonnet 5.5 point the same direction with a different instrument. The company reports 70.6% on Terminal-Bench 4.0 — the same test on which Claude Opus 5 is documented at 89.1% in an earlier Anthropic table, a figure that used a different harness revision and should not be subtracted from the newer one. That is the single most important sourcing caution in this article: Terminal-Bench moved to 4.0 partway through 2026, the index carrying it was revised to v4.3 to match, and cross-version comparisons of that benchmark are not arithmetic. Where a number has a version attached, quote the version.
What can be compared cleanly is the vendor's own Sonnet-side progression — 10.3% on Terminal-Bench 4.0 for Sonnet 5 to 70.6% for Sonnet 5.5 — and the third-party index reading, where both figures come from the same harness on the same day. On that evidence the successor has genuinely overtaken the July flagship for general reasoning, which is a stronger claim than any vendor slide makes.
The output-length trap
Here is where the arithmetic stops being kind to the newer model.
Artificial Analysis reports that evaluating Claude Sonnet 5.5 on its index suite costs $7.60 per task. Claude Opus 5 costs $5.86 per task on the same suite. The newer model, at 60% off every line of the rate card, is roughly 30% more expensive to finish a task with. The reason is in the same report: Sonnet 5.5 emitted 410 million tokens across the suite against a median of 88 million across the models tracked, which the evaluator labels "very verbose." Claude Opus 5 emitted 140 million over the same suite.
Divide it out and the mechanism is simple. Sonnet 5.5 produces roughly three times the output tokens of Claude Opus 5 on identical work, at 40% of the output price. Three times 0.4 is 1.2 — a 20% penalty before cache effects, which is the neighbourhood of the $7.60 against $5.86 result. The per-token price is not wrong; it is just not the number that governs the invoice.
This does not make the newer model the worse buy. It makes the decision depend on something most comparisons omit: whether your workload lets you bound the output. A summarisation job with a length instruction, a structured-extraction pipeline producing JSON, a classifier returning a label — in all of those the verbosity never lands, and the 60% rate-card cut is real money. An open-ended agent asked to "fix the repo" with no output discipline hands Sonnet 5.5 three times the rope.
Effort settings and the migration nobody budgets for
Both models expose reasoning depth through an effort parameter with multiple levels, and both ship a default rather than a fixed value — high on the Claude Platform, medium inside the Claude apps. The temptation on a migration is to leave the setting where it was on the old model and call the job done.
That is a mistake for a measurable reason. Anthropic's own footnote for Claude Sonnet 5.5 states that the Max effort configuration scores lower than Xhigh on FrontierCode, which means effort is not a monotonic dial and the top setting is not a free upgrade. Set effort by measuring your task, not by choosing the most expensive option available.
There is also a hard behavioural difference on the Sonnet side worth knowing before a rollout. Anthropic states that Claude Sonnet 5.5 is the first Sonnet to ship with cyber safeguards and fallbacks matching its top-tier models, so higher-risk security requests visibly fall back to the earlier Sonnet rather than being served by the model you named. If your workload sits in that domain, your logs will show a model you did not request. Claude Opus 5 pre-dates that arrangement on the Sonnet line, though the same class of routing exists across Anthropic's product for biology and cybersecurity work on the flagship tiers.
On OrcaRouter both endpoints are live today — anthropic/claude-opus-5 and anthropic/claude-sonnet-5 sit in the same catalogue as everything else, priced at the provider's list with 0% markup — and Claude Sonnet 5.5 will be reachable on that same credential as soon as it is listed. The relevant capability in the meantime is automatic failover: if one Anthropic tier is rate-limited or returning errors, the request can be re-pointed at the other without a code change, which is the cheapest possible hedge against being wrong about this comparison.
The decision, stated as a rule
If your work is bounded — you control the output shape, the prompts are structured, and the token count per task is predictable — move to Claude Sonnet 5.5 and take the 60% cut. The index agrees, the rate card agrees, and there is no capability argument left on the other side.
If your work is open-ended and agentic, run the experiment before you believe either price sheet. Measure tokens per completed task on your own traffic for a week on each model; the $7.60 against $5.86 third-party result is a specific warning that the cheaper rate card can produce the larger bill. Claude Opus 5 remains the safer default for long-horizon autonomous work until you have that number.
The honest verdict: this is the first Sonnet generation where the flagship's case rests on task shape rather than raw capability, and anyone still making the decision on the model name alone is now leaving 60% on the table or paying 30% more per finished task, depending only on which way they guess.



Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
