A generated title card reading 'Claude Sonnet 5.5' with the subtitle 'The mid-tier default, fully specified' and three badges reading 'released 2026-09-28', '$2 input / $10 output per 1M tokens' and 'Intelligence Index 56.00 vs 57.62'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Claude Sonnet 5.5: The Mid-Tier Default, Fully Specified

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Sonnet 5.5 is the mid-tier model of the Claude line, released on September 28, 2026, and the pricing is the whole reason it is worth a page: $2 per million input tokens and $10 per million output tokens, against $4 and $20 for Claude Opus 5.5, the flagship that shipped six days earlier. That is half the output price for 1.62 points of Artificial Analysis Intelligence Index — 56.00 for the mid-tier against 57.62 for the flagship, both figures measured by Artificial Analysis on index revision v4.3.2 and not by us and not by the vendor. Above both of them sits Claude Fable 5.1, at $10 / $50 and 53.35 on the same revision: dearer than the flagship, and behind it. At the same $2 / $10 list price as Claude Sonnet 5.5 sits OpenAI's GPT-6.1 Sol, at 51.83 — 417 index hundredths the other side of the line, which is the comparison a reader weighing the mid-tier actually has to make.

This page is a reference, and the dates say why rather than the freshness. Claude Sonnet 5.5 is ten days old on 2026-10-08, which puts it outside the seven-day window this blog reports news in. What is inside that window is an event six days old: two posts on X dated 2026-10-02, claiming that a successor model in this family beats a rival that its own vendor finished and withheld. That claim has no model id, no rate card, no context window and no benchmark behind it in any registry reachable from here, and it is not this page's subject — the sourcing chain, the vendor-surface absence check and the five-model index spread are all owned in full by the write-up this blog already published under the title Claude Fable 5.5 Beat GPT-6.1 Astra Before It Exists — and That Is the Interesting Part, which is dated 2026-10-03. What that claim did produce is a queue of readers asking which shipped Claude model to standardise on while they wait, and that question needs a page about the model that exists. This is it.

What shipped on September 28, 2026

The vendor's own documentation gives the specification directly, and it is a continuation of the tier rather than a new shape. Everything below was read from Anthropic's Claude Platform documentation on 2026-10-08, so it is Anthropic's specification, not an independent test of it.

• Released — September 28, 2026. The model's own documentation page opens with the line "Latest. Released September 28, 2026."

• Positioning — "The best combination of speed and intelligence", with comparative latency listed as Fast against Moderate for Claude Opus 5.5, Slower for Claude Fable 5.1 and Fastest for the small tier.

• Model ID — claude-sonnet-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS; anthropic.claude-sonnet-5-5 on Amazon Bedrock. Every Claude model ID is a pinned snapshot, so the string is not a pointer that drifts under you.

• Context and output — 1,000,000 tokens of context and 128,000 tokens of synchronous output on the Messages API, rising to 300,000 behind the output-300k-2026-03-24 beta header on the Message Batches API.

• Thinking — adaptive, always on by default, with a default effort of high. The lowest setting is between_tools, which turns off up-front thinking and is accepted at high effort or below. Setting temperature, top_p or top_k to anything other than its default returns a 400 error, so a request migrated from an older Claude model with a tuned temperature will fail loudly rather than quietly.

• Knowledge cutoff — June 2026, for both the reliable cutoff and the training-data cutoff.

• Lifecycle — Active (latest), with a retirement commitment of not sooner than September 28, 2027 on the platforms Anthropic operates. Amazon Bedrock and Google Cloud set their own dates.

The vendor's launch framing is worth quoting because it is a claim and not a measurement: Anthropic's announcement describes Claude Sonnet 5.5 as "a clear upgrade over Sonnet 5 that runs 30% faster and costs up to 30% less for most work." That sentence is the vendor's, unreproduced by us, and it is the kind of claim that depends entirely on which work you put in front of it. There is also a migration cost attached to the release that this page is not the place for: five breaking changes affect code already running on Claude Sonnet 5, and a sixth alters the response shape without failing any request. That accounting belongs to the rollout write-up this blog published in release week, and it is the first thing to read if you have Sonnet 5 in production today.

A screenshot of Anthropic's Claude Platform documentation page for Claude Sonnet 5.5, headed with the line 'Latest. Released September 28, 2026', showing the model ID claude-sonnet-5-5, pricing of $2 per input MTok and $10 per output MTok, a 1M-token context window, 128K max output, adaptive thinking with a default effort of high, a June 2026 knowledge cutoff and a retirement commitment of not sooner than September 28, 2027.

The 1.62 points, and the half price

Independent scores for this model come from one place, and they carry one number that decides the argument. Artificial Analysis measures Claude Sonnet 5.5 at 56.00 on its Intelligence Index, revision v4.3.2, in the configuration the page names as "Claude Sonnet 5.5 (Max, Default Fallback)". We read that page on 2026-10-08. On the same revision, the same day:

• Claude Opus 5.5 — 57.62, released 2026-09-22, $4 / $20

• Claude Sonnet 5.5 — 56.00, released 2026-09-28, $2 / $10

• Claude Fable 5.1 — 53.35, released 2026-09-01, $10 / $50

• GPT-6 Astra — 52.67, released 2026-09-03, $10 / $50, with a long-context tier above 272,000 prompt tokens at $20 / $75

• GPT-6.1 Sol — 51.83, released 2026-09-29, $2 / $10

Stating which revision matters, because a score from one version of the index is not comparable to a score from another: v4.3.2 incorporates ten evaluations — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1 — and the composition of that set has changed across revisions. All five figures above come from the same revision, which is why they can be read against each other, and why none of them can be read against a number quoted from an earlier version.

The price comparison is the part that decides behaviour. Claude Sonnet 5.5's output rate is half of Claude Opus 5.5's, and its input rate is half again; on the vendor's own list, the mid-tier costs $2 / $10 where the flagship costs $4 / $20. Run the two published numbers together and the ratio is stark: per dollar of output list price, Claude Sonnet 5.5 carries 5.60 index points and Claude Opus 5.5 carries 2.88 — nearly twice as much index per dollar of output, on Anthropic's own rate card and Artificial Analysis's own measurement. That is arithmetic on two published figures rather than a measurement of anything, and it is the honest form of the case for the mid-tier. It is also not a verdict: an index point is not a unit of correctness, and the point of the rest of this page is what the 1.62 is and is not made of.

A generated two-column scoreboard titled 'Claude Sonnet 5.5 vs Claude Opus 5.5 - the scoreboard'. The left column, Claude Sonnet 5.5, reads: Released September 28, 2026; Intelligence Index v4.3.2: 56.00; Price per 1M tokens: $2.00 / $10.00; Context window: 1M tokens; Max output: 128K tokens; Default effort: high. The right column, Claude Opus 5.5, reads: Released September 22, 2026; Intelligence Index v4.3.2: 57.62; Price per 1M tokens: $4.00 / $20.00; Context window: 1M tokens; Max output: 128K tokens; Default effort: medium. The footer reads 'Index figures per Artificial Analysis v4.3.2, read 2026-10-08; prices and specifications per Anthropic.' The OrcaRouter logo is composited in the bottom-right corner.

The rate card, and the two lines that move most

Anthropic's full price list, read on 2026-10-08. Every figure in this section is the vendor's.

• Input — $2.00 per million tokens. Output — $10.00 per million, and thinking tokens bill as output, which is not suppressible here because adaptive thinking is on by default.

• Cache read — $0.10 per million, which Anthropic documents as 5 percent of the base input price. That is the same proportional discount Claude Opus 5.5 carries, so caching does not change the ranking between the two, and it does mean a long agent run that keeps a stable prefix pays a tenth of the input rate on everything it re-reads.

• Cache write — $2.50 per million for the five-minute cache, $4.00 for the one-hour cache. The minimum cacheable prompt length is 512 tokens, so short prompts cannot buy the discount at all.

• Batch API — 50 percent off in both directions, so $1.00 input and $5.00 output for work that can wait.

• No long-context surcharge — Anthropic bills the full 1M-token window at the standard per-token rate on this model, so a 900,000-token request costs the same per token as a 9,000-token one. That is a real difference from GPT-6 Astra, whose long-context tier above 272,000 prompt tokens steps to $20 / $75 — worth naming because it is the one place where the two vendors' headline numbers diverge sharply once a prompt gets long.

Two lines then decide most bills rather than the headline rate: the cache-read discount across a long agent run, and the batch discount for anything asynchronous. The headline $2 / $10 is what gets compared on a leaderboard; the cache and batch lines are what a monthly invoice is actually built from.

What the gap does and does not mean

The Intelligence Index is one number over a fixed task mix, and that single property explains most of the confusion around it. The ten evaluations behind v4.3.2 include agentic coding and terminal work, graduate-level scientific reasoning, a software-engineering set, an agentic-automation set and a long-context retrieval set, weighted into a composite. A model that is 1.62 points behind another is behind on that composite — not on your workload, which may weight those ten in a completely different proportion, and may contain tasks none of them represent.

What that means in practice is a split by the shape of the work rather than by the prestige of the tier. Anthropic's own model picker makes the split explicitly, and it is worth reading because it does not flatter the mid-tier: the documentation tells readers to "start with Claude Opus 5.5 for most workloads", and to reach for Claude Fable 5.1 "for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5.5 at higher effort still fall short". Claude Sonnet 5.5 is not the model that sentence starts with. It is the model the same documentation describes as the best combination of speed and intelligence, with the fastest comparative latency outside the small tier — which is a different argument, and a narrower one.

So the honest reading is this. The mid-tier is usually the right answer for well-scoped work at volume: features and bug fixes against a known codebase, extraction and classification with a defined output shape, tool-calling loops where the hard part is the harness rather than the reasoning, and anything where a fast first token and a low cost per call matter more than the last two percent of capability. The flagship earns its rate when the work is genuinely long-horizon — multi-hour agent runs, refactors that span files the prompt cannot hold at once, and evaluations on the mid-tier that measurably fall short at higher effort. That last condition is the vendor's own, and it is a testable one: raise the effort level on the mid-tier, run your own eval, and see whether it closes the gap. If it does, the 1.62 was not load-bearing for you. If it does not, you have bought your answer with your own numbers rather than with someone else's composite.

What the gap does not mean is that 1.62 points is nothing. It is the whole margin between two models, and on a fixed harness evaluated under identical conditions it is a real ordering. It is simply an ordering of a specific task mix, at max effort with default fallback, and it says nothing about which of the two is the better buy for a workload neither of them was benchmarked on.

The name trap: which Claude Sonnet 5.5 page this is

A reader searching this model's name meets several pages on this blog that look like this one, and they are not. The one dated 2026-09-24 is the anticipation page: it was written before the model shipped, and its subject was a single X post claiming a stealth test, reconstructed against the numbers of the model it would replace. The one published in release week is the rollout write-up: its subject is the five breaking changes that affect code already running on Claude Sonnet 5, plus the sixth change that alters the response shape without raising an error. There is also a fan of matchup pages carrying this model's name against other vendors' models, and an applied page on what the release's promotion covers.

None of them is the model page, and that is the gap this piece fills. This page is the reference a reader reaches after searching the model's own name: what it is, what it costs, what it measures, and which work belongs to it. The rollout page's migration detail and the anticipation page's sourcing critique stay where they are; if you have Sonnet 5 code in production, the rollout page is your first click, and this page is what tells you whether the model you are migrating to is the one you want.

One naming habit is worth fixing while we are here, because it misleads on its own. A 5.5 suffix marks a generation, not a rank. Claude Sonnet 5.5 and Claude Opus 5.5 are the same generation of the same family, and the number in front of the dot is the tier. Nothing about the suffix tells you which one wins a benchmark.

Availability, dated

Claude Sonnet 5.5 has been on OrcaRouter since it shipped, as anthropic/claude-sonnet-5.5, at Anthropic's own list price with no markup on top: $2.00 input and $10.00 output per million tokens, the two figures read from our own model endpoint on 2026-10-08 and matching the vendor's rate card line for line. Our page for it carries the same 1M-token context window and 128K maximum output, lists text, image and file as accepted input types and text as the only output, marks it released 2026-09-28 and not deprecated, and shows it reachable both through the OpenAI-compatible chat-completions endpoint and through Anthropic's own messages endpoint — which is what makes it a drop-in change for an integration that already speaks either one.

Everything else in the same catalogue is reached from the same key: one API for 200+ models, with provider list price passed through at 0% markup so a vendor price cut reaches a customer the day it is announced, automatic failover when an endpoint degrades, and a routing DSL for the cases where you want to pin a model or compose several into one answer. That combination is what makes the question on this page cheap to answer rather than expensive. Both models in the comparison above are behind the same credential, so the price you compare against is the vendor's own published price rather than a reseller's, and the version test costs an eval run rather than a migration project: point a slice of traffic at anthropic/claude-sonnet-5.5, keep the rest on the flagship, and let your own cost-per-completed-task numbers decide which one the default should be.

A screenshot of the OrcaRouter model page for anthropic/claude-sonnet-5.5, attributed to Anthropic and dated 2026-09-28, showing a 1M-token context window, 128K maximum output, text image and file inputs with text output, an input price of $2.00 per million tokens and an output price of $10.00 per million tokens, the model id claude-sonnet-5-5, and an OrcaRouter performance panel.

The short version

Claude Sonnet 5.5 is a mid-tier model with flagship-adjacent numbers and half the flagship's output price, and the three facts that decide whether it is your default are these. It sits 1.62 points behind Claude Opus 5.5 on Artificial Analysis's v4.3.2 index, at 56.00 against 57.62, with both figures measured on the same revision and in the same max-effort configuration with default fallback. It costs $2 / $10 against the flagship's $4 / $20, with the cache and batch discounts proportional rather than different, so the price gap survives every discount either model offers. And it defaults to high effort where the flagship defaults to medium, which means a setting carried over from the flagship runs one level higher here rather than lower, and the cost of that is on you to measure rather than assume.

Anthropic's own picker still starts at Claude Opus 5.5 for most workloads, and that is worth knowing before you standardise on anything. But for most readers who arrive at this page by searching this model's name, the mid-tier is the sensible default and the flagship is the exception you buy when your own evals demand it — well-scoped work at volume, fast latency, a million-token window without a long-context surcharge, and a rate that halves the output line. Both are one key away on OrcaRouter, and the eval that decides between them costs an afternoon rather than a migration.

Compared in this article4

Detected from this article · Benchmarks: Artificial Analysis · updated daily