A hero card titled 'Claude Haiku 5.5 vs Claude Sonnet 5.5' with the subhead 'Six days and one tier apart'. Two rounded cards sit side by side: the left headed 'Claude Haiku 5.5' with rows reading Index 43, $0.21 per task and Released Oct 7, 2026; the right headed 'Claude Sonnet 5.5' with rows reading Index 56, $7.60 per task and Released Sep 28, 2026. Between them a centred badge reads '36x measured cost gap'. A footer line reads 'Cost per task and Intelligence Index per Artificial Analysis; both models 1M context.'
Guides & Insights

Claude Haiku 5.5 vs Claude Sonnet 5.5: A 20x Rate Card, a 36x Measured Bill

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Haiku 5.5 and Claude Sonnet 5.5 sit six days and one tier apart in the same family, share a one-million-token context window, and answer on the same API shape. On the published rate card the cheaper one costs a fifth of the dearer one in the band most requests land in. On the independent Artificial Analysis board the gap is wider than that: $0.21 to finish one Intelligence Index task on Claude Haiku 5.5 against $7.60 on Claude Sonnet 5.5, for an Intelligence Index of 43 against 56. The signal that started this piece put it as Claude Haiku 5.5 being "way better than GPT-6 Luna" and "close(r) to Sonnet 5.5". The first half of that is roughly what the vendor's own comparison set shows. The second half is where the measurements stop agreeing with the marketing, and it is worth being precise about which.

Two models, six days and one tier apart

Claude Haiku 5.5 shipped on October 7, 2026 under the model ID claude-haiku-5-5. It is the direct replacement for Claude Haiku 4.5, takes text and images in and returns text, and is the first Haiku-class model Anthropic has given an adjustable effort setting — Low, Medium, High, Xhigh and Max, with medium as the API default. Anthropic calls it "the cheapest, fastest, and most capable small model we've ever released", and its footnote qualifies that as fastest at each model's standard speed, still behind the Opus models when those run in Fast Mode.

Claude Sonnet 5.5 landed six days earlier, on September 28, 2026, as claude-sonnet-5-5 — the tier Anthropic positions as the best combination of speed and intelligence, and the one it recommends for complex agentic coding. Same one-million-token window. Unlike Claude Haiku 5.5, it has no prompt-length price band, which turns out to matter more than the six-day head start does.

Both are available now on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure, and both carry a retirement commitment of not sooner than one year after launch — October 7, 2027 for Claude Haiku 5.5, September 28, 2027 for Claude Sonnet 5.5. Anthropic says Claude Sonnet 5.5 is available with zero data retention, matching Claude Opus 5.5 and Claude Sonnet 5.

The rate card, both sides, in one place

Per million tokens, in US dollars, at Anthropic's published rates:

• Input — Claude Haiku 5.5 $0.10 for prompts up to 100,000 tokens, $0.50 above that line; Claude Sonnet 5.5 $2.00 flat, no step.

• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens, $2.50 above it; Claude Sonnet 5.5 $10.00 flat.

• Cache reads — Claude Haiku 5.5 $0.01 in the lower band and $0.05 above it; Claude Sonnet 5.5 $0.10. Anthropic halved Claude Sonnet 5.5 cache reads from $0.20 at the same time as the Haiku launch, and says that because cache reads are a large share of agentic token use, Claude Sonnet 5.5 now costs about 20% less on most agentic work.

• Cache writes — Claude Haiku 5.5 $0.125 for a five-minute write and $0.20 for a one-hour write in the lower band, $0.625 and $1.00 above it; Claude Sonnet 5.5 $2.50 for a five-minute write and $4.00 for one hour.

• Batch — the Batch API is 50% off both, on input and output. That puts Claude Haiku 5.5 at $0.05 and $0.25 below the 100,000-token line and $0.25 and $1.25 above it, against Claude Sonnet 5.5 at $1.00 and $5.00.

Read across those lines and the shape is consistent: in the lower band you are looking at a 20x input gap and a 20x output gap; above the line, 4x and 4x. Anthropic's headline is "around 75% less to run" on average against Claude Haiku 4.5, refined in a footnote to 90% cheaper below 100,000 tokens and 50% cheaper above it, with the tokenizer change folded into that average.

The 100,000-token step is where the arithmetic turns

Two things pull in opposite directions, and only one of them is on the rate card. Prices fall hard against Claude Haiku 4.5. At the same time, Claude Haiku 5.5 ships an updated tokenizer, similar to the ones in Claude Sonnet 5.5 and Claude Opus 5.5, which Anthropic says "uses slightly more tokens per task" — so an identical prompt arrives as a larger number of billable tokens than it would have on the previous generation. The 75% average is the net of both, and it is a vendor number.

The 100,000-token line is the part that catches people. Anthropic prices Claude Haiku 5.5 by prompt length: a request whose prompt exceeds 100,000 tokens pays the higher prices on the whole request, not just on the tokens past the threshold. Prompt length counts all input tokens, including cache reads and cache writes, and each request is priced on its own — a long request pays the higher band even when part of its prompt is a cache hit, and earlier requests keep the prices they were billed at. A retrieval pipeline sitting comfortably at 90,000 tokens pays $0.10 and $0.50. Nudge the window one notch and the entire request reprices at $0.50 and $2.50. Claude Sonnet 5.5 does not do this, because the full million-token window is billed at standard pricing.

Two consequences follow, pointing in opposite directions. First, the crossover people expect — long context making the expensive model the cheaper one — does not happen: Claude Haiku 5.5 above the line is still a quarter of Claude Sonnet 5.5 on the rate card. Second, the gap you were counting on narrows from 20x to 4x exactly when your workload gets heavy, which is exactly when absolute numbers start to matter. The step is the reason a cost-sensitive pipeline needs its own budget line rather than a per-token rate.

What each vendor reports it scores

Everything in this section is vendor-reported and unreproduced by a third party. Anthropic publishes the same comparison set on its Claude Haiku 5.5 and Claude Sonnet 5.5 pages, and where the two pages overlap they agree:

• GDPval-AA v2.1, knowledge work — 1,620 for Claude Haiku 5.5, against 1,840 for Claude Sonnet 5.5, 735 for Claude Haiku 4.5 and 1,437 for GPT-6 Luna.

• AA-Briefcase v1.1 — 1,578 for Claude Haiku 5.5, against 1,824 for Claude Sonnet 5.5, 614 for Claude Haiku 4.5 and 1,336 for GPT-6 Luna.

• OSWorld 2.1, computer use, offline subset — 72.4% for Claude Haiku 5.5, against 83.9% for Claude Sonnet 5.5, 15.7% for Claude Haiku 4.5 and 48.9% for GPT-6 Luna.

• Terminal-Bench 4.0, agentic coding — 39.2% for Claude Haiku 5.5, against 70.6% for Claude Sonnet 5.5, 0.0% for Claude Haiku 4.5 and 16.4% for GPT-6 Luna.

• FrontierCode 1.1, Main — 46.4% for Claude Haiku 5.5, against 52.1% for Claude Sonnet 5.5 at Xhigh effort, 42.4% for GPT-6 Luna. Anthropic also flags that Claude Sonnet 5.5 scores lower at Max effort than at Xhigh on this one, which it attributes to more frequent use of a code-review skill causing timeouts or out-of-scope edits.

• Chartography, visual reasoning without tools — 46.4% for Claude Haiku 5.5, against 61.6% for Claude Sonnet 5.5, 6.4% for Claude Haiku 4.5 and 29.1% for GPT-6 Luna.

• Humanity's Last Exam — 45.9% without tools and 57.4% with them for Claude Haiku 5.5; 64.5% with tools for Claude Sonnet 5.5, 18.7% for Claude Haiku 4.5, and 56.9% without tools for Claude Sonnet 5.5 on the same page. Anthropic notes the GPT-6 family figures may be stale because OpenAI fixed an image-understanding bug and third-party scores had not all been updated.

Two of those lines are worth reading twice. On FrontierCode the two models are close at all — 46.4% against 52.1% — and on Chartography, where there are no tools to hide behind, the gap is 15 points. Terminal-Bench 4.0 is not close at all: 39.2% against 70.6% is a different capability class, which is why Anthropic's Claude Sonnet 5.5 page still points at Claude Sonnet 5.5 and Claude Opus 5.5 for complex agentic coding and frames Claude Haiku 5.5 as the model for narrow, high-volume work.

What the independent board adds

Anthropic's numbers compare its own models to each other and to one competitor. An independent board puts both against the whole field, and the figure it reports that the vendors do not is cost per finished task.

Artificial Analysis scores Claude Haiku 5.5 at Max effort with an Intelligence Index of 43, well above the median of 13 among comparable models, generating 440M output tokens on the Index against a median of 100M — very verbose — at $0.10 input and $0.50 output per million, and 242 output tokens per second. Its measured cost is $0.21 per Intelligence Index task.

The same board scores Claude Sonnet 5.5 at 56, against a median of 26, generating 410M output tokens against a median of 88M, at $2.00 input and $10.00 output with a 90% cache discount. Its measured cost is $7.60 per Intelligence Index task.

A headless capture of the Artificial Analysis model page for Claude Haiku 5.5 at Max effort, showing the header 'Claude Haiku 5.5 (Max)' with the Anthropic vendor label and release month October 2026, a rank strip reading Intelligence #2/182, Speed #8/182, Cost #46/182 and Verbosity #62/182, the price strip In $0.10, Out $0.50 and $0.21 per task, a Comparison Summary stating an Artificial Analysis Intelligence Index of 43 against a median of 13 with 440M output tokens generated, and an output-token speed of 242 tokens per second described as notably fast. A specification panel confirms a 1M-token context window, reasoning enabled, text and image input and text output.A headless capture of the Artificial Analysis model page for Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback), showing the Anthropic vendor label and release month September 2026, a rank strip reading Intelligence #3/216, Cost #98/216 and Verbosity #102/216, the price strip In $2.00, Out $10.00, a 90% cache discount and $7.60 per task, a Comparison Summary stating an Intelligence Index of 56 against a median of 26 with 410M output tokens generated against a median of 88M, and a specification panel showing a 1M-token context window and text and image input.

Put those two lines side by side and the ratio is the story. $0.21 against $7.60 is a little over 36x, and the two models are nearly level on verbosity — 440M output tokens against 410M — so that multiplier is almost entirely the rate card doing what the rate card says it will. On the vendor's own rate card the same comparison is 20x. The extra comes from the things a per-token price cannot show: the cheaper model reaches its answer in fewer billed tokens per task at this effort setting, and cache reads are billed at a tenth of Claude Sonnet 5.5's rate. Same harness, same tasks, independently measured.

A generated two-column comparison scoreboard titled 'Claude Haiku 5.5 vs Claude Sonnet 5.5 - the scoreboard'. The left column, headed Claude Haiku 5.5, reads AA Intelligence Index: 43; Cost per Index task: $0.21; Input / output per 1M: $0.10 / $0.50; Output speed: 242 tok/s; Context window: 1M; Terminal-Bench 4.0: 39.2%. The right column, headed Claude Sonnet 5.5, reads AA Intelligence Index: 56; Cost per Index task: $7.60; Input / output per 1M: $2.00 / $10.00; Output speed: not published; Context window: 1M; Terminal-Bench 4.0: 70.6%. A footer line reads 'Index, cost per task and speed per Artificial Analysis; Terminal-Bench vendor-reported.'

Where the cheap tier actually runs out

The customer numbers Anthropic publishes decide the split more often than the benchmarks do, and they all point the same way. Asana reports over 30% lower latency on task completions and up to 2.5x faster inference per agent turn. Box reports a score 11 points above Claude Haiku 4.5 at about half the latency. HubSpot reports the best score it has seen on its own suite, 92.8% averaged over three runs, with the highest hit rate and lowest false-positive rate. AlphaSense runs about 8M calls a week through "Ask in Document" and reports a statistically significant improvement over Claude Haiku 4.5, 0.84 against 0.76. Cognition reports that with Claude Haiku 5.5 as the sidekick, its Fusion agent holds a FrontierCode score of 66.2.

Not one of those is a long-horizon agent. They are point solutions — one well-defined step, getting faster and cheaper. That is the shape Claude Haiku 5.5 is built for, and it is also the shape where the 36x cost advantage is real money rather than a rounding error. The advantage compounds with call volume and evaporates with call depth: a hundred thousand classification calls a day at $0.10 in and $0.50 out is a line item you notice disappearing, while a hundred agent turns per session, each re-reading accumulated context, is a line item that grows superlinearly because every turn past the threshold pays the higher band.

Claude Sonnet 5.5's counter-argument is per-attempt cost rather than per-token cost. Anthropic says it runs 30%+ faster than Claude Sonnet 5 and costs up to 30% less for most work, the saving coming from needing fewer tokens rather than from lower rates. Third-party testers put numbers on that: Slack reports roughly 14% fewer output tokens, Balyasny about 121k tokens per answer against 497k, Box 2.4x faster with 12% fewer tokens, Lovable roughly a third fewer tool calls and about half the shell runs, Atlassian up to 30% faster Rovo Agents, and Base44 3.6 iterations per build against 7.7 for Claude Opus 5. One turn that lands beats four that do not.

There is a third factor running the other way. Anthropic has moved effort to a model-level dial on this generation — Claude Haiku 5.5's effort setting is set once per request rather than sized as a per-request thinking budget, and the old budget_tokens mode is deprecated on the 4.6 models and not accepted on later ones. For a fleet calling one model ID from many services that is a simplification. For anything that sized its own reasoning per call, it is a rewrite.

Running both behind one endpoint

The reason this decision is worth getting right is that the wrong half of it is expensive to undo: rewiring a production path from one model ID to another means touching every call site that assumed the old one's behaviour. The cheaper way to find out which tier a workload belongs on is to run both behind one endpoint and move traffic between them by configuration rather than by refactor.

That is what OrcaRouter is for. Claude Sonnet 5.5 is on the catalogue as anthropic/claude-sonnet-5.5, alongside Claude Sonnet 5, Claude Opus 5.5 and Claude Haiku 4.5 — the older Haiku tier stays available as a reference point while you migrate off it. The platform passes provider list price through at zero markup, so the $2.00 and $10.00 above are the rates you pay, and the same holds in reverse: when a vendor reprices, the new price is live here the same day, which is how the Claude Sonnet 5.5 cache-read cut reached us. Two features do the work this particular decision needs. Automatic failover keeps a model you are still evaluating off your critical path on its own, and the routing DSL lets you send sub-100,000-token calls to the cheap tier and hand the long-context or long-horizon ones to the dearer one in the same request flow, without a second contract or a second SDK.

To be precise about what is and is not hosted: Claude Sonnet 5.5 is routable through us today. Claude Haiku 5.5 is not yet on the catalogue. Until it is, it lives on Anthropic's own API and the three cloud platforms listed above.

How to decide

Pick Claude Haiku 5.5 when the task is narrow, high-volume and checkable — classification, extraction, routing, summarisation, the per-turn work inside an agent you have already built. The 20x rate gap is the whole point there, the 100,000-token step is avoidable by design, and the independently measured $0.21 per task says the win survives outside the vendor's own lab. Set the effort dial per task class rather than leaving it on the default; on a Haiku-class model the difference between Low and Max is a cost multiplier, not a quality preference.

Pick Claude Sonnet 5.5 when the task is long-horizon, agentic or unverifiable — code that has to compile, computer use that has to leave the screen in a known state, anything where a wrong answer costs more than the call. Terminal-Bench 4.0 at 70.6% against 39.2% is not a nuance, it is a different capability class, and the flat rate card means a long context does not silently reprice the whole request. If your workload is genuinely in between, the honest answer is that neither model has a broad body of third-party reproduction behind it yet, the index scores come from a single harness, and the vendor figures quoted above are the vendor's own.

What would change this answer

Three things are worth watching, and none of them are benchmark releases. The first is whether independent evaluations of Claude Haiku 5.5 at Low, High and Xhigh effort — not just Max — show the same cost-per-task ratio, because the effort dial is the cheapest lever in the comparison and the least measured. The second is whether the 100,000-token step survives; a third band, or no band at all on the next generation, would change the arithmetic for every retrieval pipeline at the line more than any single benchmark score. The third is the tokenizer. If a later checkpoint tokenizes the same text at the older rate, Anthropic's 75% average becomes a floor rather than a target, and the gap against Claude Sonnet 5.5 widens without either price moving.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily