A generated hero card for 'A tenth of the rate is not a tenth of the bill', badged 'GENERATION GAP' with the kicker 'TWO HAIKU MODELS, ELEVEN MONTHS APART' and the subtitle 'Claude Haiku 5.5 against Claude Haiku 4.5 on price, tokens and the new 100,000-token step.' Four chips read '$0.10 / $0.50', '$1.00 / $5.00', '1M context' and '30% more tokens'. Two cards below are labelled 'WHAT THE RATE SAYS' ('90% lower input and output under 100,000 tokens') and 'WHAT THE INVOICE SAYS' ('around 75% less to run, on about 30% more tokens'). A footer reads 'Vendor pricing documentation read October 8, 2026.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Claude Haiku 5.5 vs Claude Haiku 4.5: The 90% Price Cut Is Not a 90% Saving

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Haiku 5.5 is priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Claude Haiku 4.5 costs $1 and $5, flat, for the whole 200,000-token window it has always had. Ant​hropic puts the difference at "90% lower" in one footnote and "around 75% less to run" in the sentence above it — and the eleven-month gap between the two models is exactly the distance between those two numbers.

That is not a marketing sleight. It is the honest shape of a model generation that changed its tokenizer, its thinking default and its context window at the same time as its price, and it means anyone swapping the model id and expecting the bill to fall by a factor of ten will be disappointed by the arithmetic rather than by the product.

Both models, as shipped

Everything below is per million tokens, in US dollars, and read from Ant​hropic's own pricing and model documentation on 2026-10-08 unless marked otherwise.

A generated scoreboard titled 'Claude Haiku 5.5 vs Claude Haiku 4.5 - as shipped'. Claude Haiku 5.5 reads input $0.10 under 100K and $0.50 above, output $0.50 and $2.50, cache read $0.01 and $0.05, context 1M tokens, maximum output 128,000 tokens, thinking adaptive with an effort dial. Claude Haiku 4.5 reads input $1.00 flat, output $5.00 flat, cache read $0.10, context 200,000 tokens, maximum output 64,000 tokens, thinking manual with no effort dial. A footer reads 'Vendors' published list rates and model documentation.'

• Input — Claude Haiku 5.5 $0.10 up to 100,000 tokens, $0.50 above it; Claude Haiku 4.5 $1.00 regardless of length.

• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens, $2.50 above it; Claude Haiku 4.5 $5.00 regardless of length.

• Cache reads — Claude Haiku 5.5 $0.01 up to 100,000 tokens and $0.05 above; Claude Haiku 4.5 $0.10. Cache writes — $0.125 and $0.625 against $1.25.

• Context — 1M tokens against 200,000. Maximum output — 128,000 tokens against 64,000.

• Thinking — Claude Haiku 5.5 is adaptive and on by default with an effort dial that defaults to medium; Claude Haiku 4.5 takes the older manual form and has no effort parameter at all.

• Independent score — Artificial Analysis Intelligence Index v4.3.2 places Claude Haiku 5.5 (Max) at 43.40, second of 182 models, and Claude 4.5 Haiku at 16.88, thirtieth of 61 — with the evaluator flagging the 4.5 score as estimated.

• Speed — the same source measures 242 output tokens per second for Claude Haiku 5.5 against 89.7 for the 4.5 generation, which is the only place the old model still looks comfortable.

Where the tenth-of-the-price story breaks

Three things eat into the cut, and only the first of them is Ant​hropic's own warning.

The tokenizer is the big one. Claude Haiku 5.5 uses the newer tokenizer introduced with Claude 4.7, and Ant​hropic's documentation says the same text "counts as approximately 30% more tokens than on Claude Haiku 4.5." You are not paying a tenth of the rate on the same number of tokens; you are paying a tenth of the rate on roughly 1.3 times the tokens. The vendor's own migration guide makes this the second item on its checklist, ahead of every code change: recount your prompts, then revisit max_tokens and your cost estimates.

Thinking tokens bill as output tokens, and Claude Haiku 5.5 thinks by default. Ant​hropic's launch page is explicit that the model is "very verbose" in its own right on the charts, and the independent measurement backs that up more sharply than the vendor does: evaluating the Intelligence Index, Artificial Analysis recorded 440 million output tokens from Claude Haiku 5.5 against a 100 million median across the board, and flagged the model as very verbose. A cheap input rate does not help a workload whose bill is mostly output.

The 100,000-token step is the third. Above it the rates are $0.50 and $2.50 — half of what Claude Haiku 4.5 charged, on the same request, which is a good deal and is not the deal most readers will have read about. Ant​hropic says prompts under the threshold accounted for around 90% of requests to the previous Haiku model, which is the number that makes the cut survivable as a headline; it is also the vendor's own traffic mix, not yours.

A screenshot of Anthropic's model documentation showing the Claude Haiku 5.5 row in the models table and the published rate card beneath it, with the input, output, cache read and cache write rates and the prompt-length thresholds set out per million tokens.

What the same work costs the two models

Cost per finished task is the metric that survives all three of the effects above, because it charges the model for everything it emitted, not just what you sent it.

Artificial Analysis puts Claude Haiku 5.5 (Max) at $0.21 per Intelligence Index task — the figure it publishes alongside a $0.10 input and $0.50 output rate. It publishes no cost-per-task figure for the 4.5 generation at all; the tile reads unknown, because the model was never run on the Index in a reasoning configuration. What the page does publish is a raw sticker of $1.00 and $5.00 and a different, lower measured score.

So the honest comparison for a buyer is not 10x on tokens but something closer to this: a tenth of the input rate on 30% more tokens, half the output rate when you cross 100K, and a model that spends more output tokens per answer than its predecessor did — against a capability jump from 16.88 to 43.40 on the same independent board. The 75% figure Ant​hropic quotes for average workloads is the sensible planning number. It is also a blended vendor estimate over a traffic mix you do not control.

The migration is not a model-id swap

Ant​hropic's migration guide runs to ten items for anyone moving off Claude Haiku 4.5, and several are hard failures rather than advice.

• Manual thinking breaks — thinking: {"type": "enabled", "budget_tokens": N} returns an error. Adaptive thinking replaces it, and thinking.display has to be set to "summarized" if you want to see the reasoning at all.

• Sampling parameters break — temperature, top_p and top_k all have to come out of the request.

• Assistant prefill breaks — a conversation ending on an assistant turn returns an error; end it on a user turn.

• Block order changes — a response can begin with thinking blocks, so code that reads the first content block as the answer has to select by type instead.

• Refusals are new — the model runs safety classifiers, can return stop_reason: "refusal", and there is no server-side fallback.

• Replay is constrained — thinking blocks work only in the account that produced them, or an account linked to it.

• Priority Tier does not carry over — organisations holding a Priority Tier commitment on Claude Haiku 4.5 have to plan capacity separately, because the tier is not supported on Claude Haiku 5.5.

Two of those deserve more attention than the rest. The refusal stop reason is a control-flow change in any loop that assumed every response has content, and the account-bound thinking blocks break any queue that stores conversations and replays them through a pool of credentials. Neither shows up until production.

Running both tiers under one key

The practical question for most teams is not which of these two models is better, because the answer depends entirely on which call you are making. It is whether the cheap leg of a cascade can be upgraded independently of everything around it.

Claude Haiku 4.5 is on the OrcaRouter catalogue today at Anthropic's list price with 0% markup, so the provider's rate is passed straight through and any change Ant​hropic makes to it appears on our side the same day. Claude Sonnet 5.5 and Claude Opus 5.5 are there too, which means a two-tier pipeline — small model for the routine majority, larger model for the escalation — can be built now with one key, one SDK and automatic failover underneath, and the small leg swapped to Claude Haiku 5.5 later as a model-id change rather than a re-write.

Claude Haiku 5.5 is not on the catalogue yet, and it is worth saying that plainly rather than leaving it implied. A launch-day gap between a model shipping and a model being routable is normal and it is not a promise about when it closes; the models callable from us this morning are the ones named above.

A screenshot of the OrcaRouter catalogue page for Claude Haiku 4.5, showing the model identifier, the provider Anthropic, the model description and a stat strip carrying the $1.00 per million input and $5.00 per million output rates alongside the context window and maximum output.

Who should switch, and which call to switch first

If your Haiku 4.5 traffic is classification, extraction, routing, compaction or summarisation with prompts under 100,000 tokens, move — the rate is a tenth, the window is five times wider, and the effort dial gives you a cost lever the old model never had. Budget about 75% lower and reconcile against your own token counts after a week.

If your prompts routinely cross 100,000 tokens, you are buying a 50% cut rather than a 90% one, and the second tier changes the comparison in ways that make shopping around worthwhile: the over-100K output rate of $2.50 is above what several competing small models charge for the whole window.

And if you are on Haiku 4.5 because it was the only cheap option that fit a 200,000-token document, note what changed. The window is now a million tokens, which is a different product, and the reason to keep the old model has quietly disappeared along with the reason it was cheap.