
Claude Haiku 5.5 vs Claude Haiku 4.5: The 90% Price Cut Is Not a 90% Saving
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 65 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Claude Haiku 5.5 is priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Claude Haiku 4.5 costs $1 and $5, flat, for the whole 200,000-token window it has always had. Anthropic puts the difference at "90% lower" in one footnote and "around 75% less to run" in the sentence above it — and the eleven-month gap between the two models is exactly the distance between those two numbers.
That is not a marketing sleight. It is the honest shape of a model generation that changed its tokenizer, its thinking default and its context window at the same time as its price, and it means anyone swapping the model id and expecting the bill to fall by a factor of ten will be disappointed by the arithmetic rather than by the product.
Both models, as shipped
Everything below is per million tokens, in US dollars, and read from Anthropic's own pricing and model documentation on 2026-10-08 unless marked otherwise.

• Input — Claude Haiku 5.5 $0.10 up to 100,000 tokens, $0.50 above it; Claude Haiku 4.5 $1.00 regardless of length.
• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens, $2.50 above it; Claude Haiku 4.5 $5.00 regardless of length.
• Cache reads — Claude Haiku 5.5 $0.01 up to 100,000 tokens and $0.05 above; Claude Haiku 4.5 $0.10. Cache writes — $0.125 and $0.625 against $1.25.
• Context — 1M tokens against 200,000. Maximum output — 128,000 tokens against 64,000.
• Thinking — Claude Haiku 5.5 is adaptive and on by default with an effort dial that defaults to medium; Claude Haiku 4.5 takes the older manual form and has no effort parameter at all.
• Independent score — Artificial Analysis Intelligence Index v4.3.2 places Claude Haiku 5.5 (Max) at 43.40, second of 182 models, and Claude 4.5 Haiku at 16.88, thirtieth of 61 — with the evaluator flagging the 4.5 score as estimated.
• Speed — the same source measures 242 output tokens per second for Claude Haiku 5.5 against 89.7 for the 4.5 generation, which is the only place the old model still looks comfortable.
Where the tenth-of-the-price story breaks
Three things eat into the cut, and only the first of them is Anthropic's own warning.
The tokenizer is the big one. Claude Haiku 5.5 uses the newer tokenizer introduced with Claude 4.7, and Anthropic's documentation says the same text "counts as approximately 30% more tokens than on Claude Haiku 4.5." You are not paying a tenth of the rate on the same number of tokens; you are paying a tenth of the rate on roughly 1.3 times the tokens. The vendor's own migration guide makes this the second item on its checklist, ahead of every code change: recount your prompts, then revisit max_tokens and your cost estimates.
Thinking tokens bill as output tokens, and Claude Haiku 5.5 thinks by default. Anthropic's launch page is explicit that the model is "very verbose" in its own right on the charts, and the independent measurement backs that up more sharply than the vendor does: evaluating the Intelligence Index, Artificial Analysis recorded 440 million output tokens from Claude Haiku 5.5 against a 100 million median across the board, and flagged the model as very verbose. A cheap input rate does not help a workload whose bill is mostly output.
The 100,000-token step is the third. Above it the rates are $0.50 and $2.50 — half of what Claude Haiku 4.5 charged, on the same request, which is a good deal and is not the deal most readers will have read about. Anthropic says prompts under the threshold accounted for around 90% of requests to the previous Haiku model, which is the number that makes the cut survivable as a headline; it is also the vendor's own traffic mix, not yours.

What the same work costs the two models
Cost per finished task is the metric that survives all three of the effects above, because it charges the model for everything it emitted, not just what you sent it.
Artificial Analysis puts Claude Haiku 5.5 (Max) at $0.21 per Intelligence Index task — the figure it publishes alongside a $0.10 input and $0.50 output rate. It publishes no cost-per-task figure for the 4.5 generation at all; the tile reads unknown, because the model was never run on the Index in a reasoning configuration. What the page does publish is a raw sticker of $1.00 and $5.00 and a different, lower measured score.
So the honest comparison for a buyer is not 10x on tokens but something closer to this: a tenth of the input rate on 30% more tokens, half the output rate when you cross 100K, and a model that spends more output tokens per answer than its predecessor did — against a capability jump from 16.88 to 43.40 on the same independent board. The 75% figure Anthropic quotes for average workloads is the sensible planning number. It is also a blended vendor estimate over a traffic mix you do not control.
The migration is not a model-id swap
Anthropic's migration guide runs to ten items for anyone moving off Claude Haiku 4.5, and several are hard failures rather than advice.
• Manual thinking breaks — thinking: {"type": "enabled", "budget_tokens": N} returns an error. Adaptive thinking replaces it, and thinking.display has to be set to "summarized" if you want to see the reasoning at all.
• Sampling parameters break — temperature, top_p and top_k all have to come out of the request.
• Assistant prefill breaks — a conversation ending on an assistant turn returns an error; end it on a user turn.
• Block order changes — a response can begin with thinking blocks, so code that reads the first content block as the answer has to select by type instead.
• Refusals are new — the model runs safety classifiers, can return stop_reason: "refusal", and there is no server-side fallback.
• Replay is constrained — thinking blocks work only in the account that produced them, or an account linked to it.
• Priority Tier does not carry over — organisations holding a Priority Tier commitment on Claude Haiku 4.5 have to plan capacity separately, because the tier is not supported on Claude Haiku 5.5.
Two of those deserve more attention than the rest. The refusal stop reason is a control-flow change in any loop that assumed every response has content, and the account-bound thinking blocks break any queue that stores conversations and replays them through a pool of credentials. Neither shows up until production.
Running both tiers under one key
The practical question for most teams is not which of these two models is better, because the answer depends entirely on which call you are making. It is whether the cheap leg of a cascade can be upgraded independently of everything around it.
Claude Haiku 4.5 is on the OrcaRouter catalogue today at Anthropic's list price with 0% markup, so the provider's rate is passed straight through and any change Anthropic makes to it appears on our side the same day. Claude Sonnet 5.5 and Claude Opus 5.5 are there too, which means a two-tier pipeline — small model for the routine majority, larger model for the escalation — can be built now with one key, one SDK and automatic failover underneath, and the small leg swapped to Claude Haiku 5.5 later as a model-id change rather than a re-write.
Claude Haiku 5.5 is not on the catalogue yet, and it is worth saying that plainly rather than leaving it implied. A launch-day gap between a model shipping and a model being routable is normal and it is not a promise about when it closes; the models callable from us this morning are the ones named above.

Who should switch, and which call to switch first
If your Haiku 4.5 traffic is classification, extraction, routing, compaction or summarisation with prompts under 100,000 tokens, move — the rate is a tenth, the window is five times wider, and the effort dial gives you a cost lever the old model never had. Budget about 75% lower and reconcile against your own token counts after a week.
If your prompts routinely cross 100,000 tokens, you are buying a 50% cut rather than a 90% one, and the second tier changes the comparison in ways that make shopping around worthwhile: the over-100K output rate of $2.50 is above what several competing small models charge for the whole window.
And if you are on Haiku 4.5 because it was the only cheap option that fit a 200,000-token document, note what changed. The window is now a million tokens, which is a different product, and the reason to keep the old model has quietly disappeared along with the reason it was cheap.
