A generated title card reading 'Claude Haiku 5.5: the 100K price step', showing two price cards — 'Under 100K tokens: $0.10 in / $0.50 out' and 'Over 100K tokens: $0.50 in / $2.50 out' — with an upward step graphic between them and a footer line 'Anthropic list prices, October 2026.' The real OrcaRouter logo is composited bottom-right.
Guides & Insights

Claude Haiku 5.5's Real Story Is the 100K Cliff: What the New Haiku Charges Past 100,000 Tokens

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Haiku 5.5 arrived on October 7, 2026 with a headline price of $0.10 per million input tokens and $0.50 per million output tokens, and the number that got repeated everywhere was the one A​nthropic itself put in the announcement: a 90 percent cut against the Haiku 4.5 rate for the same two line items. The 90 percent is real. It is also scoped to one half of one dimension of the bill, and the moment a request crosses 100,000 tokens every rate in it does something the launch coverage mostly skipped over. This piece is about that boundary line — what it costs, which workloads actually hit it, and why "90 percent cheaper" is a claim about a slice of the bill rather than about your bill.

The two prices, and where they switch

A​nthropic publishes two columns for the model, and the switch between them is a single threshold on the size of the request rather than on any model setting you choose:

• Short tier, at or under 100,000 tokens — $0.10 per million input tokens, $0.50 per million output tokens (A​nthropic list) • Long tier, over 100,000 tokens — $0.50 per million input tokens, $2.50 per million output tokens (A​nthropic list) • Cache reads — $0.01 per million (short tier) and $0.05 per million (long tier) • Batch API — 50 percent off either column, and a maximum output length of 300,000 tokens • Standard maximum output — 128,000 tokens, with a 1 million token context window in both tiers

A generated single-column scoreboard titled 'Claude Haiku 5.5 — the scoreboard' listing Input price (under 100K): $0.10 per million, Output price (under 100K): $0.50 per million, Input price (over 100K): $0.50 per million, Output price (over 100K): $2.50 per million, Context window: 1,000,000 tokens, AA Index: 43.4, with the footer 'Anthropic list prices; AA Index per Artificial Analysis.' The real OrcaRouter logo is composited bottom-right.

Those are A​nthropic's own published rates, not measured ones, and they are the current list: pricing on a model this new has not had time to move, though A​nthropic is not under any commitment to hold it.

Read those four numbers in order and the shape of the thing is obvious. Every rate in the long-context tier is exactly five times the short-context rate, and the threshold is the same 100,000 tokens for all of them at once. There is no gradual curve, no interpolation, no intermediate band — a 99,000-token request and a 101,000-token request differ by a factor of five on every token in the bill, not just on the 2,000 tokens that crossed the line. That is the part that makes this a cliff rather than a slope, and it is the part the "90 percent cheaper" framing cannot see.

A screenshot of Anthropic's Claude Platform Docs pricing page, showing the model pricing table and the per-model rate columns for the Claude line, captured in English.

Why the cut looks bigger than it is

Against Haiku 4.5, the comparison A​nthropic made, the short tier is a genuine 90 percent reduction on input and output alike — Haiku 4.5 ran $1.00 and $5.00 per million for the same two columns. The long tier is a different story: $0.50 and $2.50 against the same $1.00 and $5.00 is a 50 percent cut, not a 90 percent one. So the honest version of the launch claim is that Claude Haiku 5.5 is 90 percent cheaper than Haiku 4.5 below 100,000 tokens and 50 percent cheaper above it, which is exactly the split A​nthropic's own documentation gives once you read both columns instead of one. The single percentage is not wrong; it is the short-tier percentage, presented without its condition.

There is a second thing pulling in the opposite direction, and it is the more interesting one because it is easy to miss and hard to work around. Claude Haiku 5.5 ships with a new tokenizer, and A​nthropic's own guidance puts the token count for a given piece of text roughly 30 percent higher than Haiku 4.5 produced for the same content. Whatever you were paying per token fell and what you pay per token of *your* text rose by about a third, relative to the old model's accounting. A​nthropic's stated net figure, that the model is about 75 percent cheaper to run on average, already folds this in — a 90 percent list-price cut minus a roughly 30 percent token inflation lands in that neighbourhood on short-context traffic — but the two effects are separable and worth separating. The price move is a list-price change you can read off a page; the tokenizer move changes how many units of that price your existing prompts consume, and it shows up in your own token counts before it shows up anywhere else.

For a workload that already lived comfortably under 100,000 tokens per call, the practical effect is a straightforward reduction in cost per call, partly offset by the tokenizer. For one that did not, the arithmetic runs the other way twice on the long half.

What actually lives above 100,000 tokens

The reflex is to file long-context pricing under "agentic coding and RAG, not us." That is too quick, because the 100,000-token line is a per-request line, and per-request size is not the same thing as task size. A few patterns cross it routinely:

• A codebase-reading agent that keeps a large working set in context across a session, rather than re-retrieving per step

• Document and contract analysis where the input is a long PDF plus its own instructions and few-shot examples — the kind of prompt that was 60,000 tokens before anyone added the examples

• Ingestion pipelines that batch many small items into one call to amortise per-call overhead, and in doing so walk the whole batch over the threshold

• Chat products that keep a full conversation history inline instead of summarising it, where the request grows monotonically until something truncates it

• Audio and long-video transcription-style inputs, which are exactly the traffic that a 1 million token window is advertised to enable

The 1 million token context window is the part of the spec sheet that makes the cliff worth talking about. A​nthropic is selling the ability to hand the model a very large request, and the pricing is structured so that the last 900,000 tokens of that request cost five times what the first 100,000 did. Both facts are deliberate; they are just announced in different places. If the long-context window is why you are looking at this model at all, the long-context column is your column, and the launch percentage is somebody else's.

The pressure the price puts on the Flash tier

A​nthropic's stated intent behind the model is to hold the cheap, high-throughput end of the stack, and the price list reflects that intent plainly. Bloomberg's reporting on the launch framed it as A​nthropic defending its lower-cost position against cheaper open-weight competitors, and the vendor-side framing went further — that the new Haiku is priced to undercut its direct rivals by a wide margin.

The comparison is only clean at the short tier, where $0.10 and $0.50 are aggressively low. Above 100,000 tokens the $0.50 and $2.50 rates are no longer undercutting anybody's short tier; they are ordinary mid-tier numbers. The vendor's "wide margin" claim is a statement about the rate card's first column, and A​nthropic is entitled to make it there — it is an accurate reading of the numbers it publishes. It is not a claim about what a long-context workload pays, and it should not be quoted as one.

The rest of the model, briefly

Claude Haiku 5.5 keeps the features that shipped on the Sonnet and Opus line rather than trimming them for the size class. Adaptive thinking with a default effort setting of medium is present, and it is the first model in the Haiku class with an adjustable effort dial spanning Low to Max. The knowledge cutoff is June 2026 and the model ID is claude-haiku-5-5. A​nthropic states a retirement date no sooner than October 7, 2027 — the same date as the release, one year out — and the model is listed on Amazon Bedrock, G​oogle Cloud and Microsoft Azure alongside A​nthropic's own API.

A screenshot of the Artificial Analysis model page for Claude Haiku 5.5, headlined 'Proprietary model — Released October 2026', with an Intelligence Index of 43.395, speed #10 of 182, and a cost comparison block showing $0.10 input.

On the benchmark side, everything in the launch post carries the same provenance label and deserves it: these are vendor-reported numbers, measured by the company selling the model, with no independent reproduction published at the time of writing.

• GDPval-AA v2.1 — vendor-reported 1,620 for Claude Haiku 5.5, against 735 for Haiku 4.5 in the same harness

• AA-Briefcase v1.1 — vendor-reported 1,578, against 614 for the previous generation

• OSWorld 2.1 — vendor-reported 72.4 percent, against 15.7 percent

• Terminal-Bench 4.0 — vendor-reported 39.2, against 0.0

• Humanity's Last Exam — vendor-reported 45.9 percent, or 57.4 with tool use

The size of those deltas is the reason to flag them rather than quote them. A jump from 15.7 to 72.4 on an OS-agent benchmark measured by the model's own vendor is a number to hold loosely until a third party runs the same harness. Artificial Analysis has published its own index for the model as of early October 2026 — an independent figure of 43.395, with 0.329 on Terminal-Bench Hard and 0.444 on HLE — and that set is the one to cite when the claim needs to survive a challenge. The vendor numbers and the independent numbers are not in conflict; they are simply different harnesses, and mixing them into one sentence is how a blog post ends up asserting that a vendor eval is a fact.

The infrastructure note, since we run one

A​nthropic sells Claude Haiku 5.5 through its own API and through Bedrock, G​oogle Cloud and Azure. If you are deciding which of those to call, or you are adding it next to the other models already in your stack, note that the vendor list price is the same wherever it is sold, and OrcaRouter is built to pass that list price through at zero markup — so an A​nthropic price move on this model is live on our side the same day rather than on a delay. Our catalogue also carries qwen/qwen3.8-flash at $0.15 and $0.47 per million and z-ai/glm-5.3-flash at $0.15 and $0.50, the pair that most often ends up in the slot next to a Haiku-class model, on the same key and the same bill — the thing that makes a mixed stack cost one contract instead of three. We do not serve Claude Haiku 5.5 itself; for that, call A​nthropic directly.

One practical corollary of the cliff, if you route more than one model: the 100,000-token boundary is a place where a request that used to be cheap gets expensive without anything in the request changing meaningfully. If your routing layer can fail over, the sensible shape is to keep long-context traffic pointed at whichever model is actually cheap above the threshold and let short-context traffic fall to whichever is cheap below it, rather than assuming one model's headline rate applies to both.

What to take from the launch

The price cut is real and it is large, below 100,000 tokens. The model is the first Haiku with a shipped-through feature set and an effort dial, which matters more for agentic use than the price does. The benchmark deltas are vendor-reported and should be labelled that way every time they are repeated. And the price list has a step in it at 100,000 tokens that multiplies every rate by five at once — which is the one thing about this launch that a reader of your stack, rather than a reader of the press release, most needs to know before the next sprint's token budget gets set.

If you want to check the arithmetic against your own traffic, the two numbers to pull are the 95th percentile of your request size and your share of spend on requests above 100,000 tokens. If the first is under the line, the 90 percent applies to you and the launch coverage was right about your situation. If the second is a meaningful slice of the bill, the number that matters is the 50, and it is worth re-running the cost estimate for whichever model in the stack you picked on the assumption of the 90.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily