
Claude Haiku 5.5 Is Live at $0.10 per Million Tokens: The 75% Claim, and Its Two Footnotes
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 111 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 55 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 377 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Anthropic put Claude Haiku 5.5 into general availability on October 7, 2026 at $0.10 per million input tokens and $0.50 per million output — the same two numbers OpenAI charges for GPT-6 Luna, and a fifth of what Claude Haiku 4.5 has cost since 2025, when it launched at $1.00 and $5.00. The headline in Anthropic's launch post is that the new model costs "around 75% less to run" than its predecessor on average, and that is the figure every summary since has repeated. It travels with two footnotes most of those summaries dropped: under 100,000 tokens the cut is 90%, above it the cut is 50%, and Claude Haiku 5.5 uses the newer tokenizer shared with Claude 4.7 and later, so the same text counts as roughly 30% more tokens than on Claude Haiku 4.5. So the 75% is a statement about a bill, not about a rate card, and for anyone whose prompts are long those are different things. The vendor's own comparison table also runs GPT-6 Luna and Claude Sonnet 5.5 as columns next to it, which makes the launch documentation unusually easy to argue with.
The arithmetic behind "75% less"
Take the rate card first, because it is not flat. Input is $0.10 per million up to 100,000 tokens and $0.50 per million above that — five times. Output is $0.50 and then $2.50. Caching follows the same step: a five-minute cache write is $0.125 per million under the line and $0.625 over it, an hour-long write is $0.20 and $1.00, and a cache read is $0.01 and $0.05. The Message Batches API takes 50% off both meters, and on the batch path the model supports up to 300,000 output tokens with the output-300k-2026-03-24 beta header.
Now put the tokenizer on top of that, because it is where the headline claim gets interesting. A prompt that measured 90,000 tokens on the Claude Haiku 4.5 tokenizer measures roughly 117,000 on the new one. It has not changed size, but it has crossed the 100,000-token line, so it is billed at $0.50 per million input instead of $0.10. On Claude Haiku 4.5 that same content cost $0.09 of input ($1.00 per million × 0.09 million). On Claude Haiku 5.5 it costs $0.0585. That is a 35% cut — real, and nowhere near 90%. The 90% figure applies to requests that stay under the line after the tokenizer expands them, which is where a great deal of short-call traffic lives: a 20,000-token classification call becomes 26,000 tokens, stays in the first tier, and there the input price really is a tenth of what it was.
Three things follow from that, and they are worth separating rather than averaging:
• Short calls get the full discount. Extraction, routing, classification, summarisation of a document that fits under the line — these see the 90% figure on input and output, less the tokenizer's expansion.
• Long calls get about a third off, not three quarters. A request that lands above 100,000 tokens after re-tokenisation pays $0.50 and $2.50 per million — still half of Claude Haiku 4.5's rates, but the tier it sits in is five times the tier the sticker advertises.
• The "75% less on average" is a traffic-weighted number and it is Anthropic's. It is the vendor's own description of its own expected mix, and Anthropic is explicit that a prompt under the threshold made up around 90% of requests to the previous Haiku. If your mix does not look like that mix, your average will not look like that average — and the only way to know which is to count tokens on your own traffic, on the new tokenizer, before you set a budget.

What else shipped with it
The model is not only a price. Several of the launch details change how you write code against it, and one of them will break an existing client.
• Adaptive thinking is on by default, with an effort dial from Low to Max and a default of medium. This is the first Haiku-class model with an adjustable effort setting, and it is the main lever on both quality and cost per answer.
• Sampling parameters are rejected. Sending a non-default temperature, top_p or top_k returns a 400. Any migration that carried those arguments forward will fail loudly on the first request rather than degrade quietly, which is at least an honest failure mode.
• 1M-token context and up to 128,000 output tokens in the synchronous Messages API, with a June 2026 knowledge cutoff and a retirement commitment no sooner than October 7, 2027.
• Five platforms on day one: the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. That is a same-day rollout rather than the staged availability these launches often take.
• Tooling moved too. Anthropic's Python and TypeScript SDKs now support computer use and browser use in beta, and the company added monthly API credits for Max 5x ($100), Max 20x ($200) and Team (up to $500 pooled) subscribers.
• The safety posture is narrower than the price suggests. Anthropic says cyber safeguards on Claude Haiku 5.5 are more restrictive than Claude Haiku 4.5's but less restrictive than recent frontier models — wider defensive use than Claude Sonnet 5.5 permits, with penetration testing still blocked — and that biology safeguards match Claude Sonnet 5, Claude Sonnet 5.5 and Claude Opus 5. Alignment evaluations improved broadly over the predecessor on Anthropic's own suite.
What the vendor's table says, and what the independent board says
Anthropic published a benchmark table with three comparison columns, and it is worth reading as a document about how vendors argue. Every figure below is vendor-reported, produced on Anthropic's harnesses, and none of it has been independently reproduced; where GPT-6 Luna appears, it is Anthropic running its own evaluations against a competitor's model.
• Knowledge work — GDPval-AA v2.1 at 1620 for Claude Haiku 5.5, against 735 for Claude Haiku 4.5, 1437 for GPT-6 Luna and 1840 for Claude Sonnet 5.5. AA-Briefcase v1.1 at 1578, 614, 1336 and 1824.
• Computer use — OSWorld 2.1 offline at 72.4% against 15.7%, 48.9% and 83.9%.
• Reasoning — Humanity's Last Exam at 45.9% without tools and 57.4% with them, against 10.2% and 18.7% for Claude Haiku 4.5 and 56.9% and 64.5% for Claude Sonnet 5.5.
• Agentic coding — Terminal-Bench 4.0 at 39.2% against 0.0%, 16.4% and 70.6%. FrontierCode 1.1 (Main) at 46.4% against 42.4% for GPT-6 Luna and 52.1% for Claude Sonnet 5.5.
• Chart reading — Chartography at 46.4% without tools against 6.4%, 29.1% and 61.6%.
On that table Claude Haiku 5.5 wins every row against both the previous Haiku and GPT-6 Luna, and loses most of them to Claude Sonnet 5.5 — which is the honest shape of a small tier: a large generational jump, and still a tier below the mid model on the hardest work. Anthropic says as much in the post, recommending Claude Sonnet 5.5 and Claude Opus 5.5 for complex agentic coding while positioning the Haiku for compaction, summarisation, classification and subagent roles.
The independent picture is more nuanced and costs more attention. Artificial Analysis measures Claude Haiku 5.5 at Max effort at 43 on its Intelligence Index v4.3.2, second of 182 models, and 241.9 output tokens per second — fast, as advertised. It also measures a cost of $0.21 per completed Index task, because the model generated 440 million output tokens running the suite against a 100 million median; the evaluator describes it as very verbose. GPT-6 Luna, at 38 on the same index, costs $0.07 per task. Same sticker price, three times the bill. That is not a contradiction of the launch claim — it is the same phenomenon the launch claim is about, seen from the other end: the rate card is what you pay per token, and the number you actually pay is rate times tokens, and Claude Haiku 5.5 spends more of them per answer than its rivals do. Effort is the lever; the model's default is medium, not Max, and a team that pins it lower will see both a smaller bill and a smaller score.
Calling it from one place
The migration question this launch creates is not "is it better" but "what does my traffic cost on it", and that answer depends on your prompt lengths, your cache hit rate and the effort setting you pick. Getting there means running your own prompt set through both generations — which is cheap to do behind one endpoint and tedious to do across two.
Claude Haiku 5.5 itself is not on OrcaRouter's catalogue yet, and saying so plainly is more useful than implying otherwise. The rest of the Anthropic line is: Claude Haiku 4.5, Claude Sonnet 5.5 and Claude Opus 5.5 are all callable from us today at the provider's list price with 0% markup passed through, so the "before" leg of a migration test runs on the same key you would keep afterwards, and a vendor cut on that leg is live on our side the same day. The "after" leg is Anthropic's API until the new model reaches a runtime we route. What the shared endpoint buys you meanwhile is automatic failover underneath the call, so a new model can carry production traffic without the path depending on it, and the routing DSL for the case where the escalation from Haiku to Sonnet belongs in the route rather than in your application code.

Two customer results in the launch are worth carrying forward as leads rather than proof. Asana reports over a 30% reduction in latency and up to 2.5x faster inference per agent turn; HubSpot reports 92.8% averaged over three runs on its CRM suite; AlphaSense reports 0.84 against 0.76 on 400 queries. These are vendor-selected customer numbers on the vendor's own announcement, and their value is in telling you which workloads to test first, not in telling you what you will get.
What would change the picture
The 75% figure will be settled the same way every traffic-weighted vendor claim is settled: by your own invoice. What is genuinely open on the independent side is the cost-per-task number. If it holds as more runs accumulate, the honest summary of this launch is that Anthropic cut the rate card by 80% and built a model that spends more tokens per answer, so the effective saving is real, large, workload-dependent, and rarely the number on the poster. If it moves down with effort settings and caching, the tier looks better still. Either way, the number to check before a cutover is your own tokens per task under the new tokenizer — that is the one figure the launch post cannot give you.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
