
Claude Haiku 5.5: A 90% Price Cut, a New Tokenizer, and the Video Anthropic Didn't Post
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 65 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Somebody typed one sentence into Claude Haiku 5.5 and got a finished launch film back. The prompt, posted to X on the evening of October 7, 2026, reads in full: "make 10 times banger video for your launch that shows how good motion designer you are." The post's caption is three words long — "One shotted this video" — and it carries the clip as an attachment. No retries, no storyboard, no editor. If that is what the cheapest model in the Claude line does now, the tier stopped being a compromise.
That is the interesting half of the launch. The other half is arithmetic, and it is where the actual news sits: Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, against Claude Haiku 4.5's flat $1 and $5. Anthropic's own footnote calls that 90% lower on the common case — and then says, in the same breath, that the number a buyer should plan around is closer to 75%. Both numbers are correct, and the gap between them is the most important thing in this release.
What Anthropic put on the record on October 7
The vendor's launch post and its pricing documentation agree on the shape of the thing, and it is a larger step than a price cut usually travels with.

• Price — $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that threshold the same request is billed at $0.50 and $2.50. Cache reads cost $0.01 per million up to 100K and $0.05 above it; cache writes are $0.125 and $0.625.
• Context — 1M tokens, with a 128,000-token maximum output. That is the same window and the same output ceiling Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5.5 carry, reproduced on the cheapest tier.
• Thinking — adaptive thinking with the effort parameter, defaulting to medium. Anthropic states this is the first Haiku-class model with an adjustable effort setting, which means the same model id can be run cheap or run careful without changing anything else.
• Tokenizer — the newer tokenizer shared with Claude 4.7 and later, which "produces approximately 30% more tokens for the same text." Anthropic's documentation says the exact increase depends on the content and workload shape.
• Lifecycle — knowledge cutoff June 2026, and a retirement commitment of "not sooner than October 7, 2027."
• Availability — the Claude API, Amazon Web Services, Google Cloud, Microsoft Azure, and Claude Platform on AWS, all at once on day one.
The 10x sticker and the 75% reality
Ten times cheaper is the number that will travel furthest, and it is the one most likely to end up in somebody's budget model in the wrong place. It applies to prompts up to 100,000 tokens. Haiku 5.5 does not hold that rate for the whole window.
• Up to 100,000 tokens — $0.10 in and $0.50 out, against Haiku 4.5's $1 and $5. That is the 90% cut, and Anthropic says 90% of requests to the previous Haiku model fell into this band.
• Over 100,000 tokens — $0.50 in and $2.50 out. Still exactly half of what Haiku 4.5 charged, on the same request, with no ceiling to hit.
• The tokenizer — the same text becomes roughly 30% more tokens than it did on Haiku 4.5, so the per-token cut does not translate one-for-one into a smaller bill.
• The vendor's own average — Anthropic renders the combination as "around 75% less to run." That is its figure, not an independent measurement, and it is the one to put in a forecast.
• Cache economics — cache reads dropped from $0.10 to $0.01 per million in the common band, which matters more than the input rate for any pipeline that re-sends a long shared prefix.
The 75% figure also explains something a reader might otherwise read as a contradiction. The two headline numbers are not in conflict; one is a rate on a request shape, the other is a blended estimate over a real traffic mix that includes the over-100K minority and the tokenizer's extra tokens. If you are modelling spend, use 75% and check your own prompt-length distribution before you believe better.
What the video claim is, and what it isn't
The clip is real, the prompt is quoted verbatim above, and the timestamp — 19:49 UTC on October 7, roughly the same day the model went live — is checkable. The attribution is to an individual user, not to Anthropic.
That distinction matters here, because Anthropic's own product page for Claude Haiku 5.5 carries no embedded video at all: no player, no media file, only a link to the company's YouTube channel. If a launch film was made with the model, the vendor has not said so. So the accurate statement is narrow: a user reports generating a launch video in one shot with Claude Haiku 5.5, and published the prompt. Whether it is one shot, what it cost, and how much of it survived editing are all claims from a single source. Treat the clip as a demonstration, not as a benchmark.
The reason it is worth mentioning anyway is that video generation is not on the model's specification. Haiku 5.5 takes text and images in and returns text. A motion-design result at that quality implies the model was driving something else — a code-generating path, a rendering pipeline, a tool loop — rather than producing frames itself. That is a statement about agentic orchestration at $0.10 input, which is the actually surprising part.
The numbers Anthropic published
The vendor's launch table is a before-and-after against Haiku 4.5, with GPT-6 Luna and Claude Sonnet 5.5 in the same columns for scale. Every figure below is Anthropic's own reported evaluation result, run on Anthropic's harness, and is reproduced but not independently verified.

• Knowledge work — GDPval-AA v2.1 at 1620 against Haiku 4.5's 735, GPT-6 Luna's 1437 and Sonnet 5.5's 1840. AA-Briefcase v1.1 at 1578 against 614, 1336 and 1824.
• Computer use — OSWorld 2.1 at 72.4% on the offline subset, against 15.7% for Haiku 4.5, 48.9% for GPT-6 Luna and 83.9% for Sonnet 5.5.
• Multidisciplinary reasoning — Humanity's Last Exam at 45.9% without tools and 57.4% with them, against 10.2% and 18.7% for Haiku 4.5.
• Agentic coding — Terminal-Bench 4.0 at 39.2% against 0.0%, 16.4% and 70.6%. FrontierCode 1.1 (Main) at 46.4% against 42.4% for GPT-6 Luna and 52.1% for Sonnet 5.5.
• Visual reasoning — Chartography at 46.4% without tools, against 6.4%, 29.1% and 61.6%.
Anthropic is also unusually direct about where the small tier does not win. Its own commentary on the Terminal-Bench chart says Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks," and positions Haiku 5.5 for compaction, summarisation and subagent work instead. The customer quotes in the post point the same direction and carry the same caveat — these are early-access partners describing their own evals, not audits: Asana reports over 30% lower latency for task completions and up to 2.5x faster inference per agent turn; HubSpot reports 92.8% averaged over three runs on a CRM portal suite; AlphaSense reports 0.84 against 0.76 across 400 queries on a workload doing about 8M calls a week; Box reports 11 points higher than Haiku 4.5 at roughly half the latency; Rogo describes a Haiku 5.5 subagent pulling a segment revenue line out of a 10-K; and Cognition reports Devin Fusion holding a FrontierCode score of 66.2 with Haiku 5.5 as the sidekick.
What the independent board says
Vendor tables are the vendor's. The first outside placement is the more useful number for anyone deciding whether the tier moved far enough to change a production pipeline.

Artificial Analysis scores Claude Haiku 5.5 at 43 on its Intelligence Index v4.3.2 in the model's Max configuration, ranked second of the 182 models on that board and well above the board's median of 13. The same page measures 242 output tokens per second — ninth of 182 — and a cost of $0.21 per Intelligence Index task.
Two things on that page are worth more than the headline. The first is verbosity: evaluating the Index, Artificial Analysis recorded 440 million output tokens, against a 100 million median, and describes the model as "very verbose." A sticker price is charged per token, so a model that thinks at length pays for it — which is exactly why the cost-per-task figure sits at $0.21 rather than near zero, and why the 90% input cut is not the whole story. The second is the effort ladder: the same page exposes configurations from Max down to Low, and Anthropic's default is medium, not max. The board's 43 is the top of that ladder, not the setting you get if you do nothing.
Running the tier below the tier you can already call
Claude Haiku 5.5 is not on the OrcaRouter catalogue, and it would be worth saying so plainly rather than implying otherwise: the launch-day gap between a model existing and a model being routable is usually days, and sometimes longer.
What is on the catalogue from Anthropic today is Claude Haiku 4.5, Claude Sonnet 5.5 and Claude Opus 5.5, each at the provider's list price passed through with 0% markup, so any cut Anthropic makes lands on our side the same day it lands on theirs. That is the practical bridge for anyone who wants to test the subagent pattern now: a cascade that sends the routine turns to Haiku 4.5 and escalates the hard ones upward works today under one key and one SDK, and swapping the small leg to Haiku 5.5 later is a model-id change rather than a migration. Automatic failover sits underneath whichever legs you pick, so a single provider's bad afternoon does not take the pipeline with it.
If you would rather not wait at all, the same $0.10 input band is already occupied by GPT-6 Luna, which we do route, and which holds that rate out to 272,000 tokens before stepping up.
Who should move, and who should not
If your workload is classification, extraction, routing, compaction or summarisation at volume, and your prompts sit under 100,000 tokens, the case is simple: the rate is a tenth of what it was, the window is five times wider, and the effort dial lets you trade the other way when a request needs it. Budget for roughly 75% lower and you will not be surprised.
If your prompts routinely run past 100,000 tokens, you are buying a 50% cut rather than a 90% one, and deep-context comparison shopping is still worth the afternoon — the field has several models at or below this tier's over-100K rate.
And if your eval still fails on Terminal-Bench-shaped work, this is not the model that fixes it; Anthropic says so itself, and the 39.2% against Sonnet 5.5's 70.6% is the cleanest evidence in the post. The honest summary of October 7 is that the floor of useful intelligence dropped a long way and the ceiling did not move at all.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
