A generated hero card reading 'Daily AI Brief — October 8, 2026' with a NEWS badge, the headline 'Claude Haiku 5.5 ships, and Sonnet 5.5 cache reads halve', and four chips reading $0.10 / $0.50 per 1M, 1M-token context, Medium default effort and $0.20 to $0.10 cache read.
Engineering & Research

Daily AI Brief, Oct 8: Claude Haiku 5.5 Lands at $0.10 and Sonnet 5.5 Cache Reads Halve

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Haiku 5.5 shipped on October 7, 2026, and it is the cheapest thing the company has ever put behind an API: $0.10 per million input tokens and $0.50 per million output tokens on prompts up to 100,000 tokens, with a 1M-token context window and an effort dial running from Low to Max. The same day, with no launch post of its own, prompt-cache reads on Claude Sonnet 5.5 dropped from $0.20 to $0.10 per million tokens. One of those is a product and the other is a line item, and the line item is the one that will move more money this quarter.

The price first, because it is the part you can act on today. Then effort, which is the thing that quietly decides the price. Then the cache cut, which half the people paying for Sonnet 5.5 will not hear about until their next invoice.

What Claude Haiku 5.5 is, in the order it matters

The vendor's own framing is "the cheapest, fastest and most capable small model we've ever released," and the release date is October 7, 2026 — one day before this brief. The model ID is claude-haiku-5-5. It is text and image in, text out. The context window is 1,000,000 tokens, up from 200,000 on Claude Haiku 4.5, and the maximum output is 128,000 tokens, with a beta header on the Batch API raising that to 300,000. The training cutoff is June 2026, and Anthropic commits not to retire it before October 7, 2027.

The line that matters most for anyone routing traffic is not the context window. It is that this is the first Haiku-class model with adjustable effort. Effort runs Low, Medium, High, Xhigh and Max, the default is Medium, and adaptive thinking is on by default — a Sonnet-and-Opus feature arriving one tier down. The launch post also carries two things that are about buying rather than building: monthly API credits for Claude Max and Team plans, $100 at Max 5x and $200 at Max 20x with up to $500 pooled for Team, and beta computer-use and browser-use support in the SDKs. Safety keeps pace with the tier: the cyber safeguards are stricter than Claude Haiku 4.5's, with penetration-testing requests still blocked, and the biology safeguards match those on Claude Sonnet 5, Claude Sonnet 5.5 and Claude Opus 5.

A capture of Anthropic's Claude Haiku 5.5 announcement page showing the October 7, 2026 date, the headline framing Claude Haiku 5.5 as a performance, pricing and safety upgrade, and the on-page sections for Performance, Pricing, Safety and Further updates.

The price, and the 100,000-token cliff in the middle of it

The headline rate is $0.10 per million input tokens and $0.50 per million output tokens. Read the asterisk: that is the rate for prompts at or below 100,000 tokens. Above 100,000, both numbers multiply by five, to $0.50 input and $2.50 output. Cache writes cost $0.125 per million at the five-minute tier and $0.20 at the one-hour tier inside the cheap band, and cache reads cost $0.01 per million — a tenth of the input rate, which is the deepest cache discount Anthropic has published on any model. Batch halves everything again: $0.05 and $0.25 inside the cheap band, $0.25 and $1.25 above it.

Compare that against Claude Haiku 4.5, which is still on the price list at a flat $1.00 input and $5.00 output with no context-based tiering, a $0.10 cache read and a 200,000-token window. The new model is a tenth of the price on the same prompt shape, with five times the context. There is no migration calculus to run here; the arithmetic is not close.

The cliff is the part to design around. A 99,000-token prompt costs 99 cents per thousand calls at the input rate. A 101,000-token prompt costs $5.05 per thousand. Two thousand tokens of difference is a 5x bill, so anything you build on the cheap band should either stay comfortably under 100,000 tokens or be deliberately, consistently over it — the worst outcome is a pipeline that straddles the line and pays the high rate on half its traffic.

Effort is a pricing parameter now, and the default is not the cheapest

Because effort defaults to Medium and thinking is on by default, the sticker price and the real price are not the same number. Reasoning tokens bill as output tokens. On Anthropic's own published run of the Artificial Analysis Intelligence Index — vendor-reported, unreproduced by us — Claude Haiku 5.5 emits roughly 162,000 tokens per task, of which about 129,000 are reasoning. The index's median model emits on the order of 100 million tokens across the suite; Haiku 5.5 spends 440 million. At $0.50 per million output tokens that verbosity is affordable, and that is precisely the point: at Haiku 5.5's rate, thinking a lot is still cheaper than thinking a little on a model that charges $5.00 output.

Set effort to Low for classification, extraction and routing decisions where you want a fast answer rather than a considered one, and leave it at Medium for anything a user will read. Dropping to Low is also the reliable way to keep a verbose agent inside the 100,000-token band, because it cuts reasoning tokens before it cuts answer quality.

The quieter line: Sonnet 5.5 cache reads halved

Claude Sonnet 5.5 reads cached prompt at $0.10 per million tokens as of October 7, down from $0.20. Sonnet 5.5's input price is $2.00 and its output is $10.00, so a 0.05x cache-read multiplier is now the same multiplier Haiku 5.5 has, applied to a twenty-times-larger input rate. Anthropic's own estimate is that this cuts the cost of most agentic work by around 20%, which is a claim about workload shape rather than a benchmark: it holds only if a large share of your input is a stable prefix you are re-sending every turn. If your prompts are unique on every call, this change costs you nothing and saves you nothing.

Worth noting what the cut implies. Anthropic is pricing long-lived prefixes aggressively on the mid-tier model at exactly the moment it ships a very cheap small model, which reads as an argument for a two-model architecture: a Sonnet 5.5 leg with a hot cache for the work that needs judgement, and a Haiku 5.5 leg for the volume that does not.

What the independent numbers show, one day in

Artificial Analysis had Claude Haiku 5.5 scored the day after launch. Its Intelligence Index reads 43 overall, with the Max-effort configuration reported at 43.40, second of 182 models on the board. Cost per index task is $0.21, forty-sixth cheapest. Blended price at a 7:2:1 input-to-cached-input-to-output mix is $0.08 per million, which is the number that makes the earlier tier comparison feel unfair to everything else. Output speed is 243.4 tokens per second, ninth fastest on the board, with a reported first-token latency of about 0.3 seconds and a 90% cache discount.

That 440-million-token output figure is the caveat attached to the 43. The score is real and independently run; so is the verbosity. A model that thinks this much is cheap per token and not always cheap per task, and the gap between those two numbers is where a routing decision actually lives.

For scale, Anthropic's own published comparisons put Claude Haiku 5.5 at 1620 on GDPval-AA v2.1 against 735 for Claude Haiku 4.5, 72.4% on OSWorld 2.1 against 15.7%, 45.9% on Humanity's Last Exam without tools against 10.2%, and 39.2% on Terminal-Bench 4.0 against 0.0%. Those are vendor-reported figures on vendor-chosen evaluations and should be read as directional, but the direction is not subtle — this is not a small upgrade to a small model.

A generated single-column scoreboard titled 'Claude Haiku 5.5 — the scoreboard' listing an Artificial Analysis Intelligence Index of 43 ranked second of 182, a cost per index task of $0.21, a blended price of $0.08 per million tokens, output speed of 243.4 tokens per second, a 1M-token context window and a $0.01-per-million cache read.

Getting to these models without a second contract

Claude Haiku 5.5 is available through Anthropic's own API and, from there, wherever Anthropic's models are resold. On OrcaRouter's catalogue the Anthropic line today runs anthropic/claude-haiku-4.5, anthropic/claude-sonnet-5.5, anthropic/claude-opus-5.5 and anthropic/claude-fable-5.1 — we do not host Claude Haiku 5.5, and this brief will not pretend otherwise. What we do carry is the tier above it, and OrcaRouter passes provider list price through with 0% markup, which is why the Sonnet 5.5 cache-read cut appeared in our pricing on the same day Anthropic made it rather than on a scheduled repricing cycle.

The practical version: if you want to try Haiku 5.5 on a classification leg while your judgement leg stays on Claude Sonnet 5.5, you are holding two keys instead of one, and the routing layer is where the A/B gets decided. That is the whole argument for a gateway over a direct integration — one endpoint, model strings, and a failover path when a provider wobbles.

A capture of the OrcaRouter models catalogue showing the filterable model list with input-modality, context-length, input-price, status and series filters, the line counting the model list and providers on one API key and one bill, and model cards carrying per-million input and output prices.

What to watch from here

Three things. Whether third-party providers start reselling Haiku 5.5 at a blended rate below $0.10 blended, which is the point at which the routing layer earns its keep rather than just simplifying procurement. Whether Anthropic extends the 100,000-token tier boundary, since a cliff at exactly 100,000 is a strange place for a model with a 1M-token window to have one. And whether the 440-million-token verbosity figure comes down in a point release, because that single number is what stands between Haiku 5.5 and being the obvious default for the whole bottom half of a production stack.