
Claude Haiku 5.5 in OpenCode Go: What $10 a Month Actually Buys
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 65 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Anthropic shipped Claude Haiku 5.5 on October 7, 2026, and within a day it was listed in OpenCode's Go and Zen catalogues — the coding-agent subscriptions that meter a monthly dollar allowance instead of billing tokens. That second listing is the part worth slowing down for, because the allowance OpenCode publishes for Claude Haiku 5.5 is not a discount on the model. It is a budget meter set at Anthropic's own list price, and every figure behind it can be checked against the two pages that publish it. If you are still deciding what to call in place of Claude Haiku 4.5, the allowance is the decision, not the sticker.
The signal itself is one line: OpenCode posted "Claude 5.5 Haiku now available in OpenCode and Go" to X at 19:37 UTC on October 7, hours after the model went live. That is a platform wiring up somebody else's launch, not a launch of its own — which is exactly why it is checkable. OpenCode's marketing page carries a "Claude Haiku 5.5 — New" row at 3,850 estimated requests per five hours on the $10 Go plan and 15,380 on the $40 Go Plus plan. Its documentation lists the model id claude-haiku-5-5 against the endpoint https://opencode.ai/zen/go/v1/messages, and the same string appears in the Zen catalogue at https://opencode.ai/zen/v1/messages. Two independent surfaces, same model id, same day.
What Anthropic put on the record on October 7
The vendor did not ship this quietly. Anthropic's newsroom carries a dated announcement, a product page with a benchmark table, a system card, a pricing table and six named customer quotes. There is no repository and no weights release to inspect, because there is nothing to release: Artificial Analysis records Claude Haiku 5.5 as proprietary, with the parameter count undisclosed. Anyone describing this as a leaked or unannounced model is reading a repo that does not exist.
Anthropic's own framing is a volume model rather than a frontier one:
• Positioning — "the cheapest, fastest, and most capable small model we've ever released," aimed at summaries, compactions, database queries and classification, and pitched as a subagent that pairs with Claude Opus 5.5 and Claude Sonnet 5.5 on coding work.
• Price — $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens; $0.50 and $2.50 above that. Cache reads $0.01 / $0.05, cache writes $0.125 / $0.625.
• Context and output — a 1M-token window with a 128,000-token maximum output, the same envelope the larger 5.5-tier models carry, reproduced on the cheapest tier.
• Effort — the first Haiku-class model with an adjustable effort setting, running Low through Max, defaulting to medium. Anthropic publishes accuracy-versus-cost curves at each setting rather than a single number.
• Benchmarks, vendor-reported and not independently reproduced — GDPval-AA v2.1 1620 against Claude Haiku 4.5's 735; AA-Briefcase v1.1 1578 against 614; OSWorld 2.1 72.4% against 15.7%; Humanity's Last Exam 45.9% without tools and 57.4% with them, against 10.2% and 18.7%; Terminal-Bench 4.0 39.2% against 0.0%; Chartography 46.4% against 6.4%.
• The vendor's own ceiling — Anthropic's commentary on its Terminal-Bench chart says Claude Sonnet 5.5 and Claude Opus 5.5 "remain better choices for complex agentic coding tasks." The customer quotes in the same post are early-access partners describing their own evals, not audits: Asana reports over 30% lower latency and up to 2.5x faster inference per agent turn, HubSpot 92.8% averaged over three runs, AlphaSense 0.84 against 0.76 across 400 queries on a workload of roughly 8M calls a week, Box 11 points higher at about half the latency, and Cognition a FrontierCode score of 66.2 with Haiku 5.5 as the sidekick.

Anthropic also used the launch to move two adjacent prices: cache reads on Claude Sonnet 5.5 halved to $0.10 per million tokens, which the company says makes Sonnet 5.5 about 20% cheaper on most agentic work, and a new monthly API credit for Claude Max and Team subscribers. Both are changes to the tier above, and both matter more to a team already standardized on Sonnet than to anyone shopping for a small model.
The allowance is a dollar budget, not a discount
Here is the part the announcement tweet does not say. OpenCode's Go documentation states that "usage limits are defined as monthly dollar amounts," that token pricing is the same on both plans, and that each model's allowance tapers on a fixed schedule: 20% of the monthly limit per five hours, 50% per week, 100% per month. The documentation also publishes the per-request token profile it assumes for this model — 1,100 input tokens, 55,000 cached tokens, 240 output tokens.
Run those through Anthropic's list price and the published request counts stop looking like an arbitrary allocation:
• Cost per request at list price — 1,100 input at $0.10/M plus 55,000 cached reads at $0.01/M plus 240 output at $0.50/M comes to $0.00078.
• The monthly request count — $15 divided by $0.00078 is 19,230, which is the number the documentation prints. The $40 plan's $60 allowance divides out to 76,923, and the documentation prints 76,920.
• The taper — 20% of 19,230 is 3,846, printed as 3,850 per five hours; 50% is 9,615, printed as 9,620 per week. Go Plus follows the same arithmetic at 15,380 and 38,460.
• What that means in dollars — the allowance is worth $15 of Anthropic-priced tokens on the $10 plan and $60 on the $40 plan. A $10 subscription therefore buys at most $15 a month of this model's tokens if you spend the whole allowance on Claude Haiku 5.5 and nothing else.
That is the honest way to read the "New" badge on the Go page. The subscription is cheap, the model is cheap, and the two facts are independent: the allowance is a metered budget at the vendor's own rate, so the plan constrains how much you may spend rather than how little the model costs. Heavy users hit the ceiling and pay Anthropic directly — or route the model themselves.

Ninety percent, seventy-five percent, and the five-times cliff
Anthropic's headline is that Claude Haiku 5.5 costs about 90% less than Claude Haiku 4.5 for prompts up to 100,000 tokens, and that on average it "now costs around 75% less to run." Both numbers are the vendor's. They are also both correct, and the gap between them is the whole of the pricing risk.
• Within the band — $0.10 in and $0.50 out, against Haiku 4.5's flat $1 and $5. Anthropic says around 90% of requests to its previous Haiku model fell in this band, which is where the 90% figure comes from.
• Above the band — the same request is billed at $0.50 in and $2.50 out, a five-fold step up on input. That is still half of Haiku 4.5's rate, but it is nowhere near a 90% cut.
• The blended claim — the 75% figure is Anthropic's own estimate across a real traffic mix, not a rate card, and it is the number to put in a forecast.
• The tokenizer — the newer tokenizer shared with Claude 4.7 and later produces roughly 30% more tokens for the same text, per Anthropic's own documentation, so a per-token cut does not translate one-for-one into a smaller bill.
The effort dial compounds all of it. Because the default is medium and the published curves run up to Max, the same prompt can cost several times more depending on a parameter most callers never set — and OpenCode's request estimates assume one fixed token profile at one implicit setting. If your prompts sit near 100,000 tokens, run a bounded-use subagent, or you move the effort dial, the 90% headline is not the number you will see.
The independent board agrees on the score and flags its own maths
Artificial Analysis scores Claude Haiku 5.5 at 43 on its Intelligence Index v4.3.2 in the model's Max configuration, against a board median of 13, and measures 241.9 output tokens per second against a comparable-model median of 110.9. The same page puts the cost at $0.21 per Intelligence Index task and reports that the model generated 440 million output tokens during evaluation against a 100 million median — describing it as "very verbose."
Verbosity is where the two boards interact. The allowance arithmetic above is charged per token, so a model that reasons at length consumes the budget faster than the request count suggests, and OpenCode's per-request profile of 55,000 cached tokens is an assumption about prompt caching that a subscription user may or may not reproduce. Two disclosures on the Artificial Analysis page are worth more than its headline: the 43 is the Max configuration rather than the effort setting you get by doing nothing, and the page carries an explicit note that Claude Haiku 5.5 "has 5x higher pricing on longer-context inputs which are not currently reflected" in its cost figures. The long-context cliff is real, and the independent board has told you it is not yet in its own numbers.

The retention line most buyers will skip
OpenCode's Go documentation carries a per-model privacy table, and this is the row where Claude Haiku 5.5 differs from most of the catalogue. The vendor's own framing is zero retention with exceptions. Most models on that Go table read "0 days." This one reads "30 days," with the note that requests are retained for 30 days in accordance with Anthropic's data policies — the same treatment it gives GPT-6 Luna and GPT-5.6 Luna, and the opposite of the zero-day rows that make up the bulk of the list.
For a coding agent that means prompts, file paths and pasted source may sit with the upstream vendor for a month, on a plan whose marketing language is zero retention. That is not a reason to avoid the model; it is a reason to read the row instead of the banner. Nothing in Anthropic's launch materials, meanwhile, changes the API-side position: the Claude API has offered zero data retention, and the model card's retention terms are the ones that govern a direct integration.
What is confirmed, and what is not
The confirmed set is large and specific, which is unusual for a model this new. What is not confirmed is narrower, and worth naming rather than glossing:
• Confirmed — the model exists, is API-only under the id claude-haiku-5-5, shipped October 7 on the Claude API plus Amazon Web Services, Google Cloud and Microsoft Azure, at the prices above, with a 1M-token window, adjustable effort, and a published system card.
• Confirmed — two third-party platforms have it listed with model ids, endpoints and metered allowances the same day.
• Not confirmed — no independent benchmark we can find has reproduced the vendor's table. Every accuracy figure in Anthropic's launch post comes from Anthropic's own harness, and every customer quote is a partner's self-reported early testing.
• Not confirmed — there are no weights, no repository, no checkpoint and no licence to inspect. The model is closed, and any claim about its architecture or size is inference.
• Not confirmed — whether the platform's request estimates hold at anything other than its assumed token profile, and whether (or when) the 5x long-context pricing lands in the independent boards' cost figures.
Running it next to what you already call
The practical question this launch creates is not "is Haiku 5.5 good" — the vendor's table says it is, and nobody outside Anthropic has checked. It is "what is the bill, and what happens when the cheap tier is not the right tier for this particular call." Both halves are routing problems.
On cost, the meter is the token bill, not the sticker: a 90% input cut plus a 30% bigger tokenizer plus a verbose reasoning model plus a five-fold long-context step means the number you forecast and the number you pay diverge unless you measure your own prompt-length distribution. A model that is 5x more expensive above 100,000 tokens rewards compaction — and Haiku 5.5 is itself pitched as the compaction model, which is a loop worth noticing.
On risk, an unproven tier is exactly the case for not betting a production path on it. Claude Haiku 4.5 is already on OrcaRouter at Anthropic's list price with 0% markup, so the old tier needs no new contract, and a route that falls back from a new model to a known one is a configuration change rather than a rewrite. The effort parameter is per-call, which means the same endpoint can run a summary at low effort and a hard extraction at high effort without maintaining two integrations — the routing DSL exists for precisely this shape of decision, and automatic failover is what lets you put a days-old model in front of real traffic before you have independent numbers for it. What is not true yet is the obvious thing: Claude Haiku 5.5 is not on our routes today, and if you want it this week you are calling Anthropic directly or going through the platform that announced it.
Who should move: teams whose volume is summaries, classification, compaction and subagent work, where the 90% band applies and the tokenizer's 30% is arithmetic you can absorb. Who should wait: anyone whose prompts routinely cross 100,000 tokens, anyone who needs an independent benchmark before committing, and anyone whose code cannot sit in a 30-day retention window. The launch is real and the pricing is real. The proof is still the vendor's.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
