A generated hero card titled 'Claude Haiku 5.5 vs Sakana Namazu' with two side-by-side panels. The left panel, Claude Haiku 5.5, lists 'General purpose', '$0.10 / $0.50 per 1M' and 'Index 43'. The right panel, Sakana Namazu, lists 'Japanese specialist', '$0.95 / $4.00 per 1M' and 'Vendor benchmarks only'. A full-width strip beneath reads 'Same fast-and-cheap shape, opposite specialisms.' A footer reads 'Index per Artificial Analysis; Namazu figures per Sakana AI, unaudited.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Claude Haiku 5.5 vs Sakana Namazu: Both Cheap, Both Fast, Almost Nothing Else in Common

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

On the surface these two look like peers. Claude Haiku 5.5 is the vendor's small fast tier, released October 7, 2026 at $0.10 per million input tokens and $0.50 per million output with a one-million-token window, built for classification, extraction, routing and subagent work. Sakana Namazu is Sakana AI's Japanese-specialised API model, launched on August 3, 2026, billed at $0.95 per million input tokens and $4.00 per million output with a $0.15 cached-input rate, refined on top of the open Kimi K2.6 base with Sakana's own post-training, and available through an OpenAI-compatible endpoint.

They are both positioned as the affordable option in their respective lineups, and that is where the resemblance stops. Claude Haiku 5.5 is a general-purpose Western model you would reach for to route traffic; Sakana Namazu is a country-specialised model you would reach for when the task is in Japanese and the alternative is a general model doing it badly. The interesting question is not which is better. It is what a specialist buys you, where the boundary of that specialism actually sits, and what it costs to leave it.

What Sakana AI actually built

Sakana Namazu is worth understanding before it is compared to anything, because its construction is unusual. It is not trained from scratch. It is a refinement of an existing open-weights model — Kimi K2.6 — which Sakana post-trained further, and which it serves as a closed API. The lineage matters: the base model's general reasoning and coding capability is inherited, and what Sakana added is the Japan-specific layer.

That layer is documented, which is more than can be said for most localisation claims. Sakana reports that Namazu retains the base model's performance on general evaluations — AIME26, MMLU-Pro, LiveCodeBench v6 — while surpassing it across the Japanese and Japan-specific set, with FairPoliticsQA rising from 34.10% to 56.30%. It runs its own in-house Japanese translation benchmark, JFBench, and reports improvements across the board there too. A separate case study, an evidence-finding tool for physicians, reports 96.8% on the 119th national medical exam and 96.4% on the 120th.

Those are Sakana's numbers on Sakana's benchmarks, and JFBench is not a public standard anybody else can reproduce. The claim that survives that caveat is narrow but real: the model is a named base model plus a documented post-training step, and the reported deltas are against its own base rather than against the field.

The API and the commercial terms are the parts that will decide most business cases. Namazu is OpenAI-compatible — a change of endpoint and a sakana-namazu model name is the only code change, per Sakana's own documentation. It bills in US dollars and offers yen payment on the Enterprise plan. It also carries tooling costs that a per-token comparison will never surface: $7.00 per thousand web search calls and $0.12 per hour of code execution, which Sakana bundles into the product.

There are two constraints that matter more than any number on the rate card. Namazu is not available in the EU, EEA, UK or Switzerland pending GDPR-related compliance work, per Sakana's own page. And the default data-handling policy states that inputs may be used to train and improve Sakana's models, with opt-out available in Console settings; the company says it cannot currently guarantee that processing stays entirely within Japan. For a European or UK team, the first is disqualifying and the second is a question to answer before the first call.

The boundary, stated honestly

This is the section where a spec-sheet comparison would lie to you, so here is what is actually comparable:

• Price — Claude Haiku 5.5 $0.10 input and $0.50 output per million tokens up to 100,000 tokens of prompt, then $0.50 and $2.50. Sakana Namazu $0.95 and $4.00, with a $0.15 cached-input rate. Namazu is roughly nine and a half times more expensive on input and eight times on output at the first tier.

• Caching — Claude Haiku 5.5 cache reads are $0.01 and $0.05 per million depending on tier; Sakana Namazu's is $0.15 flat. The absolute figures are close at the top tier and twenty-nine times apart at the bottom.

• Tools — neither is a bare text model and neither charges the same way. Claude Haiku 5.5 ships adaptive thinking with an effort dial from Low to Max and server-side tool support; Sakana Namazu bundles web search at $7.00 per thousand calls and code execution at $0.12 per hour.

• Context — one million tokens for Claude Haiku 5.5. Sakana does not publish a context window for Namazu, and any figure you see quoted for it is an assumption about the Kimi K2.6 base rather than a specification.

• Independent scoring — Claude Haiku 5.5 has an Artificial Analysis Intelligence Index of 43, second of 182 in its class on that board, with a $0.21 cost per Intelligence Index task. Sakana Namazu has no Artificial Analysis entry and no third-party evaluation of any kind that I could locate. Its published figures are its own.

• Availability — Claude Haiku 5.5 is on the Claude API, Amazon Bedrock, Vertex AI, Microsoft Foundry and Claude Platform on AWS, with a published retirement commitment of no sooner than October 7, 2027. Sakana Namazu is on Sakana's own API, Sakana Chat, and the Sakana Translate product, and is unavailable in the EU, EEA, UK and Switzerland.

• Data handling — The vendor's terms are enterprise-standard, and the model's thinking blocks are scoped to the account that produced them. Sakana Namazu's default is that inputs may be used for training, with opt-out.

• Modality — Claude Haiku 5.5 takes text and images; Sakana Namazu's modality coverage is not published beyond its text-first positioning.

• Licence and lineage — Claude Haiku 5.5 is proprietary and API-only. Sakana Namazu is proprietary but built on open Kimi K2.6 weights, which means the base is inspectable and the tuned model is not.

The honest scoring problem

Read that list and notice what is missing: any line that says which model produces better Japanese output. That is not an oversight. There is no shared benchmark between them that would settle it. Claude Haiku 5.5's strong Japanese performance is not something the vendor advertises as a headline claim, and Sakana's JFBench numbers are its own instrument. A general frontier-tier small model from a vendor with strong multilingual capability, evaluated against a Kimi-based model post-trained specifically for Japanese, is a comparison nobody has actually run and published.

What can be said is structural rather than empirical. Namazu's entire commercial identity is the claim that a general model is not good enough for Japanese work — that country-specific language, cultural inference and domestic context are a distinct capability that justifies a specialised model at a nine-fold price premium. If you believe that claim, Namazu is the answer and the price is what the answer costs. If you do not, a general model at $0.10 per million tokens is what you use and the question is whether you can measure the difference on your own traffic.

The measurable version of the claim is the one Sakana provides: FairPoliticsQA going from 34.10% to 56.30% over the base model it was tuned from. That is a real improvement on a real benchmark, but it is a comparison against Kimi K2.6, not against Claude Haiku 5.5 — and it tells you that post-training for Japan-specific political context moves the needle on that task, which is a narrower and more defensible statement than "better at Japanese".

A generated two-column scoreboard titled 'Claude Haiku 5.5 vs Sakana Namazu - the scoreboard', with six rows carrying both sides. Input price: $0.10 / 1M against $0.95 / 1M. Output price: $0.50 / 1M against $4.00 / 1M. Cached input: $0.01 to $0.05 / 1M against $0.15 / 1M. Focus: general purpose against Japanese language and context. Lineage: proprietary against Kimi K2.6 post-trained. Independent index: 43 against no entry. The footer reads 'Claude Haiku 5.5 index per Artificial Analysis; Namazu figures per vendor.' The OrcaRouter logo is composited in the bottom-right corner.

What the cost gap looks like at volume

Run a modest pipeline at 10 million input and 2 million output tokens a day through both.

• Input cost, per day — $1.00 on Claude Haiku 5.5, $9.50 on Sakana Namazu.

• Output cost, per day — $1.00 on Claude Haiku 5.5, $8.00 on Sakana Namazu.

• Daily total — $2.00 against $17.50, roughly $60 a month against $525.

That $525 is before a single web search call or hour of code execution. Add 500 search calls a day and Namazu costs another $3.50 daily, taking it to $21.00 against $2.00 — a ratio of better than ten to one that no amount of token efficiency closes.

Whether that is a lot depends entirely on what the alternative is. If your choice is Namazu or a general model that produces mediocre Japanese output your reviewers have to fix by hand, the $525 is buying back reviewer hours and the comparison is not token against token, it is token against salary. If your Japanese traffic is routine and already good enough on a general model, the premium buys nothing measurable and the same dollars buy ten times the volume elsewhere.

The third option is the interesting one for a Japanese-language product: use the specialist where the specialist claim holds — translation, domestic context, regulatory and political text — and a general model for the volume work around it. That is not a compromise. It is the architecture the pricing actually implies, because Namazu's premium is paid per token and there is no reason to pay it on classification.

A headless capture of Anthropic's Claude Platform model documentation showing the Models overview comparison table with the Claude Haiku 5.5 column alongside Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5.5, including the 'From $0.10 / input MTok' and 'From $0.50 / output MTok' pricing rows, the 1M-token context window entry and the 'Fastest' comparative latency listing for Claude Haiku 5.5.

Neither model is on our catalogue

We should be plain about this, because it is the kind of comparison where an availability paragraph is easy to smuggle: OrcaRouter serves neither Claude Haiku 5.5 nor Sakana Namazu. Namazu is available through Sakana's own API, and Claude Haiku 5.5 through the vendor's and the major clouds. Neither fact changes the comparison, which rests on published rates, published capabilities and published restrictions.

The role one endpoint plays in a case like this is narrower and more honest than a plug. It is the general-model leg of a hybrid — the place you would route the non-Japanese classification and extraction traffic that does not need the specialist, without standing up a second vendor relationship for it. One API for 200-plus models, provider list rates passed through at 0% markup, automatic failover across providers, and the routing DSL when the language of the request should decide the route rather than the application. If you are running a Japanese product alongside an English one, that is the seam you would otherwise build by hand.

Which one, for whom

Choose Sakana Namazu when Japanese is the product rather than a locale — when your reviewers are correcting output, when the domain is domestic law, medicine, politics or customer service, when a nine-fold token premium is smaller than the cost of the corrections it removes, and when your users are outside the EU, EEA, UK and Switzerland or you have counsel who have cleared the data-handling question.

Choose Claude Haiku 5.5 when the work is general, the volume is high, and you want an independently measured score rather than a vendor benchmark — when $0.10 and $0.50 per million tokens is doing the arithmetic for you, when you need the one-million-token window on a long document, or when the answer to "which is better at Japanese" needs to come from your own evaluation rather than either vendor's.

The two are not really rivals. One is a general-purpose model priced to be used constantly; the other is a specialist priced to be used deliberately. The mistake is buying the specialist for work the generalist already does well, and the opposite mistake — assuming a general model handles a specific country's language and context as well as something trained for it — is the one Sakana's entire business is built on people making.

A headless capture of the OrcaRouter models catalogue page reading '207 models, 16 providers, one API key, one bill', with the filter sidebar showing input-modality counts, a pricing filter and a populated model list, and the panel that reads 'How to call any model' with an OpenAI-compatible curl sample pointing at https://api.orcarouter.ai/v1/chat/completions.