
Sakana Namazu API: Japan's Own LLM at the Base Model's Price
- metaNEWMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenNEWQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekNEWDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxNEWMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 2143 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
- grokxAI: Grok 4.52026-07-0856Intelligence72Coding
- tencentTencent: Hy32026-07-0642Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3055Intelligence72Coding
- klingKling: Kling 3.0 Turbo2026-06-1757Intelligence52Coding57Math
Sakana Namazu, Sakana AI's Japanese-business LLM API built on Moonshot's open-weight Kimi K2.6, went live on August 2, 2026 — and the number that matters most is that it costs exactly what its base model already costs. Moonshot charges $0.95 per million input tokens and $4.00 per million output tokens for Kimi K2.6 on its own API. Sakana charges the same $0.95 / $4.00 for Sakana Namazu, a model it post-trained in-house for Japanese language, keigo, and business context. The whole post-training premium is zero dollars.
What actually shipped
The API is OpenAI-compatible: keep your existing client, swap the base_url, set the model to sakana-namazu, and it just works — including function calling. What most distinguishes it from a plain Kimi K2.6 endpoint is that web search and code execution are built in as first-class tools, and the model is tuned to use them in an autonomous loop. Sakana demonstrated a weekly market-research agent that plans, searches, cross-checks, and writes a report without human intervention; a customer-support workflow that aggregates order data as it answers tickets; and a fish-school light show that reads aquarium screenshots, decides the next motif, and executes it.
Specs as published:
• Base model — Moonshot AI's Kimi K2.6, a 1T-parameter / 32B-active open-weight MoE, post-trained in-house on Sakana's own Japanese data
• Interface — OpenAI-compatible; model name sakana-namazu; change base_url and API key, no code rewrite
• Price — $0.95 per 1M input tokens, $4.00 per 1M output tokens, $0.15 per 1M cached input; thinking tokens billed at output rate
• Tools — web search at $7.00 per 1,000 calls, code execution at $0.12 per hour (session retained), billed on top of tokens
• Modality — text and image input, text output
• Context window — not disclosed by Sakana
• Availability — live via Sakana's API console; not available in the EU/EEA, UK, or Switzerland while GDPR work continues
Note the honesty gap: every quality claim here is vendor-reported. Sakana has not disclosed a context window, has not published latency or throughput numbers, and no independent lab has scored Sakana Namazu yet.

The FairPoliticsQA jump is the real story
The benchmark Sakana leads with is FairPoliticsQA, an internal evaluation of answer neutrality on politically and historically loaded questions — precisely the topics where a model trained largely on one country's internet tends to absorb that country's framing. Sakana reports Sakana Namazu rising from 34.10% to 56.30% against the base Kimi K2.6. It is a striking number and an honest one to interrogate: the evaluation is Sakana's own, its construction is not fully public, and a jump on a neutrality benchmark says as much about the base model's starting bias as it does about the fix. The direction, though, is the whole point of the product — a Japan-sovereign LLM that does not answer like a Chinese or American one.
Elsewhere the picture is "held serve plus a little": JFBench (Japanese instruction-following) up 1.5% versus the base, an internal Japanese-to-English translation eval improved on honorific-heavy text, and AIME26, MMLU-Pro, and LiveCodeBench v6 roughly unchanged. If the base model was already strong, the tuned model staying strong while improving Japanese-context behavior is exactly the trade you want from a post-training effort.

The price is the base model's price
Pricing is where Sakana Namazu is either shrewd or desperate, and it is worth taking both readings seriously. Charging the base model's list price means a developer comparing "call Kimi K2.6 directly" against "call the Japanese-tuned Sakana Namazu" sees no premium for the tuning — which is the entire adoption strategy in one number. It also makes the value obvious for Japanese-business workloads: the keigo, the jargon, the refusal-reduction, the built-in web search and code execution arrive without a per-token surcharge.
The tool fees are the part to watch, because they are where an agentic loop actually spends money. Web search at $7.00 per 1,000 calls and code execution at $0.12 per hour sound small until a weekly research agent runs a few hundred searches — the call count, not the token price, is what scales. A routing gateway that passes provider prices through without markup keeps exactly these numbers honest: no platform layer inflating the tool cost, and a vendor price cut live on your side the same day it is posted. (Sakana Namazu itself is not on OrcaRouter's catalogue yet — it runs on Sakana's own console — but the pass-through principle is why the comparison to the base model stays clean.)

The things Sakana did not publish
Three gaps matter more than any benchmark. First, no context window — Sakana's announcement and product page omit it entirely, which is unusual for an API marketed for agentic use. Second, no independent evaluation; Kimi K2.6, the base, is one of the best-measured open-weight models in the world — Artificial Analysis ranks it at the top of the open-weight field — but Sakana Namazu's deltas are entirely self-reported. Third, data handling: by default your inputs may be used for training, with an opt-out in the console, and processing is not guaranteed to stay in Japan — worth knowing before you send customer data through it.
Who should try it, who should wait
Try it if you ship Japanese-language products or internal tools and the missing piece is a model that answers like a Japanese business counterpart — keigo intact, refusals reduced, agentic tools built in — and you can live without independent benchmarks for a while. Wait if your deployment is in the EU, UK, or Switzerland (it is not available there), if your compliance regime requires guaranteed Japan-only data handling, or if you need a context window you can plan a retrieval design around.
What to watch in the next few weeks: an independent scorecard for Sakana Namazu on a public leaderboard, the disclosure of a context window, and whether the EU restriction lifts. Until those land, treat Sakana Namazu as what it is — a serious, well-priced, sovereign-Japan bet that is currently only as good as its maker's word.
