Title card for Sakana Fugu Max, the orchestration model released September 11, 2026, showing four stat cards: input $2 per 1M tokens, output $6 per 1M tokens, best overall on 6 benchmarks, and a pool of open plus specialized models including NVIDIA Nemotron, with a footer noting all figures are vendor-reported.
Guides & Insights

Sakana Fugu Max: The $2/$6 Orchestrator That Picks the Cheapest Model That Works

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Most models compete on how much they can do. Sakana Fugu Max competes on how little it can get away with doing. Sakana AI released Sakana Fugu Max on September 11, 2026, alongside Sakana Fugu Ultra v2 — two tunings of a single orchestration architecture pointed at opposite ends of the cost-performance curve. Fugu Max does not answer your request itself. It reads the task, picks the cheapest model in its pool that can plausibly handle it, routes there, and returns one answer. Sakana prices it at $2 per million input tokens and $6 per million output tokens.

The interesting part is what that price is measured against. Sakana claims Fugu Max's output rate undercuts Claude Sonnet 5, GPT-5.6 Terra and Kimi K3 by 40–60%. That claim is checkable, and unusually for a launch-day benchmark, the arithmetic holds up: against the current list prices of the three models Sakana names, $6 per million output lands almost exactly in the middle of the stated band. What is much harder to check is the performance side, because Sakana published Fugu Max's results as scatter plots rather than numbers.

What Fugu Max actually is

Fugu Max is not a checkpoint. There is no weight file, no parameter count, and nothing to download. Sakana describes it as "an architecture, not a single model" — the same core orchestration engine as Fugu Ultra v2, tuned for a different question. Ultra v2 asks what the highest capability available is. Fugu Max asks what the best output is at the lowest cost.

That distinction matters operationally. You cannot fine-tune Fugu Max, self-host it, or inspect why it chose one model over another. You also cannot see the full roster it routes across — Sakana says the pool expands to "our largest pool of open and specialized models to date" and explicitly names NVIDIA's Nemotron family, added through the NVIDIA partnership Sakana announced in August. The rest of the pool is unnamed, and Sakana does not publish a per-request routing trace.

What you can see is the shape of the thing. The vendor's own diagram shows Fugu Max at the top of a fan-out: one orchestrator, a row of interchangeable model slots behind it, and Sakana Namazu and Fugu Max itself among the targets. The pool is described as swappable by design, and the stated motivation is as much political as economic — Sakana frames orchestration as protection against vendor lock-in, API revocations and export controls, on the argument that a system which can swap its suppliers cannot be cut off by any one of them.

Sakana Fugu itself is not new. It reached general availability on June 22, 2026, and the vendor's own timeline runs April (beta), June (GA plus Fugu Ultra v1), July (Fugu-Cyber and a Claude Code interface), August (Sakana Chat plus the NVIDIA partnership), and now September. Fugu Max is the fifth monthly step, not a debut — and anyone already running Fugu upgrades to it with a one-line parameter change against the same OpenAI-compatible endpoint, no migration.

The pricing claim survives contact with the list prices

Sakana's $2 input / $6 output figure is stated plainly on the launch page. Secondary coverage adds a cached-input rate of $0.25 per million tokens, which the launch page itself does not print — treat that one as reported rather than confirmed. For comparison, Fugu Ultra v2 is $5 input, $30 output, with $0.50 cached input.

Now the 40–60% claim. The three named comparators currently sit at:

• GPT-5.6 Terra — $2 input / $12 output, after Open​AI's July 30 cut took output down from $15.

• Claude Sonnet 5 — $10 output on the launch promotional rate, $15 at the standard rate. Sources disagree on whether the scheduled September increase was scrapped or merely delayed, which is why the range is wide.

• Kimi K3 — $3 input / $15 output, with cached input at $0.30.

Against those, $6 output is 50% below Terra, 40% below Sonnet 5's promotional rate, 60% below Sonnet 5's standard rate, and 60% below Kimi K3. The claim checks out, and it checks out against published list prices rather than an internal cost model — which is not something you can say about every launch-day percentage.

One number in the same section does not survive the same treatment. Sakana also says Fugu Max delivers "performance within striking distance of elite models at two to six times lower cost." Against the three models it names for the 40–60% figure, the raw output-price gap is closer to 1.7× to 2.5×. The wider multiplier is presumably measured across the full frontier — including models that cost far more than $12 — rather than against the three in the sentence. It is a defensible claim about the Pareto curve and an indefensible one about the named rivals, and the launch page does not distinguish between the two.

This is where a routing layer earns its keep on the pricing side. OrcaRouter passes provider list prices through with no markup, so when a vendor moves list price the number you see moves that day — relevant here because Fugu Max's entire premise is that a cheaper model exists somewhere in the pool and you should reach it without thinking about it. We do not host Fugu Max; it runs on Sakana's own API. But the pass-through principle is the same one that makes a $6 output rate worth checking against a $12 incumbent rather than trusting either number on faith.

What Sakana says it beats, and what it won't show you

Screenshot of the Sakana AI Fugu Max launch page showing ten Pareto frontier scatter charts plotting benchmark score against output price in US dollars per million tokens, covering Terminal Bench 2.1, HLE, DeepSWE, CharXiv Reasoning, Chartography, GDP.pdf, GPQA-D, AutomationBench, AA-LCR and SWEFish, with Fugu Max marked as a red point at the top-left of each plot and no numeric scores printed.

The vendor's benchmark section makes three specific claims for Fugu Max: best overall score on six benchmarks — Terminal Bench 2.1, GPQA-D, AA-LCR, GDP.pdf, AutomationBench and SWEFish — an expansion of the cost-performance Pareto frontier on seven of ten benchmarks, and performance "within striking distance of elite models" at a fraction of the token spend.

Three things about that evidence are worth holding on to before you quote any of it.

The first is that Sakana published no numbers. The Fugu Max results appear as ten scatter plots — score on the vertical axis, output price in dollars per million tokens on the horizontal — with Fugu Max drawn as a red point to the northwest of the grey field. You can see that it sits up and to the left. You cannot read a score off it. Terminal Bench 2.1, GPQA-D and AA-LCR all have public leaderboards where a number would be directly comparable, and none was given.

The second is that one of the six benchmarks is Sakana's own. SWEFish is described on the launch page as "our internal benchmark reflecting Sakana AI's own coding challenges and use-cases." That is a legitimate way to measure fitness for your own product, and it is not a way to compare against a rival — a score on a benchmark the vendor designed, on tasks the vendor chose, is not evidence about anyone else's model.

The third is that the strongest numbers on the page belong to the other model. Fugu Ultra v2 gets figures where Fugu Max gets charts: Chartography 48.3 against Opus 5 at 27.3 and Fable 5 at 29.5, DeepSWE 74.3, best or joint-best on five of eight benchmarks and top-two on seven of eight. Ultra v2 also carries a stated training cutoff of 2026-08-28 — a little over two weeks before launch — and, pointedly, an explicit note that Fable 5, Fable 5.1 and GPT-6 Astra are not in its pool. Sakana is claiming it gets there without renting those specific models, which is the whole sovereignty argument in one footnote.

None of this makes the Fugu Max claims false. It makes them unverified, and it makes the six-benchmark headline safe to repeat only with the label attached. Independent evaluation of a brand-new orchestration endpoint is rare, and nothing has been published yet.

Fugu Max vs Fugu Ultra v2: which one you actually want

A two-column scoreboard comparing Sakana Fugu Max and Sakana Fugu Ultra v2 across six shared dimensions: price ($2 in / $6 out versus $5 in / $30 out), cached input ($0.25 reported versus $0.50), objective (lowest cost per task versus highest capability), benchmark evidence (charts with no numbers versus Chartography 48.3 and DeepSWE 74.3), whether Fable 5 or GPT-6 Astra are in the pool (not stated versus explicitly excluded), and best fit (high-volume simple traffic versus agentic long-horizon work).

These are not competing products, and the choice between them is not close once you see the numbers side by side. Fugu Max is the cheap one: $2 input, $6 output, built to route away from expensive models whenever a cheaper one can do the job. Fugu Ultra v2 is the strong one: $5 input, $30 output, $0.50 cached, built to route toward capability on complex multi-step work even when that costs more.

The practical split is by task shape, not by budget. If your traffic is mostly lookups, extraction, classification, short generations and straightforward tool calls, Fugu Max's whole design is to notice that and spend as little as possible — and Sakana's claim that it expands the frontier on seven of ten benchmarks is at least directionally about that. If your traffic includes long agentic runs, multi-file refactors or research tasks that fail expensively when a step goes wrong, the $30 output rate on Ultra v2 is buying something real: the top-two placement on seven of eight benchmarks is the part of the launch page that most resembles a capability claim rather than a cost one.

Both are available on the same OpenAI-compatible API through Sakana's console, and the upgrade path from existing Fugu is a parameter change. Two availability caveats apply. Fugu has been reported as unavailable in the EU and EEA pending GDPR work, which if it still holds limits where either model can be deployed. And both are closed systems: you are buying an endpoint, and the vendor chooses which models sit behind it — including, for Fugu Max, a roster it does not fully disclose.

The catch: an orchestrator that still rents its frontier

The sharpest criticism of Fugu is the one Sakana invites by framing it as sovereignty infrastructure. An orchestration layer that routes across third-party APIs does not stop those APIs from being revoked. It converts a single point of failure into a diversified one, which is a real improvement, but the underlying models are still someone else's, still reachable by the same export controls and still subject to the same sudden service cuts. The August NVIDIA partnership and the Nemotron additions are the most concrete answer to this so far — open-weight models in the pool are models nobody can switch off — but Sakana does not say how much of Fugu Max's traffic actually lands on them.

There is a second, quieter trade. Routing decisions add latency and remove determinism. A request that lands on a small model today might land on a different one tomorrow if the pool or the cost weights move, which makes behaviour harder to regression-test and harder to explain to anyone auditing the system. Sakana's answer is that the orchestration is internal and the API behaves like a single model. That is true at the interface and not true underneath, and teams with hard reproducibility requirements should test that carefully before it becomes load-bearing.

If you want the cost win without making the pool your single dependency, the same pattern is available to compose yourself. Our routing DSL lets you define a panel of models behind one endpoint and decide which one handles which class of request, and model fusion runs several models on the same prompt and reconciles the answers where agreement matters more than price. Failover across providers means that when a supplier degrades — or disappears — traffic moves to a model you already trust without a code change. That is the more conservative version of the bet Sakana is asking you to make wholesale.

Screenshot of the OrcaRouter models catalogue showing 196 models across 15 providers under one API key and one bill, an OpenAI-compatible chat completions endpoint, and per-model token rates including GPT-6 Astra at $10 input and $50 output, Claude Fable 5.1 at $10 and $50, Qwen3.8 Max at $2 and $6, and Gemini 3.8 Flash at $0.750 and $3.75.

It is worth being precise about what we do and do not offer here. Sakana Fugu Max is not in our catalogue — it is not a routable model on our side, and it runs on Sakana's own API and console. Our catalogue is 196 models across 15 providers on one key and one bill, with the per-token rates printed on each card and no markup on top of them. The relevant overlap with Fugu Max is the idea, not the inventory: both are arguments that the right answer to "which model should handle this" is usually not the biggest one.

Who should move, and who should wait

Fugu Max is one of the more interesting releases of the month precisely because it is not a model. If you are already on Sakana Fugu, the upgrade is a one-line change at a lower price and there is no reason to deliberate. If you are running high-volume, low-complexity traffic on a mid-tier model and paying $10–$15 per million output tokens for it, the arithmetic Sakana is offering is worth a benchmark of your own — take a week of real requests, run them through Fugu Max, and compare quality against your current model rather than against the scatter plots.

If you need published benchmark numbers before you commit, wait. There is nothing independent on Fugu Max yet, the vendor's own evidence is a chart rather than a table, and one of the six headline benchmarks is Sakana's. That is not a reason to dismiss it — it is a reason to test it on your own traffic, which is the only benchmark that was ever going to predict your results. And if your workload is agentic, long-horizon and expensive to get wrong, the model to look at in this pair is Fugu Ultra v2, not Fugu Max: it is the one with numbers attached.

Our routing DSL lets you define a panel of models behind one endpoint and decide which one handles which class of request, and model fusion runs several models on the same prompt and reconciles the answers where agreement matters more than price.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily