A hero title card for Fugu Ultra v2 titled 'Fugu Ultra v2' with the subtitle 'The pool got smaller and the score went up', pill badges reading 'Chartography 48.3', 'DeepSWE 74.3' and '$5 / $30 per 1M', a footer line 'Sakana AI - released September 11, 2026', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Sakana Fugu Ultra v2, explained: the pool got smaller and the score went up

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Fugu Ultra v2 is a strange kind of upgrade: Sakana AI shipped it on September 11, 2026 with a model pool that no longer contains Claude Fable 5, Claude Fable 5.1 or GPT-6 Astra — three of the strongest systems now shipping — and still claims to beat all three. The Tokyo lab's new quality-tier flagship, released the same day as a cheaper sibling called Fugu Max, is best or joint-best on five of eight benchmarks in Sakana's own evaluation, including 48.3 on Chartography against Claude Opus 5 at 27.3 and Claude Fable 5 at 29.5, and 74.3 on DeepSWE. None of those numbers has been independently reproduced, and there is still no Artificial Analysis page for any Fugu model. That gap is the story: what you are being asked to buy is a coordination layer whose entire pitch is that it does not need the best models to produce the best answers.

Sakana AI was founded in 2023 by Llion Jones, a co-author of the Transformer paper, and David Ha. The Fugu line launched in June 2026 as a multi-agent system delivered as a single model: you call one OpenAI-compatible endpoint, and behind it a trained coordinator decomposes the request, spawns agents, and synthesises the result. The research is real and published — TRINITY (arXiv 2512.04695) evolved a lightweight coordinator that assigns Thinker, Worker and Verifier roles; Conductor (arXiv 2512.04388) learned natural-language coordination between agents; the Fugu technical report sits at arXiv 2606.21228. What Sakana has never published is the thing everyone actually wants: the identity of the models in the pool, and the per-task routing decisions.

What actually changed from the June Fugu Ultra

Version 2.0 is a point-refresh roughly eleven weeks after the original, not a new architecture. The vendor's own framing is that Fugu Ultra v2 and Fugu Max are "the same core orchestration architecture" tuned for two different missions: Max pushes the cost-performance frontier down, Ultra pushes peak capability up. The model ID is fugu-ultra-v2.0, and Sakana says migrating off an earlier Fugu is a single-line parameter change with no SDK migration — which is the most consumer-friendly thing about the release.

The substantive changes are the two bookends. First, the training cutoff moved to 20260828, so the coordinator has been re-tuned against the current generation of frontier models. Second, and far more interesting, the pool composition changed in a way Sakana chose to advertise: Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra are explicitly excluded. Sakana's argument is that this proves the scores do not depend on a single proprietary frontier model, which matters if your concern is vendor lock-in, API revocations or a model disappearing behind export controls.

There is a second-order consequence Sakana does not dwell on. If the pool excludes the three strongest available models, then the benchmark comparisons against Fable 5 and Fable 5.1 are being made against systems the orchestrator could not call even if it wanted to. The comparison is legitimate — it is a claim about the architecture, not about access — but it is not a like-for-like test of who has the better model. It is a test of whether coordination beats raw capability. That is a genuinely interesting question. It is simply not the question the scoreboard implies at a glance.

A single-column scoreboard for Fugu Ultra v2 titled 'Fugu Ultra v2 - the scoreboard' listing Best or joint-best 5 of 8 benchmarks, Chartography 48.3, DeepSWE 74.3, Agent pool excludes Fable 5, Fable 5.1, GPT-6 Astra, Price $5.00 / $30.00 per 1M, and Independent evaluation none published, with a footer 'All figures vendor-reported by Sakana AI, September 2026; no independent evaluation.' and the OrcaRouter logo in the bottom-right corner.

The numbers, and who produced them

Every figure in this section is vendor-reported. Sakana has released no evaluation harness, no per-task score grid, and no reproduction recipe, and the orchestration design makes independent replication harder rather than easier: reproducing a score requires access to every model in the pool at the same versions with the same topology, and by design the topology varies per request.

• Chartography (visual reasoning over charts and structured documents) — Fugu Ultra v2 48.3, Claude Opus 5 27.3, Claude Fable 5 29.5 — vendor-reported

• DeepSWE (real-world software engineering) — Fugu Ultra v2 74.3 — vendor-reported, said to beat models costing three to five times more per token

• Best or joint-best placement — five of eight benchmarks: GDP.pdf, Chartography, SWEFish, DeepSWE and Toolathon — vendor-reported

• Top-two placement — seven of eight benchmarks — vendor-reported

• Prior Fugu Ultra (June 2026, for scale) — LiveCodeBench 93.2, GPQA-Diamond 95.5, TerminalBench 82.1, SWE-Bench Pro 73.7 — vendor-reported

The Chartography row deserves the emphasis it is getting, because it is the one category where the gap to a named, current flagship is not marginal. A 21-point gap over Claude Opus 5 on a visual-reasoning benchmark is the kind of number that usually shows up on an independent leaderboard within days. As of this writing, none has appeared. Treat the row as a hypothesis with a dollar figure attached, not as a settled result.

Pricing: unchanged from June, and flatly expensive on output

Fugu Ultra v2 costs $5 per million input tokens, $30 per million output tokens, and $0.50 per million cached input. Above 272,000 tokens of context those become $10, $45 and $1.00 respectively. Fugu Max is a different shape entirely: $2 input, $6 output, $0.25 cached, flat across context length, plus $0.007 per web search or fetch call.

Two things about that pricing table are worth reading carefully. The first is that output is six times input — an aggressive ratio that prices the model's verbosity, and orchestration is a verbose business. A coordinator that spawns agents and synthesises their answers emits a lot of tokens, and you pay for all of them at $30 per million. Sakana's efficiency claim is about the underlying pool, not about how much text the system produces on your behalf, and those are different numbers. Budget against observed token counts, not against the headline rate.

The second is the 272K cliff. Doubling the per-token rate above a threshold is a legitimate way to price long-context work, but it means the effective cost of a long-context agent run is not the number on the pricing page. If your workload sits near the boundary, the arithmetic changes materially on either side of it.

Fugu's subscription tiers — $20 Standard, $100 Pro, $200 Max per month, with 10x and 20x allowances — all include access to Fugu, Fugu Ultra and Fugu Max. Sakana also prices the base Fugu model by the top tier present in the pool for a given request, and states clearly that adding more agents does not multiply the bill. That is a real difference from the fan-out cost model you would get by calling several frontier models yourself.

What Sakana will not tell you

Ask which models answered your question and the FAQ is direct: "this routing information is not exposed by design." Ultra and Max pools are fixed, so you cannot opt a vendor out; only the base Fugu model supports per-model opt-outs in console settings. Enterprise customers can negotiate custom configurations.

That opacity is the crux of the criticism the Fugu line has attracted since June, and it is a fair criticism rather than a hostile one. When an orchestrator scores well by combining other labs' models, the score belongs to the system — which is a real and defensible thing to sell — but it does not transfer to any single model, and it cannot be audited. Reviewers have made the point bluntly: Fugu did not beat the flagship, it hired the flagship. On verifiable tasks with a cheap checker, like a coding benchmark with a test suite, a committee really can outperform its best member. On tasks without one, a panel sampling several frontier models and reporting the best result is partly reporting luck. Sakana's own category choices — Chartography, DeepSWE, Toolathon — lean toward the verifiable end, which is the right place to make the claim and also the reason the claim cannot be generalised for free.

Independent testers have also flagged latency variance as the practical cost of orchestration: on hard tasks the coordinator deliberates, and on easy tasks that deliberation is pure overhead. One researcher's shader task reportedly took around thirty minutes with quality rated only acceptable.

A screenshot of the Sakana Fugu product page at sakana.ai/fugu (captured September 11, 2026, English UI), showing the 'Sakana Fugu' wordmark with the subtitle 'One Model to Command Them All', the description that Fugu dynamically orchestrates the world's best models behind a single API, the 'Start Using Sakana Fugu' and 'Contact Us' buttons, and the notice reading 'Not yet available in the EU/EEA while we work toward compliance with GDPR and EU-specific regulations.'

Where you can actually get it

Fugu Ultra v2 is live now through Sakana's own OpenAI-compatible API — point an existing client at the endpoint with an API key and change one line — and through several third-party platforms. OrcaRouter does not carry the Fugu models; if you want to call Fugu Ultra v2, the vendor's own API is the direct route.

What is worth noting is the alternative Sakana is implicitly competing with. The pool that makes Fugu Ultra v2 work is assembled from models that are individually callable today, and the opponents it is measured against are the ones teams already run. Claude Opus 5, DeepSeek V4 Pro and Gem​ini 3.1 Pro are all on OrcaRouter at their providers' list prices with 0% markup, which means a price cut from any of those vendors is live on our side the same day it is announced. If the reason you are considering an orchestrator is resilience — not wanting a production path to depend on one vendor's uptime or one vendor's policy — then a routing layer with automatic failover across providers solves part of that problem at the model level, and Fugu Ultra v2 solves a different part at the reasoning level. One API for 200+ models with failover is not a substitute for a trained coordinator, and a trained coordinator is not a substitute for not being single-vendor-dependent. Some teams will want both.

The compliance catch

Fugu is not available in the European Union or the European Economic Area. The vendor's wording is explicit: not yet available in the EU/EEA while it works toward compliance with GDPR and EU-specific regulations, and elsewhere access may be limited by network conditions or local rules. If your organisation has EU entities in scope, this removes the Fugu line from consideration regardless of benchmark performance, which is worth knowing before the evaluation rather than after it.

A screenshot of the OrcaRouter models catalogue page (captured September 11, 2026, English UI), showing the OrcaRouter navigation with Models, Leaderboard, Offers and Developers entries, the search field, and the 'Get API key' button over the browsable model catalogue.

What to watch

Three things would move Fugu Ultra v2 from an interesting claim to a load-bearing choice. An independent evaluation — an Artificial Analysis entry, a LMArena placement, anything with a harness behind it — would settle the Chartography question in a week. Reproduced numbers from a third party on the verifiable coding benchmarks would confirm the system-level gain that the architecture predicts. And a disclosed pool, or even a partial one, would let buyers reason about exactly which single points of failure they are accepting when they route production traffic through a coordinator they cannot inspect.

Until then the honest position is this. Fugu Ultra v2 is the most serious attempt yet at selling orchestration as a product rather than a research demo, from a lab with published work behind it, at a price that makes the orchestration premium explicit rather than hidden. If your workload is long-horizon, verifiable, and confined to a region where Sakana can serve you, it is worth a scoped pilot with token accounting turned on. If it is not, the models Fugu Ultra v2 is measured against are one API call away, at list price, with the receipts to prove what they can do.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily