
Sakana Fugu Ultra v2, explained: the pool got smaller and the score went up
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Fugu Ultra v2 is a strange kind of upgrade: Sakana AI shipped it on September 11, 2026 with a model pool that no longer contains Claude Fable 5, Claude Fable 5.1 or GPT-6 Astra — three of the strongest systems now shipping — and still claims to beat all three. The Tokyo lab's new quality-tier flagship, released the same day as a cheaper sibling called Fugu Max, is best or joint-best on five of eight benchmarks in Sakana's own evaluation, including 48.3 on Chartography against Claude Opus 5 at 27.3 and Claude Fable 5 at 29.5, and 74.3 on DeepSWE. None of those numbers has been independently reproduced, and there is still no Artificial Analysis page for any Fugu model. That gap is the story: what you are being asked to buy is a coordination layer whose entire pitch is that it does not need the best models to produce the best answers.
Sakana AI was founded in 2023 by Llion Jones, a co-author of the Transformer paper, and David Ha. The Fugu line launched in June 2026 as a multi-agent system delivered as a single model: you call one OpenAI-compatible endpoint, and behind it a trained coordinator decomposes the request, spawns agents, and synthesises the result. The research is real and published — TRINITY (arXiv 2512.04695) evolved a lightweight coordinator that assigns Thinker, Worker and Verifier roles; Conductor (arXiv 2512.04388) learned natural-language coordination between agents; the Fugu technical report sits at arXiv 2606.21228. What Sakana has never published is the thing everyone actually wants: the identity of the models in the pool, and the per-task routing decisions.
What actually changed from the June Fugu Ultra
Version 2.0 is a point-refresh roughly eleven weeks after the original, not a new architecture. The vendor's own framing is that Fugu Ultra v2 and Fugu Max are "the same core orchestration architecture" tuned for two different missions: Max pushes the cost-performance frontier down, Ultra pushes peak capability up. The model ID is fugu-ultra-v2.0, and Sakana says migrating off an earlier Fugu is a single-line parameter change with no SDK migration — which is the most consumer-friendly thing about the release.
The substantive changes are the two bookends. First, the training cutoff moved to 20260828, so the coordinator has been re-tuned against the current generation of frontier models. Second, and far more interesting, the pool composition changed in a way Sakana chose to advertise: Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra are explicitly excluded. Sakana's argument is that this proves the scores do not depend on a single proprietary frontier model, which matters if your concern is vendor lock-in, API revocations or a model disappearing behind export controls.
There is a second-order consequence Sakana does not dwell on. If the pool excludes the three strongest available models, then the benchmark comparisons against Fable 5 and Fable 5.1 are being made against systems the orchestrator could not call even if it wanted to. The comparison is legitimate — it is a claim about the architecture, not about access — but it is not a like-for-like test of who has the better model. It is a test of whether coordination beats raw capability. That is a genuinely interesting question. It is simply not the question the scoreboard implies at a glance.

The numbers, and who produced them
Every figure in this section is vendor-reported. Sakana has released no evaluation harness, no per-task score grid, and no reproduction recipe, and the orchestration design makes independent replication harder rather than easier: reproducing a score requires access to every model in the pool at the same versions with the same topology, and by design the topology varies per request.
• Chartography (visual reasoning over charts and structured documents) — Fugu Ultra v2 48.3, Claude Opus 5 27.3, Claude Fable 5 29.5 — vendor-reported
• DeepSWE (real-world software engineering) — Fugu Ultra v2 74.3 — vendor-reported, said to beat models costing three to five times more per token
• Best or joint-best placement — five of eight benchmarks: GDP.pdf, Chartography, SWEFish, DeepSWE and Toolathon — vendor-reported
• Top-two placement — seven of eight benchmarks — vendor-reported
• Prior Fugu Ultra (June 2026, for scale) — LiveCodeBench 93.2, GPQA-Diamond 95.5, TerminalBench 82.1, SWE-Bench Pro 73.7 — vendor-reported
The Chartography row deserves the emphasis it is getting, because it is the one category where the gap to a named, current flagship is not marginal. A 21-point gap over Claude Opus 5 on a visual-reasoning benchmark is the kind of number that usually shows up on an independent leaderboard within days. As of this writing, none has appeared. Treat the row as a hypothesis with a dollar figure attached, not as a settled result.
Pricing: unchanged from June, and flatly expensive on output
Fugu Ultra v2 costs $5 per million input tokens, $30 per million output tokens, and $0.50 per million cached input. Above 272,000 tokens of context those become $10, $45 and $1.00 respectively. Fugu Max is a different shape entirely: $2 input, $6 output, $0.25 cached, flat across context length, plus $0.007 per web search or fetch call.
Two things about that pricing table are worth reading carefully. The first is that output is six times input — an aggressive ratio that prices the model's verbosity, and orchestration is a verbose business. A coordinator that spawns agents and synthesises their answers emits a lot of tokens, and you pay for all of them at $30 per million. Sakana's efficiency claim is about the underlying pool, not about how much text the system produces on your behalf, and those are different numbers. Budget against observed token counts, not against the headline rate.
The second is the 272K cliff. Doubling the per-token rate above a threshold is a legitimate way to price long-context work, but it means the effective cost of a long-context agent run is not the number on the pricing page. If your workload sits near the boundary, the arithmetic changes materially on either side of it.
Fugu's subscription tiers — $20 Standard, $100 Pro, $200 Max per month, with 10x and 20x allowances — all include access to Fugu, Fugu Ultra and Fugu Max. Sakana also prices the base Fugu model by the top tier present in the pool for a given request, and states clearly that adding more agents does not multiply the bill. That is a real difference from the fan-out cost model you would get by calling several frontier models yourself.
What Sakana will not tell you
Ask which models answered your question and the FAQ is direct: "this routing information is not exposed by design." Ultra and Max pools are fixed, so you cannot opt a vendor out; only the base Fugu model supports per-model opt-outs in console settings. Enterprise customers can negotiate custom configurations.
That opacity is the crux of the criticism the Fugu line has attracted since June, and it is a fair criticism rather than a hostile one. When an orchestrator scores well by combining other labs' models, the score belongs to the system — which is a real and defensible thing to sell — but it does not transfer to any single model, and it cannot be audited. Reviewers have made the point bluntly: Fugu did not beat the flagship, it hired the flagship. On verifiable tasks with a cheap checker, like a coding benchmark with a test suite, a committee really can outperform its best member. On tasks without one, a panel sampling several frontier models and reporting the best result is partly reporting luck. Sakana's own category choices — Chartography, DeepSWE, Toolathon — lean toward the verifiable end, which is the right place to make the claim and also the reason the claim cannot be generalised for free.
Independent testers have also flagged latency variance as the practical cost of orchestration: on hard tasks the coordinator deliberates, and on easy tasks that deliberation is pure overhead. One researcher's shader task reportedly took around thirty minutes with quality rated only acceptable.

Where you can actually get it
Fugu Ultra v2 is live now through Sakana's own OpenAI-compatible API — point an existing client at the endpoint with an API key and change one line — and through several third-party platforms. OrcaRouter does not carry the Fugu models; if you want to call Fugu Ultra v2, the vendor's own API is the direct route.
What is worth noting is the alternative Sakana is implicitly competing with. The pool that makes Fugu Ultra v2 work is assembled from models that are individually callable today, and the opponents it is measured against are the ones teams already run. Claude Opus 5, DeepSeek V4 Pro and Gemini 3.1 Pro are all on OrcaRouter at their providers' list prices with 0% markup, which means a price cut from any of those vendors is live on our side the same day it is announced. If the reason you are considering an orchestrator is resilience — not wanting a production path to depend on one vendor's uptime or one vendor's policy — then a routing layer with automatic failover across providers solves part of that problem at the model level, and Fugu Ultra v2 solves a different part at the reasoning level. One API for 200+ models with failover is not a substitute for a trained coordinator, and a trained coordinator is not a substitute for not being single-vendor-dependent. Some teams will want both.
The compliance catch
Fugu is not available in the European Union or the European Economic Area. The vendor's wording is explicit: not yet available in the EU/EEA while it works toward compliance with GDPR and EU-specific regulations, and elsewhere access may be limited by network conditions or local rules. If your organisation has EU entities in scope, this removes the Fugu line from consideration regardless of benchmark performance, which is worth knowing before the evaluation rather than after it.

What to watch
Three things would move Fugu Ultra v2 from an interesting claim to a load-bearing choice. An independent evaluation — an Artificial Analysis entry, a LMArena placement, anything with a harness behind it — would settle the Chartography question in a week. Reproduced numbers from a third party on the verifiable coding benchmarks would confirm the system-level gain that the architecture predicts. And a disclosed pool, or even a partial one, would let buyers reason about exactly which single points of failure they are accepting when they route production traffic through a coordinator they cannot inspect.
Until then the honest position is this. Fugu Ultra v2 is the most serious attempt yet at selling orchestration as a product rather than a research demo, from a lab with published work behind it, at a price that makes the orchestration premium explicit rather than hidden. If your workload is long-horizon, verifiable, and confined to a region where Sakana can serve you, it is worth a scoped pilot with token accounting turned on. If it is not, the models Fugu Ultra v2 is measured against are one API call away, at list price, with the receipts to prove what they can do.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
