
Fugu Ultra v2 vs Claude Opus 5: same input price, two different bets
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Fugu Ultra v2 and Claude Opus 5 cost exactly the same amount to ask a question and wildly different amounts to get an answer out of. Both list at $5.00 per million input tokens. Then Fugu Ultra v2 charges $30.00 per million output tokens against Claude Opus 5's $25.00 — and the gap between those two numbers is not a rounding difference, because the two systems emit very different quantities of text to reach an answer. Sakana AI shipped Fugu Ultra v2 on September 11, 2026 as an orchestration system that coordinates a pool of other labs' models behind one OpenAI-compatible endpoint; Anthropic's Claude Opus 5, live since late July 2026, is a single model. That difference decides most of this comparison before a benchmark is consulted, and Sakana's own scoreboard has an awkward row in it: the vendor reports 48.3 on Chartography against Claude Opus 5's 27.3, which is the largest margin in the release and also the number with the least outside evidence behind it.
The coincidence that shapes everything
Start with what the two products agree on, because it is more than the headline price.
• Input price — Fugu Ultra v2 $5.00 per 1M tokens vs Claude Opus 5 $5.00 per 1M tokens
• Output price — Fugu Ultra v2 $30.00 per 1M tokens vs Claude Opus 5 $25.00 per 1M tokens
• Cached input — Fugu Ultra v2 $0.50 per 1M above the base rate vs Claude Opus 5, which prices caching as a discount of up to 90% on cached input
• Context — Fugu Ultra v2 1M tokens, with rates doubling above 272K vs Claude Opus 5 1M tokens, flat
• Output ceiling — Fugu Ultra v2 not published as a single figure vs Claude Opus 5 128K tokens
• Availability — Fugu Ultra v2 not sold in the EU/EEA vs Claude Opus 5 generally available under standard API terms
• Evidence — Fugu Ultra v2 vendor-reported benchmarks, no independent evaluation yet vs Claude Opus 5 vendor benchmarks plus months of third-party leaderboard placements
That last row is where the matchup actually turns. At the same input price, you are choosing between a system whose scores you have to take on faith and a model whose scores other people have reproduced. The 272K cliff compounds it: past that threshold Fugu Ultra v2's input rate doubles to $10 and output to $45, while Claude Opus 5's rate does not move. For long-context agent runs — the exact workload Fugu Ultra v2 is built for — the price advantage of the cheaper headline disappears precisely where the system is supposed to shine.

What each one actually is
Claude Opus 5 is Anthropic's production flagship, and its design is legible from the outside. It runs adaptive thinking by default, deciding how long to reason before it answers, with an explicit effort dial from low to max when you want to force deeper deliberation and pay for it. It carries a 1M-token context window, a 128K output ceiling, vision, tool use and JSON mode. Anthropic reports 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro, more than doubling its predecessor's Frontier-Bench v0.1 result, and roughly 30.2% on ARC-AGI 3 against 7.8% for GPT-5.6 Sol — all vendor figures, and all of them measured on a single model you can pin by version. It also routes flagged requests to an older model rather than erroring, which is an operational detail that matters if you have compliance review in the loop.
Fugu Ultra v2 is not that, and the difference is not cosmetic. It is a coordinator — a small trained model, reported around 7B parameters, that reads your request, assembles an agent team, assigns Thinker, Worker and Verifier roles from Sakana's TRINITY work, and synthesises a final answer. The underlying models are not disclosed, the routing decisions are not exposed by design, and for the Ultra tier the pool is fixed: you cannot opt a vendor out. Sakana's own framing of version 2 is that the training cutoff moved to August 28, 2026 and the pool composition changed — specifically to exclude Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra.
That exclusion is the most revealing detail in the release. Fugu Ultra v2 is claiming benchmark wins over the three strongest models available while confirming it does not call any of them. Whatever the pool does contain, it is not the top of the market by Sakana's own accounting. The pitch is that coordination beats capability. That is a real thesis, and on verifiable tasks with a cheap checker it is a thesis with evidence behind it. It is not the same claim as "this system is smarter than Claude Opus 5," and the scoreboard is arranged so the two are easy to confuse.
The scoreboard Sakana chose
Every number below is from Sakana's own evaluation of Fugu Ultra v2. Claude Opus 5's figures are as measured by Sakana in the same comparison, except where noted. No independent party has reproduced the Fugu column, and as of this writing there is no Artificial Analysis entry for any Fugu model.
• Chartography — Fugu Ultra v2 48.3 vs Claude Opus 5 27.3 — vendor-reported, a 21-point gap
• DeepSWE — Fugu Ultra v2 74.3 vs no published Claude Opus 5 figure in the same table — vendor-reported
• Benchmark placement — Fugu Ultra v2 best or joint-best on five of eight, top-two on seven of eight — vendor-reported
• SWE-bench Verified — no Fugu Ultra v2 figure vs Claude Opus 5 96.0% — Anthropic-reported
• SWE-bench Pro — no Fugu Ultra v2 figure vs Claude Opus 5 79.2% — Anthropic-reported
• ARC-AGI 3 — no Fugu Ultra v2 figure vs Claude Opus 5 ~30.2% — Anthropic-reported
Read that as a shape rather than a ranking. Sakana published an eight-benchmark board where it wins five, and Claude Opus 5's strongest publicly documented results — the SWE-bench rows, ARC-AGI 3 — are simply not on it. That is not an accusation of bad faith; it is what a vendor scoreboard looks like, including every vendor's. But it means you cannot conclude from the release that Fugu Ultra v2 beats Claude Opus 5 at software engineering, because the two systems were not measured on the same rows often enough to say. On the one row where both appear and both numbers are public — Chartography — Fugu Ultra v2 claims a large win. On the broadest reasoning and coding rows, only one side has published.

Cost per task, not cost per token
An orchestrator's token bill is not its per-token rate. Fugu Ultra v2 spawns agents, each contributing reasoning and output tokens, and the coordinator itself produces synthesis text. Sakana's pricing FAQ makes a genuinely useful commitment here — you are charged one blended rate based on the top-tier model involved, and adding agents does not multiply the bill — which removes the worst-case fan-out multiplication that makes naive multi-agent setups expensive. But it does not remove the volume. The quantity of output tokens billed is still a function of how many agents ran and how much they wrote, and at $30 per million that volume is the variable you cannot predict from the pricing page.
Claude Opus 5's cost shape is the inverse. One call, one model, an effort dial you control, and a $25 output rate with batch processing available at 50% off for non-real-time work. You can forecast it. For a workload you run ten thousand times a day, forecastability is worth a lot more than a benchmark margin on a chart-reading task.
This is also where the routing question separates cleanly from the model question. Claude Opus 5 sits on OrcaRouter at Anthropic's list price with 0% markup — the provider's rate passed through, so a vendor price change is live on the same key the same day, with automatic failover across provider paths if one is rate-limited or down. Fugu Ultra v2 is not on our catalogue; it reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, and migrating from an earlier Fugu is a one-line parameter change. If what you actually want is Claude Opus 5's predictability with less single-provider exposure, the router is the answer and the orchestrator is not. If what you want is a system that tries several approaches to a hard problem without you writing the fan-out, Fugu Ultra v2 is the thing that does that — and you should pilot it with token accounting on.
Where Fugu Ultra v2 genuinely wins
Give the system its due. The Chartography result, if it holds, is the kind of win that single models find hard to fake: visual reasoning over structured documents rewards trying a reading, checking it against the document, and trying again — which is exactly what a Verifier role is for. The same logic applies to Toolathon and DeepSWE, where the answer can be tested rather than argued about. Sakana's published case studies lean the same way: a 123-experiment AutoResearch run over fourteen hours on one H100, a pure-Python Rubik's cube solver that finished all 300 scrambled cubes where two anonymised frontier baselines crashed entirely, a mechanical CAD iris that actually opens and closes. These are vendor-selected examples with anonymised baselines, so they prove less than they appear to. But the pattern is consistent, and it points at a real category: tasks where correctness is checkable and the hard part is persistence.
There is also a structural argument that has nothing to do with benchmarks. Claude Opus 5 is one vendor's model, subject to one vendor's pricing decisions, deprecation schedule and policy. Fugu Ultra v2 is designed to survive any single one of those changing, and Sakana is explicit that this is the point — vendor lock-in, API revocations, export controls. In a market where models have been withdrawn from regions with little notice, that is not a hypothetical. It is a hedge, and hedges cost money: $5 more per million output tokens is roughly what this hedge prices at.
The verdict
Claude Opus 5 wins the decision most teams are actually making. It has months of independent evaluation behind it, a 1M-token window that does not double in price partway through your document, a 128K output ceiling, a published price you can forecast, an effort dial that lets you buy reasoning only when the task deserves it, and third-party results on the exact benchmarks — SWE-bench Verified, SWE-bench Pro, ARC-AGI 3 — that agentic and coding teams care about most. Fugu Ultra v2 wins a narrower and more interesting question: whether a coordinator with a smaller pool can outperform a flagship on tasks where the answer can be checked. Sakana's own scoreboard says yes on five of eight rows, and the pool exclusion says the win came without the three best models in the world.
If you run verifiable, long-horizon, checkable work and you are outside the EU/EEA, Fugu Ultra v2 is worth a scoped pilot — measure output tokens per completed task, not per token, and set a ceiling before you start. If you need a production path today with evidence behind it and a bill you can forecast, Claude Opus 5 remains the better-evidenced call, and it is one API key away at Anthropic's list price with the provider's rate passed through untouched.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
