A hero title card for "Fugu Max vs Claude Opus 5" with the subtitle "Four times cheaper, and a number nobody published", three pill badges reading $6.00 vs $25.00 output", "Six wins, zero scores", "$2.00 vs $5.00 input, a footer line reading "Sakana AI vs Anthropic - September 2026", and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Fugu Max vs Claude Opus 5: four times cheaper, and a number nobody published

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Fugu Max costs $6.00 to produce a million tokens of output. Claude Opus 5 costs $25.00. Sakana AI shipped Fugu Max on September 11, 2026 as a coordinator that routes each task to the leanest model in its pool, and priced it aggressively enough that a four-fold output gap against Anth​ropic's flagship is the headline fact of the release. What the release does not contain is a single number describing how well Fugu Max actually performs. Sakana says it takes "best overall score" on six benchmarks and expands the cost-performance frontier on seven of ten — statements about placement on a chart rather than values on an axis. Claude Opus 5, by contrast, arrives with a published score for nearly everything it does. So the honest version of this comparison is not cheap-versus-expensive. It is measured-versus-asserted, and the price gap is what makes that trade worth arguing about.

The gap, and the comparison Sakana chose not to make

Sakana's release page justifies Fugu Max's output rate by comparing it to other models. The three it names are Claude Sonnet 5, GPT-5.6 Terra and Kimi K3, and the claim is that Fugu Max's output pricing runs 40 to 60 percent below them. Notice who is absent. Claude Opus 5 is not in that list, and the reason is arithmetic: at $6.00 against $25.00, Fugu Max is not 40 percent cheaper than Opus 5, it is 76 percent cheaper. Naming Opus 5 would have made the claim look implausible rather than competitive, so the comparison is drawn against the tier below the flagship.

That is not a criticism of the pricing, which is genuinely low. It is a note about what the surrounding sentence is doing. Sakana's own framing — performance "within striking distance of elite models at two to six times lower cost" — is calibrated to the models it chose to name. Against Claude Opus 5 the multiple is not two to six, it is north of four on output; whether Opus 5 is still an "elite model" by that sentence's lights is left unstated, which is a curious omission in a release whose entire pitch is cost-performance efficiency at the frontier.

• Input price — Fugu Max $2.00 per 1M tokens vs Claude Opus 5 $5.00 per 1M tokens

• Output price — Fugu Max $6.00 per 1M tokens vs Claude Opus 5 $25.00 per 1M tokens

• Cached input — Fugu Max $0.25 per 1M tokens vs Claude Opus 5 caching priced as a discount of up to 90% on cached input

• Benchmark scores — Fugu Max none published, six benchmark wins asserted without figures vs Claude Opus 5 Anthropic-reported 96.0% SWE-bench Verified, 79.2% SWE-bench Pro, roughly 30.2% on ARC-AGI 3

• Architecture — Fugu Max a trained coordinator over an undisclosed pool vs Claude Opus 5 a single model pinnable by version

• Context — Fugu Max not published vs Claude Opus 5 1M tokens, with a 128K output ceiling

• Independent evaluation — Fugu Max none exists for any Fugu model vs Claude Opus 5 months of third-party leaderboard placements

Six wins with no scores attached

The six benchmarks are Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. Five of those are recognisable public evaluations. The sixth, SWEFish, is Sakana's own coding benchmark, which means one of the six wins is scored on a test the vendor wrote and has not published.

For the other five, the release states that Fugu Max achieves the best overall score without stating what that score is, or what any competitor scored. That is an unusual shape for a model announcement. Vendors routinely overstate, but they usually overstate with numbers, because numbers are the thing that travels. A release that claims six wins and prints zero of them is asking to be taken on the strength of the claim alone.

There is a charitable reading and it is worth stating plainly. Fugu Max is a cost-optimised coordinator, not a capability push — Sakana shipped it the same day as Fugu Ultra v2, which is the version aimed at maximum capability at $5.00 input and $30.00 output. Max's job is to reach acceptable quality at $6.00 output, and an orchestrator's headline score is less meaningful than its score-per-dollar, which is exactly what a Pareto plot shows. Sakana's visual communication for this release is Pareto plots. If the argument is "look at the curve, not the peak," then a raw benchmark number is beside the point.

The uncharitable reading is that a number would invite a comparison the vendor would rather not have. Both readings are consistent with the evidence, and nothing in the public record currently distinguishes them. What can be said is that no independent party has evaluated any Fugu model, and there is no Artificial Analysis entry for Fugu Max or any of its siblings. Every figure in the release, including the six wins, originates with Sakana.

A two-column scoreboard titled "Fugu Max vs Claude Opus 5 - the scoreboard". Left column Fugu Max: Output price $6.00 / 1M, Input price $2.00 / 1M, Cached input $0.25 / 1M, Benchmarks six wins with no figures, Context not published, Evidence vendor-reported only. Right column Claude Opus 5: Output price $25.00 / 1M, Input price $5.00 / 1M, Cached input up to 90% discount, Benchmarks 96.0% SWE-bench Verified, Context 1M with a 128K output ceiling, Evidence independent evals. Footer reading "Fugu Max figures vendor-reported by Sakana AI; Opus 5 rows Anthropic-reported. No independent Fugu Max evaluation exists.", with the OrcaRouter logo bottom-right.

The purchase you are actually making

When a single-model vendor publishes a benchmark, you are buying a claim about a model. When an orchestrator publishes one, you are buying a claim about a pool — models the vendor did not train, running under a coordinator whose routing decisions are deliberately not exposed. The pool here is described as the largest Fugu has used, expanded with open-weights and specialised models including the NVIDIA Nemotron family through a collaboration with NVIDIA. Members are otherwise undisclosed.

The practical consequence is that Fugu Max's behaviour can change without its version changing. A frontier lab's deprecation, a price change, or a model being withdrawn from a region alters what the coordinator can call, and Sakana is explicit that this resilience is the point — it is selling protection against vendor lock-in and API revocations. That is a real benefit and it is not free. It is also the reason a Fugu Max benchmark, if one were published, would carry an expiry date that a Claude Opus 5 benchmark does not.

Claude Opus 5's value proposition is the inverse and it is easier to price. It is one model at one version, running adaptive thinking by default with an explicit effort dial from low to max, carrying a 1M-token context window and a 128K output ceiling, with vision, tool use and JSON mode. Anth​ropic reports 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro, and those numbers were measured on the thing you get when you call the endpoint — not on an ensemble whose composition can move underneath you. For a workload that runs in production and gets reviewed, that stability is a feature you are paying $19.00 per million output tokens to have.

Cost per finished task, not per token

Fugu Max's $6.00 output rate is genuinely attractive, and the way to test it is to stop comparing token prices. An orchestrator spawns agents; each one writes reasoning and output tokens; the coordinator writes a synthesis. Sakana's pricing FAQ for the Fugu line commits to a single blended rate based on the top-tier participating model, with multi-agent runs not stacking the bill — which removes the worst-case fan-out multiplication that makes naive multi-agent systems unaffordable. It does not remove the volume. The number of output tokens you are billed for is a function of how many agents ran, and at $6.00 per million that volume is the variable a pricing page cannot tell you.

Claude Opus 5's cost shape is one call, one model, an effort dial you control, and batch processing at half rate for work that can wait. You can forecast it before you run it. Over ten thousand executions, forecastability is worth a great deal more than a benchmark margin nobody has published.

This is where the routing question separates from the model question. Claude Opus 5 is on OrcaRouter at Anth​ropic's list price with 0% markup — the provider's rate passed straight through, so a change to Anth​ropic's pricing is live on the same key the same day, with automatic failover across provider paths when one is rate-limited or unavailable. Fugu Max is not on our catalogue; it reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, and moving to it from an earlier Fugu is a one-line parameter change. If you want Opus 5's evidence base and a bill you can predict, the router is the shorter path. If you want a system that assembles its own fan-out on hard problems, Fugu Max is the thing that does that — and you should pilot it with token accounting switched on from the first request.

A screenshot of the Sakana AI release page at sakana.ai/fugu-max-release (captured September 11, 2026, English UI), showing the sakana.ai wordmark, the headline 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' dated September 11, 2026, the opening copy arguing that the frontier which matters is two-dimensional with capability on one axis and cost on the other, a 'Try Sakana Fugu Max and Fugu Ultra v2' link, and a Pareto chart plotting performance against cost with a red 'Fugu Max' point at the low-cost end, a red 'Fugu Ultra' point at the top, a shaded 'Frontier formed by Fugu models' region and grey points labelled 'Frontier formed by single models'.

Where Fugu Max is probably the right answer

Give the design its due. If your work is high-volume, checkable, and you care about the median cost of a completed task rather than the peak capability of any single call, an orchestrator that picks the cheapest sufficient model per request is doing something a fixed flagship cannot do. Fugu Max is aimed squarely at that workload, and the $2.00 input rate with $0.25 cached input is priced for the long, repetitive contexts those workloads generate. The NVIDIA collaboration also matters more than it looks: Nemotron models are open-weights, which means the cheapest rung of the pool is not exposed to another lab's pricing decisions.

What that does not give you is a reason to believe Fugu Max will match Claude Opus 5 on the tasks where Opus 5's published numbers are strongest — repository-scale software engineering and difficult reasoning. Those are precisely the rows where Sakana's six-benchmark board offers no figure at all. If your problem is one where being the best matters more than being cheap, the $25.00 output rate buys the thing you actually need.

Bottom line

Claude Opus 5 is the better-evidenced purchase and the safer production default: published vendor benchmarks, a version you can pin, a context window that does not move, an output ceiling of 128K, an effort dial that lets you buy reasoning only when it pays, and months of third-party results behind it. Fugu Max is the more interesting bet and a genuinely cheaper one — roughly 60 percent less on input and 76 percent less on output, with a cache rate an order of magnitude below Opus 5's base input price.

If you run verifiable, repetitive, high-volume work and you are outside the EU/EEA, Fugu Max deserves a scoped pilot with an output-token ceiling set before you start and a measurement of tokens per completed task rather than tokens per call. If you need a production decision you can defend to someone else this quarter, Claude Opus 5 remains the answer — and it is available on OrcaRouter at Anth​ropic's list price with the provider's rate passed through untouched.

A screenshot of the OrcaRouter model page for Claude Opus 5 (anthropic/claude-opus-5, captured September 11, 2026, English UI), showing the Featured badge, the anthropic/claude-opus-5 model ID, the Vision, Tools, JSON and Reasoning tags, the listing date 2026-07-24, the /v1/chat/completions and /v1/messages endpoints, the pricing tiles reading INPUT $5.00 and OUTPUT $25.00 per 1M tokens with p50 TTFT 4.94s and 195.4M tokens of 7-day traffic, the 1M token context with 128K max output and text + image + file input, and the OpenAI-compatible Python sample pointing at api.orcarouter.ai/v1.