A hero title card for "Fugu Max vs DeepSeek V4 Pro" with the subtitle "The cost-optimised one is the expensive one", three pill badges reading "$6.00 vs $2.18 output", "1.6T MoE, 49B active", "384K output ceiling", a footer line reading "Sakana AI vs DeepSeek - September 2026", and the OrcaRouter logo in the bottom-right corner.
Engineering & Research

Fugu Max vs DeepSeek V4 Pro: the cost-optimised one is the expensive one

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Fugu Max is built to be the cheap option and DeepSeek V4 Pro beats it on price anyway. Sakana AI released Fugu Max on September 11, 2026 at $2.00 per million input tokens and $6.00 per million output tokens, positioning it as a coordinator that expands the cost-performance frontier by routing each task to the leanest model in its pool. DeepSeek V4 Pro, listed since April 2026, sits at $0.73 input and $2.18 output per million on OrcaRouter's metered rate — roughly a third of what Fugu Max charges on both axes, from a single open-weights mixture-of-experts model with 384K of output ceiling and no orchestration layer at all. That inversion is the whole comparison. Everything else about this pair is a question about what orchestration buys you when the baseline you are trying to undercut is already cheaper than you are.

Two ways to lower a bill, and one of them is already lower

Sakana's argument for Fugu Max is that a coordinator can cut spend by not paying frontier prices for tasks that do not need frontier capability. It is a sound argument. The problem is the model it is being made against. DeepSeek V4 Pro is a 1.6-trillion-parameter mixture-of-experts with 49 billion parameters active per token, which is the same cost-reduction idea executed one level down — you pay for the parameters a token actually activates, not for the size of the network. Both products are answers to "how do we stop paying for capacity we do not use." One of them charges $2.18 per million output tokens for the privilege.

• Input price — Fugu Max $2.00 per 1M tokens vs DeepSeek V4 Pro $0.73 per 1M tokens metered

• Output price — Fugu Max $6.00 per 1M tokens vs DeepSeek V4 Pro $2.18 per 1M tokens metered

• Cached input — Fugu Max $0.25 per 1M tokens vs DeepSeek V4 Pro cache-hit input priced far below its base rate, with the vendor's published list rate sitting under $1.00 per 1M

• Context — Fugu Max not published vs DeepSeek V4 Pro 1M tokens

• Output ceiling — Fugu Max not published vs DeepSeek V4 Pro 384K tokens

• Modality — Fugu Max not published vs DeepSeek V4 Pro text-only, with tool use, JSON mode and reasoning

• Architecture — Fugu Max a trained coordinator over an undisclosed pool vs DeepSeek V4 Pro a single MoE model, open-weights lineage

• Benchmark figures — Fugu Max six wins asserted with no scores vs DeepSeek V4 Pro vendor-reported 87.9 on Terminal Bench 2.1, 92.8 on GPQA Diamond, 96.4 on SWE-bench Vals

One number in that list deserves to be pulled out. Deep​Seek reports 87.9 on Terminal Bench 2.1. Fugu Max claims "best overall score" on Terminal Bench 2.1. Same benchmark, same release window, one side publishes a figure and the other publishes the word "best." That is the single cleanest illustration of the evidence asymmetry in this pair, and it is not a detail — it is the whole reason the price gap is not the end of the argument.

A two-column scoreboard titled "Fugu Max vs DeepSeek V4 Pro - the scoreboard". Left column Fugu Max: Output price $6.00 / 1M, Input price $2.00 / 1M, Cached input $0.25 / 1M, Context not published, Terminal Bench 2.1 best overall with no figure, Weights closed and pool undisclosed. Right column DeepSeek V4 Pro: Output price $2.18 / 1M, Input price $0.73 / 1M, Cached input below base rate, Context 1M tokens, Terminal Bench 2.1 87.9, Weights open-weights lineage. Footer reading "Fugu Max rows vendor-reported by Sakana AI with no scores published; DeepSeek rows vendor-reported, metered rate on OrcaRouter.", with the OrcaRouter logo bottom-right.

What orchestration buys when the baseline is already cheap

The case for paying more for Fugu Max has to rest on work that a single model handles badly, and there is a real category of it. Long-horizon tasks where a wrong turn compounds — multi-file refactors, repository-scale changes, anything where a verifier checking the worker's output is worth more than a stronger worker — are exactly what a Thinker/Worker/Verifier arrangement is for. Sakana's own framing for the Fugu line is that catching a mistake in a second pass is cheaper than getting it right the first time. On those tasks, a $6.00 output rate that finishes in one attempt can beat a $2.18 rate that needs three.

That argument has a testable form and Sakana has not run it in public. The comparison you would want is cost per completed task on a benchmark both systems attempt — Terminal Bench 2.1 is right there, it is agentic and terminal-based by construction, and only one of the two vendors has published a number for it. Until that gap closes, the honest position is that Fugu Max's cost advantage is asserted at the token level and unproven at the task level, while DeepSeek V4 Pro's is straightforward: the tokens themselves cost a third as much.

There is a second-order point about pool composition that matters more for this matchup than for most. Fugu Max's pool expanded this release to include open-weights and specialised models, with the NVIDIA Nemotron family named through an NVIDIA collaboration. Open weights in the pool mean the cheapest rung of Sakana's ladder is not exposed to another lab's pricing decisions. DeepSeek V4 Pro, being open-weights itself, is exposed to nobody's pricing decisions but Deep​Seek's — and it is the rung. The resilience argument Sakana makes for orchestration applies to Deep​Seek directly and to Fugu Max's underlying supply only indirectly.

Caching is where agent workloads actually get expensive

Neither vendor's base rate is the rate an agent loop pays. Long agent runs re-send a large, mostly unchanging prefix on every turn, and the cache-hit rate is what determines the real bill. Fugu Max lists cached input at $0.25 per million, which is a genuine and aggressive figure — an eighth of its own base input rate. DeepSeek V4 Pro prices cache-hit input far below its published base, and its output rate is set at roughly three times its input rather than the four-to-five-times ratio most vendors use, which compresses the cost of generation-heavy loops specifically.

What you cannot do today is compute the comparison reliably from either vendor's rate card, because Fugu Max's cached rate applies to a token count you cannot predict. An orchestrator decides how many agents run; each agent's context contributes; the coordinator adds its own. Sakana's pricing commitment for the Fugu line — one blended rate based on the top-tier participating model, with agents not stacking the bill — removes the worst case, which is a real concession and worth crediting. It does not make the volume knowable in advance. DeepSeek V4 Pro's bill is a function of tokens you send and tokens you receive, and nothing else.

That predictability is a routing property as much as a model property. DeepSeek V4 Pro is on OrcaRouter at the provider's list price with 0% markup, so a rate revision from Deep​Seek reaches your existing key the same day rather than at your next contract renewal, and automatic failover covers you when one provider path is rate-limited. That last part carries unusual weight for this model specifically: Deep​Seek has revised its pricing once already this cycle, with peak and off-peak tiers whose final figures vary between sources. When a vendor's rate card moves, the router is the difference between absorbing the change and discovering it on an invoice.

A screenshot of the Sakana AI release page at sakana.ai/fugu-max-release (captured September 11, 2026, English UI), showing the sakana.ai wordmark, the headline 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' dated September 11, 2026, the opening copy arguing that the frontier which matters is two-dimensional with capability on one axis and cost on the other, a 'Try Sakana Fugu Max and Fugu Ultra v2' link, and a Pareto chart plotting performance against cost with a red 'Fugu Max' point at the low-cost end, a red 'Fugu Ultra' point at the top, a shaded 'Frontier formed by Fugu models' region and grey points labelled 'Frontier formed by single models'.

What we could not verify

Three things in the Fugu Max release could not be checked against anything outside Sakana's own materials, and they are listed here rather than glossed over. First, the six benchmark wins carry no figures — Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish are named as wins with no scores attached, and SWEFish is Sakana's internal benchmark. Second, Fugu Max's context window, output ceiling and supported modalities are not published, so the modality and ceiling rows above compare a shipped specification against silence. Third, no independent evaluation of any Fugu model exists, and there is no Artificial Analysis entry for Fugu Max.

For DeepSeek V4 Pro the situation is the reverse. The vendor reports a long list of figures — Terminal Bench 2.1 87.9, GPQA Diamond 92.8, SWE-bench Vals 96.4, HLE 42.7 and 60.0 in two configurations, MCP Atlas 73.6, Toolathlon-Verified 74.1, CyberGym 83.3 — and all of them are vendor-reported too. The difference is not that one vendor is more trustworthy. It is that when both sides publish numbers, disagreement is possible, and a disagreement can be investigated. A claim with no number cannot be checked, only accepted or ignored.

One further caveat on the Deep​Seek pricing above: our own model page carries two figures, a metered rate of $0.73 input and $2.18 output per million in the pricing tiles and an older description of $0.44 input and $0.87 output per million. Deep​Seek has also announced peak and off-peak tiers whose final rates sources disagree on. Treat $0.73 and $2.18 as the rate you will actually be billed through us today and confirm against the published rate card before you build a budget on it.

Bottom line

DeepSeek V4 Pro wins this comparison on the terms Fugu Max chose. It is cheaper on input and output, it has a published context window and a 384K output ceiling, it reports a real number on the benchmark Fugu Max claims to lead, and its lineage is open-weights, which is the strongest available form of the vendor-independence argument Sakana is selling. If your workload is token-heavy and your priority is the lowest defensible cost per token from a model you can pin, this is not close.

Fugu Max wins a narrower question that DeepSeek V4 Pro does not answer at all: whether a coordinator that assembles its own agent team can finish hard, long-horizon work in fewer attempts than a single model. That is worth piloting if you have checkable, failure-tolerant work and you are outside the EU/EEA. Cost it on tokens per completed task, set a ceiling first, and compare it against the DeepSeek V4 Pro endpoint on OrcaRouter at the provider's list price — one key, 0% markup, and a rate that tracks the vendor's own.

A screenshot of the OrcaRouter model page for DeepSeek V4 Pro (deepseek/deepseek-v4-pro, captured September 11, 2026, English UI), showing the FLAGSHIP and Featured badges, the deepseek/deepseek-v4-pro model ID, the Tools, JSON and Reasoning tags, the listing date 2026-04-24, the /v1/chat/completions and /v1/responses endpoints, the pricing tiles reading INPUT $0.73 and OUTPUT $2.18 per 1M tokens with p50 TTFT 919ms and 6,313.7M tokens of 7-day traffic, the 1M token context with 384K max output and text-only input, and the OpenAI-compatible Python sample pointing at api.orcarouter.ai/v1.