Una card del titolo principale per Fugu Ultra v2 vs DeepSeek V4 Pro con il sottotitolo 'La questione del prezzo di output 34x', pill badge con scritto 'DeepSWE 74.3 vs 62.7', '$30.00 vs $0.87 output' e 'Pesi aperti vs chiusi', una riga a piè di pagina 'Sakana AI vs DeepSeek - settembre 2026', e il logo OrcaRouter nell'angolo in basso a destra.
Engineering & Research

Fugu Ultra v2 vs DeepSeek V4 Pro: la questione del prezzo dell'output 34x

Autore

Rowan Sterling

Data di pubblicazione

Ultimi modelli · 20Vedi tutti i modelli
Benchmark: Artificial Analysis · aggiornato ogni giorno
Torna a tutti gli articoli

Fugu Ultra v2 charges roughly thirty-four times DeepSeek V4 Pro's list price for the text it writes. That is not a typo and it is not a rounding artifact: Sakana AI's new orchestrator lists at $30.00 per million output tokens, Deep​Seek reports $0.87 per million, and on the one benchmark both vendors publish a number for, the gap in performance is nothing like thirty-four times. Sakana claims 74.3 on DeepSWE for Fugu Ultra v2, released September 11, 2026; Deep​Seek reports 62.7 for DeepSeek V4 Pro, the open-weight model it shipped in August. An 11.6-point vendor-reported edge for a 34x output price is the entire comparison in one line — and the honest question is not whether Fugu Ultra v2 is better, but whether it is nineteen dollars per million tokens better, and for whom.

Two answers to the same problem, built from opposite ends

DeepSeek V4 Pro and Fugu Ultra v2 are both responses to the same anxiety — dependence on a single closed frontier vendor — and they could hardly be more different in method.

DeepSeek V4 Pro is a model. Launch coverage dates the official version's API rollout — the DeepSeek-V4-Pro-0813 build — to around August 12 to 13, 2026, superseding the April 24 preview. It is a mixture-of-experts system reported at 1.5 to 1.6 trillion total parameters with about 49 billion activated per token, carrying a 1M-token context window and a 384K-token maximum output. It reasons in thinking and non-thinking modes with three effort levels, speaks both the Ope​nAI and Anth​ropic API protocols, and added a native Responses API aimed squarely at Codex-style agent harnesses. Critically, the weights were published under MIT, which means the resilience argument is settled by possession: nobody can revoke your copy.

Fugu Ultra v2 is not a model. It is a trained coordinator, reported around 7B parameters, that decomposes a request across a pool of undisclosed models, assigns Thinker, Worker and Verifier roles, and synthesises an answer behind one OpenAI-compatible endpoint. Its resilience argument is architectural rather than legal: because the pool is swappable, the system can survive any single vendor changing terms. Version 2.0, shipped alongside Fugu Max, moved the training cutoff to August 28, 2026 and — notably — removed Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra from the pool entirely while still claiming wins over them.

So the comparison is possession versus coordination. Deep​Seek hands you a model you can download, audit, fine-tune and serve yourself at whatever margin your own hardware allows. Sakana hands you a system that decides how to attack each problem, at a rate that funds running several frontier models on your behalf.

A two-column scoreboard for Fugu Ultra v2 vs DeepSeek V4 Pro titled 'Fugu Ultra v2 vs DeepSeek V4 Pro - the scoreboard'. Left column Fugu Ultra v2: Output price $30.00 / 1M, Input price $5.00 / 1M, DeepSWE 74.3, Weights closed, Output ceiling not published, Context 1M. Right column DeepSeek V4 Pro: Output price $0.87 / 1M, Input price $0.435 / 1M, DeepSWE 62.7, Weights open, MIT, Output ceiling 384K, Context 1M. Footer 'Both columns vendor-reported; DeepSeek API rates as reported August 2026, pricing revision pending.', with the OrcaRouter logo bottom-right.

What 34x actually buys

Benchmarks first, with the sourcing front and centre. Every Deep​Seek figure below is the vendor's own and predates the peak/off-peak pricing revision Deep​Seek announced for mid-August, whose final numbers sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against Deep​Seek's published rate card before you model a budget. Every Fugu figure is Sakana's own, and no independent party has reproduced any of them; there is still no Artificial Analysis entry for a Fugu model.

• DeepSWE — Fugu Ultra v2 74.3 (vendor-reported) vs DeepSeek V4 Pro 62.7 (vendor-reported)

• Terminal Bench 2.1 — no Fugu Ultra v2 figure published by Sakana vs DeepSeek V4 Pro 87.9 (vendor-reported)

• SWE-bench (Vals) — no Fugu Ultra v2 figure vs DeepSeek V4 Pro 96.4 (vendor-reported)

• Humanity's Last Exam — no Fugu Ultra v2 v2.0 figure vs DeepSeek V4 Pro 42.7 without tools, 60.0 with tools (vendor-reported)

• GPQA Diamond — no Fugu Ultra v2 v2.0 figure vs DeepSeek V4 Pro 92.8 (vendor-reported)

• Output price — Fugu Ultra v2 $30.00 per 1M, rising to $45.00 above 272K context vs DeepSeek V4 Pro approximately $0.87 per 1M list, with a metered rate that has run higher

Two caveats on that price row, because the whole comparison rests on it. First, Deep​Seek's figure is the list rate reported in launch coverage and restated in our own model description, but the metered rate on our listing has run at $2.18 per million output and $0.73 per million input — cache behaviour and mix move the effective number, so the honest multiple is somewhere between roughly 14x and roughly 34x depending on how much of your input is cached. Even at the floor of that range, output costs more than a dozen times as much. Second, Deep​Seek has already announced a peak/off-peak pricing revision whose final figures sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against the published rate card before you model a budget.

• Input price — Fugu Ultra v2 $5.00 per 1M, doubling above 272K vs DeepSeek V4 Pro approximately $0.435 per 1M on a cache miss, roughly 120x cheaper on a cache hit

• Weights — Fugu Ultra v2 closed, pool undisclosed, routing not exposed by design vs DeepSeek V4 Pro open under MIT, self-hostable

• Context and output — Fugu Ultra v2 1M context, output ceiling unpublished vs DeepSeek V4 Pro 1M context, 384K output ceiling

The overlap problem is obvious: the two systems barely appear on the same rows. Sakana publishes five wins out of eight benchmarks it selected; Deep​Seek publishes a wider board that includes deep agentic coverage — MCP Atlas 73.6, Toolathlon-Verified 74.1, CyberGym 83.3, NL2Repo 61.5, CorpusQA 62.0 on long context — where Fugu Ultra v2 has no published number at all. The single row you can line up, DeepSWE, is the one where Fugu claims its largest software-engineering advantage. A reader comparing the two releases is comparing two different exams.

The verbosity multiplier

Output price is not a linear tax on orchestration; it is a multiplier on a quantity orchestration increases. Fugu Ultra v2 spawns agents, each producing reasoning and output tokens, before a coordinator synthesises them. Sakana's pricing FAQ makes a genuinely consumer-friendly commitment — you pay one blended rate based on the top-tier model in the pool, and adding agents does not multiply the bill — which caps the worst case. It does not cap volume. At $30 per million output tokens, a multi-agent run that emits four times the text of a single-model call costs four times as much, and the multiplier compounds with an output rate 34 times higher.

DeepSeek V4 Pro's cost shape rewards the opposite behaviour. A cache hit on input is roughly two orders of magnitude cheaper than a miss, which makes repeated long-context work — the same repository, the same document set, the same system prompt — dramatically cheaper than the sticker rate implies. For an agent loop that re-reads the same context on every turn, that is where the effective cost lands, and it is a structural advantage no orchestrator can route around.

Put the two together and the decision becomes concrete. If Fugu Ultra v2 delivers an 11.6-point DeepSWE advantage and emits three times the output tokens of a single call, you are paying roughly a hundred times more per completed task for a ten-to-fifteen-point gain on one benchmark. That trade is defensible for research and irreducibly hard engineering problems where the marginal point is the whole game, and indefensible for a pipeline that runs a hundred thousand times a day.

A screenshot of the Sakana Fugu product page at sakana.ai/fugu (captured September 11, 2026, English UI), showing the 'Sakana Fugu' wordmark with the subtitle 'One Model to Command Them All', the description that Fugu dynamically orchestrates the world's best models behind a single API, and the notice reading 'Not yet available in the EU/EEA while we work toward compliance with GDPR and EU-specific regulations.'

Who each one is actually for

Choose DeepSeek V4 Pro if any of the following is true: your workload is high-volume and cost-sensitive; you need output longer than the ceiling an orchestrator will give you; you require a system you can audit, fine-tune or serve on your own hardware for data-residency reasons; or you need the price to be a number you can compute rather than a number you observe. The open MIT licence is not a marketing feature — it is the strongest form of the resilience argument either system is making, because it does not depend on any vendor's continued goodwill. It is also, at roughly $0.87 per million output, cheap enough that experimentation costs nothing.

Choose Fugu Ultra v2 if the marginal point is worth a multiple, the work is long-horizon, and correctness is checkable — coding with a test suite, structured document reasoning, multi-step analysis where a Verifier role can catch an error a single pass would ship. Sakana's own case studies point at exactly this: a fourteen-hour autonomous research run, a Rubik's cube solver that finished all 300 scrambles where two anonymised baselines crashed outright. Bear in mind that those baselines are anonymised and vendor-selected, which weakens them as evidence and does not make them false. Also bear in mind the constraint that ends the discussion for some teams: Fugu Ultra v2 is not sold in the EU or EEA, and Sakana states it plainly.

One practical note if you intend to run either. Fugu Ultra v2 is not on OrcaRouter — it comes through Sakana's own OpenAI-compatible API and several third-party platforms, and upgrading from an earlier Fugu is a single-line parameter change. DeepSeek V4 Pro is on OrcaRouter at the provider's list price with 0% markup, which matters here more than usual: Deep​Seek has already announced one pricing revision this cycle, and on a pass-through router a vendor price change is live on the same key the same day rather than after a renegotiation. If you want to benchmark both against your own workload before committing, the cheaper side of the comparison is the one you can call at list price with automatic failover behind it.

A screenshot of the OrcaRouter model page for DeepSeek V4 Pro (deepseek/deepseek-v4-pro, captured September 11, 2026, English UI), showing the Flagship and Featured badges, the '1.6T total / 49B active params, 1M context, top-tier reasoning + agentic tool use' summary, the 1M token context and 384K max output fields, the pricing tiles reading INPUT $0.73 and OUTPUT $2.18 per 1M tokens with p50 TTFT 919ms, and the description stating transparent pricing at $0.44 per 1 million input tokens and $0.87 per 1 million output tokens.

Il verdetto

This is not a close matchup on economics and not a decisive one on capability. DeepSeek V4 Pro is the better default by a wide margin: open weights you can hold, a 384K output ceiling, a 1M context, published long-context scores, and a price roughly one thirty-fourth of Fugu Ultra v2's on output. Fugu Ultra v2 is the better instrument for a narrow class of expensive problems, and Sakana's claim — that a fixed pool excluding the three strongest models can still beat them on verifiable work — is the most interesting architectural argument in the field right now, and the one with the least independent evidence behind it.

The practical resolution is not to pick. Run DeepSeek V4 Pro as the default because it is cheap, open and fast, and hold Fugu Ultra v2 in reserve for the tasks where a hundredfold cost increase is still smaller than the cost of being wrong. What you should not do is read Sakana's eight-benchmark board as evidence that the expensive orchestrator has replaced the cheap open model. On the one row they share, it is 11.6 points ahead for 34 times the output price. Whether that is a bargain or a trap depends entirely on which problem you point it at.

Confrontati in questo articolo1

Rilevato da questo articolo · Benchmark: Artificial Analysis · aggiornato ogni giorno