
Fugu Ultra v2 vs DeepSeek V4 Pro: la questione del prezzo dell'output 34x
- deepseekNUOVODeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligenza
- openaiNUOVOOpenAI: GPT-6 Astra2026-09-0453Intelligenza77Codice
- googleNUOVOGoogle: Gemini 3.8 Flash2026-09-0241Intelligenza76Codice
- qwenNUOVOQwen: Qwen3.8 Max (0902)2026-09-0240Intelligenza72Codice
- anthropicNUOVOAnthropic: Claude Fable 5.12026-09-0153Intelligenza82Codice
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M di token
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligenza72Codice
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M di token
- z-aiZ.ai: GLM 5.32026-08-1845Intelligenza75Codice
- obsidianQwen3.8 27B2026-08-1534Intelligenza68Codice
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligenza69Codice
- grokSpaceXAI: Grok 4.62026-08-1244Intelligenza77Codice
- metaMeta: Muse Spark 1.22026-08-0540Intelligenza72Codice
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligenza72Codice
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligenza69Codice
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M di token
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligenza78Codice
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligenza69Codice
Fugu Ultra v2 charges roughly thirty-four times DeepSeek V4 Pro's list price for the text it writes. That is not a typo and it is not a rounding artifact: Sakana AI's new orchestrator lists at $30.00 per million output tokens, DeepSeek reports $0.87 per million, and on the one benchmark both vendors publish a number for, the gap in performance is nothing like thirty-four times. Sakana claims 74.3 on DeepSWE for Fugu Ultra v2, released September 11, 2026; DeepSeek reports 62.7 for DeepSeek V4 Pro, the open-weight model it shipped in August. An 11.6-point vendor-reported edge for a 34x output price is the entire comparison in one line — and the honest question is not whether Fugu Ultra v2 is better, but whether it is nineteen dollars per million tokens better, and for whom.
Two answers to the same problem, built from opposite ends
DeepSeek V4 Pro and Fugu Ultra v2 are both responses to the same anxiety — dependence on a single closed frontier vendor — and they could hardly be more different in method.
DeepSeek V4 Pro is a model. Launch coverage dates the official version's API rollout — the DeepSeek-V4-Pro-0813 build — to around August 12 to 13, 2026, superseding the April 24 preview. It is a mixture-of-experts system reported at 1.5 to 1.6 trillion total parameters with about 49 billion activated per token, carrying a 1M-token context window and a 384K-token maximum output. It reasons in thinking and non-thinking modes with three effort levels, speaks both the OpenAI and Anthropic API protocols, and added a native Responses API aimed squarely at Codex-style agent harnesses. Critically, the weights were published under MIT, which means the resilience argument is settled by possession: nobody can revoke your copy.
Fugu Ultra v2 is not a model. It is a trained coordinator, reported around 7B parameters, that decomposes a request across a pool of undisclosed models, assigns Thinker, Worker and Verifier roles, and synthesises an answer behind one OpenAI-compatible endpoint. Its resilience argument is architectural rather than legal: because the pool is swappable, the system can survive any single vendor changing terms. Version 2.0, shipped alongside Fugu Max, moved the training cutoff to August 28, 2026 and — notably — removed Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra from the pool entirely while still claiming wins over them.
So the comparison is possession versus coordination. DeepSeek hands you a model you can download, audit, fine-tune and serve yourself at whatever margin your own hardware allows. Sakana hands you a system that decides how to attack each problem, at a rate that funds running several frontier models on your behalf.

What 34x actually buys
Benchmarks first, with the sourcing front and centre. Every DeepSeek figure below is the vendor's own and predates the peak/off-peak pricing revision DeepSeek announced for mid-August, whose final numbers sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against DeepSeek's published rate card before you model a budget. Every Fugu figure is Sakana's own, and no independent party has reproduced any of them; there is still no Artificial Analysis entry for a Fugu model.
• DeepSWE — Fugu Ultra v2 74.3 (vendor-reported) vs DeepSeek V4 Pro 62.7 (vendor-reported)
• Terminal Bench 2.1 — no Fugu Ultra v2 figure published by Sakana vs DeepSeek V4 Pro 87.9 (vendor-reported)
• SWE-bench (Vals) — no Fugu Ultra v2 figure vs DeepSeek V4 Pro 96.4 (vendor-reported)
• Humanity's Last Exam — no Fugu Ultra v2 v2.0 figure vs DeepSeek V4 Pro 42.7 without tools, 60.0 with tools (vendor-reported)
• GPQA Diamond — no Fugu Ultra v2 v2.0 figure vs DeepSeek V4 Pro 92.8 (vendor-reported)
• Output price — Fugu Ultra v2 $30.00 per 1M, rising to $45.00 above 272K context vs DeepSeek V4 Pro approximately $0.87 per 1M list, with a metered rate that has run higher
Two caveats on that price row, because the whole comparison rests on it. First, DeepSeek's figure is the list rate reported in launch coverage and restated in our own model description, but the metered rate on our listing has run at $2.18 per million output and $0.73 per million input — cache behaviour and mix move the effective number, so the honest multiple is somewhere between roughly 14x and roughly 34x depending on how much of your input is cached. Even at the floor of that range, output costs more than a dozen times as much. Second, DeepSeek has already announced a peak/off-peak pricing revision whose final figures sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against the published rate card before you model a budget.
• Input price — Fugu Ultra v2 $5.00 per 1M, doubling above 272K vs DeepSeek V4 Pro approximately $0.435 per 1M on a cache miss, roughly 120x cheaper on a cache hit
• Weights — Fugu Ultra v2 closed, pool undisclosed, routing not exposed by design vs DeepSeek V4 Pro open under MIT, self-hostable
• Context and output — Fugu Ultra v2 1M context, output ceiling unpublished vs DeepSeek V4 Pro 1M context, 384K output ceiling
The overlap problem is obvious: the two systems barely appear on the same rows. Sakana publishes five wins out of eight benchmarks it selected; DeepSeek publishes a wider board that includes deep agentic coverage — MCP Atlas 73.6, Toolathlon-Verified 74.1, CyberGym 83.3, NL2Repo 61.5, CorpusQA 62.0 on long context — where Fugu Ultra v2 has no published number at all. The single row you can line up, DeepSWE, is the one where Fugu claims its largest software-engineering advantage. A reader comparing the two releases is comparing two different exams.
The verbosity multiplier
Output price is not a linear tax on orchestration; it is a multiplier on a quantity orchestration increases. Fugu Ultra v2 spawns agents, each producing reasoning and output tokens, before a coordinator synthesises them. Sakana's pricing FAQ makes a genuinely consumer-friendly commitment — you pay one blended rate based on the top-tier model in the pool, and adding agents does not multiply the bill — which caps the worst case. It does not cap volume. At $30 per million output tokens, a multi-agent run that emits four times the text of a single-model call costs four times as much, and the multiplier compounds with an output rate 34 times higher.
DeepSeek V4 Pro's cost shape rewards the opposite behaviour. A cache hit on input is roughly two orders of magnitude cheaper than a miss, which makes repeated long-context work — the same repository, the same document set, the same system prompt — dramatically cheaper than the sticker rate implies. For an agent loop that re-reads the same context on every turn, that is where the effective cost lands, and it is a structural advantage no orchestrator can route around.
Put the two together and the decision becomes concrete. If Fugu Ultra v2 delivers an 11.6-point DeepSWE advantage and emits three times the output tokens of a single call, you are paying roughly a hundred times more per completed task for a ten-to-fifteen-point gain on one benchmark. That trade is defensible for research and irreducibly hard engineering problems where the marginal point is the whole game, and indefensible for a pipeline that runs a hundred thousand times a day.

Who each one is actually for
Choose DeepSeek V4 Pro if any of the following is true: your workload is high-volume and cost-sensitive; you need output longer than the ceiling an orchestrator will give you; you require a system you can audit, fine-tune or serve on your own hardware for data-residency reasons; or you need the price to be a number you can compute rather than a number you observe. The open MIT licence is not a marketing feature — it is the strongest form of the resilience argument either system is making, because it does not depend on any vendor's continued goodwill. It is also, at roughly $0.87 per million output, cheap enough that experimentation costs nothing.
Choose Fugu Ultra v2 if the marginal point is worth a multiple, the work is long-horizon, and correctness is checkable — coding with a test suite, structured document reasoning, multi-step analysis where a Verifier role can catch an error a single pass would ship. Sakana's own case studies point at exactly this: a fourteen-hour autonomous research run, a Rubik's cube solver that finished all 300 scrambles where two anonymised baselines crashed outright. Bear in mind that those baselines are anonymised and vendor-selected, which weakens them as evidence and does not make them false. Also bear in mind the constraint that ends the discussion for some teams: Fugu Ultra v2 is not sold in the EU or EEA, and Sakana states it plainly.
One practical note if you intend to run either. Fugu Ultra v2 is not on OrcaRouter — it comes through Sakana's own OpenAI-compatible API and several third-party platforms, and upgrading from an earlier Fugu is a single-line parameter change. DeepSeek V4 Pro is on OrcaRouter at the provider's list price with 0% markup, which matters here more than usual: DeepSeek has already announced one pricing revision this cycle, and on a pass-through router a vendor price change is live on the same key the same day rather than after a renegotiation. If you want to benchmark both against your own workload before committing, the cheaper side of the comparison is the one you can call at list price with automatic failover behind it.

Il verdetto
This is not a close matchup on economics and not a decisive one on capability. DeepSeek V4 Pro is the better default by a wide margin: open weights you can hold, a 384K output ceiling, a 1M context, published long-context scores, and a price roughly one thirty-fourth of Fugu Ultra v2's on output. Fugu Ultra v2 is the better instrument for a narrow class of expensive problems, and Sakana's claim — that a fixed pool excluding the three strongest models can still beat them on verifiable work — is the most interesting architectural argument in the field right now, and the one with the least independent evidence behind it.
The practical resolution is not to pick. Run DeepSeek V4 Pro as the default because it is cheap, open and fast, and hold Fugu Ultra v2 in reserve for the tasks where a hundredfold cost increase is still smaller than the cost of being wrong. What you should not do is read Sakana's eight-benchmark board as evidence that the expensive orchestrator has replaced the cheap open model. On the one row they share, it is 11.6 points ahead for 34 times the output price. Whether that is a bargain or a trap depends entirely on which problem you point it at.
Confrontati in questo articolo1
Rilevato da questo articolo · Benchmark: Artificial Analysis · aggiornato ogni giorno
