
Fugu Ultra v2 對比 DeepSeek V4 Pro:34 倍輸出價格的疑問
- deepseek新DeepSeek: DeepSeek V4.1 Flash2026-09-1040智能
- openai新OpenAI: GPT-6 Astra2026-09-0453智能77程式
- google新Google: Gemini 3.8 Flash2026-09-0241智能76程式
- qwen新Qwen: Qwen3.8 Max (0902)2026-09-0240智能72程式
- anthropic新Anthropic: Claude Fable 5.12026-09-0153智能82程式
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 每百萬 tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642智能72程式
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 每百萬 tokens
- z-aiZ.ai: GLM 5.32026-08-1845智能75程式
- obsidianQwen3.8 27B2026-08-1534智能68程式
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236智能69程式
- grokSpaceXAI: Grok 4.62026-08-1244智能77程式
- metaMeta: Muse Spark 1.22026-08-0540智能72程式
- qwenQwen: Qwen3.8 Max2026-08-0340智能72程式
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135智能69程式
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 每百萬 tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451智能78程式
- googleGoogle: Gemini 3.6 Flash2026-07-2134智能69程式
Fugu Ultra v2 charges roughly thirty-four times DeepSeek V4 Pro's list price for the text it writes. That is not a typo and it is not a rounding artifact: Sakana AI's new orchestrator lists at $30.00 per million output tokens, DeepSeek reports $0.87 per million, and on the one benchmark both vendors publish a number for, the gap in performance is nothing like thirty-four times. Sakana claims 74.3 on DeepSWE for Fugu Ultra v2, released September 11, 2026; DeepSeek reports 62.7 for DeepSeek V4 Pro, the open-weight model it shipped in August. An 11.6-point vendor-reported edge for a 34x output price is the entire comparison in one line — and the honest question is not whether Fugu Ultra v2 is better, but whether it is nineteen dollars per million tokens better, and for whom.
同一個問題的兩個答案,從相反的兩端構築而成
DeepSeek V4 Pro 與 Fugu Ultra v2 都是對同一種焦慮的回應——對單一封閉前沿廠商的依賴——而它們在方法上可說截然不同到了極點。
DeepSeek V4 Pro is a model. Launch coverage dates the official version's API rollout — the DeepSeek-V4-Pro-0813 build — to around August 12 to 13, 2026, superseding the April 24 preview. It is a mixture-of-experts system reported at 1.5 to 1.6 trillion total parameters with about 49 billion activated per token, carrying a 1M-token context window and a 384K-token maximum output. It reasons in thinking and non-thinking modes with three effort levels, speaks both the OpenAI and Anthropic API protocols, and added a native Responses API aimed squarely at Codex-style agent harnesses. Critically, the weights were published under MIT, which means the resilience argument is settled by possession: nobody can revoke your copy.
Fugu Ultra v2 不是一個模型。它是一個受過訓練的協調器,據報導約有 7B 參數,會將一項請求拆解到一組未公開的模型池中,指派 Thinker、Worker 與 Verifier 角色,並在單一 OpenAI 相容端點背後綜合成一個答案。其韌性論證是架構性的,而非法律性的:因為模型池可抽換,該系統能夠承受任何單一供應商更改條款。2.0 版與 Fugu Max 一同推出,將訓練截止時間移至 2026 年 8 月 28 日,而且——值得注意的是——將 Claude Fable 5、Claude Fable 5.1 與 GPT-6 Astra 從模型池中完全移除,卻仍宣稱能勝過它們。
So the comparison is possession versus coordination. DeepSeek hands you a model you can download, audit, fine-tune and serve yourself at whatever margin your own hardware allows. Sakana hands you a system that decides how to attack each problem, at a rate that funds running several frontier models on your behalf.

34x 究竟能換來什麼
Benchmarks first, with the sourcing front and centre. Every DeepSeek figure below is the vendor's own and predates the peak/off-peak pricing revision DeepSeek announced for mid-August, whose final numbers sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against DeepSeek's published rate card before you model a budget. Every Fugu figure is Sakana's own, and no independent party has reproduced any of them; there is still no Artificial Analysis entry for a Fugu model.
• DeepSWE — Fugu Ultra v2 74.3(廠商回報)對 DeepSeek V4 Pro 62.7(廠商回報)
• Terminal Bench 2.1 — Sakana 未公布 Fugu Ultra v2 的分數,相較於 DeepSeek V4 Pro 的 87.9(廠商提供)
• SWE-bench(Vals)— 沒有 Fugu Ultra v2 的數據,對比 DeepSeek V4 Pro 96.4(廠商回報)
• Humanity's Last Exam — 無 Fugu Ultra v2 v2.0 數據 vs DeepSeek V4 Pro 42.7(不使用工具)、60.0(使用工具)(廠商報告)
• GPQA Diamond — 沒有 Fugu Ultra v2 v2.0 的數字,對比 DeepSeek V4 Pro 的 92.8(廠商回報)
• 輸出價格 — Fugu Ultra v2 每 1M 為 $30.00,在超過 272K 上下文時升至 $45.00;相較之下,DeepSeek V4 Pro 的標價約為每 1M $0.87,而按量計費的費率曾更高
Two caveats on that price row, because the whole comparison rests on it. First, DeepSeek's figure is the list rate reported in launch coverage and restated in our own model description, but the metered rate on our listing has run at $2.18 per million output and $0.73 per million input — cache behaviour and mix move the effective number, so the honest multiple is somewhere between roughly 14x and roughly 34x depending on how much of your input is cached. Even at the floor of that range, output costs more than a dozen times as much. Second, DeepSeek has already announced a peak/off-peak pricing revision whose final figures sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against the published rate card before you model a budget.
• 輸入價格 — Fugu Ultra v2 每 1M 為 $5.00,超過 272K 後加倍;相較之下,DeepSeek V4 Pro 在快取未命中時約每 1M 為 $0.435,在快取命中時便宜約 120 倍
• 權重 — Fugu Ultra v2 為閉源,池未披露,路由依設計不公開,對比 DeepSeek V4 Pro 以 MIT 授權開源、可自架設
• 上下文與輸出 — Fugu Ultra v2 1M 上下文,輸出上限未公布 vs DeepSeek V4 Pro 1M 上下文,384K 輸出上限
The overlap problem is obvious: the two systems barely appear on the same rows. Sakana publishes five wins out of eight benchmarks it selected; DeepSeek publishes a wider board that includes deep agentic coverage — MCP Atlas 73.6, Toolathlon-Verified 74.1, CyberGym 83.3, NL2Repo 61.5, CorpusQA 62.0 on long context — where Fugu Ultra v2 has no published number at all. The single row you can line up, DeepSWE, is the one where Fugu claims its largest software-engineering advantage. A reader comparing the two releases is comparing two different exams.
詳細程度倍數
輸出價格並不是對編排課徵的線性稅;它是對編排所增加之數量的乘數。Fugu Ultra v2 會生成代理,每個代理都會產生推理與輸出 token,之後再由協調器綜合彙整。Sakana 的定價常見問題提出了一項真正對消費者友善的承諾——你根據池中最高階模型支付單一混合費率,而增加代理不會讓帳單倍增——這為最壞情況設下了上限。但它並未限制用量。以每百萬個輸出 token 30 美元計算,一次多代理執行若輸出單一模型呼叫四倍的文字量,成本就會是四倍,而這個乘數還會與高出 34 倍的輸出費率相互疊加。
DeepSeek V4 Pro 的成本結構獎勵的卻是相反的行為。輸入端的快取命中比未命中便宜大約兩個數量級,這使得重複性的長上下文工作——同一個儲存庫、同一份文件集、同一個系統提示——遠比標價所暗示的便宜。對於每一回合都重新讀取相同上下文的代理迴圈而言,實際成本就落在這裡,而這是任何協調器都無法繞過的結構性優勢。
把這兩者放在一起看,決策就變得具體了。如果 Fugu Ultra v2 能帶來 11.6 點的 DeepSWE 優勢,而且輸出的 token 數量是單次呼叫的三倍,那麼為了在某個基準測試上換得十到十五點的提升,你每完成一項任務所付出的成本大約是原來的百倍。這樣的取捨,在研究領域,以及那些邊際的那一點就是勝負全部的、無法再簡化的艱難工程問題上,說得過去;但對於每天要跑十萬次的管線來說,則站不住腳。

每一個實際上適合誰
若符合以下任一情況,請選擇 DeepSeek V4 Pro:你的工作負載屬於高用量且對成本敏感;你需要比編排器所能給你的上限更長的輸出;你需要一個可稽核、可微調,或基於資料駐留理由而能在自有硬體上提供服務的系統;或者你需要價格是一個你能計算出來的數字,而不是一個你只能觀察到的數字。開放的 MIT 授權不是行銷噱頭——它是這兩個系統所提出的韌性論據中最強的一種形式,因為它不依賴任何供應商持續的善意。它也很便宜,每百萬輸出約 $0.87,便宜到實驗不花什麼成本。
若那最後一點邊際效益值得付出數倍代價、工作屬於長時程性質,而且正確性是可以檢驗的,就選擇 Fugu Ultra v2——例如搭配測試套件的程式開發、結構化文件推理,或是由「驗證者」角色能抓出單次執行就會直接交出去的錯誤的多步驟分析。Sakana 自家的案例研究正指向這一點:一次長達十四小時的自主研究執行,以及一個完成全部 300 種打亂狀態的魔方求解器,而兩個匿名化的基準方法則直接崩潰。請記住,那些基準方法是匿名化且由廠商挑選的,這削弱了它們作為證據的效力,但並不代表它們是假的。也請記住這項對某些團隊而言直接結束討論的限制:Fugu Ultra v2 不在歐盟或歐洲經濟區販售,而 Sakana 也明白說明了這一點。
One practical note if you intend to run either. Fugu Ultra v2 is not on OrcaRouter — it comes through Sakana's own OpenAI-compatible API and several third-party platforms, and upgrading from an earlier Fugu is a single-line parameter change. DeepSeek V4 Pro is on OrcaRouter at the provider's list price with 0% markup, which matters here more than usual: DeepSeek has already announced one pricing revision this cycle, and on a pass-through router a vendor price change is live on the same key the same day rather than after a renegotiation. If you want to benchmark both against your own workload before committing, the cheaper side of the comparison is the one you can call at list price with automatic failover behind it.

裁決
這在經濟性上並非勢均力敵的對決,在能力上也不是決定性的勝負。DeepSeek V4 Pro 則是明顯更佳的預設選擇:可實際持有的開放權重、384K 輸出上限、1M 上下文、已公布的長上下文分數,以及輸出價格約為 Fugu Ultra v2 的三十四分之一。Fugu Ultra v2 則是針對一小類昂貴問題的更好工具,而 Sakana 的說法——一個排除掉三個最強模型的固定模型池,仍可在可驗證的工作上勝過它們——是當前這個領域最有趣的架構論點,也是背後獨立證據最少的一個。
實際的解決方案不是二選一。把 DeepSeek V4 Pro 當作預設選項,因為它便宜、開放且快速,並將 Fugu Ultra v2 保留備用,用於那些即使成本增加一百倍,仍小於判斷錯誤代價的任務。你不該做的,是把 Sakana 的八項基準測試榜單解讀為昂貴的協調器已經取代了便宜的開源模型。在它們唯一共同的那一列上,它領先 11.6 分,而輸出價格是 34 倍。這究竟是划算還是陷阱,完全取決於你把它指向哪個問題。
本文中的比較1
根據本文內容識別 · 基準測試:Artificial Analysis · 每日更新
