
Fugu Ultra v2 與 Claude Opus 5:相同輸入價格,兩種不同的押注
- deepseek新DeepSeek: DeepSeek V4.1 Flash2026-09-1040智能
- openai新OpenAI: GPT-6 Astra2026-09-0453智能77程式
- google新Google: Gemini 3.8 Flash2026-09-0241智能76程式
- qwen新Qwen: Qwen3.8 Max (0902)2026-09-0240智能72程式
- anthropic新Anthropic: Claude Fable 5.12026-09-0153智能82程式
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 每百萬 tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642智能72程式
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 每百萬 tokens
- z-aiZ.ai: GLM 5.32026-08-1845智能75程式
- obsidianQwen3.8 27B2026-08-1534智能68程式
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236智能69程式
- grokSpaceXAI: Grok 4.62026-08-1244智能77程式
- metaMeta: Muse Spark 1.22026-08-0540智能72程式
- qwenQwen: Qwen3.8 Max2026-08-0340智能72程式
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135智能69程式
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 每百萬 tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451智能78程式
- googleGoogle: Gemini 3.6 Flash2026-07-2134智能69程式
Fugu Ultra v2 and Claude Opus 5 cost exactly the same amount to ask a question and wildly different amounts to get an answer out of. Both list at $5.00 per million input tokens. Then Fugu Ultra v2 charges $30.00 per million output tokens against Claude Opus 5's $25.00 — and the gap between those two numbers is not a rounding difference, because the two systems emit very different quantities of text to reach an answer. Sakana AI shipped Fugu Ultra v2 on September 11, 2026 as an orchestration system that coordinates a pool of other labs' models behind one OpenAI-compatible endpoint; Anthropic's Claude Opus 5, live since late July 2026, is a single model. That difference decides most of this comparison before a benchmark is consulted, and Sakana's own scoreboard has an awkward row in it: the vendor reports 48.3 on Chartography against Claude Opus 5's 27.3, which is the largest margin in the release and also the number with the least outside evidence behind it.
形塑一切的巧合
先從兩個產品的共同點開始,因為它比表面上的價格更有價值。
• 輸入價格 — Fugu Ultra v2 每 100 萬個 token 5.00 美元,對比 Claude Opus 5 每 100 萬個 token 5.00 美元
• 輸出價格 — Fugu Ultra v2 每 100 萬個 token 為 $30.00,相較於 Claude Opus 5 每 100 萬個 token 為 $25.00
• 快取輸入 — Fugu Ultra v2 每 1M 較基本費率高 $0.50,相較之下,Claude Opus 5 將快取定價為快取輸入最高 90% 的折扣
• 上下文 — Fugu Ultra v2 為 100 萬個 token,超過 272K 之後費率加倍;相較之下,Claude Opus 5 為 100 萬個 token,費率維持不變
• 輸出上限 — Fugu Ultra v2 未公布為單一數字 vs Claude Opus 5 128K tokens
• 可用性 — Fugu Ultra v2 未在 EU/EEA 銷售,對比 Claude Opus 5 在標準 API 條款下普遍可用
• 證據 — Fugu Ultra v2 廠商自行公布的基準測試,尚無獨立評估;相較之下,Claude Opus 5 有廠商提供的基準測試,加上數個月來在第三方排行榜上的排名
最後一列正是這場對決真正出現轉折的地方。在相同的輸入價格下,你得在兩種系統之間做選擇:一種是你只能憑信心接受其分數的系統,另一種是其分數已被其他人重現的模型。272K 的斷崖讓情況雪上加霜:一旦超過這個門檻,Fugu Ultra v2 的輸入費率翻倍至 $10,輸出則為 $45,而 Claude Opus 5 的費率則維持不變。對於長脈絡的代理執行作業——正是 Fugu Ultra v2 所量身打造的工作負載——表面上較便宜的那個價格優勢,恰恰就在這套系統理應大放異彩的地方消失無蹤。

每一個實際上是什麼
Claude Opus 5 is Anthropic's production flagship, and its design is legible from the outside. It runs adaptive thinking by default, deciding how long to reason before it answers, with an explicit effort dial from low to max when you want to force deeper deliberation and pay for it. It carries a 1M-token context window, a 128K output ceiling, vision, tool use and JSON mode. Anthropic reports 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro, more than doubling its predecessor's Frontier-Bench v0.1 result, and roughly 30.2% on ARC-AGI 3 against 7.8% for GPT-5.6 Sol — all vendor figures, and all of them measured on a single model you can pin by version. It also routes flagged requests to an older model rather than erroring, which is an operational detail that matters if you have compliance review in the loop.
Fugu Ultra v2 不是那樣,而差異並非表面上的。它是一個協調器——一個小型訓練模型,據報導約有 7B 參數,會讀取你的請求、組建一個代理團隊、依據 Sakana 的 TRINITY 研究成果指派 Thinker、Worker 與 Verifier 角色,並綜合出最終答案。底層模型並未公開,路由決策在設計上也不對外揭露,而就 Ultra 層級而言,池是固定的:你無法選擇將某個供應商排除在外。Sakana 自己對第 2 版的說法是,訓練截止日期移至 2026 年 8 月 28 日,且池的組成有所變動——具體而言是排除 Claude Fable 5、Claude Fable 5.1 與 GPT-6 Astra。
那個排除條款是這次發布中最能說明問題的細節。Fugu Ultra v2 宣稱在基準測試中勝過目前可用的三大最強模型,同時又確認它並未呼叫其中任何一個。無論這個池子裡實際包含什麼,按 Sakana 自己的說法,它都不是市場頂端。這個賣點是:協調勝過能力。這是一個真實的論點,而且在有廉價檢查器的可驗證任務上,這個論點是有證據支持的。這與「這個系統比 Claude Opus 5 更聰明」不是同一種主張,而計分板的編排方式讓這兩者很容易被混淆。
Sakana 選擇的計分板
以下每個數字都來自 Sakana 自己對 Fugu Ultra v2 的評估。Claude Opus 5 的數據是 Sakana 在同一項比較中測得的,除非另有註明。沒有任何獨立第三方重現過 Fugu 那一欄,且截至本文撰寫時,Artificial Analysis 沒有任何 Fugu 模型的條目。
• Chartography — Fugu Ultra v2 48.3 對比 Claude Opus 5 27.3 — 廠商自行回報,差距 21 分
• DeepSWE — Fugu Ultra v2 74.3,相較於同一份表格中並未公布 Claude Opus 5 的數據 — 由廠商回報
• 基準測試排名 — Fugu Ultra v2 在八項中的五項為最佳或並列最佳,八項中的七項排名前二 — 廠商提供
• SWE-bench Verified — 無 Fugu Ultra v2 數據 vs Claude Opus 5 96.0% — Anthropic 回報
• SWE-bench Pro — 未提供 Fugu Ultra v2 的數據,對比 Claude Opus 5 的 79.2% — 由 Anthropic 回報
• ARC-AGI 3 — Fugu Ultra v2 無數據,對比 Claude Opus 5 約 30.2% — Anthropic 提供
把這讀成一種格局,而不是一份排名。Sakana 發布了一份八項基準的榜單,其中它贏了五項,而 Claude Opus 5 最強、公開有紀錄的結果——SWE-bench 的幾列、ARC-AGI 3——根本不在上面。這不是在指控對方心懷不軌;廠商的記分板就是這個樣子,每一家的都一樣。但這意味著你不能從這次發布推論出 Fugu Ultra v2 在軟體工程上勝過 Claude Opus 5,因為這兩個系統在同一批列上被量測的次數不夠多,無法這麼說。在唯一雙方都有出現、且雙方數字都公開的那一列——Chartography——Fugu Ultra v2 宣稱大幅勝出。在最廣泛的推理與程式編寫列上,只有一方公布了結果。

按任務計費,而非按代幣計費
協調器的 Token 帳單並非其每 Token 費率。Fugu Ultra v2 會生成代理,每個代理都會貢獻推理與輸出 Token,而協調器本身則產生合成文字。Sakana 的定價常見問題在此做出了一項真正有用的承諾——您會根據所涉及的頂級模型按單一混合費率計費,且增加代理不會使帳單倍增——這消除了最壞情況下的扇出倍增,而這種倍增會讓樸素的多代理設定變得昂貴。但它並未消除用量。計費的輸出 Token 數量仍然是取決於多少代理執行以及它們寫了多少,而在每百萬 30 美元的情況下,該用量是您無法從定價頁面預測的變數。
Claude Opus 5 的成本結構恰恰相反。一次呼叫、一個模型、一個由你掌控的投入程度調節旋鈕,以及 $25 的輸出費率,非即時工作還可享批次處理五折優惠。你可以預測它。對於每天執行一萬次的工作負載,可預測性的價值遠高於在圖表判讀任務上的基準分數優勢。
This is also where the routing question separates cleanly from the model question. Claude Opus 5 sits on OrcaRouter at Anthropic's list price with 0% markup — the provider's rate passed through, so a vendor price change is live on the same key the same day, with automatic failover across provider paths if one is rate-limited or down. Fugu Ultra v2 is not on our catalogue; it reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, and migrating from an earlier Fugu is a one-line parameter change. If what you actually want is Claude Opus 5's predictability with less single-provider exposure, the router is the answer and the orchestrator is not. If what you want is a system that tries several approaches to a hard problem without you writing the fan-out, Fugu Ultra v2 is the thing that does that — and you should pilot it with token accounting on.
Fugu Ultra v2 真正勝出之處
該給這套系統應有的肯定。Chartography 的成果若能成立,正是單一模型難以偽裝的那種勝利:對結構化文件進行視覺推理,需要有個讀法去嘗試、對照文件檢查,然後再試一次——這正是 Verifier 角色的用途。同樣的邏輯也適用於 Toolathon 和 DeepSWE,在那裡答案是可以被檢驗的,而非靠爭論。Sakana 發表的案例研究也指向同一方向:一次在一台 H100 上歷時十四小時、進行 123 項實驗的 AutoResearch 執行,一個純 Python 的魔術方塊求解器完成了全部 300 個打亂的方塊,而兩個匿名化的前沿基準則完全崩潰,一個機械 CAD 虹膜,能夠實際開合。這些是廠商挑選的範例,且基準經過匿名化,因此它們證明的東西比表面上看起來的要少。但這個模式是一致的,而且指向一個真實的類別:正確性可被檢驗、而困難之處在於持久堅持的任務。
此外還有一個與基準測試無關的結構性論點。Claude Opus 5 是單一供應商的模型,受制於該供應商的定價決策、淘汰時程與政策。Fugu Ultra v2 的設計初衷,就是在上述任何一項出現變化時仍能存續,而 Sakana 也明確表示這正是重點——供應商鎖定、API 撤銷、出口管制。在一個模型可能在幾乎沒有預警下就被撤出某些地區的市場裡,這並非假設。這是一種避險,而避險要付出代價:每百萬輸出詞元多付 5 美元,大致就是這項避險的價格。
裁決
Claude Opus 5 在多數團隊實際上正在做的決策中勝出。它背後有數個月的獨立評估、一個不會在你的文件進行到一半時價格翻倍的 1M token 視窗、128K 的輸出上限、可讓你預測的公開價格、一個努力程度調節旋鈕,讓你只在任務值得時才購買推理,以及第三方在確切基準測試——SWE-bench Verified、SWE-bench Pro、ARC-AGI 3——上的結果,而這些正是代理式與程式開發團隊最重視的。Fugu Ultra v2 則在一個更狹窄也更有趣的問題上勝出:一個擁有較小模型池的協調者,能否在答案可被檢驗的任務上超越旗艦模型。Sakana 自家的計分板顯示,在八列中有五列答案是肯定的,而模型池的排除條件則顯示,這項勝利是在未使用世界上最好的三個模型之下達成的。
If you run verifiable, long-horizon, checkable work and you are outside the EU/EEA, Fugu Ultra v2 is worth a scoped pilot — measure output tokens per completed task, not per token, and set a ceiling before you start. If you need a production path today with evidence behind it and a bill you can forecast, Claude Opus 5 remains the better-evidenced call, and it is one API key away at Anthropic's list price with the provider's rate passed through untouched.

本文中的比較1
根據本文內容識別 · 基準測試:Artificial Analysis · 每日更新
