
Fugu Ultra v2 vs Claude Opus 5:相同的输入价格,两种不同的押注
- deepseek新DeepSeek: DeepSeek V4.1 Flash2026-09-1040智能
- openai新OpenAI: GPT-6 Astra2026-09-0453智能77代码
- google新Google: Gemini 3.8 Flash2026-09-0241智能76代码
- qwen新Qwen: Qwen3.8 Max (0902)2026-09-0240智能72代码
- anthropic新Anthropic: Claude Fable 5.12026-09-0153智能82代码
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 每百万 tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642智能72代码
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 每百万 tokens
- z-aiZ.ai: GLM 5.32026-08-1845智能75代码
- obsidianQwen3.8 27B2026-08-1534智能68代码
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236智能69代码
- grokSpaceXAI: Grok 4.62026-08-1244智能77代码
- metaMeta: Muse Spark 1.22026-08-0540智能72代码
- qwenQwen: Qwen3.8 Max2026-08-0340智能72代码
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135智能69代码
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 每百万 tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451智能78代码
- googleGoogle: Gemini 3.6 Flash2026-07-2134智能69代码
Fugu Ultra v2 and Claude Opus 5 cost exactly the same amount to ask a question and wildly different amounts to get an answer out of. Both list at $5.00 per million input tokens. Then Fugu Ultra v2 charges $30.00 per million output tokens against Claude Opus 5's $25.00 — and the gap between those two numbers is not a rounding difference, because the two systems emit very different quantities of text to reach an answer. Sakana AI shipped Fugu Ultra v2 on September 11, 2026 as an orchestration system that coordinates a pool of other labs' models behind one OpenAI-compatible endpoint; Anthropic's Claude Opus 5, live since late July 2026, is a single model. That difference decides most of this comparison before a benchmark is consulted, and Sakana's own scoreboard has an awkward row in it: the vendor reports 48.3 on Chartography against Claude Opus 5's 27.3, which is the largest margin in the release and also the number with the least outside evidence behind it.
塑造一切的巧合
先看这两款产品的共同点,因为它比标价更重要。
• 输入价格 — Fugu Ultra v2 每 100 万 tokens 5.00 美元,对比 Claude Opus 5 每 100 万 tokens 5.00 美元
• 输出价格 — Fugu Ultra v2 每 100 万 token 30.00 美元,对比 Claude Opus 5 每 100 万 token 25.00 美元
• 缓存输入 — Fugu Ultra v2 每百万 $0.50,高于基础费率;相比之下,Claude Opus 5 将缓存定价为对缓存输入最高 90% 的折扣
• 上下文 — Fugu Ultra v2 为 100 万 tokens,超过 272K 后费率翻倍;而 Claude Opus 5 为 100 万 tokens,费率固定
• 输出上限 — Fugu Ultra v2 未以单一数字公布,对比 Claude Opus 5 的 128K tokens
• 可用性 — Fugu Ultra v2 未在欧盟/欧洲经济区销售,而 Claude Opus 5 依据标准 API 条款普遍可用
• 证据 — Fugu Ultra v2 厂商报告的基准测试,尚无独立评估;对比 Claude Opus 5 的厂商基准测试,外加数月第三方排行榜排名。
最后那一行才是这场对比真正发生转折的地方。在相同的输入价格下,你要在两种选择之间做出取舍:一个是其得分只能靠你盲目信任的系统,另一个是其得分已被他人复现的模型。272K 的断崖式涨价更是雪上加霜:一旦超过这一阈值,Fugu Ultra v2 的输入费率翻倍至 $10,输出费率升至 $45,而 Claude Opus 5 的费率则纹丝不动。对于长上下文智能体运行——正是 Fugu Ultra v2 所针对的那类工作负载——更低的标价所带来的价格优势,恰恰在这个系统本应大放异彩的地方消失殆尽。

每一个到底是什么
Claude Opus 5 is Anthropic's production flagship, and its design is legible from the outside. It runs adaptive thinking by default, deciding how long to reason before it answers, with an explicit effort dial from low to max when you want to force deeper deliberation and pay for it. It carries a 1M-token context window, a 128K output ceiling, vision, tool use and JSON mode. Anthropic reports 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro, more than doubling its predecessor's Frontier-Bench v0.1 result, and roughly 30.2% on ARC-AGI 3 against 7.8% for GPT-5.6 Sol — all vendor figures, and all of them measured on a single model you can pin by version. It also routes flagged requests to an older model rather than erroring, which is an operational detail that matters if you have compliance review in the loop.
Fugu Ultra v2 并非如此,而且这一差别并不只是表面上的。它是一个协调器——一个小型训练模型,据报道约有 7B 参数,会读取你的请求,组建一个 agent 团队,根据 Sakana 的 TRINITY 工作分配 Thinker、Worker 和 Verifier 角色,并综合出最终答案。底层模型未披露,路由决策在设计上就不对外暴露,而对于 Ultra 档,模型池是固定的:你无法将某个供应商排除在外。Sakana 自己对第 2 版的表述是,训练截止时间已移至 2026 年 8 月 28 日,模型池构成也发生了变化——具体就是将 Claude Fable 5、Claude Fable 5.1 和 GPT-6 Astra 排除在外。
那次排除是本次发布中最能说明问题的细节。Fugu Ultra v2 声称在基准测试中胜过现有的三个最强模型,同时确认自己并未调用其中任何一个。无论这个池子里究竟包含什么,按照 Sakana 自己的说法,它都不是市场的顶尖。其卖点是协调胜过能力。这是一个真实的论点,而且在有廉价检查器的可验证任务上,这是一个有证据支持的论点。这与“这个系统比 Claude Opus 5 更聪明”不是同一种主张,而记分牌的设计方式让两者很容易被混淆。
Sakana 选择的记分牌
下文中的每一个数字都来自 Sakana 对 Fugu Ultra v2 自己的评估。Claude Opus 5 的数据是由 Sakana 在同一比较中测得的,除非另有说明。没有任何独立第三方复现过 Fugu 那一列,而且截至本文撰写时,Artificial Analysis 中还没有任何 Fugu 模型的条目。
• Chartography — Fugu Ultra v2 48.3 对比 Claude Opus 5 27.3 — 厂商报告,相差21分
• DeepSWE — Fugu Ultra v2 74.3,对比同一表格中未公布的 Claude Opus 5 数据 — 厂商报告
• 基准排名 — Fugu Ultra v2 在八项中五项最佳或并列最佳,在八项中七项位列前二 — 由厂商提供
• SWE-bench Verified — 没有 Fugu Ultra v2 的数据,对比 Claude Opus 5 的 96.0% — Anthropic 报告
• SWE-bench Pro — 未提供 Fugu Ultra v2 的数值,而 Claude Opus 5 为 79.2% — Anthropic 报告
• ARC-AGI 3 — Fugu Ultra v2 无数据,对比 Claude Opus 5 约 30.2% — 据 Anthropic 报告
把它当作一种格局而非一份排名来解读。Sakana 发布了一个包含八项基准的榜单,其中它赢了五项,而 Claude Opus 5 最有力的公开记录结果——SWE-bench 的各行、ARC-AGI 3——根本不在榜单上。这并不是在指控对方动机不纯;这恰恰就是厂商计分板的模样,每一家的都是如此。但这意味着你不能从这次发布中得出 Fugu Ultra v2 在软件工程上击败 Claude Opus 5 的结论,因为这两个系统在相同项目上被测量的次数还不够多,无法下此判断。在唯一一行两者都出现、且双方数字都公开的项目——Chartography——上,Fugu Ultra v2 声称取得了大幅领先。而在最广泛的推理与编程项目上,只有一方公布了结果。

按任务计费,而非按令牌计费
编排器的 token 账单并不等于它的每 token 费率。Fugu Ultra v2 会派生多个 agent,每个 agent 都会贡献推理 token 和输出 token,而协调器本身还会生成综合文本。Sakana 的定价 FAQ 在这里做出了一个确实有用的承诺——你会根据所涉及的顶级模型按一个混合费率计费,增加 agent 并不会让账单成倍增长——这消除了使朴素多 agent 配置变得昂贵的最坏情况扇出倍增。但它并没有消除体量。计费的输出 token 数量仍然取决于运行了多少个 agent 以及它们写了多少,而在每百万 $30 的价格下,这个体量是你无法从定价页面上预测的变量。
Claude Opus 5 的成本形态正好相反。一次调用、一个模型、一个你可控制的 effort 调节旋钮,以及 25 美元的输出费率,非实时工作还可使用五折的批处理。你可以预测它。对于每天运行一万次的工作负载,可预测性的价值远高于在一项读图表任务上的基准测试优势。
This is also where the routing question separates cleanly from the model question. Claude Opus 5 sits on OrcaRouter at Anthropic's list price with 0% markup — the provider's rate passed through, so a vendor price change is live on the same key the same day, with automatic failover across provider paths if one is rate-limited or down. Fugu Ultra v2 is not on our catalogue; it reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, and migrating from an earlier Fugu is a one-line parameter change. If what you actually want is Claude Opus 5's predictability with less single-provider exposure, the router is the answer and the orchestrator is not. If what you want is a system that tries several approaches to a hard problem without you writing the fan-out, Fugu Ultra v2 is the thing that does that — and you should pilot it with token accounting on.
Fugu Ultra v2 真正胜出之处
该给这套系统应有的认可。Chartography 的结果如果站得住,就是单一模型难以伪造的那类胜利:对结构化文档进行视觉推理,奖励的是尝试一种读法、对照文档加以核查、然后再试一次——这正是 Verifier 角色存在的意义。同样的逻辑也适用于 Toolathon 和 DeepSWE,在这些任务里,答案可以被检验,而不是靠争论。Sakana 公布的案例研究也指向同一方向:一次在单块 H100 上历时十四小时、包含 123 次实验的 AutoResearch 运行;一个纯 Python 魔方求解器完成了全部 300 个打乱的魔方,而两个匿名化的前沿基线则彻底崩溃;一个真正能开合的机械 CAD 光圈。这些是厂商挑选的例子,基线也经过匿名化处理,所以它们所能证明的没有表面上那么多。但这个模式是一致的,并且指向一个真实的类别:正确性可检验、而难点在于能否坚持下去的任务。
还有一个与基准测试无关的结构性论点。Claude Opus 5 是某一家供应商的模型,受制于该供应商的定价决策、弃用时间表和政策。Fugu Ultra v2 的设计目标,是在上述任何一项发生变化时依然能够存活,而 Sakana 明确表示这正是其用意所在——供应商锁定、API 撤销、出口管制。在一个模型可能在几乎没有预警的情况下从某些地区被下架的市场里,这并非假设。这是一种对冲,而对冲是有成本的:每百万输出 token 多花 5 美元,大致就是这种对冲的定价。
裁决
Claude Opus 5 赢下的是大多数团队实际正在做的决定。它背后有数月的独立评测,一个不会在你文档处理到一半时价格翻倍的 100 万 token 窗口,128K 的输出上限,可预测的公示价格,一个让你只在任务值得时才购买推理的投入程度调节旋钮,以及在智能体与编程团队最关心的那些基准——SWE-bench Verified、SWE-bench Pro、ARC-AGI 3——上的第三方结果。Fugu Ultra v2 赢下的则是一个更狭窄也更有趣的问题:在答案可被核验的任务上,一个模型池更小的协调者能否胜过旗舰模型。Sakana 自己的记分板显示,八行中有五行答案是肯定的,而模型池的排除条件则说明,这一胜利是在没有全球最好的三个模型的情况下取得的。
If you run verifiable, long-horizon, checkable work and you are outside the EU/EEA, Fugu Ultra v2 is worth a scoped pilot — measure output tokens per completed task, not per token, and set a ceiling before you start. If you need a production path today with evidence behind it and a bill you can forecast, Claude Opus 5 remains the better-evidenced call, and it is one API key away at Anthropic's list price with the provider's rate passed through untouched.

本文中的对比1
根据本文内容识别 · 基准测试:Artificial Analysis · 每日更新
