一张“Fugu Max vs GPT-5.6 Sol”的主视觉标题卡,副标题为“该定价声明针对的是该系列的中端”,三个胶囊徽章分别写着“$6.00 vs $20.00 输出”、“40-60% 说法 vs Terra”、“六胜,无比分”,页脚一行写着“Sakana AI vs OpenAI - 2026 年 9 月”,右下角是 OrcaRouter 标志。
Guides & Insights

Fugu Max 对比 GPT-5.6 Sol:该定价主张瞄准的是产品系列的中端

作者

Magnus Corvin

发布日期

最新模型 · 20查看全部模型
基准测试:Artificial Analysis · 每日更新
返回全部文章

Fugu Max's price justification names a G​PT-5.6 model, and it is not the flagship. Sakana AI released Fugu Max on September 11, 2026 at $2.00 per million input tokens and $6.00 per million output tokens, and stated that its output pricing runs 40 to 60 percent below Claude Sonnet 5, GPT-5.6 Terra and Kimi K3. GPT-5.6 Sol — the flagship of the same Ope​nAI family, listed at $4.00 input and $20.00 output per million — does not appear in that sentence. That omission is not an oversight; it is the arithmetic. Sol is more than three times Fugu Max's output rate, a gap that would have made the surrounding "within striking distance of elite models at two to six times lower cost" framing read as a different kind of claim entirely. So this matchup is worth running on Sol's terms rather than the vendor's chosen ones.

正在对照该主张所指向的族进行核查

Sakana 所说的 40% 至 60% 这个数字是可以验证的,因为它点名的模型有已公布的费率。GPT-5.6 Terra 的标价是每百万 token 输入 2.00 美元、输出 12.00 美元。Fugu Max 的标价是输入 2.00 美元、输出 6.00 美元。一对比,就会得出两点。在输出端,6.00 美元对 12.00 美元正好是降价 50%——恰好落在所称 40% 至 60% 区间的正中心,这表明该区间是围绕这一比较划定的,而不是从中发现的。在输入端,两项费率精确到美分都相同。

That is a well-constructed claim, and constructing it well is what a release page is for. But it is calibrated to the middle of the G​PT-5.6 ladder. Terra is the balanced tier of a three-model family — Sol above it for hard complex work, Luna below it for speed and cost. Naming Terra sets the "elite model" benchmark at the middle rung and lets Fugu Max claim half its output price honestly. Naming Sol would have required a different sentence, because the honest version is that Fugu Max costs 70 percent less on output than Sol, and 50 percent less on input — a much larger gap that raises the obvious follow-up question of what capability you are giving up for it.

输入价格 — Fugu Max 每 100 万 token 2.00 美元,而 GPT-5.6 Sol 每 100 万 token 4.00 美元

• 输出价格 — Fugu Max 每 100 万 token 6.00 美元,而 GPT-5.6 Sol 每 100 万 token 20.00 美元

• 缓存输入 — Fugu Max 每 100 万 tokens $0.25,而 GPT-5.6 Sol 的缓存按基础输入费率的折扣计价

• 背景 — Fugu Max 未发布 对比 GPT-5.6 Sol 1M tokens

• 输出上限 — Fugu Max 未公布 vs GPT-5.6 Sol 128K tokens

• 输入模态 — Fugu Max 未公布,对比 GPT-5.6 Sol 的文本、图像和文件

• 能力标签 — Fugu Max 未公布任何标签,而 GPT-5.6 Sol 则涵盖视觉、工具调用、JSON 模式、推理

• 端点 — Fugu Max 单一 OpenAI 兼容接口 vs GPT-5.6 Sol 的 chat completions 与 responses API

• 基准测试数据——Fugu Max 声称取得六项胜利,却未给出任何分数;而 GPT-5.6 Sol 公布了厂商数据,包括 Terminal-Bench 2.0 的 91.9% 和 SWE-bench Pro 的 64.6%

为什么中间阶梯才是值得命名的合适阶梯

Sakana 的这一选择有一个站得住脚的版本,而在提出批评之前,值得先把它说清楚。一个把每项任务路由到最精简但足以胜任的模型的编排器,并不是在旗舰级工作上与旗舰模型竞争。它是在中端工作上与中端模型竞争,并提出以更低成本完成其中的一个子集。如果卖点是“你的大部分流量并不需要 Sol”,那么正确的比较对象就是 Terra,因为当大多数团队不想支付 Sol 的价格时,他们实际会选择的正是 Terra。以客户本来会购买的档位作为衡量基准,比以可用的最昂贵模型作为基准更诚实。

难点在于,同一句话还宣称性能“与精英模型相差无几”。这两半宣传说辞彼此拉扯,方向相反。如果 Fugu Max 是 Terra 的替代品,那么精英模型这个框架就在承担它无法支撑的营销作用,因为 Terra 明确定位并非精英——它被定位为“够用就好”的层级。如果 Fugu Max 确实与 Sol 相差无几,那么价格对比本应针对 Sol 来做,因为在那里差距最大,也最能抬高自己。一篇发布文案在同一段里既作出激进的能力宣称,又进行保守的价格对比,就是在两头占便宜,而读者根本无从判断哪一半才是承重的那一半。

A two-column scoreboard titled "Fugu Max vs GPT-5.6 Sol - the scoreboard". Left column Fugu Max: Output price $6.00 / 1M, Input price $2.00 / 1M, Cached input $0.25 / 1M, Context not published, Modality not published, Benchmarks six wins with no figures. Right column GPT-5.6 Sol: Output price $20.00 / 1M, Input price $4.00 / 1M, Cached input discount on base rate, Context 1M tokens with 128K output, Modality text, image and file, Benchmarks 91.9% on Terminal-Bench 2.0. Footer reading "Fugu Max figures vendor-reported by Sakana AI. Sakana's 40-60% output claim names GPT-5.6 Terra at $2.00 / $12.00, not Sol.", with the OrcaRouter logo bottom-right.

六项基准,没有数字,还有一个理应引起疑问的名字

这次推介的能力部分全靠一份清单支撑。Fugu Max 在 Terminal Bench 2.1、GPQAD、AA-LCR、GDP.pdf、AutomationBench 和 SWEFish 上“取得最佳总体得分”。这些都没有给出任何分数。SWEFish 是 Sakana 自家的内部编码基准,这意味着六项胜利中有一项是在供应商自己编写且尚未公开的测试上评出的。

GPT-5.6 Sol 公布了相关数据。在第三方对比数据中,Sol 在 Terminal-Bench 2.0 上录得 91.9%,在 SWE-bench Pro 上录得 64.6%。请注意终端基准测试上的版本不一致问题:Sol 公布的数字对应的是 Terminal-Bench 2.0,而 Sakana 宣称的胜出是在 2.1 上,这是另一项不同的评测,不能与之并列比较。这一点值得指出,因为这两项主张在标题中看起来相邻,但实际上并不可比。

对于这场对比,相关的比较应该是 Sol 与某个 Fugu 模型在共享评测上的表现。第三方对比页面确实收录了 GPT-5.6 Sol 与基础版 Sakana Fugu 代的对比——Sol 在 Terminal-Bench 2.0 上以 91.9% 对 80.2% 领先,在 SWE-bench Pro 上以 64.6% 对 59% 领先——但那些是较早一代的数据,并非 Fugu Max,而且这些来源明确拒绝给出总体胜者,因为共享基准的覆盖范围太薄。Fugu Max 与 GPT-5.6 Sol 之间没有任何已发布的对比。任何断言存在此类对比的人都是在做外推。

没有任何独立第三方评估过 Fugu Max,Artificial Analysis 中也没有它或任何 Fugu 模型的条目。该发布内容中的每一个数字,包括六项胜绩,都出自厂商。相比之下,Sol 自 2026 年 7 月 9 日起已全面可用,并出现在第三方排行榜中——这种程度的外部审查,Fugu Max 还没有机会经历。

你在每一方实际买到的是什么

配置 GPT-5.6 Sol 不是什么未知数。它具备 1M token 上下文窗口、128K 输出上限,以及文本、图像和文件输入、视觉、工具使用、JSON 模式和推理等已记录的能力,还有两个端点——聊天补全和 responses API——因此现有集成有处可去。在我们所服务的流量中,它显示出 6.68 秒的 p50 首 token 时间,这比我们目录中最快的模型更慢,在你基于它构建交互式产品之前值得了解。

配置 Fugu Max 更像是购买一项服务级别,而不是购买一个模型。你会获得一个 OpenAI 兼容的端点、只需改一行参数即可从更早的 Fugu 迁移到它,以及一个决定调用哪些模型的协调器。从发布页面上,你得不到的是上下文窗口、输出上限、模态列表、延迟数值或基准测试分数。协调器的路由决策按设计不予公开。本次发布中,该池已扩展,纳入了开放权重模型和专用模型,其中 NVIDIA Nemotron 系列是通过与 NVIDIA 的合作被点名列出的,除此之外则未披露。

Sakana 对这种不透明性的公开理由站得住脚——一个可替换的、不依赖特定供应商的资源池,可以防范弃用、价格变动和区域撤出。这是一种真正的对冲,而 $0.25 的缓存输入费率表明,Sakana 预期的是长上下文、高重复的工作负载,这种对冲在这些场景中最重要。但这确实意味着,你购买的规格是对供应链的承诺,而不是对产品的描述。

A screenshot of the Sakana AI release page at sakana.ai/fugu-max-release (captured September 11, 2026, English UI), showing the sakana.ai wordmark, the headline 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' dated September 11, 2026, the opening copy arguing that the frontier which matters is two-dimensional with capability on one axis and cost on the other, a 'Try Sakana Fugu Max and Fugu Ultra v2' link, and a Pareto chart plotting performance against cost with a red 'Fugu Max' point at the low-cost end, a red 'Fugu Ultra' point at the top, a shaded 'Frontier formed by Fugu models' region and grey points labelled 'Frontier formed by single models'.

如何运行测试

解决这个问题的正确方式不是去读任何一方的宣传页面。从你自己的流量中取样——几百个真实请求,最好包含你们团队目前会升级处理的那些——开启 token 统计,把它们分别打到两个端点上。按每完成一个任务所消耗的 token 来计成本,而不是按每次调用消耗的 token,并且在开始之前就设定好输出预算,这样即使出现扇出也不会让你措手不及;质量则用第二个人是否认可该答案可以接受来衡量。这个实验一周之内就能给出答案,而且花费很少。

The friction is that Fugu Max is not on OrcaRouter. It reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, which means a second integration, a second credential set and a second invoice. GPT-5.6 Sol is on our catalogue at Ope​nAI's list price with 0% markup — the provider's rate passed through rather than marked up, so a change to Ope​nAI's pricing is live on your existing key the same day, with automatic failover across provider paths when one is rate-limited or unavailable. It sits behind the same key as more than 200 other models, so when the pilot ends you keep the integration and change one string.

底线

GPT-5.6 Sol 在 Sakana 拒绝做出的那项对比中胜出。它的输入成本是两倍,输出成本超过三倍,作为交换,你得到的是一个公开的上下文窗口、128K 的输出上限、有文档记录的多模态输入、具名的能力标签、两个端点,以及第三方已有数月时间审视的厂商基准测试。如果你的工作正是 Sol 为之而生的那类——深度多步推理、大规模软件工程、长周期智能体任务——那么价格差就是为买到一件你能明确指定之物所付出的代价。

Fugu Max 赢下的赌注更窄,却真正有趣,而且价格便宜得多:输出价格比 GPT-5.6 Terra 低 50%,比 Sol 低 70%,缓存费率在两张费率表中都是最低的。如果你的流量规模大、纯文本、可核查,且大多不需要旗舰级推理,那么一个把请求路由到最精简的够用模型的编排器就是一种自洽的设计,价值正落在 6.00 美元的输出费率上。按发布方自己的说法照单全收,把它当作 Terra 的替代品而非 Sol 的替代品来做试点——如果你想要旗舰模型,同时希望供应商的标价原样传递,GPT-5.6 Sol 在 OrcaRouter 上只差一个密钥。

A screenshot of the OrcaRouter model page for GPT-5.6 Sol (openai/gpt-5.6-sol, captured September 11, 2026, English UI), showing the Featured badge, the openai/gpt-5.6-sol model ID, the Vision, Tools, JSON and Reasoning tags, the listing date 2026-07-09, the /v1/chat/completions and /v1/responses endpoints, the pricing tiles reading INPUT $4.00 and OUTPUT $20.00 per 1M tokens with p50 TTFT 6.68s and 153.0M tokens of 7-day traffic, the 1M token context with 128K max output and text + image + file input, and the OpenAI-compatible Python sample pointing at api.orcarouter.ai/v1.

本文中的对比1

根据本文内容识别 · 基准测试:Artificial Analysis · 每日更新