Fugu Ultra v2 vs DeepSeek V4 Pro のヒーロータイトルカード。サブタイトルは「34倍の出力価格という問い」、ピル型バッジは「DeepSWE 74.3 vs 62.7」「$30.00 vs $0.87 の出力価格」「オープンウェイト vs クローズド」、フッター行は「Sakana AI vs DeepSeek - 2026年9月」、右下隅に OrcaRouter のロゴ。
Engineering & Research

Fugu Ultra v2 vs DeepSeek V4 Pro:出力価格34倍の疑問

著者

Rowan Sterling

公開日

最新モデル · 20すべてのモデルを見る
ベンチマーク:Artificial Analysis · 毎日更新
すべての記事に戻る

Fugu Ultra v2 charges roughly thirty-four times DeepSeek V4 Pro's list price for the text it writes. That is not a typo and it is not a rounding artifact: Sakana AI's new orchestrator lists at $30.00 per million output tokens, Deep​Seek reports $0.87 per million, and on the one benchmark both vendors publish a number for, the gap in performance is nothing like thirty-four times. Sakana claims 74.3 on DeepSWE for Fugu Ultra v2, released September 11, 2026; Deep​Seek reports 62.7 for DeepSeek V4 Pro, the open-weight model it shipped in August. An 11.6-point vendor-reported edge for a 34x output price is the entire comparison in one line — and the honest question is not whether Fugu Ultra v2 is better, but whether it is nineteen dollars per million tokens better, and for whom.

同じ問題に対する2つの答え、正反対の端から組み立てられた

DeepSeek V4 ProとFugu Ultra v2はどちらも、単一の閉鎖的なフロンティアベンダーへの依存という同じ不安への応答であり、その手法はこれ以上ないほど異なっている。

DeepSeek V4 Pro is a model. Launch coverage dates the official version's API rollout — the DeepSeek-V4-Pro-0813 build — to around August 12 to 13, 2026, superseding the April 24 preview. It is a mixture-of-experts system reported at 1.5 to 1.6 trillion total parameters with about 49 billion activated per token, carrying a 1M-token context window and a 384K-token maximum output. It reasons in thinking and non-thinking modes with three effort levels, speaks both the Ope​nAI and Anth​ropic API protocols, and added a native Responses API aimed squarely at Codex-style agent harnesses. Critically, the weights were published under MIT, which means the resilience argument is settled by possession: nobody can revoke your copy.

Fugu Ultra v2はモデルではない。約7Bパラメータと報じられている訓練済みコーディネーターであり、リクエストを非公開モデルのプール全体に分解し、Thinker、Worker、Verifierの役割を割り当て、1つのOpenAI互換エンドポイントの背後で回答を統合する。その回復力に関する主張は法的なものではなくアーキテクチャ上のものだ。プールは差し替え可能なので、単一ベンダーが条件を変更してもシステムは生き残れる。Fugu Maxとともに提供されたバージョン2.0は、トレーニングカットオフを2026年8月28日に移し、そして——特筆すべきことに——Claude Fable 5Claude Fable 5.1GPT-6 Astraをプールから完全に排除しながらも、それらに対する勝利を依然として主張している。

So the comparison is possession versus coordination. Deep​Seek hands you a model you can download, audit, fine-tune and serve yourself at whatever margin your own hardware allows. Sakana hands you a system that decides how to attack each problem, at a rate that funds running several frontier models on your behalf.

A two-column scoreboard for Fugu Ultra v2 vs DeepSeek V4 Pro titled 'Fugu Ultra v2 vs DeepSeek V4 Pro - the scoreboard'. Left column Fugu Ultra v2: Output price $30.00 / 1M, Input price $5.00 / 1M, DeepSWE 74.3, Weights closed, Output ceiling not published, Context 1M. Right column DeepSeek V4 Pro: Output price $0.87 / 1M, Input price $0.435 / 1M, DeepSWE 62.7, Weights open, MIT, Output ceiling 384K, Context 1M. Footer 'Both columns vendor-reported; DeepSeek API rates as reported August 2026, pricing revision pending.', with the OrcaRouter logo bottom-right.

34xが実際にもたらすもの

Benchmarks first, with the sourcing front and centre. Every Deep​Seek figure below is the vendor's own and predates the peak/off-peak pricing revision Deep​Seek announced for mid-August, whose final numbers sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against Deep​Seek's published rate card before you model a budget. Every Fugu figure is Sakana's own, and no independent party has reproduced any of them; there is still no Artificial Analysis entry for a Fugu model.

• DeepSWE — Fugu Ultra v2 74.3(ベンダー報告)対 DeepSeek V4 Pro 62.7(ベンダー報告)

• Terminal Bench 2.1 — Sakana が公表した Fugu Ultra v2 の数値はなし 対 DeepSeek V4 Pro 87.9(ベンダー報告)

• SWE-bench(Vals) — Fugu Ultra v2の数値なし 対 DeepSeek V4 Pro 96.4(ベンダー報告)

• Humanity's Last Exam — Fugu Ultra v2 v2.0 の数値はなく、DeepSeek V4 Pro はツールなしで 42.7、ツールありで 60.0(ベンダー報告)

• GPQA Diamond — DeepSeek V4 Pro の 92.8(ベンダー報告)に対し、Fugu Ultra v2 v2.0 の数値はなし

• 出力価格 — Fugu Ultra v2 は100万あたり30.00ドル、272Kコンテキストを超えると45.00ドルに上昇。一方、DeepSeek V4 Pro はおおよそ100万あたり0.87ドルのリスト価格で、従量制レートはそれより高くなった実績がある

Two caveats on that price row, because the whole comparison rests on it. First, Deep​Seek's figure is the list rate reported in launch coverage and restated in our own model description, but the metered rate on our listing has run at $2.18 per million output and $0.73 per million input — cache behaviour and mix move the effective number, so the honest multiple is somewhere between roughly 14x and roughly 34x depending on how much of your input is cached. Even at the floor of that range, output costs more than a dozen times as much. Second, Deep​Seek has already announced a peak/off-peak pricing revision whose final figures sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against the published rate card before you model a budget.

• 入力価格 — Fugu Ultra v2 は100万トークンあたり$5.00、272Kを超えると倍額。一方、DeepSeek V4 Pro はキャッシュミス時で100万トークンあたり約$0.435、キャッシュヒット時は約120倍安い

• 重み — Fugu Ultra v2 はクローズド、プールは非公開、ルーティングは設計上公開されていない vs DeepSeek V4 Pro は MIT の下でオープン、セルフホスト可能

• コンテキストと出力 — Fugu Ultra v2:1Mコンテキスト、出力上限は非公開 対 DeepSeek V4 Pro:1Mコンテキスト、384K出力上限

The overlap problem is obvious: the two systems barely appear on the same rows. Sakana publishes five wins out of eight benchmarks it selected; Deep​Seek publishes a wider board that includes deep agentic coverage — MCP Atlas 73.6, Toolathlon-Verified 74.1, CyberGym 83.3, NL2Repo 61.5, CorpusQA 62.0 on long context — where Fugu Ultra v2 has no published number at all. The single row you can line up, DeepSWE, is the one where Fugu claims its largest software-engineering advantage. A reader comparing the two releases is comparing two different exams.

冗長性の乗数

出力価格はオーケストレーションに対する線形の税ではない。それは、オーケストレーションが増大させる量に対する乗数である。Fugu Ultra v2 はエージェントを生成し、各エージェントが推論トークンと出力トークンを生成した後、コーディネーターがそれらを統合する。Sakana の料金 FAQ は、真に消費者に優しい約束をしている——プール内の最上位モデルに基づく単一のブレンド料率を支払い、エージェントを追加しても請求額は倍増しない——これにより最悪のケースには上限がつく。ただし、量には上限をつけない。出力トークン100万あたり30ドルでは、単一モデル呼び出しの4倍のテキストを出力するマルチエージェント実行は4倍のコストがかかり、その乗数は34倍高い出力レートと相まって複合する。

DeepSeek V4 Proのコスト構造は、正反対の行動に報いる。入力のキャッシュヒットはミスよりもおよそ2桁安く、そのため長いコンテキストを繰り返し扱う作業——同じリポジトリ、同じ文書セット、同じシステムプロンプト——は、表示価格が示唆するよりも劇的に安くなる。毎ターン同じコンテキストを読み直すエージェントループにとって、実効コストはそこに落ち着く。そしてこれは、どんなオーケストレーターも迂回できない構造的な優位性である。

この2つを合わせて考えると、判断は具体的になる。Fugu Ultra v2がDeepSWEで11.6ポイントの優位をもたらし、1回の呼び出しの3倍の出力トークンを生成するなら、1つのベンチマークで10〜15ポイントの向上を得るために、完了したタスクあたりおよそ100倍のコストを支払っていることになる。そのトレードオフは、限界的な1ポイントがすべてを決める研究や、本質的に難解な工学問題では正当化できるが、1日に10万回実行されるパイプラインでは正当化できない。

A screenshot of the Sakana Fugu product page at sakana.ai/fugu (captured September 11, 2026, English UI), showing the 'Sakana Fugu' wordmark with the subtitle 'One Model to Command Them All', the description that Fugu dynamically orchestrates the world's best models behind a single API, and the notice reading 'Not yet available in the EU/EEA while we work toward compliance with GDPR and EU-specific regulations.'

それぞれが実際に誰向けなのか

次のいずれかに当てはまるなら、DeepSeek V4 Pro を選んでください。ワークロードが大規模でコストに敏感である。オーケストレーターが与える上限より長い出力が必要である。データ所在地の理由から、監査・ファインチューニング・自社ハードウェアでの提供が可能なシステムが必要である。または、価格が観測する数値ではなく計算できる数値である必要がある。オープンな MIT ライセンスはマーケティング上の機能ではない。それは、どちらのシステムが掲げる回復力の主張の中でも最も強力な形である。なぜなら、それはいかなるベンダーの継続的な好意にも依存しないからだ。また、出力100万トークンあたり約$0.87で、実験にコストがかからないほど安い。

Fugu Ultra v2 を選ぶべきなのは、わずかな性能差に何倍ものコストを払う価値があり、作業が長期にわたり、正しさを検証できる場合だ — テストスイートを伴うコーディング、構造化された文書の推論、単一パスなら見逃して出荷してしまう誤りを Verifier 役が捉えられる多段階の分析などがそうだ。Sakana 自身のケーススタディはまさにこの点を指し示している。14時間に及ぶ自律研究実行、そして2つの匿名化されたベースラインが完全にクラッシュしたなかで300通りのスクランブルをすべて解き切ったルービックキューブソルバーだ。ただし、これらのベースラインは匿名化され、ベンダーが選んだものであることを念頭に置いておこう。それは証拠としての強さを弱めるが、虚偽にするものではない。また、一部のチームにとって議論をそこで終わらせる制約も念頭に置いておこう。Fugu Ultra v2 は EU でも EEA でも販売されておらず、Sakana はそれを明言している。

One practical note if you intend to run either. Fugu Ultra v2 is not on OrcaRouter — it comes through Sakana's own OpenAI-compatible API and several third-party platforms, and upgrading from an earlier Fugu is a single-line parameter change. DeepSeek V4 Pro is on OrcaRouter at the provider's list price with 0% markup, which matters here more than usual: Deep​Seek has already announced one pricing revision this cycle, and on a pass-through router a vendor price change is live on the same key the same day rather than after a renegotiation. If you want to benchmark both against your own workload before committing, the cheaper side of the comparison is the one you can call at list price with automatic failover behind it.

A screenshot of the OrcaRouter model page for DeepSeek V4 Pro (deepseek/deepseek-v4-pro, captured September 11, 2026, English UI), showing the Flagship and Featured badges, the '1.6T total / 49B active params, 1M context, top-tier reasoning + agentic tool use' summary, the 1M token context and 384K max output fields, the pricing tiles reading INPUT $0.73 and OUTPUT $2.18 per 1M tokens with p50 TTFT 919ms, and the description stating transparent pricing at $0.44 per 1 million input tokens and $0.87 per 1 million output tokens.

評決

これは経済性では接戦ではなく、能力でも決定的な差があるわけではない。DeepSeek V4 Pro は大差でより優れたデフォルトだ。手元に保持できるオープンウェイト、384K の出力上限、1M のコンテキスト、公開された長文コンテキストのスコア、そして出力あたりの価格は Fugu Ultra v2 の約34分の1である。Fugu Ultra v2 は、限られたクラスの高コストな問題に対してはより優れた道具であり、Sakana の主張——最強の3モデルを除いた固定プールでも、検証可能な作業ではそれらを上回れる——は、現在この分野で最も興味深いアーキテクチャ上の議論であり、同時に最も独立した証拠が少ないものでもある。

現実的な解決策は、どちらかを選ぶことではない。DeepSeek V4 Pro をデフォルトとして使う。安価でオープンで高速だからだ。そして Fugu Ultra v2 は、コストが100倍になってもなお、間違えることのコストより小さいタスクのために取っておく。してはいけないのは、Sakana の8ベンチマークのボードを、高価なオーケストレーターが安価なオープンモデルを置き換えた証拠として読むことだ。両者が共通して載っている唯一の行では、それは出力価格34倍で11.6ポイント上回っている。それがお買い得か罠かは、それをどの問題に当てるかによって完全に決まる。

この記事で比較したモデル1

この記事から検出 · ベンチマーク:Artificial Analysis · 毎日更新