Fugu Ultra v2 vs DeepSeek V4 Pro의 히어로 타이틀 카드로, 부제는 '34배 출력 가격에 관한 질문', 알약 배지는 'DeepSWE 74.3 vs 62.7', '$30.00 vs $0.87 출력', '오픈 웨이트 vs 클로즈드'를 표시하며, 하단 문구는 'Sakana AI vs DeepSeek - 2026년 9월', 오른쪽 아래 모서리에는 OrcaRouter 로고가 있습니다.
Engineering & Research

Fugu Ultra v2 vs DeepSeek V4 Pro: 34배 출력 가격에 대한 질문

작성자

Rowan Sterling

게시일

최신 모델 · 20모든 모델 보기
벤치마크: Artificial Analysis · 매일 업데이트
모든 게시물로 돌아가기

Fugu Ultra v2 charges roughly thirty-four times DeepSeek V4 Pro's list price for the text it writes. That is not a typo and it is not a rounding artifact: Sakana AI's new orchestrator lists at $30.00 per million output tokens, Deep​Seek reports $0.87 per million, and on the one benchmark both vendors publish a number for, the gap in performance is nothing like thirty-four times. Sakana claims 74.3 on DeepSWE for Fugu Ultra v2, released September 11, 2026; Deep​Seek reports 62.7 for DeepSeek V4 Pro, the open-weight model it shipped in August. An 11.6-point vendor-reported edge for a 34x output price is the entire comparison in one line — and the honest question is not whether Fugu Ultra v2 is better, but whether it is nineteen dollars per million tokens better, and for whom.

같은 문제에 대한 두 가지 답, 서로 반대쪽 끝에서 만들어졌다

DeepSeek V4 Pro와 Fugu Ultra v2는 모두 동일한 불안, 즉 단일 폐쇄형 프런티어 벤더에 대한 의존에 대한 대응이지만, 방식에 있어서는 그 차이가 극에 달한다.

DeepSeek V4 Pro is a model. Launch coverage dates the official version's API rollout — the DeepSeek-V4-Pro-0813 build — to around August 12 to 13, 2026, superseding the April 24 preview. It is a mixture-of-experts system reported at 1.5 to 1.6 trillion total parameters with about 49 billion activated per token, carrying a 1M-token context window and a 384K-token maximum output. It reasons in thinking and non-thinking modes with three effort levels, speaks both the Ope​nAI and Anth​ropic API protocols, and added a native Responses API aimed squarely at Codex-style agent harnesses. Critically, the weights were published under MIT, which means the resilience argument is settled by possession: nobody can revoke your copy.

Fugu Ultra v2는 모델이 아니다. 그것은 약 7B 매개변수로 알려진 훈련된 코디네이터로, 비공개 모델 풀에 걸쳐 요청을 분해하고 Thinker, Worker, Verifier 역할을 할당하며, 하나의 OpenAI 호환 엔드포인트 뒤에서 답변을 종합한다. 그것의 복원력 논거는 법적이라기보다 아키텍처적이다: 풀이 교체 가능하므로 시스템은 어느 한 공급업체가 약관을 바꾸더라도 버틸 수 있다. Fugu Max와 함께 출시된 버전 2.0은 학습 컷오프를 2026년 8월 28일로 옮겼고 — 특히 — Claude Fable 5, Claude Fable 5.1GPT-6 Astra를 풀에서 완전히 제거했으면서도 여전히 그들에 대한 승리를 주장한다.

So the comparison is possession versus coordination. Deep​Seek hands you a model you can download, audit, fine-tune and serve yourself at whatever margin your own hardware allows. Sakana hands you a system that decides how to attack each problem, at a rate that funds running several frontier models on your behalf.

A two-column scoreboard for Fugu Ultra v2 vs DeepSeek V4 Pro titled 'Fugu Ultra v2 vs DeepSeek V4 Pro - the scoreboard'. Left column Fugu Ultra v2: Output price $30.00 / 1M, Input price $5.00 / 1M, DeepSWE 74.3, Weights closed, Output ceiling not published, Context 1M. Right column DeepSeek V4 Pro: Output price $0.87 / 1M, Input price $0.435 / 1M, DeepSWE 62.7, Weights open, MIT, Output ceiling 384K, Context 1M. Footer 'Both columns vendor-reported; DeepSeek API rates as reported August 2026, pricing revision pending.', with the OrcaRouter logo bottom-right.

34x가 실제로 사주는 것

Benchmarks first, with the sourcing front and centre. Every Deep​Seek figure below is the vendor's own and predates the peak/off-peak pricing revision Deep​Seek announced for mid-August, whose final numbers sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against Deep​Seek's published rate card before you model a budget. Every Fugu figure is Sakana's own, and no independent party has reproduced any of them; there is still no Artificial Analysis entry for a Fugu model.

• DeepSWE — Fugu Ultra v2 74.3(공급업체 보고) vs DeepSeek V4 Pro 62.7(공급업체 보고)

• Terminal Bench 2.1 — Sakana가 발표한 Fugu Ultra v2 수치 없음 vs DeepSeek V4 Pro 87.9(공급업체 보고)

• SWE-bench (Vals) — Fugu Ultra v2 수치 없음 vs DeepSeek V4 Pro 96.4(공급업체 보고)

• Humanity's Last Exam — Fugu Ultra v2 v2.0 수치는 없음 vs DeepSeek V4 Pro는 도구 없이 42.7, 도구 사용 시 60.0 (공급업체 보고)

• GPQA Diamond — Fugu Ultra v2 v2.0 수치 없음 vs DeepSeek V4 Pro 92.8 (공급업체 보고)

• 출력 가격 — Fugu Ultra v2는 1M당 $30.00이며, 272K 컨텍스트를 초과하면 $45.00로 오르고, DeepSeek V4 Pro는 약 1M당 $0.87 정가인데, 종량제 요율은 더 높게 적용되어 왔습니다

Two caveats on that price row, because the whole comparison rests on it. First, Deep​Seek's figure is the list rate reported in launch coverage and restated in our own model description, but the metered rate on our listing has run at $2.18 per million output and $0.73 per million input — cache behaviour and mix move the effective number, so the honest multiple is somewhere between roughly 14x and roughly 34x depending on how much of your input is cached. Even at the floor of that range, output costs more than a dozen times as much. Second, Deep​Seek has already announced a peak/off-peak pricing revision whose final figures sources still disagree on — one report puts peak output near ¥27 per million, another near ¥12, against a ¥6 base. Verify against the published rate card before you model a budget.

• 입력 가격 — Fugu Ultra v2는 1M당 $5.00이며, 272K를 초과하면 두 배가 되고, DeepSeek V4 Pro는 캐시 미스 시 1M당 약 $0.435로, 캐시 적중 시 약 120배 더 저렴합니다

• 가중치 — Fugu Ultra v2는 비공개, 풀 미공개, 라우팅은 설계상 노출되지 않음 vs DeepSeek V4 Pro는 MIT 라이선스로 공개, 자체 호스팅 가능

• 맥락 및 출력 — Fugu Ultra v2 1M 컨텍스트, 출력 상한 미공개 vs DeepSeek V4 Pro 1M 컨텍스트, 384K 출력 상한

The overlap problem is obvious: the two systems barely appear on the same rows. Sakana publishes five wins out of eight benchmarks it selected; Deep​Seek publishes a wider board that includes deep agentic coverage — MCP Atlas 73.6, Toolathlon-Verified 74.1, CyberGym 83.3, NL2Repo 61.5, CorpusQA 62.0 on long context — where Fugu Ultra v2 has no published number at all. The single row you can line up, DeepSWE, is the one where Fugu claims its largest software-engineering advantage. A reader comparing the two releases is comparing two different exams.

상세도 승수

출력 가격은 오케스트레이션에 매기는 선형 세금이 아니라, 오케스트레이션이 늘리는 양에 곱해지는 승수다. Fugu Ultra v2는 에이전트를 생성하는데, 각 에이전트는 추론 및 출력 토큰을 만들어 내고, 그다음 코디네이터가 이를 종합한다. Sakana의 가격 FAQ는 진정으로 소비자 친화적인 약속을 한다 — 풀에서 최상위 모델을 기준으로 하나의 혼합 요금을 지불하며, 에이전트를 추가해도 청구서가 배로 늘지 않는다는 것 — 이는 최악의 경우를 상한으로 막아 준다. 그것이 볼륨을 상한하지는 않는다. 출력 토큰 100만 개당 $30이라면, 단일 모델 호출보다 네 배 많은 텍스트를 내보내는 멀티 에이전트 실행은 비용이 네 배 더 든다. 그리고 그 승수는 34배 더 높은 출력 속도와 맞물려 복리로 불어난다.

DeepSeek V4 Pro의 비용 구조는 정반대 행동에 보상을 줍니다. 입력에 대한 캐시 히트는 미스보다 대략 100배 저렴하므로, 반복적인 장문 컨텍스트 작업 — 같은 저장소, 같은 문서 세트, 같은 시스템 프롬프트 — 은 표시 요금이 시사하는 것보다 훨씬 저렴합니다. 매 턴마다 같은 컨텍스트를 다시 읽는 에이전트 루프의 경우, 실효 비용이 바로 그 지점에서 결정되며, 이는 어떤 오케스트레이터도 우회할 수 없는 구조적 이점입니다.

두 가지를 함께 놓으면 결정은 구체적으로 드러난다. Fugu Ultra v2가 DeepSWE에서 11.6점 우위를 제공하고 단일 호출의 세 배에 달하는 출력 토큰을 내보낸다면, 하나의 벤치마크에서 10~15점 향상을 얻기 위해 완료된 작업당 대략 100배 더 많은 비용을 지불하는 셈이다. 그런 맞바꿈은 그 한 점의 차이가 전부인 연구와 환원 불가능하게 어려운 엔지니어링 문제에서는 정당화될 수 있지만, 하루에 10만 번 실행되는 파이프라인에서는 정당화될 수 없다.

A screenshot of the Sakana Fugu product page at sakana.ai/fugu (captured September 11, 2026, English UI), showing the 'Sakana Fugu' wordmark with the subtitle 'One Model to Command Them All', the description that Fugu dynamically orchestrates the world's best models behind a single API, and the notice reading 'Not yet available in the EU/EEA while we work toward compliance with GDPR and EU-specific regulations.'

각각이 실제로 누구를 위한 것인지

다음 중 하나라도 해당한다면 DeepSeek V4 Pro를 선택하세요: 워크로드가 대용량이고 비용에 민감하다; 오케스트레이터가 제공하는 상한보다 더 긴 출력이 필요하다; 데이터 상주 요건 때문에 직접 감사하고, 미세 조정하거나, 자체 하드웨어에서 서빙할 수 있는 시스템이 필요하다; 또는 가격이 관측하는 숫자가 아니라 계산할 수 있는 숫자이기를 원한다. 오픈 MIT 라이선스는 마케팅 기능이 아니다 — 그것은 두 시스템 중 어느 쪽이든 내세우는 복원력 논변의 가장 강력한 형태다. 왜냐하면 그것은 어떤 공급업체의 지속적인 선의에 의존하지 않기 때문이다. 또한 백만 출력당 약 $0.87로, 실험이 아무 비용도 들지 않을 만큼 저렴하다.

한계적 이점이 몇 배의 가치가 있고, 작업이 장기적이며, 정확성을 검증할 수 있다면 Fugu Ultra v2를 선택하라 — 테스트 스위트를 갖춘 코딩, 구조화된 문서 추론, 단일 패스로는 그대로 내보냈을 오류를 Verifier 역할이 잡아낼 수 있는 다단계 분석. Sakana 자체의 사례 연구가 바로 이것을 가리킨다: 14시간의 자율 연구 실행, 두 개의 익명화된 베이스라인이 완전히 크래시한 가운데 300개의 스크램블을 모두 완료한 루빅스 큐브 솔버. 그 베이스라인들은 익명화되었고 벤더가 선택한 것이라는 점을 명심하라. 이는 증거로서 약화시키지만 거짓이 되게 하지는 않는다. 또한 일부 팀에게는 논의를 끝내는 제약을 명심하라: Fugu Ultra v2는 EU나 EEA에서 판매되지 않으며, Sakana는 이를 분명히 밝힌다.

One practical note if you intend to run either. Fugu Ultra v2 is not on OrcaRouter — it comes through Sakana's own OpenAI-compatible API and several third-party platforms, and upgrading from an earlier Fugu is a single-line parameter change. DeepSeek V4 Pro is on OrcaRouter at the provider's list price with 0% markup, which matters here more than usual: Deep​Seek has already announced one pricing revision this cycle, and on a pass-through router a vendor price change is live on the same key the same day rather than after a renegotiation. If you want to benchmark both against your own workload before committing, the cheaper side of the comparison is the one you can call at list price with automatic failover behind it.

A screenshot of the OrcaRouter model page for DeepSeek V4 Pro (deepseek/deepseek-v4-pro, captured September 11, 2026, English UI), showing the Flagship and Featured badges, the '1.6T total / 49B active params, 1M context, top-tier reasoning + agentic tool use' summary, the 1M token context and 384K max output fields, the pricing tiles reading INPUT $0.73 and OUTPUT $2.18 per 1M tokens with p50 TTFT 919ms, and the description stating transparent pricing at $0.44 per 1 million input tokens and $0.87 per 1 million output tokens.

평결

경제성에서 접전이 아니며, 역량에서도 결정적이지 않다. DeepSeek V4 Pro는 훨씬 더 나은 기본값이다: 직접 보유할 수 있는 오픈 웨이트, 384K 출력 상한, 1M 컨텍스트, 공개된 장문맥 점수, 그리고 출력 기준으로 Fugu Ultra v2 가격의 약 1/34에 불과한 가격. Fugu Ultra v2는 좁은 부류의 값비싼 문제에 더 나은 도구이며, Sakana의 주장 — 가장 강력한 세 모델을 제외한 고정 풀이 검증 가능한 작업에서 여전히 그들을 이길 수 있다는 — 은 현재 이 분야에서 가장 흥미로운 아키텍처 논거이자, 독립적 증거가 가장 적게 뒷받침되는 논거다.

실용적인 해법은 둘 중 하나를 고르는 것이 아니다. DeepSeek V4 Pro를 기본으로 실행하라. 저렴하고 개방적이며 빠르기 때문이다. 그리고 Fugu Ultra v2는 100배의 비용 증가가 틀렸을 때의 비용보다 여전히 작은 작업을 위해 예비로 남겨 두어라. 해서는 안 되는 것은 Sakana의 8개 벤치마크 표를 비싼 오케스트레이터가 저렴한 오픈 모델을 대체했다는 증거로 읽는 것이다. 두 모델이 공유하는 단 하나의 행에서, 그것은 출력 가격 34배로 11.6점 앞선다. 그것이 이득인지 함정인지는 전적으로 그것을 어느 문제에 겨냥하느냐에 달려 있다.

이 글에서 비교한 모델1

이 글에서 자동 인식 · 벤치마크: Artificial Analysis · 매일 업데이트