
Fugu Max vs Gemini 3.1 Pro: 동일한 입력 가격, 절반의 출력 가격, 그리고 아무도 볼 수 없는 풀
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040지능
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453지능77코딩
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241지능76코딩
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240지능72코딩
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153지능82코딩
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 100만 토큰당
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642지능72코딩
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 100만 토큰당
- z-aiZ.ai: GLM 5.32026-08-1845지능75코딩
- obsidianQwen3.8 27B2026-08-1534지능68코딩
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236지능69코딩
- grokSpaceXAI: Grok 4.62026-08-1244지능77코딩
- metaMeta: Muse Spark 1.22026-08-0540지능72코딩
- qwenQwen: Qwen3.8 Max2026-08-0340지능72코딩
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135지능69코딩
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 100만 토큰당
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451지능78코딩
- googleGoogle: Gemini 3.6 Flash2026-07-2134지능69코딩
Fugu Max and Gemini 3.1 Pro charge identically to be asked a question. Both list at $2.00 per million input tokens. They part company on the way out: Sakana AI's Fugu Max writes for $6.00 per million output tokens, and Google's Gemini 3.1 Pro writes for $12.00. Sakana released Fugu Max on September 11, 2026 as a coordinator that routes each task to the cheapest sufficient model in a pool it does not disclose. Gemini 3.1 Pro is a named model — callable by ID, versioned, multimodal — which is the thing an orchestrator is designed not to be. So the interesting question here is not which one is better at $2.00 input. It is whether a comparison between a coordinator and a model is even the right shape, given that Sakana's earlier Fugu generation reportedly counted Gemini 3.1 Pro among the models it routed to.
풀 질문을 정확히 서술하면
When the first Fugu Ultra shipped in June 2026, reporting on its architecture described a pool that included Gemini 3.1 Pro alongside models from other frontier labs. Sakana has not published the membership of Fugu Max's pool. The release says the pool expanded to include an unprecedented number of open-weights and specialised models, naming the NVIDIA Nemotron family through an NVIDIA collaboration, and stops there. It also does not say Gemini 3.1 Pro was removed.
That leaves two possibilities and no public way to choose between them. If Gemini 3.1 Pro is in the Fugu Max pool, then this matchup is partly a "versus" and partly a "through" — a Fugu Max answer on a hard reasoning task may be a Gemini answer that a coordinator selected, checked, and rewrote, in which case the $6.00 output rate is buying Gemini's reasoning plus orchestration overhead at a discount to calling Gemini directly. If it is not in the pool, the two products are genuinely different instruments and the price difference reflects different underlying capability.
릴리스에는 이 문제를 해결하는 내용이 없다. 그것이 어느 것인지 말해 주는 사람은 그것을 읽은 것이 아니라 추론한 것이다. 자신 있게 말할 수 있는 것은 더 좁지만 여전히 유용하다: 오케스트레이터가 공개한 점수는 코디네이터가 아니라 풀에 속하며, 구성원이 공개되지 않은 풀은 그 점수들을 귀속시키는 것을 불가능하게 만든다.
• Input price — Fugu Max $2.00 per 1M tokens vs Gemini 3.1 Pro $2.00 per 1M tokens
• Output price — Fugu Max $6.00 per 1M tokens vs Gemini 3.1 Pro $12.00 per 1M tokens
• Cached input — Fugu Max $0.25 per 1M tokens vs Gemini 3.1 Pro caching offered at a discount on the base input rate
• Context — Fugu Max not published vs Gemini 3.1 Pro 1M tokens
• Output ceiling — Fugu Max not published vs Gemini 3.1 Pro 65K tokens
• Input modality — Fugu Max not published vs Gemini 3.1 Pro audio, file, image, text and video
• Output modality — Fugu Max not published vs Gemini 3.1 Pro text, with a separate image-generation line in the 3.1 family
• Benchmark figures — Fugu Max six wins asserted with no scores vs Gemini 3.1 Pro Google-reported ARC-AGI-2 77.1%, GPQA Diamond 94.3%, SWE-bench Verified 80.6%, BrowseComp 85.9%
모델이 약속할 수 있지만 코디네이터는 할 수 없는 것
The single largest practical difference in that list is not the output price. It is the modality row. Gemini 3.1 Pro accepts audio, video and images as input and is documented as such, which makes it usable for work that has nothing to do with text — reading a screen recording, transcribing and reasoning over a call, parsing a scanned form. Fugu Max publishes no modality specification at all. It may well handle some of those inputs, depending on what its pool contains and whether the coordinator forwards multimodal content or flattens it to text first. Nobody outside Sakana knows, and that is a different kind of uncertainty from "we haven't benchmarked it yet." It means you cannot decide fitness for a multimodal workload from the documentation; you have to test it.
There is a second asymmetry in the output ceiling. Gemini 3.1 Pro documents 65K maximum output tokens. Fugu Max documents nothing. For a coordinator, the effective ceiling is the minimum across whatever models it routes to, so the figure is genuinely harder to state — but a system with an unstated ceiling is awkward to plan around if your application generates long artefacts.

벤치마크 행을 솔직하게 읽기
Sakana names six benchmarks on which Fugu Max achieves "best overall score": Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. No figures accompany the claim, and SWEFish is Sakana's own internal coding benchmark. Google, meanwhile, publishes real numbers for Gemini 3.1 Pro including 77.1% on ARC-AGI-2, 94.3% on GPQA Diamond, 80.6% on SWE-bench Verified and 85.9% on BrowseComp.
Be careful with one near-miss. Google reports Terminal-Bench 2.0 at 68.5%. Sakana claims a best overall score on Terminal Bench 2.1. Those are adjacent version numbers on a benchmark that changes between them, which means the two results cannot be placed side by side even in principle. A reader scanning headlines would see "both claim Terminal Bench leadership" and conclude they are comparable. They are not.
GPQA is the closer call — Sakana lists GPQAD as one of its six wins; Google reports GPQA Diamond 94.3%. Same underlying evaluation family, both vendor-reported, and only one gives you a number. If Sakana's GPQAD result were published, this would be the one row where a genuine comparison was possible. It is not published, so the row reads as an assertion against a figure.
No independent party has evaluated Fugu Max, and there is no Artificial Analysis entry for it or any Fugu model. Gemini 3.1 Pro, by contrast, has been in public circulation long enough to appear in third-party leaderboards and evaluation write-ups — which is what "months of independent scrutiny" buys you over a same-week release.
The reason this matters more than usual for this specific matchup is that Fugu Max's own architecture makes its scores harder to interpret. Gemini 3.1 Pro's 94.3% on GPQA Diamond is a fact about one model at one version. A Fugu Max score on the same test would be a fact about a routing policy, a pool whose membership is undisclosed, and a coordinator trained to pick among them. Change any one of those and the number moves without the product name changing. Gemini 3.1 Pro's score has a fixed referent. Fugu Max's would not.
그들을 호출하는 것, 그리고 어떤 라우팅 변경이 있는지
Gemini 3.1 Pro is on OrcaRouter under the model ID google/gemini-3.1-pro-preview, at Google's list price with 0% markup — the provider's rate passed through, so a Google price change lands on your existing key the same day rather than at renewal. It sits alongside more than 200 other models behind one API key, which matters here for a specific reason: if you want to find out whether Fugu Max's $6.00 output rate is actually better for your workload, the clean experiment is to run both against the same task set, and Fugu Max is not on our catalogue. It reaches you through Sakana's own OpenAI-compatible API and several third-party platforms. Running your own A/B means two integrations, two bills, and two sets of credentials.
Automatic failover is the other piece worth naming. Gemini 3.1 Pro runs on multiple provider paths through the router, and if one is rate-limited or degraded, traffic moves without a code change. That is a partial, practical version of the resilience Sakana sells as the entire premise of orchestration — you get it at the transport layer rather than the reasoning layer, and you get it for a model whose behaviour you can pin.

그것들 중에서 어떻게 결정할까요
Choose Gemini 3.1 Pro if your work involves non-text input, if you need a documented output ceiling, if you want a model whose benchmark figures can be checked against other people's results, or if you simply need a component you can pin by version and reason about. At $2.00 input it costs nothing to try, and 65K of output is enough for most generation tasks. Its published numbers put it among the stronger general-purpose models available, and its multimodal input set is not something an orchestration layer substitutes for.
Choose Fugu Max if your work is text-heavy, verifiable, high-volume, and you care about the median cost of a completed task rather than the ceiling on any one call. The $6.00 output rate is half of Gemini 3.1 Pro's and the $0.25 cache rate is the most aggressive number in either rate card. Run it on problems where a wrong answer is detectable, set an output-token budget before you begin, and measure tokens per finished task. Do not select it for a multimodal pipeline on the strength of the release page, because the release page does not tell you whether it accepts images.
결정
Gemini 3.1 Pro is the safer and better-documented choice: identical input price, a published context window, a stated 65K output ceiling, documented multimodal input, and vendor benchmarks that third parties have had a chance to examine. Its output rate is double Fugu Max's, and that is the entire case for the other side.
Fugu Max is the cheaper way to generate text and the more opaque way to do anything else. Its six benchmark wins carry no figures, its context window and modality support are unpublished, and the composition of the pool that produces its results is undisclosed — which means you cannot rule out that you are paying $6.00 for a discounted version of the reasoning you could be buying from Gemini 3.1 Pro directly at $2.00 input on OrcaRouter with the provider's list price passed through and 0% markup on top.

