Cartão de título principal para "Fugu Max vs Gemini 3.1 Pro" com o subtítulo "Mesmo preço de entrada, metade do preço de saída", três selos em forma de pílula com os textos "$2.00 de entrada em ambos", "$6.00 vs $12.00 de saída", "Pool não divulgado", uma linha de rodapé com "Sakana AI vs Google - setembro de 2026", e o logotipo do OrcaRouter no canto inferior direito.
Guides & Insights

Fugu Max vs Gemini 3.1 Pro: mesmo preço de input, metade do preço de output, e um pool que ninguém consegue ver

Autor

Elias Hawthorne

Data de publicação

Modelos mais recentes · 20Ver todos os modelos
Benchmarks: Artificial Analysis · atualizado diariamente
Voltar para todas as publicações

Fugu Max and Gem​ini 3.1 Pro charge identically to be asked a question. Both list at $2.00 per million input tokens. They part company on the way out: Sakana AI's Fugu Max writes for $6.00 per million output tokens, and Goo​gle's Gem​ini 3.1 Pro writes for $12.00. Sakana released Fugu Max on September 11, 2026 as a coordinator that routes each task to the cheapest sufficient model in a pool it does not disclose. Gem​ini 3.1 Pro is a named model — callable by ID, versioned, multimodal — which is the thing an orchestrator is designed not to be. So the interesting question here is not which one is better at $2.00 input. It is whether a comparison between a coordinator and a model is even the right shape, given that Sakana's earlier Fugu generation reportedly counted Gem​ini 3.1 Pro among the models it routed to.

The pool question, stated precisely

When the first Fugu Ultra shipped in June 2026, reporting on its architecture described a pool that included Gem​ini 3.1 Pro alongside models from other frontier labs. Sakana has not published the membership of Fugu Max's pool. The release says the pool expanded to include an unprecedented number of open-weights and specialised models, naming the NVIDIA Nemotron family through an NVIDIA collaboration, and stops there. It also does not say Gem​ini 3.1 Pro was removed.

That leaves two possibilities and no public way to choose between them. If Gem​ini 3.1 Pro is in the Fugu Max pool, then this matchup is partly a "versus" and partly a "through" — a Fugu Max answer on a hard reasoning task may be a Gem​ini answer that a coordinator selected, checked, and rewrote, in which case the $6.00 output rate is buying Gem​ini's reasoning plus orchestration overhead at a discount to calling Gem​ini directly. If it is not in the pool, the two products are genuinely different instruments and the price difference reflects different underlying capability.

Nothing in the release resolves this. Anyone who tells you which one it is has inferred it rather than read it. What can be said with confidence is narrower and still useful: an orchestrator's published scores belong to the pool, not to the coordinator, and a pool whose members are undisclosed makes those scores impossible to attribute.

• Input price — Fugu Max $2.00 per 1M tokens vs Gem​ini 3.1 Pro $2.00 per 1M tokens

• Output price — Fugu Max $6.00 per 1M tokens vs Gem​ini 3.1 Pro $12.00 per 1M tokens

• Cached input — Fugu Max $0.25 per 1M tokens vs Gem​ini 3.1 Pro caching offered at a discount on the base input rate

• Context — Fugu Max not published vs Gem​ini 3.1 Pro 1M tokens

• Output ceiling — Fugu Max not published vs Gem​ini 3.1 Pro 65K tokens

• Input modality — Fugu Max not published vs Gem​ini 3.1 Pro audio, file, image, text and video

• Output modality — Fugu Max not published vs Gem​ini 3.1 Pro text, with a separate image-generation line in the 3.1 family

• Benchmark figures — Fugu Max six wins asserted with no scores vs Gem​ini 3.1 Pro Google-reported ARC-AGI-2 77.1%, GPQA Diamond 94.3%, SWE-bench Verified 80.6%, BrowseComp 85.9%

What a model can promise that a coordinator cannot

The single largest practical difference in that list is not the output price. It is the modality row. Gem​ini 3.1 Pro accepts audio, video and images as input and is documented as such, which makes it usable for work that has nothing to do with text — reading a screen recording, transcribing and reasoning over a call, parsing a scanned form. Fugu Max publishes no modality specification at all. It may well handle some of those inputs, depending on what its pool contains and whether the coordinator forwards multimodal content or flattens it to text first. Nobody outside Sakana knows, and that is a different kind of uncertainty from "we haven't benchmarked it yet." It means you cannot decide fitness for a multimodal workload from the documentation; you have to test it.

There is a second asymmetry in the output ceiling. Gem​ini 3.1 Pro documents 65K maximum output tokens. Fugu Max documents nothing. For a coordinator, the effective ceiling is the minimum across whatever models it routes to, so the figure is genuinely harder to state — but a system with an unstated ceiling is awkward to plan around if your application generates long artefacts.

A two-column scoreboard titled "Fugu Max vs Gemini 3.1 Pro Preview - the scoreboard". Left column Fugu Max: Input price $2.00 / 1M, Output price $6.00 / 1M, Cached input $0.25 / 1M, Context not published, Input modality not published, Evidence vendor-reported only. Right column Gemini 3.1 Pro Preview: Input price $2.00 / 1M, Output price $12.00 / 1M, Cached input discount on base rate, Context 1M tokens with 65K output, Input modality audio, file, image, text and video, Evidence independent evals. Footer reading "Fugu Max figures vendor-reported by Sakana AI; Gemini 3.1 Pro rows Google-reported. Pool membership is not disclosed.", with the OrcaRouter logo bottom-right.

Reading the benchmark rows honestly

Sakana names six benchmarks on which Fugu Max achieves "best overall score": Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. No figures accompany the claim, and SWEFish is Sakana's own internal coding benchmark. Goo​gle, meanwhile, publishes real numbers for Gem​ini 3.1 Pro including 77.1% on ARC-AGI-2, 94.3% on GPQA Diamond, 80.6% on SWE-bench Verified and 85.9% on BrowseComp.

Be careful with one near-miss. Goo​gle reports Terminal-Bench 2.0 at 68.5%. Sakana claims a best overall score on Terminal Bench 2.1. Those are adjacent version numbers on a benchmark that changes between them, which means the two results cannot be placed side by side even in principle. A reader scanning headlines would see "both claim Terminal Bench leadership" and conclude they are comparable. They are not.

GPQA is the closer call — Sakana lists GPQAD as one of its six wins; Goo​gle reports GPQA Diamond 94.3%. Same underlying evaluation family, both vendor-reported, and only one gives you a number. If Sakana's GPQAD result were published, this would be the one row where a genuine comparison was possible. It is not published, so the row reads as an assertion against a figure.

No independent party has evaluated Fugu Max, and there is no Artificial Analysis entry for it or any Fugu model. Gem​ini 3.1 Pro, by contrast, has been in public circulation long enough to appear in third-party leaderboards and evaluation write-ups — which is what "months of independent scrutiny" buys you over a same-week release.

The reason this matters more than usual for this specific matchup is that Fugu Max's own architecture makes its scores harder to interpret. Gem​ini 3.1 Pro's 94.3% on GPQA Diamond is a fact about one model at one version. A Fugu Max score on the same test would be a fact about a routing policy, a pool whose membership is undisclosed, and a coordinator trained to pick among them. Change any one of those and the number moves without the product name changing. Gem​ini 3.1 Pro's score has a fixed referent. Fugu Max's would not.

Calling them, and what routing changes

Gem​ini 3.1 Pro is on OrcaRouter under the model ID google/gemini-3.1-pro-preview, at Goo​gle's list price with 0% markup — the provider's rate passed through, so a Goo​gle price change lands on your existing key the same day rather than at renewal. It sits alongside more than 200 other models behind one API key, which matters here for a specific reason: if you want to find out whether Fugu Max's $6.00 output rate is actually better for your workload, the clean experiment is to run both against the same task set, and Fugu Max is not on our catalogue. It reaches you through Sakana's own OpenAI-compatible API and several third-party platforms. Running your own A/B means two integrations, two bills, and two sets of credentials.

Automatic failover is the other piece worth naming. Gem​ini 3.1 Pro runs on multiple provider paths through the router, and if one is rate-limited or degraded, traffic moves without a code change. That is a partial, practical version of the resilience Sakana sells as the entire premise of orchestration — you get it at the transport layer rather than the reasoning layer, and you get it for a model whose behaviour you can pin.

A screenshot of the Sakana AI release page at sakana.ai/fugu-max-release (captured September 11, 2026, English UI), showing the sakana.ai wordmark, the headline 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' dated September 11, 2026, the opening copy arguing that the frontier which matters is two-dimensional with capability on one axis and cost on the other, a 'Try Sakana Fugu Max and Fugu Ultra v2' link, and a Pareto chart plotting performance against cost with a red 'Fugu Max' point at the low-cost end, a red 'Fugu Ultra' point at the top, a shaded 'Frontier formed by Fugu models' region and grey points labelled 'Frontier formed by single models'.

How to decide between them

Choose Gem​ini 3.1 Pro if your work involves non-text input, if you need a documented output ceiling, if you want a model whose benchmark figures can be checked against other people's results, or if you simply need a component you can pin by version and reason about. At $2.00 input it costs nothing to try, and 65K of output is enough for most generation tasks. Its published numbers put it among the stronger general-purpose models available, and its multimodal input set is not something an orchestration layer substitutes for.

Choose Fugu Max if your work is text-heavy, verifiable, high-volume, and you care about the median cost of a completed task rather than the ceiling on any one call. The $6.00 output rate is half of Gem​ini 3.1 Pro's and the $0.25 cache rate is the most aggressive number in either rate card. Run it on problems where a wrong answer is detectable, set an output-token budget before you begin, and measure tokens per finished task. Do not select it for a multimodal pipeline on the strength of the release page, because the release page does not tell you whether it accepts images.

A decisão

Gem​ini 3.1 Pro is the safer and better-documented choice: identical input price, a published context window, a stated 65K output ceiling, documented multimodal input, and vendor benchmarks that third parties have had a chance to examine. Its output rate is double Fugu Max's, and that is the entire case for the other side.

Fugu Max is the cheaper way to generate text and the more opaque way to do anything else. Its six benchmark wins carry no figures, its context window and modality support are unpublished, and the composition of the pool that produces its results is undisclosed — which means you cannot rule out that you are paying $6.00 for a discounted version of the reasoning you could be buying from Gem​ini 3.1 Pro directly at $2.00 input on OrcaRouter with the provider's list price passed through and 0% markup on top.

A screenshot of the OrcaRouter model page for Gemini 3.1 Pro Preview (google/gemini-3.1-pro-preview, captured September 11, 2026, English UI), showing the FLAGSHIP and Featured badges, the google/gemini-3.1-pro-preview model ID, the Vision, Audio, Tools, JSON and Reasoning tags, the listing date 2026-02-19, the /v1/chat/completions and /vibeta/models/{model}:generateContent endpoints, the pricing tiles reading INPUT $2.00 and OUTPUT $12.00 per 1M tokens with p50 TTFT 3.82s and 70.1M tokens of 7-day traffic, the 1M token context with 65K max output and audio + file + image + text + video input, and the OpenAI-compatible Python sample pointing at api.orcarouter.ai/v1.