A hero title card for Fugu Ultra v2 vs Gemini 3.1 Pro with the subtitle 'The opponent that may already be inside the machine', pill badges reading '$30.00 vs $12.00 output', 'CharXiv 86.6 vs 80.2 directional' and 'Undisclosed pool', a footer line 'Sakana AI vs Google DeepMind - September 2026', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Fugu Ultra v2 vs Gemini 3.1 Pro: the opponent that may already be inside the machine

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

There is a version of the Fugu Ultra v2 vs Gem​ini 3.1 Pro comparison in which Gem​ini wins whichever way the benchmark goes — because it may be doing some of the work. Sakana AI's orchestration system, released September 11, 2026, does not disclose which models it coordinates; the June Fugu Ultra's reported pool included Gem​ini 3.1 Pro among others, and version 2.0 tells us only what it removed, naming Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra as excluded. Goo​gle's Gem​ini 3.1 Pro is a single multimodal frontier model with a 1M-token context that handles text, images, PDFs, audio and video, generally available since May 2026 at $2.00 per million input and $12.00 per million output. One of these is a model you can inspect. The other is a system whose benchmark wins may route through the first one, and neither vendor will tell you whether they do.

What each system actually is

Gem​ini 3.1 Pro is legible in a way that has become rare among frontier releases. It is a sparse mixture-of-experts transformer from Goo​gle DeepMind, first shipped in preview on February 19, 2026, with one source dating general availability to May 2026 — our own listing carries it under the preview model ID. It carries a 1M-token input window and a 65K-token output ceiling, accepts text, images, PDFs, audio, video and code, and exposes a three-tier reasoning control — low, medium or high — with high as the API default. Goo​gle reports 77.1% on ARC-AGI-2, more than double the previous generation's 31.1%; 94.3% on GPQA Diamond; 80.6% on SWE-bench Verified; 68.5% on Terminal-Bench 2.0; 85.9% on BrowseComp; and 44.4% on Humanity's Last Exam without tools, rising to 51.4% with search and code execution. Those are vendor figures. Its knowledge cutoff is January 2025, which is worth flagging if your work depends on recent events.

Fugu Ultra v2 is a coordinator. It is a trained model, reported around 7B parameters, that reads a request, assembles an agent team with Thinker, Worker and Verifier roles drawn from Sakana's TRINITY research, and synthesises an answer behind an OpenAI-compatible endpoint. It has no published modality list of its own — anything it can see, it sees because a model in its pool can see it. Its context is 1M tokens with a sharp pricing cliff at 272K, where rates double to $10 input and $45 output. Its output ceiling is not published. Its pool is undisclosed, its routing decisions are "not exposed by design," and for the Ultra tier the pool is fixed, so you cannot exclude a provider.

That asymmetry is the whole matchup. Gem​ini 3.1 Pro is a component you can buy directly and reason about. Fugu Ultra v2 is an assembly whose parts you cannot see and whose strongest benchmark claims may be, in part, Gem​ini's own results combined with someone else's.

A screenshot of the Sakana Fugu product page at sakana.ai/fugu (captured September 11, 2026, English UI), showing the 'Sakana Fugu' wordmark with the subtitle 'One Model to Command Them All' and the description that Fugu dynamically orchestrates the world's best models to tackle complex, multi-step tasks behind a single API, plus the notice reading 'Not yet available in the EU/EEA while we work toward compliance with GDPR and EU-specific regulations.'

The comparison where the two overlap

Very little of the published evidence lines the two systems up on the same row. Gem​ini 3.1 Pro's figures below are Goo​gle's own; Fugu Ultra v2's are Sakana's own; where a third-party aggregator has published a shared row, it is marked and comes with a caveat about being directional rather than decisive.

• Multimodal reasoning (CharXiv, shared row) — Fugu Ultra v2 86.6% vs Gem​ini 3.1 Pro 80.2%, per BenchLM, which labels the category directional and declines to name a winner

• TerminalBench 2.1 — no Fugu Ultra v2 v2.0 figure vs Gem​ini 3.1 Pro no comparable published figure; the June Fugu Ultra reported 82.1% against Gem​ini 3.1 Pro's 70.3% in third-party aggregation

• SWE-Bench Pro — no Fugu Ultra v2 v2.0 figure vs Gem​ini 3.1 Pro no comparable published figure; the June Fugu Ultra reported 73.7% against Gem​ini 3.1 Pro's 54.2%

• ARC-AGI-2 — no Fugu Ultra v2 figure vs Gem​ini 3.1 Pro 77.1%, vendor-reported

• Humanity's Last Exam — no Fugu Ultra v2 v2.0 figure vs Gem​ini 3.1 Pro 44.4% without tools, 51.4% with, vendor-reported

• Modality coverage — Fugu Ultra v2 inherits whatever its pool supports, undisclosed vs Gem​ini 3.1 Pro text, image, PDF, audio, video and code, documented

• Input / output price — Fugu Ultra v2 $5.00 / $30.00 per 1M, doubling above 272K vs Gem​ini 3.1 Pro $2.00 / $12.00 per 1M up to 200K, $4.00 / $18.00 beyond, cached input $0.20 / $0.40

• Context and output — Fugu Ultra v2 1M context, output ceiling unpublished vs Gem​ini 3.1 Pro 1M input, 65K output

• Weights and access — Fugu Ultra v2 closed, EU/EEA excluded vs Gem​ini 3.1 Pro proprietary but broadly available across Goo​gle surfaces

The multimodal row is the one to handle carefully, because it is the most quoted and the least solid. BenchLM's CharXiv aggregation puts Fugu Ultra v2 ahead by 6.4 normalized points — and the same page states plainly that directional rows never receive a winner, that Gem​ini had no measured results in the agentic, coding, knowledge or multilingual categories, and that Fugu had none in math or multilingual. A category where only one side has been measured in most rows cannot produce a verdict. If you take nothing else from this comparison, take that: on multimodal work specifically, the published evidence does not support a conclusion either way, and the one row that exists is explicitly not a settled result.

The possibility that changes the frame

Sakana will not say what is in the pool. That is a documented design choice, not an oversight — the FAQ says routing information is not exposed by design, and there is no mechanism to inspect it. What we know from the June generation is that reported pool members included Gem​ini 3.1 Pro alongside other frontier models. Version 2.0 changed the pool, and Sakana described the change only by subtraction: Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra are out. Nothing in the release says Gem​ini is still in.

If Gem​ini 3.1 Pro is in the pool, three consequences follow, and all three are practical rather than philosophical. First, some share of Fugu Ultra v2's multimodal performance is Gem​ini's performance, combined with other models' — which means you are paying a coordination premium on top of a per-token rate you could pay directly. Second, Fugu Ultra v2 at $30 output per million against Gem​ini 3.1 Pro at $12 is a premium for assembly, not for raw capability; if your task is a single multimodal question, Gem​ini answers it cheaper and you can see exactly which model answered. Third, and most usefully, the honest benchmark of an orchestrator is not "does it beat its components" but "does it beat its best component when you use it alone." Sakana's architecture is built on the claim that it does, on verifiable tasks. That claim is worth testing on your own workload before you accept it on a chart-reading benchmark.

There is a version where the answer is no and the orchestrator still earns its place. If Gem​ini is in the pool and Fugu Ultra v2's Chartography score of 48.3 against Claude Opus 5's 27.3 holds up, then what Sakana built is a coordination layer that reliably gets more out of the same model than you get calling it once. That is a real product, and the open question is only whether the margin covers the price. Nobody has published the experiment.

A two-column scoreboard for Fugu Ultra v2 vs Gemini 3.1 Pro titled 'Fugu Ultra v2 vs Gemini 3.1 Pro - the scoreboard'. Left column Fugu Ultra v2: Input price $5.00 / 1M, Output price $30.00 / 1M, Modalities inherited, undisclosed, CharXiv 86.6 (directional), Weights closed, pool hidden, EU/EEA availability not served. Right column Gemini 3.1 Pro: Input price $2.00 / 1M, Output price $12.00 / 1M, Modalities text, image, PDF, audio, video, CharXiv 80.2 (directional), Weights proprietary, documented, EU/EEA availability available. Footer 'CharXiv row directional only per BenchLM; Fugu figures vendor-reported by Sakana.', with the OrcaRouter logo bottom-right.

Where the honest edges are

Gem​ini 3.1 Pro's advantages are concrete. It is roughly two-fifths the price of Fugu Ultra v2 on input and output, and unlike Fugu its rates do not double past a context threshold — the step at 200K is gentler than the 272K cliff. It documents its modalities instead of inheriting them, which means audio and video work is a supported feature rather than a hope. It has a published output ceiling. It is available through Goo​gle's own surfaces and, like the other models Fugu Ultra v2 is measured against, it is callable on OrcaRouter — as google/gemini-3.1-pro-preview, at Goo​gle's list price with 0% markup, so a Goo​gle price change lands on the same key the same day it is announced.

Fugu Ultra v2's advantages are narrower and real. It attacks a hard problem more than once, with a Verifier role checking the work, which is why its strongest claims sit on benchmarks where correctness is checkable — Chartography, DeepSWE, Toolathon — rather than on knowledge recall. Sakana's published case studies lean the same way: a mechanical CAD iris that opens and closes correctly where other models left gaps and weak linkages; a pure-Python Rubik's cube solver finishing all 300 scrambled cubes where two anonymised frontier baselines crashed entirely. Those baselines are anonymised and vendor-chosen, so treat them as illustration rather than proof. But the pattern is consistent across the release, and it describes a genuine capability that a single model call does not have: persistence on a problem until the answer passes a check.

Weigh the constraints too. Fugu Ultra v2 is not sold in the EU or the EEA, and Sakana is explicit that it is working toward GDPR compliance. Gem​ini 3.1 Pro has no such restriction and a much longer track record of independent evaluation behind it.

A screenshot of the OrcaRouter model page for Gemini 3.1 Pro Preview (google/gemini-3.1-pro-preview, captured September 11, 2026, English UI), showing the Flagship and Featured badges, the capability tags Vision, Audio, Tools, JSON and Reasoning, the 1M token context and 65K max output fields, the input and output modality line reading 'audio + file + image + text + video' to 'text', the pricing tiles reading INPUT $2.00 and OUTPUT $12.00 per 1M tokens with p50 TTFT 3.82s, and the OpenAI-compatible Python code sample pointing at api.orcarouter.ai/v1.

The verdict

For most multimodal work, Gem​ini 3.1 Pro is the right call and it is not close. It is cheaper per token, cheaper per document, documented for audio and video, capped at a context threshold that does not double your bill mid-run, and — the part that matters most here — you can see what answered you. If you need one model to read a PDF, watch a clip and answer a question, there is no argument for routing that through an undisclosed pool at two and a half times the output price.

Fugu Ultra v2 is the right call for a specific and much narrower shape of work: long-horizon tasks with a checkable answer, where trying a problem three ways and verifying the result beats trying it once, and where the marginal point of quality is worth the multiple. It is outside consideration entirely if you have EU or EEA users in scope.

The uncomfortable question the matchup leaves open is the one nobody has measured. If Gem​ini 3.1 Pro is inside the pool — and the only honest answer today is that Sakana has not said it is not — then part of what you are buying with Fugu Ultra v2 is Gem​ini 3.1 Pro with a coordination layer on top, at a price that implies the layer is worth roughly two and a half times the model. That may well be true. Sakana's architecture was built to make it true. But it is a hypothesis with a vendor scoreboard behind it and no independent test, and until someone runs it, the model you can see is the one you can defend.