
Fugu Ultra v2 vs Gemini 3.1 Pro: the opponent that may already be inside the machine
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
There is a version of the Fugu Ultra v2 vs Gemini 3.1 Pro comparison in which Gemini wins whichever way the benchmark goes — because it may be doing some of the work. Sakana AI's orchestration system, released September 11, 2026, does not disclose which models it coordinates; the June Fugu Ultra's reported pool included Gemini 3.1 Pro among others, and version 2.0 tells us only what it removed, naming Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra as excluded. Google's Gemini 3.1 Pro is a single multimodal frontier model with a 1M-token context that handles text, images, PDFs, audio and video, generally available since May 2026 at $2.00 per million input and $12.00 per million output. One of these is a model you can inspect. The other is a system whose benchmark wins may route through the first one, and neither vendor will tell you whether they do.
What each system actually is
Gemini 3.1 Pro is legible in a way that has become rare among frontier releases. It is a sparse mixture-of-experts transformer from Google DeepMind, first shipped in preview on February 19, 2026, with one source dating general availability to May 2026 — our own listing carries it under the preview model ID. It carries a 1M-token input window and a 65K-token output ceiling, accepts text, images, PDFs, audio, video and code, and exposes a three-tier reasoning control — low, medium or high — with high as the API default. Google reports 77.1% on ARC-AGI-2, more than double the previous generation's 31.1%; 94.3% on GPQA Diamond; 80.6% on SWE-bench Verified; 68.5% on Terminal-Bench 2.0; 85.9% on BrowseComp; and 44.4% on Humanity's Last Exam without tools, rising to 51.4% with search and code execution. Those are vendor figures. Its knowledge cutoff is January 2025, which is worth flagging if your work depends on recent events.
Fugu Ultra v2 is a coordinator. It is a trained model, reported around 7B parameters, that reads a request, assembles an agent team with Thinker, Worker and Verifier roles drawn from Sakana's TRINITY research, and synthesises an answer behind an OpenAI-compatible endpoint. It has no published modality list of its own — anything it can see, it sees because a model in its pool can see it. Its context is 1M tokens with a sharp pricing cliff at 272K, where rates double to $10 input and $45 output. Its output ceiling is not published. Its pool is undisclosed, its routing decisions are "not exposed by design," and for the Ultra tier the pool is fixed, so you cannot exclude a provider.
That asymmetry is the whole matchup. Gemini 3.1 Pro is a component you can buy directly and reason about. Fugu Ultra v2 is an assembly whose parts you cannot see and whose strongest benchmark claims may be, in part, Gemini's own results combined with someone else's.

The comparison where the two overlap
Very little of the published evidence lines the two systems up on the same row. Gemini 3.1 Pro's figures below are Google's own; Fugu Ultra v2's are Sakana's own; where a third-party aggregator has published a shared row, it is marked and comes with a caveat about being directional rather than decisive.
• Multimodal reasoning (CharXiv, shared row) — Fugu Ultra v2 86.6% vs Gemini 3.1 Pro 80.2%, per BenchLM, which labels the category directional and declines to name a winner
• TerminalBench 2.1 — no Fugu Ultra v2 v2.0 figure vs Gemini 3.1 Pro no comparable published figure; the June Fugu Ultra reported 82.1% against Gemini 3.1 Pro's 70.3% in third-party aggregation
• SWE-Bench Pro — no Fugu Ultra v2 v2.0 figure vs Gemini 3.1 Pro no comparable published figure; the June Fugu Ultra reported 73.7% against Gemini 3.1 Pro's 54.2%
• ARC-AGI-2 — no Fugu Ultra v2 figure vs Gemini 3.1 Pro 77.1%, vendor-reported
• Humanity's Last Exam — no Fugu Ultra v2 v2.0 figure vs Gemini 3.1 Pro 44.4% without tools, 51.4% with, vendor-reported
• Modality coverage — Fugu Ultra v2 inherits whatever its pool supports, undisclosed vs Gemini 3.1 Pro text, image, PDF, audio, video and code, documented
• Input / output price — Fugu Ultra v2 $5.00 / $30.00 per 1M, doubling above 272K vs Gemini 3.1 Pro $2.00 / $12.00 per 1M up to 200K, $4.00 / $18.00 beyond, cached input $0.20 / $0.40
• Context and output — Fugu Ultra v2 1M context, output ceiling unpublished vs Gemini 3.1 Pro 1M input, 65K output
• Weights and access — Fugu Ultra v2 closed, EU/EEA excluded vs Gemini 3.1 Pro proprietary but broadly available across Google surfaces
The multimodal row is the one to handle carefully, because it is the most quoted and the least solid. BenchLM's CharXiv aggregation puts Fugu Ultra v2 ahead by 6.4 normalized points — and the same page states plainly that directional rows never receive a winner, that Gemini had no measured results in the agentic, coding, knowledge or multilingual categories, and that Fugu had none in math or multilingual. A category where only one side has been measured in most rows cannot produce a verdict. If you take nothing else from this comparison, take that: on multimodal work specifically, the published evidence does not support a conclusion either way, and the one row that exists is explicitly not a settled result.
The possibility that changes the frame
Sakana will not say what is in the pool. That is a documented design choice, not an oversight — the FAQ says routing information is not exposed by design, and there is no mechanism to inspect it. What we know from the June generation is that reported pool members included Gemini 3.1 Pro alongside other frontier models. Version 2.0 changed the pool, and Sakana described the change only by subtraction: Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra are out. Nothing in the release says Gemini is still in.
If Gemini 3.1 Pro is in the pool, three consequences follow, and all three are practical rather than philosophical. First, some share of Fugu Ultra v2's multimodal performance is Gemini's performance, combined with other models' — which means you are paying a coordination premium on top of a per-token rate you could pay directly. Second, Fugu Ultra v2 at $30 output per million against Gemini 3.1 Pro at $12 is a premium for assembly, not for raw capability; if your task is a single multimodal question, Gemini answers it cheaper and you can see exactly which model answered. Third, and most usefully, the honest benchmark of an orchestrator is not "does it beat its components" but "does it beat its best component when you use it alone." Sakana's architecture is built on the claim that it does, on verifiable tasks. That claim is worth testing on your own workload before you accept it on a chart-reading benchmark.
There is a version where the answer is no and the orchestrator still earns its place. If Gemini is in the pool and Fugu Ultra v2's Chartography score of 48.3 against Claude Opus 5's 27.3 holds up, then what Sakana built is a coordination layer that reliably gets more out of the same model than you get calling it once. That is a real product, and the open question is only whether the margin covers the price. Nobody has published the experiment.

Where the honest edges are
Gemini 3.1 Pro's advantages are concrete. It is roughly two-fifths the price of Fugu Ultra v2 on input and output, and unlike Fugu its rates do not double past a context threshold — the step at 200K is gentler than the 272K cliff. It documents its modalities instead of inheriting them, which means audio and video work is a supported feature rather than a hope. It has a published output ceiling. It is available through Google's own surfaces and, like the other models Fugu Ultra v2 is measured against, it is callable on OrcaRouter — as google/gemini-3.1-pro-preview, at Google's list price with 0% markup, so a Google price change lands on the same key the same day it is announced.
Fugu Ultra v2's advantages are narrower and real. It attacks a hard problem more than once, with a Verifier role checking the work, which is why its strongest claims sit on benchmarks where correctness is checkable — Chartography, DeepSWE, Toolathon — rather than on knowledge recall. Sakana's published case studies lean the same way: a mechanical CAD iris that opens and closes correctly where other models left gaps and weak linkages; a pure-Python Rubik's cube solver finishing all 300 scrambled cubes where two anonymised frontier baselines crashed entirely. Those baselines are anonymised and vendor-chosen, so treat them as illustration rather than proof. But the pattern is consistent across the release, and it describes a genuine capability that a single model call does not have: persistence on a problem until the answer passes a check.
Weigh the constraints too. Fugu Ultra v2 is not sold in the EU or the EEA, and Sakana is explicit that it is working toward GDPR compliance. Gemini 3.1 Pro has no such restriction and a much longer track record of independent evaluation behind it.

The verdict
For most multimodal work, Gemini 3.1 Pro is the right call and it is not close. It is cheaper per token, cheaper per document, documented for audio and video, capped at a context threshold that does not double your bill mid-run, and — the part that matters most here — you can see what answered you. If you need one model to read a PDF, watch a clip and answer a question, there is no argument for routing that through an undisclosed pool at two and a half times the output price.
Fugu Ultra v2 is the right call for a specific and much narrower shape of work: long-horizon tasks with a checkable answer, where trying a problem three ways and verifying the result beats trying it once, and where the marginal point of quality is worth the multiple. It is outside consideration entirely if you have EU or EEA users in scope.
The uncomfortable question the matchup leaves open is the one nobody has measured. If Gemini 3.1 Pro is inside the pool — and the only honest answer today is that Sakana has not said it is not — then part of what you are buying with Fugu Ultra v2 is Gemini 3.1 Pro with a coordination layer on top, at a price that implies the layer is worth roughly two and a half times the model. That may well be true. Sakana's architecture was built to make it true. But it is a hypothesis with a vendor scoreboard behind it and no independent test, and until someone runs it, the model you can see is the one you can defend.
