
Fugu Ultra v2 vs Grok 4.6: Five Times the Output Price, One Shared Benchmark
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
There is exactly one benchmark on which Fugu Ultra v2 and Grok 4.6 both publish a number, and it is DeepSWE: Sakana AI reports 74.3 for Fugu Ultra v2 in its 11 September 2026 announcement, while xAI's launch material for Grok 4.6 puts its own DeepSWE v1.1 score at 65.9%. That is an 8.4-point gap, and it is the only clean comparison the two companies offer. Everything else you might line up — Sakana's Chartography and SWEFish against Grok's GPQA Diamond and CursorBench — is measured on different instruments by different labs. So the honest question is not "which is smarter." It is whether Fugu Ultra v2's claimed edge is worth paying $30 per million output tokens for, when Grok 4.6 charges $6.
Two different answers to the same problem
Grok 4.6 is a single model, and it is the current flagship of the Grok line — listed as "Latest" in the vendor's own developer documentation, succeeding Grok 4.5 with the same 500,000-token context window, the same text/image/file input surface, and the same base pricing. It carries a knowledge cutoff of 1 February 2026 and exposes reasoning effort across low, medium, high and xhigh, with high as the default. One detail that surprises people: reasoning cannot be switched off. Thinking tokens bill on every request, which makes Grok 4.6's real output cost depend on how hard it decides to think.
Fugu Ultra v2 is not a model at all in the usual sense. It is the peak-capability tier of Sakana AI's orchestration line — a single OpenAI-compatible endpoint that decomposes a task and dispatches it across a pool of other models, some open-weights, some specialised. Sakana's framing is that it reaches frontier-level output "without the indispensable reliance on the frontier models it orchestrates," and it states plainly that Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra are excluded from that pool. You are paying for the scheduling layer, and for every token the sub-agents burn.
That distinction is not cosmetic. It is the entire reason the price gap exists, and it is the only thing that makes the gap defensible.
The rate cards, including the part that repeats
• Input — Fugu Ultra v2 $5.00 per 1M vs Grok 4.6 $2.00 per 1M. Fugu is 2.5× before any sub-calls happen.
• Output — Fugu Ultra v2 $30.00 per 1M vs Grok 4.6 $6.00 per 1M. A 5× gap, and the one that decides most bills.
• Cached input — $0.50 per 1M on both. Identical. Two vendors with nothing else in common landed on the same cache read price, which makes cache-heavy agent loops the one workload where the two are genuinely comparable line for line.
• Long-context repricing — Grok 4.6 doubles to $4.00 / $12.00 once a prompt crosses 200,000 tokens, and the repricing applies to the whole request, not just the overage. Fugu Ultra v2's reported tier above roughly 272,000 input tokens moves to about $10.00 / $45.00, also on the whole request. Both punish long prompts; Grok starts punishing 72K tokens earlier.
• Context window — Grok 4.6 500K tokens vs Fugu Ultra v2 1M as reported by third-party listings. Sakana's announcement page does not state a context figure at all, so treat its 1M as unconfirmed by the vendor.
The long-context row cuts both ways and is worth doing the arithmetic on. A 300K-token prompt costs you $4.00 per million input on Grok 4.6 and around $10.00 on Fugu Ultra v2. A 150K-token prompt costs $2.00 on Grok and $5.00 on Fugu. Sakana's model is more expensive at every prompt length — the only question is how often it needs to be called.

The argument for paying 5× on output
An orchestrator's case is not that it is cheaper per token. It is that it needs fewer attempts, and that the tokens it spends on coordination are cheaper than the tokens you would spend failing.
That is a real argument, and Sakana is making it explicitly: Fugu Ultra v2 is aimed at complex multi-step work — autonomous research, full-stack development — where a single model's failure mode is drifting off the rails partway through a long run. If a $30 output rate buys you a task completed on the first pass instead of the third, the per-token premium is irrelevant.
There is supporting evidence from the vendor's own side of the fence, though it is vendor-favourable and should be read as such. xAI's launch material for Grok 4.6 cites roughly 53 turns and about 0.5 billion input tokens to resolve long-horizon tasks, against approximately 103 turns and 2.0 billion input tokens for a competing frontier model — an argument that Grok itself is economical per task even at a low per-token price. Both companies are making the same move: pushing you away from the rate card and toward cost per finished job.
Which means the deciding number is one neither vendor publishes, and that you have to measure yourself: tokens burned, start to finish, on a task that looks like yours.
Where each one is genuinely weak
Grok 4.6's published weak spot is terminal work. Its Terminal-Bench v3.0 score sits at 26%, well behind the top of the field, even though the same model posts a much stronger 88.4% on the older Terminal-Bench v2.1 — those are different tests and blending them is a mistake you will see in a lot of coverage. It also carries the highest GPQA Diamond score of its cohort at 94.9%, but GPQA Diamond has since been dropped from the Artificial Analysis index as saturated, so a high score there no longer discriminates between frontier models the way it did a year ago.
On the independent side, Artificial Analysis scores Grok 4.6 at 44 on its v4.3 Intelligence Index, ranked #20 of 200, at a cost per task of $1.86. The often-circulated figure of 61 is from a superseded index scale and should not be quoted without saying which version produced it.
Fugu Ultra v2's weak spot is verification. As of publication, nothing Sakana has claimed about it — the five best-or-joint-best results, the 48.3 Chartography score against Opus 5's 27.3 and Fable 5's 29.5, the 74.3 DeepSWE — has been independently reproduced. There is also no published latency or throughput figure, which matters more for an orchestrator than for a single model, because its response time depends entirely on what it decides to call. For the previous Fugu Ultra, testers publicly reported runs stretching to around 30 minutes; that is the June-2026 model, not v2, but it is the failure mode to test for before you commit a production path.

A note on the vendor name
One wrinkle worth knowing before you go looking for Grok 4.6 in a developer console: the model is consistently called Grok 4.6, but the vendor branding around it is in flux. xAI's own documentation now carries the SpaceXAI name, a change most launch coverage has not caught up with. The model identifiers and the API surface are unchanged — Grok 4.6 is a first-class OpenAI Responses model and drops into existing tool-calling loops without a translation layer — but if you are searching for a model card and the branding looks wrong, it is not your mistake.
Calling either of these in practice
Grok 4.6 is on OrcaRouter at the provider's own rate — $2.00 per million input and $6.00 per million output, passed through with zero markup, so a vendor price change reaches here the same day rather than on a migration schedule. It is OpenAI-compatible on the same base URL as everything else in the catalogue, which means you can put it behind a fallback chain or a routing rule instead of hard-coding it, and you can A/B it against another model without a second contract or a code change.
Fugu Ultra v2 is not routed by us. It is available from Sakana AI's own OpenAI-compatible API, and the company says an existing Fugu integration moves to it with a single parameter change. If you want to compare the two on your workload, the cheapest arrangement is to leave Grok 4.6 on your gateway and call Fugu Ultra v2 directly from Sakana behind a feature flag — both speak the same request format, so the same test harness works on both.

The verdict
Grok 4.6 is the rational default for most teams, and it is not close on price. At $2.00 / $6.00 with a $0.50 cache read and a 500K window, it handles long-document reasoning, coding agents and multimodal file work at a fifth of Fugu Ultra v2's output rate, and it is independently scored and independently benchmarked. Its exposure is prompt length — the 200K cliff doubles the whole request — and terminal-heavy agent loops, where its published scores are not competitive.
Fugu Ultra v2 is worth its premium only if you have a specific kind of workload: long, multi-step, autonomy-heavy tasks where a single model tends to derail, and where you also want the architectural property Sakana is selling — a capability tier that does not depend on any frontier vendor staying available or staying cheap. That is a real thing to want. But it is a claim, not yet a measurement, and the DeepSWE gap that makes it look credible is a single benchmark from a single vendor on the day of release.
If you cannot say which of your tasks currently fails on the second or third attempt, you do not yet have the number that would justify the price difference. Measure that first. The rate cards already tell you everything else.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
