A hero title card reading 'Fugu Max vs Grok 4.6' with the subtitle 'Identical rate cards, one published score', three pill badges reading '$2.00 vs $2.00 input', '$6.00 vs $6.00 output' and 'Six wins, zero figures', a footer line reading 'Sakana AI, September 2026 vs SpaceXAI, August 2026', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Fugu Max vs Grok 4.6: Identical Rate Cards, One Published Score

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Fugu Max and Grok 4.6 cost exactly the same. Sakana AI shipped Fugu Max on September 11, 2026, priced at $2.00 per million input tokens and $6.00 per million output tokens — and that is line-for-line the rate card SpaceXAI has been charging for Grok 4.6 since it launched on August 12. The coincidence is the most useful fact in this matchup, because it deletes price from the argument. When two products bill identically, the only question left is what arrives for the money, and there the two part company completely. Grok 4.6 is a single model you can pin by version, read a benchmark table for, and check against an independent index. Fugu Max is a trained coordinator that dispatches each task to the cheapest model in its pool, and Sakana's announcement claims it takes best overall score on six benchmarks without printing a figure from any of them. One of these purchases is measurable before you make it. The other is not.

Two products, one rate card

• Input — Fugu Max $2.00 per 1M tokens vs Grok 4.6 $2.00 per 1M tokens

• Output — Fugu Max $6.00 per 1M tokens vs Grok 4.6 $6.00 per 1M tokens

• Cached input — Fugu Max $0.25 per 1M tokens vs Grok 4.6 $0.50 per 1M tokens, the only line where the two prices differ at all

• Context window — Fugu Max not published by the vendor vs Grok 4.6 500,000 tokens, unchanged from Grok 4.5

• Long-prompt repricing — Fugu Max no tier stated in the release vs Grok 4.6 doubles to $4.00 / $12.00 once a single request passes 200,000 tokens, applied to the whole request rather than the overage

• Reasoning control — Fugu Max not exposed vs Grok 4.6 four effort levels, low through xhigh, defaulting to high and not switchable off

• Architecture — Fugu Max a trained coordinator over an undisclosed pool that includes open-weights and specialised models, expanded for Max with the NVIDIA Nemotron family through a collaboration with NVIDIA vs Grok 4.6 one model at one version

• Independent evaluation — Fugu Max none published, and no Artificial Analysis page exists for any Fugu model vs Grok 4.6 a live Artificial Analysis Intelligence Index of 44, ranked #20 of 200

The cheap column is the unforecastable one

Look at where Fugu Max actually wins on price: the cache read, at half of Grok 4.6's. It is a real advantage and it is exactly the wrong place to build a cost model, because cache pricing only pays off in proportion to your cache hit rate — and Sakana does not publish one for Fugu Max. The release does not state a context window either, nor a long-context tier, nor an output ceiling. So the one column where Fugu Max undercuts Grok 4.6 is a discount of unknown size applied to an unknown share of your tokens.

Grok 4.6's weaknesses are at least visible. Its 200,000-token threshold is aggressive — it reprices the entire request, so a 250,000-token prompt is billed at double rate from the first token, not the 200,001st. But you can see the cliff, measure your prompts against it, and decide whether to chunk. That is the difference in kind here: one system publishes a price schedule with a shape you can plan around, and the other publishes a headline rate with no schedule attached.

Six wins and not one score

The benchmarks Sakana names are Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. Five of those are recognised public evaluations. The sixth, SWEFish, is Sakana's own coding benchmark, which means one of the six claimed wins is scored on a test the vendor wrote and has not released. For the other five, the release says Fugu Max achieves the best overall score and stops there — no score, no margin, no competitor's number to sit beside it.

Set that against the table SpaceXAI published for Grok 4.6, which is vendor-reported throughout but is at least a table: GDPVal-AA v2 at 1,753, AA-Briefcase at 1,577, DeepSWE v1.1 at 65.9%, CursorBench v3.2 at 69.9% and GPQA Diamond at 94.9%. The launch material also reports roughly 53 turns and about half a billion input tokens to resolve long-horizon tasks on AA-Briefcase, against roughly 103 turns and two billion tokens for a competing frontier model — an argument, again vendor-reported, that Grok is cheap per completed job even at a modest per-token rate.

Two of those Grok numbers need handling before you quote them. The first is that its headline GPQA Diamond score of 94.9% comes from a benchmark Artificial Analysis has since retired as saturated, so it no longer separates frontier models the way it once did. The second is the number you will see in older coverage: an Artificial Analysis Intelligence Index of 61 for Grok 4.6. That was the pre-restatement scale. Artificial Analysis recalibrated the index in September 2026, and the live figure is 44 at #20 of 200, with a cost of $1.86 per Intelligence Index task. Anyone ranking Grok 4.6 against Claude Fable 5 or GPT-5.6 Sol on a 61 is comparing readings from two different instruments.

A two-column scoreboard titled 'Fugu Max vs Grok 4.6 - the scoreboard' contrasting six dimensions. Fugu Max: output price $6.00 per 1M, input price $2.00 per 1M, cached input $0.25 per 1M, context not published, benchmarks six wins with no figures, evidence vendor-reported only. Grok 4.6: output price $6.00 per 1M, input price $2.00 per 1M, cached input $0.50 per 1M, context 500K with a 200K repricing cliff, benchmarks GDPVal 1,753 and DeepSWE 65.9%, evidence Artificial Analysis Index 44 ranked #20 of 200. Footer reads 'Fugu Max figures vendor-reported by Sakana AI; Grok 4.6 rows SpaceXAI-reported and per Artificial Analysis. No independent Fugu Max evaluation exists.' OrcaRouter logo bottom-right.

What an orchestrator's token price does not include

Sakana's own coverage of the release concedes the structural problem. Because Fugu Max switches between models inside a single task, it "may reduce context cache reuse and also generate extra orchestration tokens" — and the company did not disclose the resulting cache hit rate or a total token cost per task. Every agent the coordinator spawns writes reasoning and output text, and all of it is billed at $6.00 per million. The rate is identical to Grok 4.6's. The volume is not, and volume is the variable a pricing page cannot show you.

Both vendors are pushing you toward the same correction — stop comparing rates, start comparing cost per finished job — but only one of them hands you a number to start from. Grok 4.6 arrives with a published turns-and-tokens profile. Fugu Max arrives with an instruction to measure it yourself.

There is a fair reading of Fugu Max's silence and it deserves stating. Max is the cost-optimised tier of the Fugu line, shipped the same day as Fugu Ultra v2, which is the capability push at $5.00 input and $30.00 output. Sakana's visual argument for Max is a Pareto plot — capability against cost — and on a Pareto plot a raw benchmark score is beside the point; score-per-dollar is the whole claim. If that is the argument, printing six winning scores without their costs would be the misleading chart, not the missing one. What the release still owes a buyer is a denominator, and it does not provide one.

Where Grok 4.6 is genuinely exposed

Terminal work is Grok 4.6's weakest published surface. Its Terminal-Bench v3.0 score sits at 26%, well behind the front of the field. Be careful with the neighbouring figure: the same model posts 88.4% on Terminal-Bench v2.1, and those are different tests. You will see the two blended in coverage, and the blend flatters Grok considerably.

The second exposure is that reasoning cannot be turned off. Thinking tokens bill on every request, so Grok 4.6's effective output cost depends on how hard the model decides to think about your prompt. The effort dial gives you control at the low end, but there is no zero setting.

Against that, Grok 4.6 has something Fugu Max cannot currently offer at any price: a number. Every row above can be checked by a third party, and the Artificial Analysis entry means one already has been. Fugu Max's six wins were published on the same day as the model, by the company that built it, and as of writing nobody outside Sakana has run it.

Calling one, piloting the other

Grok 4.6 is on OrcaRouter at SpaceXAI's list price — $2.00 per million input, $0.50 on a cache hit, $6.00 per million output — passed through with 0% markup, so a vendor price change reaches our gateway the same day rather than on a migration schedule. It is OpenAI-compatible on the same base URL as the rest of the catalogue, which means you can put it in a failover chain or behind a conditional routing rule instead of hard-coding a provider into a production path, and you can A/B it against another model without a second contract. It moved 46.6 million tokens through the gateway in the last seven days.

Fugu Max is not on our catalogue. It reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, and the company says an existing Fugu integration upgrades to it with a single parameter change. Because both speak the same request format, an evaluation that puts them behind one flag is a base-URL swap rather than a rewrite — which is the right way to settle this particular argument, since the argument is about tokens per finished task and only your own workload can produce that number.

A screenshot of the Sakana AI announcement page (English UI, captured September 11, 2026) headed 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' and dated September 11, 2026, showing the introductory argument that a system deploying a multi-trillion-parameter model for a simple lookup is wasteful, and a cost-versus-performance chart placing Fugu Max on a frontier formed by Fugu models above the frontier formed by single models.

Which one to buy

Grok 4.6 is the defensible default. At a rate card identical to Fugu Max's it gives you a 500K context window you can plan against, a version you can pin, four levels of reasoning control, an independent index placement at 44 with a cost-per-task figure beside it, and months of third-party measurement. Its costs are specific and known: the 200K repricing cliff, reasoning that always bills, and a terminal-work profile that trails the field.

Fugu Max is the more interesting bet and, on the evidence available, the less inspectable one. The design case is sound — if a coordinator reliably picks the cheapest model that can finish your task, your median cost per completed job falls even when your rate card does not. But at $2.00 / $6.00 the token rate is not where Fugu Max's advantage lives; that rate is exactly Grok 4.6's. The advantage has to come from spending fewer tokens, and Sakana has not published the measurement that would show it does.

So run the comparison the way both vendors are asking you to. Put Grok 4.6 on your gateway, call Fugu Max behind a feature flag, and count output tokens per completed task on work that looks like yours. If Fugu Max finishes in fewer tokens than a single model at the same rate, the cache discount and the coordination are real money and the switch is justified. If it does not, you are paying identical rates for a system whose composition can change underneath you without its version number moving.

For everything you can decide this quarter without that measurement, start with the model that has a scoreboard.

A screenshot of the OrcaRouter model page for Grok 4.6 (English UI, captured September 11, 2026), showing the NEW and Featured badges, the identifier grok/grok-4.6, by SpaceXAI dated 2026-08-12, vision, tools, JSON and reasoning capability tags, a 500K-token context window with text, image and file input, a p50 TTFT of 10.00 seconds, list pricing of $2.00 per 1M input tokens and $6.00 per 1M output tokens, traffic of 46.6 million tokens over 7 days, and an OpenAI-compatible base URL with Python and cURL code samples.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily