A hero title card for "Fugu Max vs GPT-5.6 Sol" with the subtitle "The pricing claim was aimed at the middle of the family", three pill badges reading "$6.00 vs $20.00 output", "40-60% claim vs Terra", "Six wins, no scores", a footer line reading "Sakana AI vs OpenAI - September 2026", and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Fugu Max vs GPT-5.6 Sol: the pricing claim was aimed at the middle of the family

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Fugu Max's price justification names a G​PT-5.6 model, and it is not the flagship. Sakana AI released Fugu Max on September 11, 2026 at $2.00 per million input tokens and $6.00 per million output tokens, and stated that its output pricing runs 40 to 60 percent below Claude Sonnet 5, GPT-5.6 Terra and Kimi K3. GPT-5.6 Sol — the flagship of the same Ope​nAI family, listed at $4.00 input and $20.00 output per million — does not appear in that sentence. That omission is not an oversight; it is the arithmetic. Sol is more than three times Fugu Max's output rate, a gap that would have made the surrounding "within striking distance of elite models at two to six times lower cost" framing read as a different kind of claim entirely. So this matchup is worth running on Sol's terms rather than the vendor's chosen ones.

Checking the claim against the family it points at

Sakana's 40-to-60-percent figure is testable, because the model it names has a published rate. GPT-5.6 Terra lists at $2.00 input and $12.00 output per million tokens. Fugu Max lists at $2.00 input and $6.00 output. Run the comparison and two things fall out. On output, $6.00 against $12.00 is a cut of exactly 50 percent — dead centre of the stated 40-to-60 window, which is a sign the window was drawn around this comparison rather than discovered from it. On input, the two rates are identical to the cent.

That is a well-constructed claim, and constructing it well is what a release page is for. But it is calibrated to the middle of the G​PT-5.6 ladder. Terra is the balanced tier of a three-model family — Sol above it for hard complex work, Luna below it for speed and cost. Naming Terra sets the "elite model" benchmark at the middle rung and lets Fugu Max claim half its output price honestly. Naming Sol would have required a different sentence, because the honest version is that Fugu Max costs 70 percent less on output than Sol, and 50 percent less on input — a much larger gap that raises the obvious follow-up question of what capability you are giving up for it.

• Input price — Fugu Max $2.00 per 1M tokens vs GPT-5.6 Sol $4.00 per 1M tokens

• Output price — Fugu Max $6.00 per 1M tokens vs GPT-5.6 Sol $20.00 per 1M tokens

• Cached input — Fugu Max $0.25 per 1M tokens vs GPT-5.6 Sol caching priced as a discount on the base input rate

• Context — Fugu Max not published vs GPT-5.6 Sol 1M tokens

• Output ceiling — Fugu Max not published vs GPT-5.6 Sol 128K tokens

• Input modality — Fugu Max not published vs GPT-5.6 Sol text, image and file

• Capability tags — Fugu Max none published vs GPT-5.6 Sol vision, tool use, JSON mode, reasoning

• Endpoints — Fugu Max one OpenAI-compatible surface vs GPT-5.6 Sol chat completions and the responses API

• Benchmark figures — Fugu Max six wins asserted with no scores vs GPT-5.6 Sol published vendor figures including Terminal-Bench 2.0 at 91.9% and SWE-bench Pro at 64.6%

Why the middle rung is the right rung to name

There is a defensible version of Sakana's choice, and it deserves stating before the criticism. An orchestrator that routes each task to the leanest capable model is not competing with a flagship on flagship work. It is competing with a mid-tier model on mid-tier work, and offering to do a subset of it more cheaply. If the pitch is "most of your traffic does not need Sol," then Terra is the correct comparator, because Terra is what most teams actually reach for when they do not want to pay Sol prices. Measuring against the tier your customer would otherwise buy is more honest than measuring against the most expensive model available.

The difficulty is that the same sentence also claims performance "within striking distance of elite models." Those two halves of the pitch pull in opposite directions. If Fugu Max is a Terra substitute, then the elite-model framing is doing marketing work it cannot support, because Terra is explicitly not positioned as elite — it is positioned as the good-enough tier. If Fugu Max genuinely sits within striking distance of Sol, then the pricing comparison should have been drawn against Sol, where the gap is largest and most flattering. A release that makes the aggressive capability claim and the conservative price comparison in the same paragraph is having it both ways, and the reader has no way to tell which half is load-bearing.

A two-column scoreboard titled "Fugu Max vs GPT-5.6 Sol - the scoreboard". Left column Fugu Max: Output price $6.00 / 1M, Input price $2.00 / 1M, Cached input $0.25 / 1M, Context not published, Modality not published, Benchmarks six wins with no figures. Right column GPT-5.6 Sol: Output price $20.00 / 1M, Input price $4.00 / 1M, Cached input discount on base rate, Context 1M tokens with 128K output, Modality text, image and file, Benchmarks 91.9% on Terminal-Bench 2.0. Footer reading "Fugu Max figures vendor-reported by Sakana AI. Sakana's 40-60% output claim names GPT-5.6 Terra at $2.00 / $12.00, not Sol.", with the OrcaRouter logo bottom-right.

Six benchmarks, no numbers, and one name that should raise a question

The capability half of the pitch rests on a list. Fugu Max "achieves best overall score" on Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. No scores are printed for any of them. SWEFish is Sakana's own internal coding benchmark, which means one of the six wins is scored on a test the vendor wrote and has not released.

GPT-5.6 Sol reports figures. Sol is listed at 91.9% on Terminal-Bench 2.0 and 64.6% on SWE-bench Pro in third-party comparison data. Note the version mismatch on the terminal benchmark: Sol's published number is for Terminal-Bench 2.0 and Sakana claims a win on 2.1, which is a different evaluation and cannot be set against it. That is worth flagging because the two claims look adjacent in a headline and are not comparable in fact.

The relevant comparison for this matchup would be Sol against a Fugu model on a shared evaluation. Third-party comparison pages do carry GPT-5.6 Sol against the base Sakana Fugu generation — Sol ahead on Terminal-Bench 2.0 at 91.9% to 80.2% and on SWE-bench Pro at 64.6% to 59% — but those are figures for the earlier generation, not Fugu Max, and the same sources explicitly decline to name an overall winner because the shared-benchmark coverage is too thin. There is no published comparison of any kind between Fugu Max and GPT-5.6 Sol. Anyone asserting one is extrapolating.

No independent party has evaluated Fugu Max, and no Artificial Analysis entry exists for it or for any Fugu model. Every number in the release, including the six wins, originates with the vendor. Sol, by contrast, has been generally available since July 9, 2026 and appears in third-party leaderboards — a level of external scrutiny Fugu Max has not had a chance to receive.

What you are actually buying on each side

Configuring GPT-5.6 Sol is a known quantity. It carries a 1M-token context window, a 128K output ceiling, text, image and file input, vision, tool use, JSON mode and reasoning as documented capabilities, and two endpoints — chat completions and the responses API — so an existing integration has somewhere to go. It shows a p50 time-to-first-token of 6.68 seconds across the traffic we serve, which is slower than the fastest models on our catalogue and worth knowing before you build an interactive product on it.

Configuring Fugu Max is closer to buying a service level than a model. You get one OpenAI-compatible endpoint, a single-line parameter change to move to it from an earlier Fugu, and a coordinator that decides which models to call. What you do not get, from the release page, is a context window, an output ceiling, a modality list, a latency figure, or a benchmark score. The coordinator's routing decisions are not exposed by design. The pool expanded this release to include open-weights and specialised models with the NVIDIA Nemotron family named through an NVIDIA collaboration, and is otherwise undisclosed.

Sakana's stated rationale for that opacity is defensible — a vendor-agnostic pool that can be swapped is protection against deprecations, price changes and regional withdrawals. It is a real hedge, and the $0.25 cached input rate suggests Sakana expects the long-context, high-repetition workloads where that hedge would matter most. But it does mean the specification you are purchasing is a promise about a supply chain rather than a description of a product.

A screenshot of the Sakana AI release page at sakana.ai/fugu-max-release (captured September 11, 2026, English UI), showing the sakana.ai wordmark, the headline 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' dated September 11, 2026, the opening copy arguing that the frontier which matters is two-dimensional with capability on one axis and cost on the other, a 'Try Sakana Fugu Max and Fugu Ultra v2' link, and a Pareto chart plotting performance against cost with a red 'Fugu Max' point at the low-cost end, a red 'Fugu Ultra' point at the top, a shaded 'Frontier formed by Fugu models' region and grey points labelled 'Frontier formed by single models'.

How to run the test

The right way to settle this is not to read either vendor's page. Take a sample of your own traffic — a few hundred real requests, ideally including the ones your team currently escalates — and run them against both endpoints with token accounting on. Cost them on tokens per completed task rather than tokens per call, set an output budget before you start so a fan-out cannot surprise you, and measure quality by whether a second person agrees the answer was acceptable. That experiment answers the question in a week and costs little.

The friction is that Fugu Max is not on OrcaRouter. It reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, which means a second integration, a second credential set and a second invoice. GPT-5.6 Sol is on our catalogue at Ope​nAI's list price with 0% markup — the provider's rate passed through rather than marked up, so a change to Ope​nAI's pricing is live on your existing key the same day, with automatic failover across provider paths when one is rate-limited or unavailable. It sits behind the same key as more than 200 other models, so when the pilot ends you keep the integration and change one string.

Bottom line

GPT-5.6 Sol wins the comparison Sakana declined to draw. It costs twice as much on input and more than three times as much on output, and in exchange you get a published context window, a 128K output ceiling, documented multimodal input, named capability tags, two endpoints, and vendor benchmarks that third parties have had months to examine. If your work is the kind Sol was built for — deep multi-step reasoning, large-scale software engineering, long-horizon agentic tasks — the price difference is the cost of buying something you can specify.

Fugu Max wins a narrower and genuinely interesting bet, and it is the cheaper one by a wide margin: 50 percent below GPT-5.6 Terra on output, 70 percent below Sol, with the lowest cache rate in either rate card. If your traffic is high-volume, text-only, checkable, and mostly does not need flagship reasoning, an orchestrator that routes to the leanest sufficient model is a coherent design and the $6.00 output rate is where the value sits. Take the release's own framing at face value and pilot it as a Terra substitute rather than a Sol one — and if you want the flagship with the provider's list price passed through untouched, GPT-5.6 Sol is a key away on OrcaRouter.

A screenshot of the OrcaRouter model page for GPT-5.6 Sol (openai/gpt-5.6-sol, captured September 11, 2026, English UI), showing the Featured badge, the openai/gpt-5.6-sol model ID, the Vision, Tools, JSON and Reasoning tags, the listing date 2026-07-09, the /v1/chat/completions and /v1/responses endpoints, the pricing tiles reading INPUT $4.00 and OUTPUT $20.00 per 1M tokens with p50 TTFT 6.68s and 153.0M tokens of 7-day traffic, the 1M token context with 128K max output and text + image + file input, and the OpenAI-compatible Python sample pointing at api.orcarouter.ai/v1.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily