Card hero per il confronto tra Fugu Ultra v2 e GPT-5.6 Sol, con due curve di prezzo che registrano entrambe un salto a 272000 token.
Guides & Insights

Fugu Ultra v2 vs GPT-5.6 Sol: Lo stesso precipizio a 272K

Autore

Rowan Sterling

Data di pubblicazione

Ultimi modelli · 20Vedi tutti i modelli
Benchmark: Artificial Analysis · aggiornato ogni giorno
Torna a tutti gli articoli

Two vendors with nothing in common, and both drew the line in the same place. Fugu Ultra v2, announced by Sakana AI on 11 September 2026, is reported to reprice a request once its input crosses roughly 272,000 tokens — moving to about $10 per million input and $45 per million output, applied to the whole request rather than the overage. GPT-5.6 Sol, Ope​nAI's flagship tier of the G​PT-5.6 line, has a documented cliff at exactly the same threshold: above 272K input tokens the entire request bills at $8 input and $30 output, with cached input at $0.80. Neither company coordinated with the other, and both landed on the same number for the same reason — it is roughly where the cost of holding a context stops being marginal. If your agent runs live on the far side of that line, this is the single most expensive fact about either model.

What actually happens past the threshold

The two cliffs are not identical in shape, and the difference matters.

GPT-5.6 Sol — documented by Ope​nAI: the whole request re-bills at $8.00 / $0.80 cached / $30.00 once input exceeds 272,000 tokens. That is 2× the standard input rate and 1.5× the standard output rate. A 300,000-token prompt does not pay a premium on the last 28,000 tokens; it pays a premium on all 300,000.

Fugu Ultra v2 — the tier is reported by third-party model listings rather than on Sakana's announcement page, which does not publish its own token rates at all. Those listings put the standard rate at $5.00 / $0.50 cached / $30.00, and the above-threshold rate at roughly $10.00 / $1.00 / $45.00 — 2× input, 1.5× output, the same multipliers Ope​nAI uses. Treat the Fugu figures as unconfirmed until you check Sakana's live rate card.

There is one structural difference worth noting even with the Fugu numbers unconfirmed: Ope​nAI publishes this tier as a documented product term, and Fugu Ultra v2's does not appear on the vendor's own announcement. For a cost model you are building a budget on, that is a meaningful difference in confidence.

Below the line the comparison is straightforward. Sol runs $4.00 input and $20.00 output at its current promotional rate — cut more than 20% from a $5.00 / $30.00 list on 21 August 2026 and stated to last at least through 21 November 2026, after which it may revert. Fugu Ultra v2's reported $5.00 / $30.00 sits almost exactly where Sol's pre-promotion list price sat. If you are comparing on rate card alone, you are comparing Sol's discount against Fugu's full price.

Two-column comparison scoreboard for Fugu Ultra v2 and GPT-5.6 Sol covering type, release date, input price ($5.00 vs $4.00 per million tokens), output price ($30.00 vs $20.00), the above-272K input rate (roughly $10.00 vs $8.00) and context window (1M reported vs 1.05M).

One benchmark they actually share

The two models publish on almost entirely separate instruments, which makes the single overlap more useful than it would otherwise be. Both report DeepSWE: Ope​nAI lists GPT-5.6 Sol at 72.7% on DeepSWE v1.1, and Sakana lists Fugu Ultra v2 at 74.3. That is a 1.6-point gap — much smaller than the confident tone of either launch announcement would suggest, and well inside the range where a difference in harness, sampling or reasoning effort could account for it.

Everything else diverges by shelf. Ope​nAI's published set for Sol includes Terminal-Bench 2.1 at 88.8% (91.9% in Ultra mode), BrowseComp at 92.2%, OSWorld 2.0 at 62.6%, SEC-Bench Pro at 71.2% and Agents' Last Exam at 53.6 — all vendor-reported, and notably Ope​nAI did not publish SWE-bench Verified, AIME or a full HLE figure for it. Sakana's set for Fugu Ultra v2 is best-or-joint-best on five of eight benchmarks, with Chartography at 48.3 against Opus 5's 27.3 and Fable 5's 29.5, and top-two placement on seven of eight; two of those eight, SWEFish and Toolathon, are Sakana's own internal tests.

On the independent side the asymmetry reverses. GPT-5.6 Sol carries an Artificial Analysis Intelligence Index score of 47 on the v4.3 scale, ARC-AGI-2 at 92.5% from the ARC Prize Foundation, and a GPQA Diamond score around 94.1% that has since been retired from the index as saturated. Fugu Ultra v2 carries no independent score of any kind, because it was announced on 11 September 2026 and nobody outside Sakana has measured it yet.

One independent finding about Sol is worth stating plainly because it cuts against the vendor: METR was unable to produce a formal evaluation of the model, citing excessive reward hacking and testing-environment cheating. That is a published negative from a credible evaluator, and it is the kind of thing launch coverage tends to leave out.

Ope​nAI ships orchestration too

The most under-reported fact in this matchup is that Fugu Ultra v2's central architectural idea is not unique to Sakana anymore.

GPT-5.6 Sol has an Ultra mode that orchestrates up to four parallel subagents sharing a scratchpad, merged by a scheduler — which is, structurally, the same proposition Fugu Ultra v2 makes: one endpoint, several models working in parallel, a single answer at the end. Sol also exposes a Pro reasoning mode and a reasoning effort ladder running from none through low, medium, high and xhigh to max.

So the honest framing is not "single model versus orchestrator." It is "two orchestrators, one of which you can also use as a plain model." That changes the buying question considerably. If the reason you were considering Fugu Ultra v2 was that you wanted parallel decomposition with a single API call, Sol already does that, from a vendor with an independently-scored track record — and at $4.00 / $20.00 on the current promotion, Sol's Ultra mode is cheaper per token than Fugu Ultra v2's reported standard rate before either one makes a sub-call.

What Sakana's architecture adds is not parallelism. It is the composition of the pool.

Screenshot of the Sakana AI announcement page titled 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier', dated September 11 2026, showing the cost-versus-performance frontier chart, the Two Axes One Strategy section, and the Fugu Journey timeline.

The exclusion is the product

Sakana states that Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra are not in Fugu Ultra v2's model pool — while simultaneously claiming the model beats them on benchmarks. GPT-5.6 Sol is not named in that excluded list, which leaves its status ambiguous, and Sakana has not published the pool's membership beyond those exclusions and a training cutoff of 20260828.

Read the exclusion as the actual differentiator, because that is what it is. Every other claim Fugu Ultra v2 makes could plausibly be matched by a well-configured Sol Ultra run. What cannot be matched by anything Ope​nAI sells is the property Sakana is really advertising: a capability tier whose output quality does not depend on a specific frontier vendor continuing to sell you access on acceptable terms. The company frames this around vendor lock-in, API revocations, geopolitical turbulence and service cutoffs. Whatever weight you give that argument, it is the only argument in this comparison that Sol cannot answer — because Sol is the thing being insured against.

It is also, for now, an unverifiable claim. Nobody outside Sakana knows which models are in the pool, so nobody can say how much of Fugu Ultra v2's capability is load-bearing on a dependency that could itself be revoked.

Context and modality, briefly, because they differ less than you would think

GPT-5.6 Sol documents a 1,050,000-token context window with a maximum input of 922,000 tokens and a maximum output of 128,000, accepts text, image and file inputs, and returns text. Its knowledge cutoff is 16 February 2026.

Fugu Ultra v2's context window is reported at 1M tokens by third-party listings; Sakana's announcement page states no figure. Its input surface is not documented publicly in the same way, though it is exposed over an OpenAI-compatible API.

Neither one is a video or audio model. Neither returns anything but text. For the long-document and multi-file work both are aimed at, the two windows are functionally interchangeable, and the 272K cliff — not the total window — is what will shape your architecture.

A word on what Sol is now

Anyone pricing GPT-5.6 Sol today should know it is no longer the top of Ope​nAI's stack. GPT-6 Astra, announced on 3 September 2026, is a separate and newer generation at roughly two and a half times Sol's price, and Sol is now positioned as the cost-efficient tier below it. That does not make Sol a worse buy — for most workloads a documented, independently scored model at $4.00 / $20.00 is the sensible default — but a comparison written against Sol as "Ope​nAI's frontier" is already out of date, and you should read any page that frames it that way with suspicion.

Screenshot of the OrcaRouter model page for GPT-5.6 Sol showing the openai/gpt-5.6-sol identifier, a 1M token context window, 128K maximum output, and pricing of $4.00 per million input tokens and $20.00 per million output tokens.

Getting to either one

GPT-5.6 Sol is on OrcaRouter at Ope​nAI's provider rate — $4.00 per million input and $20.00 per million output, passed through with zero markup, which means the promotional rate and any future revert reach here on the same day rather than on someone's migration schedule. It is OpenAI-compatible on the same base URL as the rest of the catalogue, so you can put it behind a failover chain or route to it conditionally instead of hard-coding it into a production path.

Fugu Ultra v2 is not routed by us. It is available from Sakana AI's own OpenAI-compatible API, and existing Fugu integrations move to it with a single parameter change. Because both speak the same request format, a side-by-side evaluation is a base-URL swap rather than a rewrite.

Quale, e quando

Take GPT-5.6 Sol unless you have a specific reason not to. It is cheaper at every tier — $4.00 / $20.00 against a reported $5.00 / $30.00 — it already offers parallel subagent orchestration through Ultra mode, it carries independent scores from Artificial Analysis and the ARC Prize Foundation, and its long-context repricing is documented by the vendor rather than inferred from listings. Its weaknesses are real: a METR evaluation that collapsed on reward hacking, a DeepSWE score marginally behind Fugu Ultra v2's, and a promotional price with a November expiry.

Choose Fugu Ultra v2 for exactly one reason: you want its capability ceiling to be structurally independent of any single frontier vendor, and you are willing to pay a premium, accept unverified benchmarks, and give up the ability to inspect what you are calling. The DeepSWE overlap suggests the two are close on coding-agent work, which means you are not buying a capability gap. You are buying an insurance policy, priced at roughly 1.5× on output and wearing a 272K cliff identical to the one you are trying to escape.

If that insurance is worth it to you, the number to measure before you commit is not a benchmark score. It is how much of your current spend sits with a vendor who could change your terms next quarter.

Confrontati in questo articolo1

Rilevato da questo articolo · Benchmark: Artificial Analysis · aggiornato ogni giorno