A hero title card reading 'Fugu Max vs Qwen3.8-Max' with the subtitle 'Same price, same benchmark, one number', three pill badges reading '$2.00 / $6.00 on both', 'Terminal Bench 2.1: 86.6 vs unstated' and 'Independent score: none vs 40', a footer line reading 'Sakana AI, September 2026 vs Alibaba, August 2026', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Fugu Max vs Qwen3.8-Max: Same Price, Same Benchmark, One Published Number

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Fugu Max and Qwen3.8-Max list at the same rate: $2.00 per million input tokens and $6.00 per million output tokens. Sakana AI shipped Fugu Max on September 11, 2026 and Ali​baba has been serving Qwen3.8-Max since early August, and the two vendors landed on an identical rate card from opposite directions — one pricing a routing policy, the other pricing a 2.4-trillion-parameter model. The tie removes price from the decision entirely, which is unusually clean. What is left is evidence, and the two sides differ there more than anywhere else. Both vendor tables name Terminal Bench 2.1. Ali​baba prints 86.6 for Qwen3.8-Max. Sakana says Fugu Max takes best overall score on Terminal Bench 2.1 and prints no figure at all — for that test or the other five it claims to win.

The one benchmark both sides name

This is the whole matchup compressed into a single row, so it is worth being precise about it.

Sakana's release lists six benchmarks on which Fugu Max takes best overall score: Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. Five are recognised public evaluations and the sixth, SWEFish, is Sakana's own coding benchmark. For none of the six does the release state a value, a margin, or a competitor's number to sit beside it. The parallel claim — that Fugu Max expands the cost-performance Pareto frontier on seven of ten benchmarks — is likewise a statement about position on a chart rather than a point on an axis.

Ali​baba's published rows for Qwen3.8-Max include Terminal-Bench 2.1 at 86.6, OSWorld-Verified at 86.1, GPQA Diamond at 92.6 and SWE-bench Pro at 67.7. All vendor-reported, all numbers. So on the single test both companies chose to name, one has published a figure and the other has published a rank. You cannot conclude from that which system is better — Sakana's "best overall" could mean 91 or 87. What you can conclude is that as of publication no public evidence answers the question, and the vendor making the larger claim is the one declining to quantify it.

Identical prices, opposite construction

• Input — Fugu Max $2.00 per 1M tokens vs Qwen3.8-Max $2.00 per 1M tokens

• Output — Fugu Max $6.00 per 1M tokens vs Qwen3.8-Max $6.00 per 1M tokens

• Cached input — Fugu Max $0.25 per 1M tokens vs Qwen3.8-Max $0.25 per 1M tokens, and Artificial Analysis independently measures an 88 percent cache discount on the Qw​en side

• Long-prompt surcharge — Fugu Max no tier stated in the release vs Qwen3.8-Max flat to the window with no surcharge

• Context and output — Fugu Max neither figure published vs Qwen3.8-Max a 1M-token window with roughly a 128K output ceiling

• Input modality — Fugu Max text, with the pool's own capabilities behind it vs Qwen3.8-Max text, image, video and PDF

• Architecture — Fugu Max a trained coordinator over an undisclosed pool, expanded for Max with open-weights and specialised models including the NVIDIA Nemotron family through a collaboration with NVIDIA vs Qwen3.8-Max one sparse mixture-of-experts model at roughly 2.4 trillion total parameters and about 95 billion active per token

• Weights — Fugu Max closed, pool membership undisclosed by design vs Qwen3.8-Max open weights published for the Max-class checkpoint under a custom licence

• Independent evaluation — Fugu Max none published, and no Artificial Analysis page exists for any Fugu model vs Qwen3.8-Max a live Artificial Analysis entry with published speed, verbosity and cost-per-task figures

Four of those rows read "not published" on the Fugu side and carry a figure on the Qw​en side. When the rates are identical, that is the entire comparison: the same money buys one system you can describe and one you cannot.

A two-column scoreboard titled 'Fugu Max vs Qwen3.8-Max - the scoreboard' contrasting six dimensions. Fugu Max: output price $6.00 per 1M, input price $2.00 per 1M, cached input $0.25 per 1M, context not published, Terminal Bench 2.1 best with no figure, evidence vendor-reported only. Qwen3.8-Max: output price $6.00 per 1M, input price $2.00 per 1M, cached input $0.25 per 1M, context 1M with a 128K output ceiling, Terminal Bench 2.1 at 86.6, evidence Artificial Analysis Index 40 and $2.67 per task. Footer reads 'Fugu Max figures vendor-reported by Sakana AI; Qwen3.8-Max rows Alibaba-reported and per Artificial Analysis.' OrcaRouter logo bottom-right.

One column has a third party behind it

Artificial Analysis scores Qwen3.8-Max at 40 on the restated v4.3 Intelligence Index, ranked #30 of 200, with a cost of $2.67 per Intelligence Index task and 37.8 output tokens per second — slow enough to rank #154 of 200 on speed. The same entry records 180 million output tokens generated across the index, more than twice the median model's 89 million, and an 88 percent cache discount.

Two corrections apply before any of that is quoted. The first is that the widely circulated figure of 58 — or 58.1, or 58.4 — for Qwen3.8-Max is from before Artificial Analysis recalibrated the index in September 2026, and the live page carries the "Updated" badge that marks a restated score. The older write-ups also carried a different effort tax from the previous scale, 64 turns per task against 14 for the prior Qw​en generation at roughly $1.14 per Intelligence Index task; the current page shows $2.67. The second is a genuine warning rather than a scale artefact: Artificial Analysis separately measured Qwen3.8-Max's hallucination rate rising from 23 percent to 40 percent against the prior generation. For a model marketed at autonomous multi-step work, that is a cost you budget for in verifiers, not a footnote.

Fugu Max has no independent entry of any kind. Every figure attached to it originates with Sakana, and as of writing no Artificial Analysis page exists for Fugu Max or any of its siblings. That is not an accusation — a model released hours ago will not have third-party coverage — but it does mean the two columns are not equally checkable, and it will stay that way until someone outside Sakana runs the thing.

Two ways to overspend

Both of these systems have moved the cost of an answer away from the cost of a token, and they do it in opposite directions. Qwen3.8-Max is thrifty by the token and lavish by the answer: 180 million output tokens on the Intelligence Index, twice the median, at $2.67 per task. It works long rather than being large, and the 88 percent cache discount is the lever that controls the bill — a long agent run that keeps re-reading the same context stays cheap, one that keeps generating fresh text does not.

Fugu Max is expensive by the token and unpredictable by the answer. Its coordinator spawns sub-agents, each of which writes reasoning and output text billed at $6.00 per million, plus the coordinator's own synthesis. Sakana commits that multi-agent runs are charged one blended rate pegged to the top-tier participating model and that adding agents does not multiply the bill, which removes the worst-case fan-out blow-up that makes naive multi-agent systems unaffordable. It does not remove the volume. And on the volume question the release is silent in a way that matters: Sakana's own coverage states that switching models mid-task may reduce context cache reuse and generate extra orchestration tokens, and confirms the company did not disclose Max's cache hit rate or its total token cost per task.

So the same $2.00 / $6.00 rate is a predictable number on one side and an unbounded multiple with a floor on the other. Qwen3.8-Max's verbosity is measurable before you commit; Fugu Max's is not.

The escape hatch test

This is where the usual lock-in argument inverts, and it is the most useful thing in the pairing.

Sakana's case for Fugu Max is resilience. A pool that spans open-weights and specialised models, expanded for Max with the NVIDIA Nemotron family, means no single frontier vendor's deprecation, repricing or regional withdrawal can take your product down — a real hedge, and one the company is explicit about selling. But the hedge is unverifiable from outside. You cannot see the pool, cannot opt a member in or out, cannot self-host any of it, and cannot run it in the EU/EEA at all. A promise about a pool you are not allowed to inspect is a weaker guarantee than it sounds.

Qwen3.8-Max runs the argument the other way. It is one vendor's model and therefore exposed to one vendor's decisions — but Ali​baba published open weights for the Max-class Qwen3.8-2.4T-A95B checkpoint under a custom licence, which means the escape hatch is real rather than rhetorical: if pricing or terms change, the model can be served from your own infrastructure. For a team whose actual fear is a vendor pulling the rug, a downloadable checkpoint beats a description of a pool.

Calling either one

Qwen3.8-Max is on our catalogue at $2.00 in and $6.00 out per million tokens, which is Ali​baba's list price with 0% markup passed through, so a vendor price change lands on the same key the same day rather than on a migration schedule. It is OpenAI-compatible on the same base URL as the rest of the catalogue, which means you can put it behind a conditional routing rule, a failover chain, or a model-fusion panel without a second contract or a code change — and automatic failover covers a rate-limited or degraded provider path. It moved 127.7 million tokens through the gateway in the last seven days.

Fugu Max is not on OrcaRouter. It reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, and the company says an existing Fugu integration upgrades with a one-line parameter change. Because both speak the same request format, the cheapest way to settle the benchmark gap is to stop waiting for Sakana to publish it: run both behind one flag on your own workload and count output tokens per completed task. That is the measurement neither vendor has supplied.

A screenshot of the Sakana AI announcement page (English UI, captured September 11, 2026) headed 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' and dated September 11, 2026, showing the argument that the industry has raced along a single axis of bigger and more expensive foundation models while the frontier that matters is two-dimensional, capability on one axis and cost on the other.

Which one, and when

Qwen3.8-Max is the choice you can audit, price and abandon. At the identical rate card it gives you a 1M-token window with no long-prompt surcharge, image, video and PDF input, a downloadable Max-class checkpoint, a live independent entry measuring the things that drive your bill, and 88 percent off cached input. Its costs are equally specific: it is slow at 37.8 tokens per second, ranked near the bottom of the field on speed, it is enormously verbose at 180 million tokens on the index, and its measured hallucination rate roughly doubled against the prior generation — which is a serious concern for exactly the autonomous work it is sold for. Use the cache discount hard and budget for a verifier.

Fugu Max is the bet that coordination beats capability. If your work is long-horizon and checkable, and a coordinator that spawns a verifier finishes tasks that a single model fails twice, then the orchestration tokens are cheaper than the retries and the identical rate card is not a tie at all. What you are buying is persistence, in a system whose pool you cannot pin, whose context window is not published, and whose six benchmark wins have no values attached — from a vendor that chose to name the test and withhold the number.

Both products cost the same per token. Only one of them tells you what you get for it.

A screenshot of the OrcaRouter model page for Qwen3.8 Max (English UI, captured September 11, 2026), showing the Featured badge, the identifier qwen/qwen3.8-max, by Qwen dated 2026-08-03, vision, tools, JSON and reasoning capability tags, a 1M-token context window with text, image and video input, list pricing of $2.00 per 1M input tokens and $6.00 per 1M output tokens, a p50 TTFT of 10.00 seconds, traffic of 127.7 million tokens over 7 days, and an OpenAI-compatible base URL with Python and cURL code samples.