
Fugu Max vs Qwen3.8-Max: Same Price, Same Benchmark, One Published Number
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.24 / $0.73 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Fugu Max and Qwen3.8-Max list at the same rate: $2.00 per million input tokens and $6.00 per million output tokens. Sakana AI shipped Fugu Max on September 11, 2026 and Alibaba has been serving Qwen3.8-Max since early August, and the two vendors landed on an identical rate card from opposite directions — one pricing a routing policy, the other pricing a 2.4-trillion-parameter model. The tie removes price from the decision entirely, which is unusually clean. What is left is evidence, and the two sides differ there more than anywhere else. Both vendor tables name Terminal Bench 2.1. Alibaba prints 86.6 for Qwen3.8-Max. Sakana says Fugu Max takes best overall score on Terminal Bench 2.1 and prints no figure at all — for that test or the other five it claims to win.
The one benchmark both sides name
This is the whole matchup compressed into a single row, so it is worth being precise about it.
Sakana's release lists six benchmarks on which Fugu Max takes best overall score: Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. Five are recognised public evaluations and the sixth, SWEFish, is Sakana's own coding benchmark. For none of the six does the release state a value, a margin, or a competitor's number to sit beside it. The parallel claim — that Fugu Max expands the cost-performance Pareto frontier on seven of ten benchmarks — is likewise a statement about position on a chart rather than a point on an axis.
Alibaba's published rows for Qwen3.8-Max include Terminal-Bench 2.1 at 86.6, OSWorld-Verified at 86.1, GPQA Diamond at 92.6 and SWE-bench Pro at 67.7. All vendor-reported, all numbers. So on the single test both companies chose to name, one has published a figure and the other has published a rank. You cannot conclude from that which system is better — Sakana's "best overall" could mean 91 or 87. What you can conclude is that as of publication no public evidence answers the question, and the vendor making the larger claim is the one declining to quantify it.
Identical prices, opposite construction
• Input — Fugu Max $2.00 per 1M tokens vs Qwen3.8-Max $2.00 per 1M tokens
• Output — Fugu Max $6.00 per 1M tokens vs Qwen3.8-Max $6.00 per 1M tokens
• Cached input — Fugu Max $0.25 per 1M tokens vs Qwen3.8-Max $0.25 per 1M tokens, and Artificial Analysis independently measures an 88 percent cache discount on the Qwen side
• Long-prompt surcharge — Fugu Max no tier stated in the release vs Qwen3.8-Max flat to the window with no surcharge
• Context and output — Fugu Max neither figure published vs Qwen3.8-Max a 1M-token window with roughly a 128K output ceiling
• Input modality — Fugu Max text, with the pool's own capabilities behind it vs Qwen3.8-Max text, image, video and PDF
• Architecture — Fugu Max a trained coordinator over an undisclosed pool, expanded for Max with open-weights and specialised models including the NVIDIA Nemotron family through a collaboration with NVIDIA vs Qwen3.8-Max one sparse mixture-of-experts model at roughly 2.4 trillion total parameters and about 95 billion active per token
• Weights — Fugu Max closed, pool membership undisclosed by design vs Qwen3.8-Max open weights published for the Max-class checkpoint under a custom licence
• Independent evaluation — Fugu Max none published, and no Artificial Analysis page exists for any Fugu model vs Qwen3.8-Max a live Artificial Analysis entry with published speed, verbosity and cost-per-task figures
Four of those rows read "not published" on the Fugu side and carry a figure on the Qwen side. When the rates are identical, that is the entire comparison: the same money buys one system you can describe and one you cannot.

One column has a third party behind it
Artificial Analysis scores Qwen3.8-Max at 40 on the restated v4.3 Intelligence Index, ranked #30 of 200, with a cost of $2.67 per Intelligence Index task and 37.8 output tokens per second — slow enough to rank #154 of 200 on speed. The same entry records 180 million output tokens generated across the index, more than twice the median model's 89 million, and an 88 percent cache discount.
Two corrections apply before any of that is quoted. The first is that the widely circulated figure of 58 — or 58.1, or 58.4 — for Qwen3.8-Max is from before Artificial Analysis recalibrated the index in September 2026, and the live page carries the "Updated" badge that marks a restated score. The older write-ups also carried a different effort tax from the previous scale, 64 turns per task against 14 for the prior Qwen generation at roughly $1.14 per Intelligence Index task; the current page shows $2.67. The second is a genuine warning rather than a scale artefact: Artificial Analysis separately measured Qwen3.8-Max's hallucination rate rising from 23 percent to 40 percent against the prior generation. For a model marketed at autonomous multi-step work, that is a cost you budget for in verifiers, not a footnote.
Fugu Max has no independent entry of any kind. Every figure attached to it originates with Sakana, and as of writing no Artificial Analysis page exists for Fugu Max or any of its siblings. That is not an accusation — a model released hours ago will not have third-party coverage — but it does mean the two columns are not equally checkable, and it will stay that way until someone outside Sakana runs the thing.
Two ways to overspend
Both of these systems have moved the cost of an answer away from the cost of a token, and they do it in opposite directions. Qwen3.8-Max is thrifty by the token and lavish by the answer: 180 million output tokens on the Intelligence Index, twice the median, at $2.67 per task. It works long rather than being large, and the 88 percent cache discount is the lever that controls the bill — a long agent run that keeps re-reading the same context stays cheap, one that keeps generating fresh text does not.
Fugu Max is expensive by the token and unpredictable by the answer. Its coordinator spawns sub-agents, each of which writes reasoning and output text billed at $6.00 per million, plus the coordinator's own synthesis. Sakana commits that multi-agent runs are charged one blended rate pegged to the top-tier participating model and that adding agents does not multiply the bill, which removes the worst-case fan-out blow-up that makes naive multi-agent systems unaffordable. It does not remove the volume. And on the volume question the release is silent in a way that matters: Sakana's own coverage states that switching models mid-task may reduce context cache reuse and generate extra orchestration tokens, and confirms the company did not disclose Max's cache hit rate or its total token cost per task.
So the same $2.00 / $6.00 rate is a predictable number on one side and an unbounded multiple with a floor on the other. Qwen3.8-Max's verbosity is measurable before you commit; Fugu Max's is not.
The escape hatch test
This is where the usual lock-in argument inverts, and it is the most useful thing in the pairing.
Sakana's case for Fugu Max is resilience. A pool that spans open-weights and specialised models, expanded for Max with the NVIDIA Nemotron family, means no single frontier vendor's deprecation, repricing or regional withdrawal can take your product down — a real hedge, and one the company is explicit about selling. But the hedge is unverifiable from outside. You cannot see the pool, cannot opt a member in or out, cannot self-host any of it, and cannot run it in the EU/EEA at all. A promise about a pool you are not allowed to inspect is a weaker guarantee than it sounds.
Qwen3.8-Max runs the argument the other way. It is one vendor's model and therefore exposed to one vendor's decisions — but Alibaba published open weights for the Max-class Qwen3.8-2.4T-A95B checkpoint under a custom licence, which means the escape hatch is real rather than rhetorical: if pricing or terms change, the model can be served from your own infrastructure. For a team whose actual fear is a vendor pulling the rug, a downloadable checkpoint beats a description of a pool.
Calling either one
Qwen3.8-Max is on our catalogue at $2.00 in and $6.00 out per million tokens, which is Alibaba's list price with 0% markup passed through, so a vendor price change lands on the same key the same day rather than on a migration schedule. It is OpenAI-compatible on the same base URL as the rest of the catalogue, which means you can put it behind a conditional routing rule, a failover chain, or a model-fusion panel without a second contract or a code change — and automatic failover covers a rate-limited or degraded provider path. It moved 127.7 million tokens through the gateway in the last seven days.
Fugu Max is not on OrcaRouter. It reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, and the company says an existing Fugu integration upgrades with a one-line parameter change. Because both speak the same request format, the cheapest way to settle the benchmark gap is to stop waiting for Sakana to publish it: run both behind one flag on your own workload and count output tokens per completed task. That is the measurement neither vendor has supplied.

Which one, and when
Qwen3.8-Max is the choice you can audit, price and abandon. At the identical rate card it gives you a 1M-token window with no long-prompt surcharge, image, video and PDF input, a downloadable Max-class checkpoint, a live independent entry measuring the things that drive your bill, and 88 percent off cached input. Its costs are equally specific: it is slow at 37.8 tokens per second, ranked near the bottom of the field on speed, it is enormously verbose at 180 million tokens on the index, and its measured hallucination rate roughly doubled against the prior generation — which is a serious concern for exactly the autonomous work it is sold for. Use the cache discount hard and budget for a verifier.
Fugu Max is the bet that coordination beats capability. If your work is long-horizon and checkable, and a coordinator that spawns a verifier finishes tasks that a single model fails twice, then the orchestration tokens are cheaper than the retries and the identical rate card is not a tie at all. What you are buying is persistence, in a system whose pool you cannot pin, whose context window is not published, and whose six benchmark wins have no values attached — from a vendor that chose to name the test and withhold the number.
Both products cost the same per token. Only one of them tells you what you get for it.

