A hero title card reading 'Fugu Ultra v2 vs Qwen3.8-Max' with the subtitle 'Why the price tag is the wrong number', three pill badges reading 'Output: $30.00 vs $6.00', 'Above 272K: rate doubles' and 'Independent Fugu score: none', a footer line reading 'Sakana AI, September 11 2026 vs Alibaba, August 2026', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Fugu Ultra v2 vs Qwen3.8-Max: why the price tag is the wrong number

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Fugu Ultra v2 costs $30.00 per million output tokens. Qwen3.8-Max costs $6.00. That is a five-to-one gap, and if you pick between these two agentic systems on that ratio you will pick wrong — because neither of them is actually billed the way its price list suggests. On Artificial Analysis's own measurement, Qwen3.8-Max generated 180 million output tokens to complete the Intelligence Index, more than twice the median model's 89 million, at a cost of $2.67 per task. It is a thrifty model by the token and a lavish one by the answer. Sakana AI's Fugu Ultra v2, shipped September 11, 2026, has the opposite shape: an undisclosed pool of other labs' models that it spawns and coordinates per request, billed at one blended rate that Sakana says does not multiply as agents are added. One is a single 2.4-trillion-parameter model that thinks out loud for a long time. The other is a small coordinator that hires help. Comparing their token prices tells you almost nothing about which is cheaper to run.

Two opposite ways to be expensive

Start with the mechanism, because it explains every other difference in this comparison.

Qwen3.8-Max is a sparse mixture-of-experts model — 2.4 trillion total parameters, roughly 95 billion active per token — that reaches frontier-adjacent results by working longer rather than by being larger. Artificial Analysis measured it at 37.8 output tokens per second, ranking it #154 of 200 models on speed, with 180 million output tokens emitted across the Intelligence Index (#70 of 200 on verbosity). Its cost per Intelligence Index task reads $2.67, and its cache discount is unusually deep at 88%, which is the lever that actually controls the bill here: a long agent run that keeps re-reading the same context is cheap, and one that keeps generating fresh text is not.

Fugu Ultra v2 approaches the same problem from the other end. It is an orchestrator — a small trained coordinator, reported around 7 billion parameters, that reads a request, assembles an agent team along Thinker, Worker and Verifier roles from Sakana's TRINITY work, and then synthesises an answer. Its token bill is the sum of what its sub-agents produce plus the coordinator's own synthesis text. Sakana commits in its pricing FAQ that you are charged one blended rate pegged to the top-tier model involved and that adding agents does not multiply the bill, which removes the classic multi-agent blow-up where fan-out multiplies cost. What it does not remove is volume. At $30.00 per million output tokens, the number that matters is how many agents ran and how much they wrote, and that is a number neither you nor Sakana can predict from the pricing page.

That is the honest framing of the price gap: it is a rate, not a total. A five-times-higher rate applied to a workload whose output volume you cannot forecast is not five times the cost — it is an unbounded multiple with a predictable floor.

A two-column scoreboard titled 'Fugu Ultra v2 vs Qwen3.8-Max Scoreboard' contrasting six dimensions. Fugu Ultra v2: price $5.00/$30.00 per 1M, context 1M tokens, above 272K input doubles to $10, AA Index none published, speed not measured, weights closed with pool undisclosed. Qwen3.8-Max: price $2.00/$6.00 per 1M, context 1M tokens, above 272K no surcharge, AA Index 40 ranked #30 of 200, speed 37.8 tokens/sec, weights open checkpoint published. Footer reads 'Fugu Ultra v2 figures vendor-reported; no independent evaluation. Qwen3.8-Max figures per Artificial Analysis, recalibrated scale.' OrcaRouter logo bottom-right.

The spec sheet, and the one row that decides it

• Price — Fugu Ultra v2 $5.00 in / $30.00 out per 1M tokens vs Qwen3.8-Max $2.00 in / $6.00 out

• Cached input — Fugu Ultra v2 $0.50 per 1M vs Qwen3.8-Max $0.25 per 1M read, and an 88% cache discount measured by Artificial Analysis

• Long-context surcharge — Fugu Ultra v2 input doubles to $10.00 and output to $45.00 above 272K tokens vs Qwen3.8-Max no long-prompt surcharge, flat to the window

• Context window — Fugu Ultra v2 1M tokens vs Qwen3.8-Max 1M tokens

• Output ceiling — Fugu Ultra v2 not published as a single figure vs Qwen3.8-Max roughly 128K tokens

• Input modality — Fugu Ultra v2 text, with the pool's own capabilities behind it vs Qwen3.8-Max text, image, video and PDF

• Weights — Fugu Ultra v2 closed, pool membership undisclosed by design vs Qwen3.8-Max open weights published for the Max-class checkpoint under a custom licence

• Regional availability — Fugu Ultra v2 not sold in the EU/EEA vs Qwen3.8-Max available without a regional restriction

• Independent evaluation — Fugu Ultra v2 none published, and no Artificial Analysis page exists for any Fugu model vs Qwen3.8-Max a live Artificial Analysis entry with published speed, verbosity and cost-per-task figures

Read the last two rows together and the matchup sharpens. Qwen3.8-Max is a model you can pin by version, read a licence for, download a checkpoint of, and check against a third party's scoreboard. Fugu Ultra v2 is a system you cannot inspect, cannot self-host, cannot buy in Europe, and cannot verify against anyone outside Sakana. The 272K surcharge compounds it: past that threshold Fugu Ultra v2's input rate doubles and its output rate rises by half, while Qwen3.8-Max's rate does not move. For the long-horizon document and repository work both systems are built for, the cheaper headline is also the flatter one.

The Qwen3.8-Max number everyone quotes is out of date

If you have read about this model anywhere in the last two weeks you have probably seen an Artificial Analysis Intelligence Index of 58, or 58.1, or 58.4. Those figures were real and they are no longer current. Artificial Analysis recalibrated its index to v4.3 on September 7, 2026, compressing scores across the board, and the live Qwen3.8-Max page now reads 40, placing it #30 of 200 models — with the "Updated" badge that marks a restated score. The older write-ups also carried the effort tax measured on the previous scale: 64 turns per task against 14 for the prior Qw​en generation, and roughly $1.14 per Intelligence Index task. The current page shows $2.67 per task, 37.8 output tokens per second, and 180 million output tokens on the index.

Two things follow. First, on the recalibrated scale Qwen3.8-Max sits in the same band as the rest of the frontier's second tier rather than level with it, and any comparison you read that ranks it against Claude Opus 5 or Kimi K3 on a 58 is comparing snapshots from different scales. Second, and more usefully for this matchup, the correction is one-directional: Qwen3.8-Max has a number that can be wrong and then fixed. Fugu Ultra v2 has no number. Every Fugu figure in this article comes from Sakana's own evaluation, and as of September 11 there is no Artificial Analysis page for Fugu Ultra v2 or any other Fugu model — the URL 404s. When a vendor publishes a scoreboard and no independent party has run the model, the scoreboard is the entire record.

Where they actually overlap, and where they don't

Sakana reports Fugu Ultra v2 as best or joint-best on five of eight benchmarks — GDP.pdf, Chartography, SWEFish, DeepSWE and Toolathon — and top-two on seven of eight. The two concrete figures in the release are Chartography at 48.3, against 27.3 for Claude Opus 5 and 29.5 for Claude Fable 5 in the same table, and DeepSWE at 74.3. Ali​baba's own published rows for Qwen3.8-Max include Terminal-Bench 2.1 at 86.6, OSWorld-Verified at 86.1, GPQA Diamond at 92.6 and SWE-bench Pro at 67.7.

Notice what those two lists have in common: nothing. Not one benchmark appears in both vendor tables. There is no row where you can hold a Sakana figure and an Ali​baba figure side by side and read off a winner, which means the honest answer to "which is better at coding agents" is that no published evidence answers it. What exists is one system with eight vendor-chosen rows and no outside verification, and one system with vendor rows plus an independent entry that measures cost and speed rather than capability head-to-head. That is a gap in the evidence, not a gap in the models, and it should change how you buy: you cannot pick between these on benchmarks, so pick on the properties you can actually verify.

A screenshot of the Sakana AI announcement page (English UI, captured September 11, 2026) headed 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' dated September 11, 2026, with the opening paragraphs arguing that the frontier which matters is two-dimensional — capability against cost — and stating that Fugu Ultra v2 pushes peak performance higher without indispensable reliance on the frontier models it orchestrates.

Open weights versus an undisclosed pool

This is where the usual lock-in argument inverts, and it is the most interesting thing about the pairing.

Sakana's pitch for Fugu Ultra v2 is sovereignty. A swappable pool of open and specialised models means no single vendor's pricing decision, deprecation schedule or policy shift can take your product down, and the release makes the point explicitly by noting the pool's composition changed in v2. The company also states plainly that Fugu Ultra v2's scores were achieved without Claude Fable 5, Claude Fable 5.1 or GPT-6 Astra in the agent pool — claiming wins over the three strongest models available while confirming it calls none of them. Whatever the pool contains, it is not the top of the market by Sakana's own accounting.

But the hedge has a cost that is not on the price list. You cannot see the pool, you cannot opt a member in or out of the Ultra tier, you cannot self-host any of it, and you cannot run it in the EU/EEA at all — Singapore-based orchestration over other labs' weights is a different risk profile from owning the weights. Qwen3.8-Max inverts that exactly. It is one vendor's model and therefore exposed to one vendor's decisions, but Ali​baba published open weights for the Max-class Qwen3.8-2.4T-A95B checkpoint under a custom licence, which means the escape hatch is real rather than rhetorical: if pricing or terms change, the model can be served from your own infrastructure. For a team whose actual fear is a vendor pulling the rug, a downloadable checkpoint is a stronger guarantee than a promise about a pool you are not allowed to inspect.

There is a second-order point for anyone running agents across both. Qwen3.8-Max is on our catalogue at $2.00 in and $6.00 out, which is Ali​baba's list price with 0% markup passed through — a vendor price change lands on the same key the same day, and automatic failover covers a rate-limited or degraded provider path. Fugu Ultra v2 is not on OrcaRouter. It reaches you through Sakana's own OpenAI-compatible API and several third-party platforms, and moving from an earlier Fugu is a single-line parameter change. A router is not a substitute for a coordinator, and a coordinator is not a substitute for the ability to fail over; if you want both properties you are running two vendors' systems and should budget for two bills.

Which one you should pick

Pick Qwen3.8-Max if you need a model you can forecast, verify and escape from. It is a third of Fugu Ultra v2's output rate at list, it has no long-context cliff, it accepts images, video and PDFs where the Fugu pool's modality is whatever the underlying agents happen to support, it publishes a checkpoint you can fall back to, it sells everywhere, and it has a live independent entry measuring the things that actually drive your bill — 37.8 tokens per second and a cost-per-task figure you can model before you commit. Its weaknesses are equally specific and worth naming: it is slow, ranked #154 of 200 on speed, it is enormously verbose at 180 million tokens on the index, and Artificial Analysis separately measured its hallucination rate rising from 23% to 40% against the prior generation, which is a serious concern for exactly the autonomous multi-step work it is marketed for. Budget for a verifier, and use the deep cache discount aggressively.

Pick Fugu Ultra v2 if your work is long-horizon and checkable, your budget is exploratory rather than production, and you are outside the EU/EEA. The structural case for a coordinator is real and the vendor's own examples point at it consistently — a 123-experiment automated research run over fourteen hours on a single H100, a pure-Python Rubik's cube solver that completed all 300 scrambled cubes where two anonymised frontier baselines crashed outright. Those are vendor-selected examples with anonymised baselines and they prove less than they appear to, but the pattern is coherent: on tasks where correctness is checkable and the hard part is persistence, spawning a verifier beats asking one model harder. What you are buying is that persistence, at a rate five times Qwen3.8-Max's, in a system whose quality you cannot audit and whose pool you cannot pin. Pilot it scoped, instrument output tokens per completed task rather than per token, and set a ceiling before the first run.

A screenshot of the OrcaRouter model page for Qwen3.8 Max (English UI, captured September 11, 2026), showing the model title, the 'Featured' badge, the identifier qwen/qwen3.8-max, vision, tools, JSON and reasoning capability tags, a 1M-token context window, text and image and video input, a p50 TTFT of 10.00 seconds, list pricing of $2.00 per 1M input tokens and $6.00 per 1M output tokens, and an OpenAI-compatible base URL with Python and cURL code samples.

The reason the price tag is the wrong number is that both of these products have moved the cost of an answer away from the cost of a token. Fugu Ultra v2 does it by fanning out to models it will not name. Qwen3.8-Max does it by thinking for 180 million tokens. If you already know your token volume per task, the arithmetic is trivial and you should do it. Most teams do not know that number, and for them the decision reduces to something simpler: Qwen3.8-Max is the choice you can audit, price and abandon, and Fugu Ultra v2 is the choice that bets coordination beats capability. Both are defensible. Only one of them comes with the receipts.