A hero title card for "Fugu Max vs GLM-5.3" with the subtitle "The value model is undercut by the volume model", three pill badges reading "$6.00 vs $3.96 output", "7,962.7M tokens per 7d", "$2.00 vs $1.26 input", a footer line reading "Sakana AI vs Z.ai - September 2026", and the OrcaRouter logo in the bottom-right corner.
Engineering & Research

Fugu Max vs GLM-5.3: the value model is undercut by the volume model

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Fugu Max was priced to win on cost and GLM-5.3 is cheaper on both axes. Sakana AI's coordinator launched on September 11, 2026 at $2.00 per million input tokens and $6.00 per million output tokens, built to route each task to the least expensive model that can handle it. Z.ai's GLM-5.3, listed since August 18, 2026, sits at $1.26 input and $3.96 output per million on OrcaRouter — between a third and two-fifths below Fugu Max on both, from a single text-only model with a published 1M context window and a 128K output ceiling. GLM-5.3 also happens to be the busiest model on our catalogue by a wide margin. That combination — cheaper per token, and demonstrably carrying production traffic at scale — makes this the matchup where the orchestration premium is hardest to justify.

Cheaper on both axes, with the specification published

The comparison is uncomfortable for Fugu Max before any benchmark enters the picture, because it is not a case of trading capability for price. GLM-5.3 is a frontier-adjacent flagship in its own right — Z.ai describes it as delivering roughly a 50% coding improvement over GLM-5.2 and matching Mythos 5 on selected cybersecurity capabilities. It is not a budget model being offered as a cheap alternative. It is a flagship that happens to be priced below a cost-optimised orchestrator.

• Input price — Fugu Max $2.00 per 1M tokens vs GLM-5.3 $1.26 per 1M tokens

• Output price — Fugu Max $6.00 per 1M tokens vs GLM-5.3 $3.96 per 1M tokens

• Cached input — Fugu Max $0.25 per 1M tokens vs GLM-5.3 caching priced as a discount on the base input rate

• Context — Fugu Max not published vs GLM-5.3 1M tokens

• Output ceiling — Fugu Max not published vs GLM-5.3 128K tokens

• Modality — Fugu Max not published vs GLM-5.3 text-only, documented as such

• Capability tags — Fugu Max none published vs GLM-5.3 tool use, JSON mode, reasoning

• Benchmark figures — Fugu Max six wins asserted with no scores vs GLM-5.3 vendor-reported DeepSWE 46.2 to 66.9 over the prior generation, plus a published Artificial Analysis intelligence score

• Weights — Fugu Max closed, pool undisclosed vs GLM-5.3 from a lab with an open-weights lineage

• Observed traffic on OrcaRouter — Fugu Max none, not carried vs GLM-5.3 7,962.7M tokens over the last seven days

Notice what the last two rows do together. GLM-5.3 comes from a lineage with open weights, which is the strongest available form of the vendor-independence argument Sakana is selling as the entire premise of Fugu Max. And it has an observable traffic figure an order of magnitude beyond most models we serve. If the question is "what do I run for text-heavy production work at a low rate," the evidence points somewhere other than the orchestrator.

A two-column scoreboard titled "Fugu Max vs GLM-5.3 - the scoreboard". Left column Fugu Max: Output price $6.00 / 1M, Input price $2.00 / 1M, Context not published, Output ceiling not published, Modality not published, Observed traffic not carried. Right column GLM-5.3: Output price $3.96 / 1M, Input price $1.26 / 1M, Context 1M tokens, Output ceiling 128K tokens, Modality text-only and documented, Observed traffic 7,962.7M tokens over 7 days. Footer reading "Fugu Max figures vendor-reported by Sakana AI; GLM-5.3 pricing and traffic from OrcaRouter, September 11 2026.", with the OrcaRouter logo bottom-right.

The volume argument, and why it is not just a popularity contest

A traffic number is not a quality number, and it should not be read as one. What 7.96 billion tokens in seven days does demonstrate is that GLM-5.3 has been run at production scale by many independent teams, on real workloads, and has not generated a mass migration away from it. That is a weak signal about quality and a strong signal about operational fitness: latency is stable, throughput holds, the API behaves under load, and the failure modes are known to a large enough population that they surface publicly. A model released the same week as Fugu Max has none of that behind it.

The latency row sharpens the point. GLM-5.3 shows a p50 time-to-first-token of 4.33 seconds across the traffic we serve. For agentic work — the workload both of these products are aimed at — time-to-first-token compounds across every turn of every subagent the coordinator spawns. A cheaper orchestrator that spawns four agents is not cheaper if each one waits several seconds before emitting anything, and that cost never appears on a rate card.

Fugu Max's counter-argument is that its coordinator routes to the leanest capable model per request, which in principle means many tasks never touch an expensive backend at all. Sakana's pool expansion this release added open-weights and specialised models including the NVIDIA Nemotron family through an NVIDIA collaboration, so the cheap rungs of that ladder are real. The question is whether the rungs are cheaper than $3.96 per million output tokens. For most text tasks, they are not — that is a low bar, and open-weights models running at hosted rates generally clear it only by a small margin.

Questions worth asking before you pick

Is an orchestrator cheaper than a single open-weights model? Not by default, and the intuition runs the wrong way. Orchestration adds a coordinator's synthesis tokens and, on multi-agent runs, the output of every agent in the team. Sakana's Fugu pricing removes the worst case by billing one blended rate based on the top-tier participating model rather than stacking agents, which is a genuine concession. But the base rate you are blending is $6.00 output, and the single model you would otherwise call is $3.96. Orchestration saves money when it prevents retries, not when it adds parallelism for its own sake.

Does a text-only limitation matter for either one? For GLM-5.3 it is documented, so it is a constraint you can plan around — no vision, no audio, no video, and the catalogue says so. For Fugu Max, modality is simply not published. That is worse than a limitation, because you cannot tell whether a pipeline that sends a PDF will work until you try it. An unstated constraint is not the same as an absent one.

What happens to each one when the underlying models change? GLM-5.3 is one model at one version, trained and served by Z.ai. If Z.ai changes its pricing, the change lands on your key the same day through OrcaRouter at 0% markup, and you can see it and respond. Fugu Max's behaviour is a function of a pool whose members Sakana does not disclose, so a frontier lab's deprecation or price change silently alters what you are buying without the product name changing. Sakana presents this as resilience. It is resilience for the vendor and opacity for the buyer.

A screenshot of the Sakana AI release page at sakana.ai/fugu-max-release (captured September 11, 2026, English UI), showing the sakana.ai wordmark, the headline 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' dated September 11, 2026, the opening copy arguing that the frontier which matters is two-dimensional with capability on one axis and cost on the other, a 'Try Sakana Fugu Max and Fugu Ultra v2' link, and a Pareto chart plotting performance against cost with a red 'Fugu Max' point at the low-cost end, a red 'Fugu Ultra' point at the top, a shaded 'Frontier formed by Fugu models' region and grey points labelled 'Frontier formed by single models'.

Where routing decisions actually get made

Both of these are, in effect, routing decisions — Fugu Max makes them inside an opaque coordinator, and you make them at the API layer. OrcaRouter is the second kind: one API for over 200 models with a routing DSL that lets you express the policy explicitly, automatic failover when a provider path degrades, and model fusion when you want to combine outputs deliberately rather than trusting a black box to do it. GLM-5.3 sits on that catalogue at Z.ai's list price with 0% markup, which is worth more than usual for a model with this much traffic behind it: the rate you see is the rate Z.ai publishes, passed through untouched.

The practical difference is auditability. When Fugu Max picks a model for your request, you find out nothing about which one it was or why. When you write the routing policy yourself, the decision is in your repository, reviewable by whoever reviews your code, and changeable without waiting on a vendor. For teams under any kind of compliance or cost-accounting obligation, that is not a philosophical difference.

The verdict

GLM-5.3 wins this comparison on the terms that matter most to a production team: it is cheaper on input and output, its context window and output ceiling are published, its modality limits are stated, it carries tool use, JSON mode and reasoning as documented capabilities, it comes from an open-weights lineage, and it has a seven-day traffic figure approaching eight billion tokens. There is no reading of the public evidence on which Fugu Max is the better-evidenced purchase at this price point.

Fugu Max keeps one argument, and it is a real one: for tasks where a coordinator's second pass catches a mistake the first pass would have shipped, cost per completed task can beat cost per token. That is worth a scoped pilot if your work is checkable, high-volume, and you are outside the EU/EEA — with an output-token ceiling set in advance and a measurement of tokens per finished task. Otherwise, the single model that costs a third less and has been running under production load for a month is the answer, and it is on OrcaRouter at Z.ai's list price with the provider's rate passed through.

A screenshot of the OrcaRouter model page for GLM 5.3 (z-ai/glm-5.3, captured September 11, 2026, English UI), showing the NEW and Featured badges, the z-ai/glm-5.3 model ID, the Tools, JSON and Reasoning tags over a text-only model, the listing date 2026-08-18, the /v1/chat/completions endpoint, the pricing tiles reading INPUT $1.26 and OUTPUT $3.96 per 1M tokens with p50 TTFT 4.33s and 7,962.7M tokens of 7-day traffic, the 1M token context with 128K max output, and the OpenAI-compatible Python sample pointing at api.orcarouter.ai/v1.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily