A hero title card reading 'Fugu Max vs Kimi K3' with the subtitle 'Sakana chose this yardstick', three pill badges reading '$6.00 vs $15.00 output', '60% below, on output' and '1M flat vs window unpublished', a footer line reading 'Sakana AI, September 2026 vs Moonshot AI, July 2026', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Fugu Max vs Kimi K3: Sakana Chose This Yardstick, So Let's Read It

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Sakana AI's release for Fugu Max names exactly three models to justify its pricing, and Moonshot AI's Kimi K3 is one of them. The claim is that Fugu Max's output rate runs 40 to 60 percent below those three, which makes Kimi K3 — not Grok 4.6, not Qwen3.8-Max — the benchmark Sakana itself picked for value. That is an unusually generous thing for a vendor to do: it hands you a yardstick and tells you where to hold it. So this page does the obvious thing and holds it. Fugu Max shipped on September 11, 2026 at $2.00 per million input and $6.00 per million output tokens. Kimi K3 lists at $3.00 and $15.00. The 60 percent claim on output is arithmetically correct. What the sentence does not mention is that the model it is undercutting publishes a full record of independently measured results, and the model doing the undercutting publishes none.

The 40-to-60-percent sentence, examined

• Input — Fugu Max $2.00 per 1M tokens vs Kimi K3 $3.00 per 1M tokens, a 33 percent saving rather than the 40 to 60 percent the release leads with

• Output — Fugu Max $6.00 per 1M tokens vs Kimi K3 $15.00 per 1M tokens, which is 60 percent below, the top of Sakana's stated range

• Cached input — Fugu Max $0.25 per 1M tokens vs Kimi K3 $0.30 per 1M tokens, a gap small enough that cache-heavy loops will not notice it

• Context window — Fugu Max no figure published in the release vs Kimi K3 1,048,576 tokens

• Long-prompt repricing — Fugu Max no tier stated vs Kimi K3 flat across the entire window with no surcharge at any length

• Deployment — Fugu Max vendor API only, pool membership undisclosed vs Kimi K3 downloadable weights under the Kimi K3 License

• Independent evaluation — Fugu Max none, and no Artificial Analysis page exists for any Fugu model vs Kimi K3 an Artificial Analysis Intelligence Index of 44 and a standardised 93.40% on SWE-bench Verified from Vals AI

Note the shape of that first row. The headline range, 40 to 60 percent, is the range across three named competitors, and Kimi K3 is the one that produces the 60. On input alone the comparison is 33 percent, which is still a real saving but a different sentence. That is not a criticism of the pricing — $2.00 and $6.00 are genuinely low numbers for a system that claims frontier-adjacent results. It is a note about which number in the release is doing the persuading, and about how the paragraph is arranged around it.

Kimi K3 is on our catalogue at Moonshot's own rate — $3.00 per million input, $0.30 on a cache hit, $15.00 per million output — passed through with 0% markup, so if Moonshot moves its pricing the new rate is live on the same key the same day rather than on a migration cycle. That matters more here than usual, because a price cut upstream is precisely what erodes or widens the 60 percent gap this whole matchup rests on.

What the yardstick publishes

Kimi K3's record is the opposite of sparse. Moonshot's own model card reports Terminal-Bench 2.1 at 88.3, GPQA Diamond at 93.5, HLE-Full at 43.5 rising to 56.0 with tools, SWE-Marathon at 42.0, OSWorld-Verified at 84.8, MCPMark-Verified at 94.5 and DeepSWE at 67.5. All vendor-reported, and there is a lot of it.

The independent layer exists too, which is the part that matters for this comparison. Artificial Analysis scores Kimi K3 at 44 on the v4.3 Intelligence Index — second among open-weight models, behind GLM-5.3 at 45 — with recorded throughput of 37.9 output tokens per second and a time to first token of 3.80 seconds. Vals AI's standardised agentic harness puts it at 93.40% on SWE-bench Verified, behind Claude Opus 5 and DeepSeek V4 Pro but close. It also leads the Frontend Code Arena at 1679 Elo, ahead of Claude Fable 5 at 1631 and GPT-5.6 Sol at 1618. One caution: if you have seen Kimi K3 quoted at 57 or 60 on the Intelligence Index, those are pre-restatement figures from before Artificial Analysis recalibrated the scale in September 2026 and should not be repeated alongside current numbers.

There is a usage signal as well, and it comes from our own gateway rather than anyone's marketing: Kimi K3 moved roughly 3.27 billion tokens through OrcaRouter in the last seven days, the heaviest traffic of any model in this comparison. That is not a quality measurement. It is a large sample of what developers reach for when the only cost of choosing is the token price, and it says the flat 1M window and the open weights have already won a substantial share of real long-horizon work.

A two-column scoreboard titled 'Fugu Max vs Kimi K3 - the scoreboard' contrasting six dimensions. Fugu Max: output price $6.00 per 1M, input price $2.00 per 1M, cached input $0.25 per 1M, context not published, benchmarks six wins with no figures, evidence vendor-reported only. Kimi K3: output price $15.00 per 1M, input price $3.00 per 1M, cached input $0.30 per 1M, context 1M flat with no surcharge, benchmarks Terminal-Bench 2.1 at 88.3, evidence Artificial Analysis Index 44 and SWE-bench Verified 93.40%. Footer reads 'Fugu Max figures vendor-reported by Sakana AI; Kimi K3 rows Moonshot-reported and per Artificial Analysis and Vals AI.' OrcaRouter logo bottom-right.

The scoreboard with no numbers on it

Sakana reports Fugu Max as taking best overall score on six benchmarks: Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. It also says Max expands the cost-performance Pareto frontier on seven of ten benchmarks. Both are statements about rank on a chart. Neither comes with a value, a margin, or a competitor's figure beside it.

The two lists barely touch, which is the practical problem. Kimi K3 reports Terminal-Bench 2.1 at 88.3; Sakana says Fugu Max takes best overall on Terminal Bench 2.1 and prints nothing to compare against it. That single row is the cleanest illustration of the whole matchup: both vendors name the same test, one prints a number, and the other prints the word "best". Until Sakana publishes the figure, no published evidence answers whether Fugu Max beats Kimi K3 on the benchmark Sakana chose to lead with.

There is a fair reading of the release worth stating. Fugu Max is the cost tier — shipped the same day as Fugu Ultra v2, which is the capability push at $5.00 input and $30.00 output — and its argument is drawn on a Pareto plot where score-per-dollar, not score, is the claim. On that framing, a bare benchmark number would be the less honest statistic. The problem is that the release also omits the denominator: Sakana's own coverage concedes that switching models mid-task may reduce context cache reuse and generate extra orchestration tokens, and confirms the company did not disclose Max's cache hit rate or its total token cost per task. So the one measurement that would settle a cost-tier argument is the one not supplied.

Weights you can hold, or a pool you cannot see

This is where the two products stop being comparable at all.

Kimi K3 ships downloadable weights, which is a genuine escape hatch: if pricing or terms change, the model can be served from your own infrastructure. The licence is narrower than "open weights" usually implies. It is the Kimi K3 License rather than an OSI-approved grant, and the frequently repeated description of it as modified MIT is disputed. The terms require a separate agreement with Moonshot for model-as-a-service businesses above $20M in revenue over any twelve consecutive months, plus attribution for products above 100M monthly users or $20M monthly revenue, with internal use exempt. And the practical scale is real: the checkpoint is roughly 1.5 terabytes and Moonshot recommends 64 or more accelerators, so self-hosting is an inference-cluster decision rather than a laptop one.

Fugu Max offers none of that — no weights, no inspectable pool, no self-hosting option — and in exchange imposes no licence condition, because there is nothing to license. Sakana's stated benefit is the inverse of control: a pool that spans open-weights and specialised models, expanded for Max with the NVIDIA Nemotron family, means no single upstream vendor's deprecation or repricing can take your product down. That is a real hedge. It is also unverifiable from outside, because the membership is deliberately not published, and it means a Fugu Max result carries an expiry date that a Kimi K3 result does not.

Open weights and an undisclosed pool are two different answers to the same fear. One says you can leave. The other says there is nothing to leave.

Where the flat window decides it

Kimi K3's 1,048,576-token context is flat — no long-context tier, no surcharge at any length. Fugu Max's release states no window at all, so there is nothing to compare it to and nothing to plan against. For the long-horizon agentic coding both products are aimed at, where sessions spend most of their life deep in the context, that asymmetry matters more than the rate card does.

Speed is the mirror image of the same problem. Kimi K3's 37.9 tokens per second with a 3.80-second time to first token is measured, and it is slow — genuinely a constraint for a model built around long loops, and Artificial Analysis flags it as notably verbose. Sakana publishes no latency or throughput figure for Fugu Max, which for an orchestrator is arguably honest rather than evasive: the response time depends on which models it picks and how many passes it makes. But it means you cannot model the wall-clock cost of switching without measuring it yourself. For the earlier Fugu Ultra, users publicly reported runs stretching toward 30 minutes. That is the June-2026 model and not evidence about Max, but it is the number to beat when you benchmark.

A screenshot of the Sakana AI announcement page (English UI, captured September 11, 2026) headed 'Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier' and dated September 11, 2026, showing the opening argument about the two-dimensional frontier of capability against cost and the statement that Fugu Max expands the Pareto efficiency frontier by orchestrating Sakana's largest pool of open and specialized models to date.

The verdict, in Sakana's own terms

Sakana picked Kimi K3 as one of the three models Fugu Max undercuts, and on the output rate the claim holds: $6.00 against $15.00 is 60 percent below, exactly the top of the stated range. Take the release at its word and Fugu Max is the cheaper product.

Now apply the yardstick the release avoided. Kimi K3 is 33 percent more expensive on input and 150 percent more expensive on output — and for that money it gives you a 1M-token window with no repricing cliff, a downloadable checkpoint under a revenue-gated licence, an independent Intelligence Index placement of 44, a standardised SWE-bench Verified score of 93.40%, first place on the Frontend Code Arena, and the heaviest real traffic of any model in this comparison. Fugu Max gives you a lower rate on both sides of the token, six unquantified benchmark wins, a cache rate half of K3's with no published hit rate, and a pool you cannot inspect.

The honest conclusion is that Sakana chose a yardstick that flatters it on price and gave you nothing to check on capability. If your workload is high-volume and your tasks are the kind a coordinator can decompose reliably, the rate difference is real money and Fugu Max deserves a scoped pilot with token accounting switched on from the first request. If you need a production decision you can defend this quarter — a window you can plan against, a score you can point at, and a model you could still run if the vendor disappeared — Kimi K3 is the answer, and it is on OrcaRouter at Moonshot's list price with the provider's rate passed through untouched.

A screenshot of the OrcaRouter model page for Kimi K3 (English UI, captured September 11, 2026), showing the Featured badge, the identifier kimi/kimi-k3, by MoonshotAI dated 2026-07-15, vision, tools, JSON and reasoning capability tags, a 2.8-trillion-parameter Mixture-of-Experts description, a 1M-token context window with text and image input, list pricing of $3.00 per 1M input tokens and $15.00 per 1M output tokens, traffic of 3,274.3 million tokens over 7 days, and an OpenAI-compatible base URL with Python and cURL code samples.