A generated hero card titled 'Qwen 4 Max vs Kimi K3' with the subtitle 'The open-weights promise meets the open-weights delivery' and three pill badges reading 'Weights promised, not shipped', 'K3 weights out July 27, 2026' and 'Modified MIT licence'. A footer line reads 'Alibaba Yunqi Conference, Hangzhou'. Flat vector editorial styling on a white background with a blue-to-cyan gradient wash and the OrcaRouter logo composited bottom right.
Engineering & Research

Qwen 4 Max vs Kimi K3: the open-weights promise meets the open-weights delivery

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Moonshot AI released Kimi K3 on July 16, 2026 and then did the thing most labs only talk about: it published the weights. The 2.8-trillion-parameter checkpoint landed on July 27, 2026 under a Modified MIT licence, eleven days after launch. Meanwhile, the Yunqi conference on September 22, 2026 was used to announce Qwen 4 Max as the flagship of a four-tier Qwen 4 family that includes an open-weight Qwen 4 27B — a tier that is promised, not shipped. Those two facts are the comparison. Everything else about Qwen 4 Max is currently unknown, and the gap between a promised open tier and a delivered one is where this matchup actually has something to say.

What Kimi K3 put on the table in July

Kimi K3 is a 2.8-trillion-parameter mixture-of-experts that activates 16 of 896 experts per token, which puts roughly 104 billion parameters to work on any given step. It carries a 1,048,576-token context window on both the input and the output side and accepts text, image and video input with text output. Its first-party API lists at $3.00 per million input tokens, $15.00 per million output tokens and $0.30 per million cached input tokens, and third-party rates run below that.

The architecture is the part Alibaba's own preview echoes. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which Moonshot claims deliver roughly 2.5 times better scaling efficiency than Kimi K2 and up to 6.3 times faster decoding at million-token contexts. That is the same problem Qwen's QSA is aimed at — making attention affordable at extreme length — which means the two families are not competing on scale so much as on whose sparse-attention design ages better.

The scores Moonshot published at launch, vendor-reported and unreproduced at the time:

• Artificial Analysis Intelligence Index — 57, third at the time behind Claude Fable 5 and GPT-5.6 Sol. That figure was recorded under an earlier index revision; the index has since been rebuilt twice, so it is a historical reading rather than a current one

• Frontend Code Arena — first place at 1,679 Elo, ahead of both models that outscored it on the general index

• Terminal-Bench 2.1 — 88.3

• SWE-Marathon — 42, ahead of Opus 4.8 and GPT-5.6 Sol

• DeepSearchQA — 0.95, first place, and MATH-Vision 0.98, also first

• GPQA-Diamond — 0.94

Moonshot's own launch material concedes that K3 still trails Claude Fable 5 and GPT-5.6 Sol on overall capability, which is a more useful sentence than any of the first-place claims above it. The interesting shape of the numbers is that a model third on the general index finished first on frontend code and first on long-horizon search — a profile that says the aggregate index is hiding a specialised tool rather than describing a general one.

A generated scoreboard titled 'Qwen 4 Max vs Kimi K3 - the scoreboard'. Left column 'Qwen 4 Max' reads announced not released, Qwen 4 27B promised, not published, not published, not published, not published. Right column 'Kimi K3' reads GA since July 16 2026, published July 27 2026, Modified MIT, $3.00 per 1M tokens, $15.00 per 1M tokens, 1,048,576 tokens. Footer reads 'Kimi K3 figures from Moonshot's launch material; Qwen 4 Max status per Alibaba's September 22 2026 Yunqi conference.'

What Alibaba has promised, and what a promise is worth

The Qwen 4 announcement named four tiers. Qwen 4 Max as flagship, Qwen 4 Flash for throughput, Qwen 4 Plus as the balanced middle, and Qwen 4 27B as the open-weight local tier. Qwen LLM lead Liu Da Yiheng described the family as training on a new-generation architecture and said it would arrive soon. There is no date, no model card, no API identifier, no cloud listing, no price and no published score for any of the four.

Two specific things about that announcement get misreported, and both matter for anyone trying to plan around it:

• The 5-trillion-to-10-trillion-parameter figure attached to Qwen 4 coverage is not a Qwen 4 specification. It belongs to the Qwen 4.5 and Qwen 5 roadmap. No parameter count of any kind has been published for Qwen 4 Max

• The tier names continuing does not carry the specifications forward. Qwen 4 Max's context window, price and licence are unknown; the shipping Qwen3.8-Max numbers — 1,000,000 tokens at $2.00 and $6.00 per million — describe a different model

The reason the 27B tier is the interesting one has nothing to do with benchmarks. It is the only tier in the Qwen 4 line that has historically shipped under open weights, and the Qwen 3.8 generation showed what that unlocks: within days of the 27B weights landing, the community had produced MLX conversions, NVFP4 quantisations and fine-tunes of every description. An announced 27B is a commitment to a downstream ecosystem, and it is the most predictable delivery on the roadmap.

But predictable is not delivered. Kimi K3's weights are a file you can download today. Qwen 4 27B's weights are a line in a conference talk.

Open weights and runnable weights are different claims

There is a trap in treating Kimi K3 as the open-weights answer, and it is worth stating plainly because it is the most common overclaim in coverage of Chinese frontier models. A 2.8-trillion-parameter mixture-of-experts is a datacenter-scale artefact. Published weights mean you are permitted to run it, can inspect it, can fine-tune it, and can quantise it. They do not mean you can run it on a workstation. Efficient serving of a checkpoint that size needs tens of accelerators, and the quantisation work that follows an open release — the NVFP4 and REAP-style variants that circulate within weeks — is as much about making the model servable as about making it smaller.

So the honest reading of Kimi K3's licence is narrower and still significant: it is an auditable, modifiable, self-hostable frontier model, under a Modified MIT licence, for organisations with the hardware to serve it. That is a real strategic asset and a poor fit for anyone who read "open weights" as "runs on my laptop".

The same caveat will apply to Qwen 4 27B if and when it ships — and it applies less sharply, because a 27B open-weight tier is genuinely local-scale in a way that a 2.8T flagship is not. That is precisely why the Qwen 4 27B tier deserves more attention than the Max tier in this comparison, even though the Max tier is what the announcement led with.

A screenshot of the OrcaRouter model page for Kimi K3 (kimi/kimi-k3), captured September 22 2026, showing the model name, the 1M token context, text and image input with text output, vision, tool and reasoning badges, and a p50 time-to-first-token figure.

The commercial question underneath the licence question

Kimi K3's first-party rate of $3.00 input and $15.00 output per million tokens is a premium position, and Moonshot has been explicit that it is deliberately priced above the cheap end of the Chinese market. Against Qwen3.8-Max at $2.00 and $6.00 that is one and a half times the input rate and two and a half times the output rate — a gap large enough that a workload has to care about Kimi K3's specific strengths, frontend code and long-horizon search among them, for the premium to be worth paying.

OrcaRouter routes kimi/kimi-k3 alongside the Qwen 3.8 family on one OpenAI-compatible endpoint at 0% markup, which means the vendor's list rate is what you are billed rather than a marked-up equivalent. In a comparison where the price difference is a factor of 2.5 on output, the absence of a platform margin is not a rounding detail — a ten percent layer on a $15.00 output rate is $1.50 per million tokens, which is more than the entire input cost of the cheaper model in this matchup. The same endpoint carries automatic failover and a routing DSL for expressing policy explicitly, which is how you try a model with Kimi K3's profile on a slice of production traffic without betting a whole path on a vendor you have not run before.

What OrcaRouter does not carry is Qwen 4 Max or any Qwen 4 tier, because none of them exist as products yet.

A screenshot of the OrcaRouter models catalogue, captured September 22 2026, showing the filter sidebar (input modalities, context length, input price, status, series), the model card grid and a code sample posting to the OpenAI-compatible endpoint at api.orcarouter.ai.

What to watch, and what to ignore

Three signals would move this comparison, and they are specific enough to watch for:

• The Qwen 4 27B weights. Not the Max tier's benchmark table — the 27B checkpoint, its licence and its size. That is the release that will tell you whether Alibaba's open-weight commitment in this generation is as real as the last one

• Qwen 4 Max's licence, if any. If Alibaba publishes weights for its flagship the way it did for Qwen3.8-Max, the open-weights argument in this matchup stops being Kimi's advantage and becomes a tie

• A fresh independent reading on Kimi K3. Moonshot's launch scores were self-reported and the Artificial Analysis index has been rebuilt twice since, so the 57 and the Frontend Code Arena Elo both need re-basing before they are compared to anything current

What to ignore is any Qwen 4 Max specification table that appears before Alibaba publishes a model card, and any claim that Kimi K3's open weights make it locally runnable. Both errors are already circulating. In the meantime the comparison you can act on is between two models that exist: Kimi K3 at a premium rate with delivered weights and a genuinely differentiated coding profile, and Qwen3.8-Max at less than half the output rate with a flat million-token window and weights of its own. That is a real choice, priced on both sides, and it does not require waiting for a conference slide to become a product.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily