A generated title card reading "Kimi K4" with the subtitle "what the leak says, and what it does not", a LEAK / UNVERIFIED badge, three stacked lines reading "No model card", "No weights" and "No price or route", and a footer line "As of 25 September 2026", with the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

Kimi K4: what the leak actually claims, and why "fewer active params" would be the real story

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

A single post has put four unannounced model names into circulation at once. The list starts with K​imi K4, the supposed next generation from Moonshot AI, sitting alongside GLM-5.5 Flash and GLM-5.4 from Z.ai and D​eepSeek V4.1P — and the author's own emphasis is not on any of the flagship names but on a suspicion about the arithmetic underneath K4: that it may activate fewer parameters than the model it replaces. That claim is unverified. There is no K​imi K4 model card, no weights, no API identifier, no price and no route, and the same is true of every Z.ai name in the list. What is worth doing now is working out which of these claims would be checkable, what the shipped baseline looks like, and what a sparser K4 would actually mean.

The signal, and its tier

This is an early signal, and it should be read as one. The post that carries it is a read of model naming conventions rather than a report of a sighting: it flags "GLM-5.5 Flash, GLM-5.4" as interesting nomenclature and calls K​imi K4 "very strange", then offers a speculative mechanism — fewer activated experts — without claiming to have seen a config, a checkpoint listing, or a staging endpoint.

Nothing in the public record contradicts the names, and nothing corroborates them either. That asymmetry is normal for this stage of a release cycle and it is exactly why the useful move is to fix the baseline and the tests, not to speculate further. A name in a post is a hypothesis about a product, not evidence of one.

The baseline a Kimi K4 would be measured against

Kimi K3 is real, shipped, and unusually well documented, which makes it a firm reference point. Moonshot's own model card describes a 2.8-trillion-parameter Mixture-of-Experts model that activates 104B parameters per token, spread over 93 layers — one dense layer and 92 sparse ones. The sparsity is aggressive and is stated plainly: 16 of 896 experts fire per token, on top of 2 shared experts, which Moonshot credits with roughly a 2.5× improvement in scaling efficiency over Kimi K2.

The attention stack is where K3 breaks with the previous generation. Of its 92 sparse layers, 69 use K​imi Delta Attention — a hybrid linear-attention mechanism — and 24 use gated MLA, a 3:1 split that keeps a minority of full-attention layers and pushes the rest onto the cheaper linear path. Layers carry an attention residual every 12 blocks, the activation function is SiTU-GLU, and the whole model is trained with MXFP4 weights and MXFP8 activations under quantization-aware training. Context length is 1,048,576 tokens, vocabulary is 163,840, and vision runs through a 401M-parameter MoonViT-V2 encoder, making K3 natively image-capable rather than multimodal by bolt-on.

Screenshot of the Artificial Analysis model page for Kimi K3 (max), headed "Released July 2026", whose comparison summary reads In $3.00 / Out $15.00 and describes the model as leading on intelligence but expensive next to other open-weight models of similar size, notably slow and somewhat verbose, with text and image input and a 1M-token context window.

Those are the numbers a successor has to move. Note what they imply about the "fewer active params" idea: K3 is already an extraordinarily sparse model. It activates about 3.7% of its total parameter count per token. A K4 that activates fewer than 104B would not be trimming fat from an over-provisioned design — it would be walking back a sparsity ratio Moonshot has publicly framed as a win, and it would do so in the generation where the rest of the field is going the other way.

What "fewer active params" would mean if it is true

Two readings are available and they point in opposite directions. The charitable one is architectural: Moonshot could be extending the KDA hybrid further, replacing more full-attention layers with linear ones and re-balancing the expert budget so that a lower activation count buys a longer context, a cheaper serving footprint, or a better quality-per-flop ratio. Under that reading, fewer active parameters is a refinement of the same thesis K3 already states — that sparsity is a scaling strategy, not a compromise.

The other reading is that K4 is not a uniform successor at all. K​imi's line has already branched once this year, and reports have variously gestured at K3.1, K3.2 and K4 as three different products. A K4 that activates less than K3, in a family that ships a separate coding variant, is at least as consistent with a repositioning as it is with a straight upgrade. Both readings are speculation. Neither is settled by a post.

There is also a licence dimension nobody has addressed. K3's weights are released under the Kimi K3 License, a custom "other" licence rather than a standard open-source one, and it is the first open 3T-class model. Whether a sparser K4 inherits that licence, or narrows it, matters more to anyone building on the weights than a few points of index score.

A generated two-column scoreboard titled "Kimi K3 vs Kimi K4 — the scoreboard", with six shared rows: Status (K3 "released 15 July 2026" vs K4 "unreleased"), Total parameters ("2.8T" vs "unverified"), Activated parameters ("104B" vs "unverified"), Experts per token ("16 of 896" vs "unverified"), Context ("1M tokens" vs "unverified") and Price ("$3 / $15" vs "none"), with the footer line "Kimi K3 figures per Moonshot's model card; Kimi K4 has no model card, weights or price."

The other three names in the same post

GLM-5.5 Flash and GLM-5.4 are the more surprising pair, and not because of the numbers. Z.ai's published ceiling is GLM-5.3, which shipped on 14 August 2026 at 743B parameters, with GLM-5.3-Flash following on 26 August and the BF16 and Flash-BF16 weight releases landing on 25 August. Z.ai's own documentation lists exactly three current model identifiers: glm-5.3, glm-5.3-flash and glm-5.2. There is no GLM-5.4, no GLM-5.5 and no GLM-5.5 Flash — and the naming direction is the odd part, since a 5.5 generation preceding a 5.4 one is not how Z.ai has ever numbered a family. Version numbers appearing out of order in a leak are a familiar failure mode: they tend to be adjacent products, internal codenames, or a serving-tier label rather than a new base model.

D​eepSeek V4.1P is the name the signal's author says they are most interested in, and it is the easiest of the four to check. D​eepSeek's API documentation names exactly two current model strings: deepseek-flash and deepseek-v4-pro. The "P" suffix has no counterpart in any D​eepSeek identifier, and the most recent weight release in D​eepSeek's own organisation is DeepSeek-V4.1-Flash, published 10 September 2026. A V4.1P would most plausibly be a Pro-tier sibling of that model — but "most plausibly" is a guess about a string, not a product.

How you would know any of this is real

The four names fail the same test in the same way, which makes the wait straightforward rather than agonising. A real release leaves artefacts, and each of them is cheap to check:

• A model card in the vendor's own Hugging Face organisation. Moonshot's newest entry is Kimi-K3, created 13 June 2026; Z.ai's newest are the GLM-5.3 family from 25 August; D​eepSeek's newest is DeepSeek-V4.1-Flash from 10 September. Nothing newer exists in any of the three.

• A config that settles the argument. For K4 specifically, the fields to look for are num_experts and num_experts_per_token — K3's read 896 and 16. A K4 config with a lower per-token figure is the claim in this leak being confirmed or refuted in a single file.

• A routable model. A name that appears in a catalogue is a name with a price, a context window and a provider behind it; a name that is only in a post has none of those things.

Screenshot of the OrcaRouter model page for kimi/kimi-k3 by MoonshotAI, dated 2026-07-15, with Vision, Tools, JSON and Reasoning capability chips and pricing of $3.00 per million input tokens and $15.00 per million output tokens.

What you can route today

The practical version of this leak is that the generation it describes is not what you can build on. What is live is the previous one, and the router's job is to make the gap between generations a routing decision rather than a migration project. Kimi K3 is available now with its full 1M-token context and native vision, priced at $3.00 per million input tokens and $15.00 per million output tokens, with cache reads at $0.30 — billed at the provider rate with zero markup, which is why a vendor price change on K3 reaches your bill the same day rather than at the next repricing cycle. Kimi K3 on OrcaRouter carries the full spec sheet and current rate.

The same applies across the rest of the board. GLM-5.3, GLM-5.3-Flash and DeepSeek V4.1 Flash are all routable alongside Kimi K3, over one OpenAI-compatible endpoint covering 200+ models, with automatic failover when a provider degrades and a routing DSL for expressing preferences — cost ceilings, quality floors, latency budgets — as policy rather than per-call logic. When a K​imi K4, GLM-5.5 Flash or D​eepSeek V4.1P does appear, the migration is a string change in the routing policy and nothing else. Model fusion lets you combine outputs from several models on the same endpoint when a task benefits from it.

To be explicit about what that is not: none of the four leaked names is a route on OrcaRouter today. We do not host them and cannot serve them.

Bottom line

The interesting claim here is not the version numbers. It is the mechanism — a next-generation flagship that activates fewer parameters than its predecessor would be a statement that Moonshot believes sparsity, not raw activation count, is the axis worth pushing, and K3's 16-of-896 expert split is already an argument in that direction. Every part of the claim is unverified: K​imi K4, GLM-5.5 Flash, GLM-5.4 and D​eepSeek V4.1P have no model cards, no weights, no API identifiers and no prices, and the vendors' own catalogues stop one generation earlier in each case. Treat the names as a preview of a question, not an answer — and if you need frontier-class open weights at 1M context today, Kimi K3 is the model that already exists.

When a K​imi K4 does arrive, automatic failover is what keeps the migration from turning into an incident