A generated title card comparing Xiaomi MiMo-V2.6-Pro and Kimi K3 at maximum effort. The left card reads Intelligence Index v4.3.2: 46, 1.02T total parameters, $206.66 to run the index, and 134.3 tokens per second; the right card reads Intelligence Index v4.3.2: 44, 2.8T total parameters, $3,658.07 to run the index, and 42.8 tokens per second. A footer reads: Both scores on Artificial Analysis Intelligence Index v4.3.2, read 22 September 2026. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

MiMo-V2.6-Pro vs Kimi K3: Two Trillion-Parameter Open Weights, Fourteen Times Apart on Cost per Task

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Both are open-weights mixture-of-experts models from Chinese labs with a 1M-token context window. Both are available to download. On the Artificial Analysis Intelligence Index v4.3.2, read on 22 September 2026, MoonshotAI's Kimi K3 scores 44 and Xiaomi's MiMo-V2.6-Pro scores 46. Two points. The measured cost of running that same evaluation suite against each is $3,658.07 for Kimi K3 and $206.66 for MiMo-V2.6-Pro — a factor of seventeen. If you have already decided you want open weights in this class, that ratio, not the two points, is the decision.

Kimi K3 is the established model in this pair: released 16 July 2026, with full weights following on 27 July, and months of independent evaluation behind it. Xiaomi MiMo-V2.6-Pro went generally available on 22 September 2026, with its reinforcement-learning checkpoint repositories appearing the day before. The comparison is therefore between a proven open model at premium pricing and a brand-new one at commodity pricing, and the interesting part is how narrow the measured capability gap turned out to be.

What each model actually is

Kimi K3 is a 2.8-trillion-parameter open-weight multimodal reasoning model, described at release as the first open model in the three-trillion class. It activates 16 of 896 experts per token — roughly 104 billion active parameters — and uses Kimi Delta Attention and Attention Residuals to hold a 1M-token context efficiently, with output of up to 1M tokens per request. It is always-on reasoning with native vision. Moonshot's first-party API charges $3.00 per million input tokens, $0.30 per million cached input and $15.00 per million output, flat with no context-length tiering, and Artificial Analysis reports a 90% cache discount and a blended rate of $2.31. It measured 42.8 output tokens per second — notably slow against a median of 77.9 for models of comparable size — and 3.93 seconds to first token.

Xiaomi MiMo-V2.6-Pro is a sparse mixture-of-experts model at 1.02 trillion total parameters and 42 billion activated per token, with a 1M-token context window and native multimodality across text, image, video and audio. Its weights carry an MIT licence tag. Xiaomi kept the previous generation's pricing rather than repricing the new checkpoint upward, which is why it lists at $0.43 per million input tokens and $0.87 per million output with a 99% cache discount and a blended $0.18. It measured 134.3 output tokens per second — eleventh of 114 models on speed — and 2.15 seconds to first token.

• Active parameters — MiMo-V2.6-Pro 42B per token; Kimi K3 about 104B, activating 16 of 896 experts

• Total parameters — MiMo-V2.6-Pro 1.02T; Kimi K3 2.8T

• Index score — MiMo-V2.6-Pro 46; Kimi K3 44, both Artificial Analysis v4.3.2

• List price — MiMo-V2.6-Pro $0.43 / $0.87 per 1M; Kimi K3 $3.00 / $15.00, cached input $0.30

• Blended rate — MiMo-V2.6-Pro $0.18 per 1M; Kimi K3 $2.31

• Cost to run the index — MiMo-V2.6-Pro $206.66; Kimi K3 $3,658.07

• Speed — MiMo-V2.6-Pro 134.3 output tokens/second; Kimi K3 42.8, against a 77.9 median for the size class

• Time to first token — MiMo-V2.6-Pro 2.15 seconds; Kimi K3 3.93 seconds

• Context — 1M tokens both; Kimi K3 publishes output up to 1M tokens per request

• Weights — both downloadable; MiMo-V2.6-Pro carries an MIT tag

A two-column scoreboard titled MiMo-V2.6-Pro vs Kimi K3, subtitled Two open-weight MoEs tied on the only index both were run against. The left column, Xiaomi MiMo-V2.6-Pro, reads Intelligence Index v4.3.2 46; total parameters 1.02T; list price $0.43 / $0.87 per 1M; cost to run the index $206.66; speed 134.3 tokens per second; time to first token 2.15 s. The right column, Kimi K3 (max), reads Intelligence Index v4.3.2 44; total parameters 2.8T; list price $3.00 / $15.00 per 1M; cost to run the index $3,658.07; speed 42.8 tokens per second; time to first token 3.93 s. A footer notes both models are open-weights, that MiMo-V2.6-Pro carries an MIT licence tag, and that blended rates per 1M tokens are $0.18 and $2.31. The OrcaRouter logo is composited in the bottom-right corner.

Speed is the difference the index hides

The two-point composite gap is the least interesting number here, and the 134.3 against 42.8 tokens per second is the most. A model that generates three times faster is not three times better, but it is the difference between a class of application that works and one that does not — and Kimi K3's own Artificial Analysis page flags it as notably slow for its size class. For a long-horizon agent loop that makes dozens of sequential calls, throughput compounds the way latency does: a three-fold per-call difference is not a three-fold end-to-end difference, but it is the difference between a session a user waits through and one they abandon.

Kimi K3's advantage is on the input side of the ledger. Its 3.93 seconds to first token is slower than MiMo-V2.6-Pro's 2.15 but not by the order of magnitude that separates either of them from a max-effort frontier reasoning model. Where K3 pays for its size is in generation, and generation is what you are billed for.

The vendor numbers, and the one that gets misread

Both labs published DeepSWE v1.1 figures from their own harnesses, and the two are in circulation as if they were comparable. They are not, and this pair is where that error does the most damage because the numbers look so close together.

Xiaomi reports MiMo-V2.6-Pro moving from 58.4 to 72.57 on DeepSWE v1.1 across its reinforcement-learning run, evaluated with mini-swe-agent at average-of-three on Xiaomi's own grader. The run is documented as 30 steps and roughly 750,000 trajectories per model, completed in under six days at a published spend of about $2.62 million for the Pro model and $854,044 for the smaller Flash. A widely circulated reading of that run places Kimi K3 at 68.51 on the same benchmark, which would make the Xiaomi model the winner by four points. That reading is a community figure, not a Moonshot number, and the two measurements use different metrics — average-of-three against a pass rate — so subtracting them is arithmetic on mismatched scales. Neither vendor has submitted a configuration to Datacurve's public board for these models.

The number that is genuinely comparable is the index, because Artificial Analysis ran both itself: 46 for MiMo-V2.6-Pro and 44 for Kimi K3, on v4.3.2, with the per-evaluation breakdown published for neither. That two-point margin is small enough to be within the noise of a ten-evaluation composite, and the honest conclusion is that on measured capability these two models are effectively tied.

What Kimi K3's premium buys

It would be wrong to read the cost gap as pure margin. Kimi K3 is a 2.8-trillion-parameter model against Xiaomi's 1.02 trillion, and it activates roughly two and a half times as many parameters per token. It was the first open model in its size class, and it holds a set of independent results that MiMo-V2.6-Pro simply has not had time to accumulate: on Artificial Analysis's own reporting it placed third on the Intelligence Index behind Claude Fable 5 and GPT-5.6 Sol at the time of its release, with GDPval-AA v2 at Elo 1668, first place on AutomationBench-AA at 53%, second on AA-Briefcase at Elo 1547, and first on the Frontend Code Arena at 1679. Those are independent measurements, not launch claims, and they describe a model with breadth that a two-point composite margin does not capture.

There is also a capacity story attached. Demand for Kimi K3 reportedly overwhelmed serving capacity at launch, causing temporary API subscription pauses and upstream capacity limits on some gateways. A model that is hard to get is a model that is hard to build on, and that is an availability problem rather than a quality one.

Screenshot of the Artificial Analysis model page for Xiaomi MiMo-V2.6-Pro, captured 22 September 2026. The header reads 46 Artificial Analysis Intelligence Index, ranked first of 114 in its class; 134.3 output tokens per second, ranked eleventh of 114; $0.435 input and $0.87 output per 1M tokens with a 99% cache discount; $0.13 cost per Intelligence Index task, ranked twelfth of 114; and 140M output tokens from the Intelligence Index. Technical specifications list reasoning Yes, input modality text, image, speech and video, output modality text, a 1M-token context window, 1.0T total parameters, 42B active parameters, an MIT licence and Hugging Face model weights, under a breadcrumb reading Xiaomi, Open weights model, Released September 2026. A comparison summary notes it cost $206.66 to evaluate MiMo-V2.6-Pro on the Intelligence Index.

One endpoint, and a straight answer about what we route

Kimi K3 is on OrcaRouter as kimi/kimi-k3 — one OpenAI-compatible key, provider list price passed through with 0% markup, so a Moonshot price change is live on our side the same day, and automatic failover across providers covers the capacity limits that have dogged this model since launch. Xiaomi MiMo-V2.6-Pro is not on OrcaRouter. No Xiaomi model is, and the route to it today is Xiaomi's own platform or one of the third-party catalogues carrying it. That is worth stating plainly rather than letting a model page imply otherwise.

Which makes this pairing a good example of what the routing layer is for. Two open models, tied on the only independent measure both were run against, seventeen times apart on measured cost per task — that is not a choice, it is a two-lane configuration. Put the bulk of your traffic on the cheap lane and route the tail to the other on a predicate you define, or fail over when one route degrades. With the models behind one key, the experiment that tells you which share each should take is a config change on your own prompts rather than a procurement conversation — and if Xiaomi reprices MiMo-V2.6-Pro once the launch window closes, or Moonshot cuts K3's rates in response to the competition, both moves land on our side the same day rather than at the next billing cycle.

Screenshot of the OrcaRouter model page for Kimi K3, captured 22 September 2026, showing the breadcrumb Home > Models > MoonshotAI > Kimi K3, the model id kimi/kimi-k3 dated 2026-07-15, the Vision, Tools, JSON and Reasoning capability badges, and the opening of the model description, which calls Kimi K3 MoonshotAI's flagship and most capable release to date and a 2.8-trillion-parameter mixture-of-experts model. The rate card sits below the visible area of the capture.

Which one to reach for

Pick Kimi K3 when you need the breadth that comes with a model this size and this independently measured — vision, long-horizon coding across large repositories, agentic tool use — or when you are already running it and the migration cost is real. It is the more thoroughly characterised model of the two, and its 1M-token output ceiling is unusual. Budget for the generation speed: at 42.8 tokens per second it is a batch and offline-workload model in interactive terms, whatever its reasoning quality suggests.

Pick MiMo-V2.6-Pro when throughput and unit economics are the constraint, when the loop is context-heavy and repetitive, when first-token latency matters, or when you want the weights on your own hardware under a permissive licence. It is the rare case of a cheap model that is also the fast one, and on the only independent score both models share, it is ahead by a margin small enough to be a tie.

Pick both if you have a router. The failure mode to avoid is standardising on either one because of a headline ratio: paying Kimi K3 rates for a summarisation job a 46-index model finishes identically, or discovering that a repository-scale coding task needs the 2.8T model after you have moved everything to the cheap lane. Neither is a model problem, and both are configuration.

OrcaRouter puts 200+ models behind one OpenAI-compatible endpoint at provider list price with 0% markup, so a vendor price change is live the same day.