OrcaRouter model radar hero card for Xiaomi MiMo-V2.6-Pro, titled with the model name and the line 'Open weights, 46 on the index, $0.43 / $0.87', with two panels labelled What shipped and What it costs.
Guides & Insights

Xiaomi MiMo-V2.6-Pro Tops the Open-Weights Index at 46 — and Publishes the Training Bill

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Xiaomi MiMo-V2.6-Pro is now the highest-ranked open-weights model on the Artificial Analysis Intelligence Index, at 46 on v4.3.2 — one point clear of GLM-5.3 and two clear of Kimi K3, and twenty points above the 26 that its predecessor MiMo-V2.5-Pro recorded on the earlier index scale. That number was produced by Artificial Analysis running its own harness, not by Xiaomi, and it is the single most useful fact in this release. The second most useful fact is the price printed beside it: $0.43 per million input tokens, $0.87 per million output, against $5.00 and $25.00 for Claude Opus 5 — a model that sits five index points higher. The third is the one nobody expected Xiaomi to hand over: a public accounting of what the reinforcement-learning run that produced MiMo-V2.6-Pro actually cost.

The score is a measurement, not a claim

Almost everything published about a Chinese model launch arrives as a vendor number, and readers have learned to discount it. This one is different, and the difference is worth being precise about because the launch coverage has already blurred it.

Artificial Analysis evaluated MiMo-V2.6-Pro itself, on the v4.3.2 composite — ten evaluations covering reasoning, knowledge, mathematics and coding, including AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. The 46 is what that suite returned. It also recorded the operating characteristics that decide whether a model is usable rather than merely good: 134.3 output tokens per second, 2.15 seconds to first token, a 99% cache discount, and a measured cost of $0.13 per Intelligence Index task.

Two of those deserve emphasis. The 134.3 tokens per second puts MiMo-V2.6-Pro at eleventh of 114 models on speed — for a trillion-parameter mixture-of-experts model, that is a serving result, not just a training result. And the $0.13 per task is the number that survives contact with a budget: it is what the index actually cost to run against this model, measured the same way across every model on the board, which makes it comparable in a way that headline per-token rates are not.

Scoreboard card headed 'Xiaomi MiMo-V2.6-Pro — the scoreboard', listing the index score of 46 on Intelligence Index v4.3.2 as the highest of any open-weights model, the open-weights field behind it at GLM-5.3 45 and Kimi K3 44, the top of the same index at Claude Fable 5.1 53, GPT-6 Astra 53 and Claude Opus 5 51, serving at 134.3 output tokens per second and 2.15 seconds to first token, a price of $0.43 in and $0.87 out per million tokens with a 99% cache discount and $0.13 per index task, MIT-licensed weights at 1.02T total and 42B active parameters with a 1M-token context window and text, image, video and audio input, and a route count of one API provider.

What the index does and does not cover

The open-weights crown is real but narrower than it reads. At 46, MiMo-V2.6-Pro leads the open-weights category. It does not lead the index. The top of v4.3 is Claude Fable 5.1 at maximum effort with default fallback and GPT-6 Astra at max effort, both at 53, with Claude Opus 5 at 51 and Claude Fable 5 at 50. So the honest framing is a five-point gap to the frontier on a composite that rewards breadth, and a twenty-point improvement over Xiaomi's own previous generation.

• Open-weights leader — Xiaomi MiMo-V2.6-Pro 46, GLM-5.3 (max) 45, Kimi K3 (max) 44

• Top of the same index — Claude Fable 5.1 (max, fallback) 53, GPT-6 Astra (max) 53, Claude Opus 5 (max) 51

• Speed — 134.3 output tokens/second, eleventh of 114 models measured; median for the size class is 77.9

• Time to first token — 2.15 seconds, against a 2.34-second median

• Cost per Intelligence Index task — $0.13, with the whole evaluation costing $206.66 to run

• Listed rate — $0.43 per million input tokens, $0.87 per million output, a 99% cache discount, and a blended $0.18 at a 7:2:1 cache-hit/input/output mix

One number on the Artificial Analysis page is a genuine caveat rather than a boast: it counted a single API provider for MiMo-V2.6-Pro at the time of writing. That is the practical bottleneck for this model right now, and it is not a quality problem — it is an availability one. A model reachable through one route has no failover. If that route degrades, your application degrades with it, and the failover story is exactly what a router exists to provide. The relevant observation for anyone planning around this release is that we do not host MiMo-V2.6-Pro, and we will not pretend otherwise; the route to it today is Xiaomi's own platform and the handful of third-party catalogues that list it.

Screenshot of the Artificial Analysis model page for Xiaomi MiMo-V2.6-Pro, showing the Intelligence Index score of 46 on version 4.3.2, the listed price of $0.43 per million input tokens and $0.87 per million output tokens, a 99% cache discount, output speed of 134.3 tokens per second, time to first token of 2.15 seconds, and the open-weights attribution with the MIT licence.

The two numbers that must not be mixed

Xiaomi also published benchmark figures of its own, and they are a different kind of object from the 46. The company's dashboard reports MiMo-V2.6-Pro moving from 58.4 to 72.57 on DeepSWE v1.1 across its reinforcement-learning run, evaluated with mini-swe-agent at average-of-three. That is Xiaomi's harness, Xiaomi's grader, Xiaomi's offline run. It has not been submitted to the public DeepSWE leaderboard and it has not been reproduced by anyone.

Read the two side by side and the temptation is to treat 72.57 as a score on a par with the 74% that the public DeepSWE board lists at the top tier. They are not on the same scale. The leaderboard figure is a submitted configuration, Pass@1, with a published confidence interval and an average cost per task; the vendor figure is an internal evaluation with none of that attached. Our own September 21 write-up made this point when the numbers were still dashboard readings, and it holds now that the weights are out. A vendor harness can be entirely honest and still not be a leaderboard entry, because the leaderboard entry is a claim about a specific configuration under specific conditions that a third party can re-run.

The same discipline applies to the training spend. Xiaomi reports roughly $3,474,715 across the Pro and Flash runs — $2,620,670 for Pro, $854,044 for Flash — over thirty reinforcement-learning steps each and about 750,000 trajectories apiece, completed in under six days. Those are the company's own compute-accounting figures. They are interesting because labs almost never publish them, and they are not an API price, not a cost you will pay, and not a number any third party has audited.

What Xiaomi actually shipped

The release is a sparse mixture-of-experts model at 1.02 trillion total parameters with 42 billion activated per token, a 1M-token context window, and native multimodality across text, image, video and audio. Its sibling Xiaomi MiMo-V2.6-Flash is the same architecture at a smaller scale, 309 billion total and 15 billion active. The published weights carry an MIT licence tag, and the release bundle includes a technical report, deployment instructions, and — per Xiaomi — the reinforcement-learning training environments and code needed to examine and reproduce the post-training.

That last item is the substantive difference between this launch and a typical API announcement, and it is worth separating from the marketing. Publishing the RL environment is a claim that the training setup is examinable. Whether the published environments are sufficient to reproduce a 72.57 is a separate question that will be settled by people attempting it, not by the release notes. Xiaomi's own training notes describe a run that mixed coding, general-agent, visual and cybersecurity tasks in a single pass, with graders that rank successful trajectories and redistribute reward toward better paths — and they also record failures: a GPU out-of-memory restart driven by expert load imbalance, a network fault between the training cluster and the grader deployment, and a Flash restart after an infrastructure error went undetected for roughly three hours. Publishing the faults alongside the wins is the part that makes the rest of the disclosure credible.

What it costs to call, and where the routing question sits

At $0.43 and $0.87 per million tokens, MiMo-V2.6-Pro is priced like a mid-tier model and benchmarked like a frontier open-weights model. Xiaomi kept the standard pricing of the previous MiMo-V2.5 generation rather than repricing the new checkpoint upward, which is why the rate looks low relative to the parameter count. A 99% cache discount means a cache-heavy agent loop reads its context for effectively nothing, and that is where a long-horizon agent's bill actually accumulates.

There is a version of this trade that routes cleanly and a version that does not. The version that does: the frontier models you would compare MiMo-V2.6-Pro against — Claude Opus 5 at $5.00/$25.00, GPT-6 Astra, Gemini 3.8 Flash, GLM-5.3 — are all reachable through one OrcaRouter key at provider list price with 0% markup, so a vendor price cut on any of them is live on our side the same day, and a routing rule can send a request to a second model when the first is slow or unavailable. That matters more than usual here, because the model with the best open-weights score is also the one with the thinnest route to it. Composing the two — a cheap open-weights lane for bulk work and a frontier lane for the hard tail — is a routing decision, and it is the kind of decision that stops being a bet the moment both lanes are on the same key.

One more product to keep straight, because it is easy to conflate with the model: Xiaomi is also offering MiMo-V2.6-Pro-UltraSpeed, a hosted serving tier built from the same checkpoint. Xiaomi's own material claims up to twenty times the output speed of the standard Pro service at the same quality; third-party catalogue listings describe roughly ten times. Those are two different numbers describing the same tier, nobody has published per-request latency distributions that reconcile them, and the tier carries a materially higher rate. Nothing about UltraSpeed changes what the weights can do — it changes how fast someone else's servers will run them.

What would change this picture

Three things, in order of how much they would matter.

A second and third serving route. The single-provider count on Artificial Analysis is the constraint on everything else. More routes mean failover, mean price competition at the serving layer, and mean the model can be put in a production path rather than an experiment.

Independent reproduction of the DeepSWE figures. The 46 is already independent; the 72.57 is not. If outside runs land near it, the case for the model strengthens considerably. If they land materially below, the Artificial Analysis number remains the one to trust and the vendor figure joins the long list of launch claims that did not survive contact.

Timeline card headed 'MiMo-V2.6-Pro — what was disclosed, and by whom', with six dated rows: 21 September 2026 weights, technical report, deployment instructions and RL environments published; 22 September 2026 Xiaomi's launch page carrying the DeepSWE v1.1 figures and the training-spend accounting; the independent Artificial Analysis score of 46 on v4.3.2; the vendor harness result moving 58.4 to 72.57 on DeepSWE v1.1; the vendor accounting of $2,620,670 for Pro and $854,044 for Flash across the RL runs; and a route count of one API provider.

Per-evaluation breakdowns. The composite is a composite. Which of the ten evaluations MiMo-V2.6-Pro wins and which it loses is the difference between a model you deploy for agentic coding and a model you deploy for retrieval-heavy question answering, and the index alone does not tell you. That detail is what to watch for next, and it is what a routing configuration needs before it is worth writing down.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily