
MiMo-V2.6-Pro Puts Open Weights on the Pareto Line at $0.13 a Task
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Xiaomi MiMo-V2.6-Pro went up on Hugging Face on 21 September 2026, and the number that has travelled furthest since is 46 — its score on the Artificial Analysis Intelligence Index v4.3.2, the highest any open-weights model has recorded on that composite. The number that decides whether you can actually run it is smaller. Artificial Analysis measured $0.1332 of compute per Intelligence Index task while evaluating MiMo-V2.6-Pro on its own harness, and that figure places the model on the Pareto line of the site's Intelligence-versus-Cost-per-Task chart. Above the rounding floor, it is the only open-weights model on that line.
That is a narrower and more useful claim than "top open-weights model", which is how most of the launch coverage framed it. Being first on the intelligence axis tells you the ceiling. Being on the cost curve tells you whether anything else in the market is strictly better at that price. For Xiaomi MiMo-V2.6-Pro, as of this writing, nothing is.
What the Pareto line actually measures
The chart plots every model Artificial Analysis has evaluated on two axes: the Intelligence Index on one, the weighted average dollar cost of producing one index task on the other. A model on the Pareto line cannot be beaten on both axes by any single other model — you cannot find something smarter for less. Points below and left of the line are dominated and drop off it.
We recomputed the frontier from the data embedded in the chart rather than reading the line by eye. Working from the 326 models that carry both an index score and a measured cost per task, the current frontier is fourteen points long:
• Qwen3.5 2B (Non-reasoning) — index 6.2, cost per task rounds to $0.000, open weights
• Devstral 2 — index 8.6, cost per task rounds to $0.000, open weights
• North Mini Code — index 9.9, cost per task rounds to $0.000, open weights
• Claude Sonnet 4.6 (Non-reasoning, High Effort) — index 24.7, cost per task rounds to $0.000, closed
• Grok Build 0.1 0616 — index 27.2, cost per task $0.006, closed
• GPT-5.6 Luna (high) — index 32.1, cost per task $0.044, closed
• GPT-5.6 Luna (xhigh) — index 34.6, cost per task $0.085, closed
• Xiaomi MiMo-V2.6-Pro — index 46.3, cost per task $0.133, MIT-licensed open weights
• GPT-6 Astra (medium) — index 49.6, cost per task $1.541, closed
• GPT-6 Astra (high) — index 50.9, cost per task $1.725, closed
• GPT-6 Astra (xhigh) — index 52.4, cost per task $2.309, closed
• GPT-6 Astra (max) — index 52.7, cost per task $3.258, closed
• Claude Fable 5.1 (xhigh, default fallback) — index 53.2, cost per task $5.978, closed
• Claude Fable 5.1 (max, default fallback) — index 53.4, cost per task $7.630, closed
Three things fall out of that list. The four cheapest points are models whose measured cost rounds to zero at four decimal places, so the interesting part of the curve starts at Grok Build 0.1 0616. Xiaomi MiMo-V2.6-Pro is a genuine knee in it, not a rounding artifact: the next point up costs 11.6 times as much per task for 3.3 extra index points. And of the ten frontier points with a non-zero cost, exactly one is open weights — which is the fact the "top open-weights model" framing buries. The nearest open-weights model on the intelligence axis, GLM-5.3 (max) at 44.8, is nowhere near the line: it costs $2.006 per index task, fifteen times MiMo-V2.6-Pro's figure. Kimi K3 (max) sits at 43.6 and $2.000 per task.

The price list did not move. The bill did.
Xiaomi did not raise prices for this generation. MiMo-V2.6-Pro lists at $0.435 per million input tokens and $0.87 per million output, with a cache-hit input price of $0.0036 — a 99.17% discount — which is the same schedule the previous flagship carried. Blended at a 7:2:1 cache-hit/input/output mix, that works out to $0.1765 per million tokens.
And yet the cost of running the Intelligence Index against MiMo-V2.6-Pro came to $206.66, or $0.1332 per task, where the same index against MiMo-V2.5-Pro cost $0.054 per task. Artificial Analysis recorded 140 million output tokens for the newer model against 110 million for the older one, on the same suite. The rate card is unchanged; the model simply writes more to answer the same question. That is the honest qualification to the Pareto position — Xiaomi MiMo-V2.6-Pro earns its place on the cost axis in spite of getting more verbose, not because it got cheaper.
For anyone sizing a budget, the practical reading is that per-token pricing is now a weak predictor of what a workload costs. Two models on identical rate cards can differ by 2.5x in cost per completed task. The cache discount is the lever that moves furthest: at $0.0036 per million cached input tokens, the difference between a prompt that reuses context and one that does not is larger than the difference between most pairs of models on this list.
What the score does and does not cover
The 46 is a third-party measurement, and that matters because almost everything else published about MiMo-V2.6-Pro is not. The index version matters too: v4.3.2 is a composite of ten evaluations covering reasoning, knowledge, mathematics and coding, and a score is only comparable against other scores on the same version. The vendor figures Xiaomi published alongside the release — DeepSWE v1.1 at 71.9, AutomationBench v1.0.6 at 53.1, Toolathlon-Verified at 76.9, Terminal Bench 2.1 at 89.9, CyberGym at 94.0 — come from Xiaomi's own harness and grader and have not been reproduced by anyone. They are useful as directional evidence of what the model was tuned for; they are not leaderboard entries.
The other half of the release is the open-weights package. Artificial Analysis lists the model at 1.0T total parameters with 42B active, a 1M-token context window, text, image, speech and video input with text output, and an MIT licence with weights on Hugging Face. The repository itself reports a Safetensors model size of 524B parameters — a third figure that matches neither the 1.0T total nor the 42B active count, and one Xiaomi has not reconciled. If you are planning a self-hosted deployment, that discrepancy is worth resolving before you provision anything, because a 524B FP8 checkpoint and a 1T one are different hardware conversations.

A frontier point you cannot currently route
The availability picture is the part of this release that most limits it. When Artificial Analysis evaluated Xiaomi MiMo-V2.6-Pro it counted a single API provider for the model. The Hugging Face repository states that Xiaomi MiMo-V2.6-Pro is not deployed by any inference provider. We do not host it either, and we are not going to imply otherwise — the route to this model today is Xiaomi's own platform and the third-party catalogues that list it.
That is a failover problem rather than a quality problem, and it is the specific problem a router exists to solve. One API for hundreds of models, with automatic failover between providers, is what turns "one provider currently lists this" from an architectural risk into a configuration detail. It is also how you would stage a migration onto a model this new: send a share of traffic to Xiaomi MiMo-V2.6-Pro, keep a proven model behind it, and let the router move the request rather than your on-call rotation. That works for the models we carry. For this one, until it is routable, the honest advice is to keep it off a critical path.
What would move it off the line
A Pareto position is a snapshot, not a property of the model. Three things would change it. A price cut from Xiaomi would move Xiaomi MiMo-V2.6-Pro further left and make it harder to beat — and because we pass provider list price through at 0% markup, a cut like that is live on our side the same day it lands anywhere. A cheaper model scoring above 46.3 would knock it off entirely; GPT-6 Astra (low) already sits at 45.8 for $0.818 a task, more expensive but uncomfortably close on the intelligence axis. And independent benchmarking of MiMo-V2.6-Flash — which currently has no Artificial Analysis page at all — could reveal that the cheaper sibling occupies a better spot on the curve for most workloads.
For now, the defensible statement is the arithmetic one. On the current index version, at the current published prices, there is no model in the market that is both cheaper per task and more intelligent than Xiaomi MiMo-V2.6-Pro. Whether that survives its first independent replication is the question worth watching.

