
Xiaomi MiMo-V2.6-Pro vs MiniMax M3: The Better Score Has the Thinner Route
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Xiaomi MiMo-V2.6-Pro went generally available on 22 September 2026 and took the top open-weights position on the Artificial Analysis Intelligence Index at 46 on v4.3.2. MiniMax M3 has held a version of that title since it shipped on 31 May 2026, and on the same index today it scores 29 — seventeen points behind. The score is the least interesting thing about the pair. MiniMax M3 serves at 168.9 output tokens per second against 134.3, answers in 0.92 seconds against 2.15, is cheaper per input token, and is reachable through fifteen API providers to Xiaomi's one. The newer, higher-scoring model is the harder one to actually call, and that inversion is what this comparison is about.
Both open weights, both multimodal, fourteen weeks apart
The two models are close relatives in design intent. Both are sparse mixture-of-experts models from Chinese labs, both carry a 1M-token context window, both accept images and video natively rather than through a bolted-on adapter, and both publish weights. Fourteen weeks separate their releases, and the index gap between them is the distance the open-weights frontier travelled in that window.
• Total parameters — Xiaomi MiMo-V2.6-Pro 1.02T; MiniMax M3 roughly 427B, both sparse MoE
• Active per token — Xiaomi MiMo-V2.6-Pro 42B; MiniMax M3 roughly 23–25B
• Context — 1M tokens each, with MiniMax M3 guaranteeing a 512K minimum
• Modalities in — text, image and video for both; Xiaomi adds audio input
• Weights — MIT licence on Xiaomi MiMo-V2.6-Pro; MiniMax M3 ships under the MiniMax Community Licence, which is not permissive and requires a separate agreement for commercial use
• Released — Xiaomi MiMo-V2.6-Pro 21–22 September 2026; MiniMax M3 31 May 2026
• Reachability — one API provider counted for Xiaomi MiMo-V2.6-Pro on Artificial Analysis; fifteen for MiniMax M3
That last line is the one to keep in view. Everything else on the sheet favours the newer model. That line does not, and it is the line that decides whether a model can sit in a production path.
What the index measures, and who measured it
Artificial Analysis ran both models itself on v4.3.2 — a ten-evaluation composite covering reasoning, knowledge, mathematics and coding, including AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. Neither the 46 nor the 29 is a vendor number, which matters here because both labs also published benchmark figures of their own that are vendor numbers, and the two kinds must not be mixed.
The 46 puts Xiaomi MiMo-V2.6-Pro at the head of the open-weights field on that composite. It does not put it at the head of the index: Claude Fable 5.1 and GPT-6 Astra both sit at 53, and Claude Opus 5 at 51. MiniMax M3's 29 places it sixteenth of 114 models measured, well above the median for its size class but a long way from the frontier it was being compared to in June.
MiniMax M3 has the receipts Xiaomi does not
This is the part of the comparison that gets skipped, and it is the part a buyer should weigh most heavily. MiniMax M3 is on the public DeepSWE v1.1 leaderboard with an official entry: 20.4% Pass@1 and 48.7% Pass@4, published 8 June 2026 after a MiniMax inference upgrade fixed abnormal token generation and improved long-context cache behaviour. That entry came with an efficiency accounting attached, and the accounting is not flattering: a median of 311 agent steps, roughly 91,000 output tokens, and about $5.04 per task. It passes long-horizon engineering tasks, and it does so expensively.
Xiaomi's headline benchmark figure for its new model is 72.57 on DeepSWE v1.1, up from a 58.4 baseline, measured with mini-swe-agent at average-of-three on Xiaomi's own harness. It is not a leaderboard entry and has not been submitted as one. Putting 72.57 next to MiniMax's 20.4 and reading off a 52-point win is arithmetic across two different measurements: the board publishes Pass@1 for a specific configuration, with a confidence interval and an average cost per task that a third party can re-run, and the vendor figure has none of that attached to it. The 72.57 may well be honest. It is simply not the same kind of object.
MiniMax also published launch figures of its own, and the same caveat applies to every one: SWE-Bench Pro at 59.0, BrowseComp at 83.5, SVG-Bench at 63.7 and Terminal-Bench 2.1 at 66.0, run on MiniMax's own infrastructure with its own scaffolding. Treat them as claims about the model, not measurements of it.
Speed, latency, and the two numbers that decide a bill
The throughput gap runs opposite to the quality gap, and it is wide enough to matter. MiniMax M3 generates 168.9 output tokens per second against 134.3 for Xiaomi MiMo-V2.6-Pro, and its 0.92 seconds to first token is roughly two and a half times faster than Xiaomi's 2.15. For an interactive product, first-token latency is the number a user feels; for a batch pipeline, throughput is the number that sets the wall clock.
On cost the two are closer than the headline rates suggest, and the direction depends on your traffic. MiniMax M3 lists at $0.30 per million input tokens and $1.20 per million output; Xiaomi MiMo-V2.6-Pro lists at $0.43 and $0.87. M3 is cheaper to feed, Xiaomi is cheaper to generate from. Once cache discounts are applied — 80% for M3, 99% for Xiaomi — the blended rates land at $0.22 and $0.18 per million on a 7:2:1 cache-hit/input/output mix. Artificial Analysis also records a cost of $0.13 per Intelligence Index task for Xiaomi MiMo-V2.6-Pro, measured the same way across every model on the board, which is the most comparable single cost figure available for the two.

Fifteen routes against one, and what failover is worth
A seventeen-point index advantage does not help you if the endpoint is down. Artificial Analysis counted a single API provider for Xiaomi MiMo-V2.6-Pro at the time of writing, and that count is a fact about the release rather than about the model: a checkpoint published three days ago has had no time to propagate. MiniMax M3, four months old, is served by fifteen.
That difference is what a router is for, and it is where the two models diverge commercially. MiniMax M3 is on OrcaRouter as minimax/minimax-m3: one OpenAI-compatible endpoint, provider list price passed through with 0% markup, so the $0.30 and $1.20 above are MiniMax's own numbers rather than ours, and a MiniMax price change is live on our side the same day. Its fifteen upstreams are what makes an automatic failover chain meaningful — a request that stalls on one provider retries onto another before the response starts, which is a different guarantee from a model that has one place to go. Xiaomi MiMo-V2.6-Pro is not on OrcaRouter; no Xiaomi model is, and reaching it today means Xiaomi's own platform and the third-party catalogues that list it.
If you want both in the same application, the practical shape is one key holding the models you have already decided on, with the newcomer evaluated against them rather than adopted on the strength of a three-day-old score. That is a routing decision, and the useful property is that a fallback lane costs nothing until the day it saves you.

Which one to pick
If you need an open-weight, multimodal, million-token model in production this week, the answer is MiniMax M3, and it is not close — fifteen providers, sub-second first token, a permissive-enough commercial path only if you have read the community licence, and an official long-horizon benchmark entry you can inspect. The efficiency warning is real: budget for the 91,000 output tokens per task, not just the per-token rate.
If you are choosing a model for work you will start next month, Xiaomi MiMo-V2.6-Pro is the more capable model by a wide measured margin, at a lower blended rate and with an MIT licence that removes the licensing conversation entirely. What it does not yet have is a second place to run. The thing to watch is the provider count: when it moves off one, the seventeen-point gap stops being a research result and becomes a deployment decision.
Until then, the honest summary is that the better model is the one you cannot fail over from, and the weaker model is the one you can put in front of customers tonight.

Where the numbers in this piece come from
Index scores, throughput, latency, cache discounts and provider counts for both models are Artificial Analysis measurements on Intelligence Index v4.3.2, read on 22 September 2026, and the 29 for MiniMax M3 sits on the same version of the composite as the 46 for Xiaomi MiMo-V2.6-Pro. The DeepSWE v1.1 figures for MiniMax M3 are from the public leaderboard entry of 8 June 2026. The 58.4-to-72.57 DeepSWE movement, the reinforcement-learning spend of roughly $2.62 million for the Pro run over thirty steps, and the SWE-Bench Pro, BrowseComp, SVG-Bench and Terminal-Bench numbers for MiniMax M3 are all vendor-reported on the respective labs' own harnesses and have not been independently reproduced. The per-token rates are the vendors' own list prices.
