
MiMo-V2.6-Pro vs MiMo-V2.6-Flash: A 3x Price Gap That Two Benchmarks Do Not Explain
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
On Xiaomi's own evaluation tables, MiMo-V2.6-Flash beats MiMo-V2.6-Pro. The row is CyberGym, and the margin is 95.1 to 94.0. It is one row out of fifteen, and it is the reason to read the rest of this comparison carefully rather than assuming the flagship is always the answer — because MiMo-V2.6-Flash costs roughly one third of what MiMo-V2.6-Pro costs, and on most of the rows the two are within a few points of each other.
Both models were published as Hugging Face repositories on 21 September 2026, under the MIT licence, eighteen seconds apart, as separate training runs rather than rungs on one ladder. Xiaomi MiMo-V2.6-Pro is the trillion-parameter tier. MiMo-V2.6-Flash is the efficiency tier. The question this page answers is which of them belongs on your production path, and the honest answer depends on a number that does not exist in either release.
What each one is
• Parameters — Xiaomi MiMo-V2.6-Pro 1.02T total with 42B active per token, against Xiaomi MiMo-V2.6-Flash at roughly 309B total with 15B active
• Architecture — both sparse mixture-of-experts with native omnimodal input; Flash is documented at 48 layers, 39 sliding-window and 9 full attention, hidden size 4096, 256 routed experts with 8 active and no shared experts, alongside a 681M vision transformer, a 308M audio tokenizer and a 127M audio patch encoder
• Context — 1M tokens on both, with up to 128,000 output tokens on both
• Licence — MIT on both, weights ungated on both
• Price — $0.435 in and $0.87 out per million tokens for Pro, against $0.14 and $0.28 for Flash: a 3.1x gap on both sides of the card
• Cache-hit input — $0.0036 per million on Pro, $0.0028 on Flash
• Training spend — Xiaomi reports roughly $2.62M for the Pro RL run and $854K for Flash, each completed in under six days across 30 reinforcement-learning steps and about 750,000 trajectories
• Independent index score — 46 for Xiaomi MiMo-V2.6-Pro on Artificial Analysis Intelligence Index v4.3.2; none published for MiMo-V2.6-Flash
That last line is the one to keep hold of. Everything above it except the Pro score is Xiaomi-reported.
Where three times the price buys almost nothing
Xiaomi's Table 3 puts the two side by side across the harnesses the series was trained on, and on long-horizon agentic work the margins are thin:
• DeepSWE v1.1 — Pro 71.9, Flash 67.9
• MiMo Code Bench — Pro 63.2, Flash 61.2
• AutomationBench v1.0.6 — Pro 53.1, Flash 52.3
• Toolathlon-Verified — Pro 76.9, Flash 73.6
• Terminal Bench 2.1 — Pro 89.9, Flash 87.6
• OSWorld-Verified — Pro 82.0, Flash 80.8
• JobBench — Pro 62.0, Flash 61.2
• MiMo Visual Coding — Pro 72.3, Flash 71.5
• CyberGym — Pro 94.0, Flash 95.1
• MiMo Cyber Bench — Pro 80.2, Flash 77.2
Read down that column and the flagship's advantage is mostly between one and four points, with one row lost outright. On a workflow where the difference between a 67.9 and a 71.9 is the difference between a task completing and a task failing, three times the price is cheap. On a workflow where it is the difference between two acceptable answers, it is not. Vendor tables with decimal places like these invite false precision — nobody outside Xiaomi has run them, and a two-point gap on an internally graded harness is well inside the range that a different grader or a different retry policy can move.

Where it buys a lot
There is a cluster of rows where the gap stops being a rounding argument. Every one of them is security work:
• ExploitGym — Pro 17.8, Flash 6.0
• ExploitBench — Pro 47.9, Flash 25.3
• SEC Bench Pro — Pro 66.3, Flash 47.5
• Agents' Last Exam — Pro 31.6, Flash 27.6
• Terminal Bench 4.0 — Pro 34.9, Flash 28.8
On ExploitGym the flagship is three times the cheaper model, which is the same factor as the price, and on SEC Bench Pro it is 1.4x. That pattern is what you would expect from a model with 42B active parameters against one with 15B: a task that needs a long chain of reasoning with no partial credit is the kind of task where the extra capacity shows up, and CyberGym sitting the other way is the exception that proves the two runs learned different things rather than the same thing at different sizes.
So the split maps onto a real workload boundary rather than a marketing one. High-volume agent work with a forgiving success criterion — retrieval, classification, routine edits, calls that a retry can rescue — is Flash territory. Work where a wrong answer is expensive and a partial answer is worthless — exploit development, security review, terminal sessions with multi-step state — is where the Pro tier earns the multiple.
The number that does not exist
MiMo-V2.6-Flash has no Artificial Analysis model page. Its sibling does. That asymmetry is the single most important thing on this page, because it means the only independently measured figure anywhere in the MiMo-V2.6 series is the 46 next to Xiaomi MiMo-V2.6-Pro. Every Flash number above — including the CyberGym row it wins — comes from a Xiaomi harness, a Xiaomi grader and a Xiaomi offline run.
There is a reasonable case that this is temporary: the model is a day old at the time of writing and third-party evaluation takes time. There is also a reasonable case that it is structural, since Flash is the volume tier and volume tiers attract less benchmarking attention. Either way, the practical consequence today is that you cannot compare Flash to anything outside its own family with the same confidence you can compare Pro. If you are choosing between Flash and a non-Xiaomi model on price-per-quality grounds, you are comparing a vendor figure against somebody else's measured one.
Both model cards carry the same second caveat. Neither repository is deployed by any inference provider — Hugging Face says so explicitly on each page, and Artificial Analysis counted a single API provider for Pro. We do not host either checkpoint, so neither is routable through us. The route to both is Xiaomi's own platform and the third-party catalogues that carry them.

The price you pay in hardware, not tokens
Per-token pricing is only half of what a checkpoint costs you, and the two differ far more sharply on disk than on the rate card. MiMo-V2.6-Flash publishes 172.9 GB of FP8 weights across 65 shards. Xiaomi MiMo-V2.6-Pro carries roughly three times the parameters, which puts its raw download in the half-terabyte range before any quantization — and its own repository reports a Safetensors model size of 524B parameters, a figure that reconciles with neither the 1.02T total nor the 42B active count.
That gap changes the shape of the decision. Flash at 172.9 GB is a model a small team can self-host on rented capacity and treat as a commodity. Pro at roughly three times that is a cluster decision — the coverage around the launch put the two training runs at roughly 4,000 and 8,000 H200-equivalent GPUs respectively. If your reason for looking at open weights is that you want the option to stop paying per token, the two checkpoints offer that option at very different scales of commitment.
Choosing between them
The rule that survives the caveats above is to pick by workload class rather than by benchmark average. Flash for anything you would be comfortable retrying; Pro for anything where a retry does not rescue you. Then measure, because neither model has an independent score on the benchmarks you actually care about.
Measuring two checkpoints side by side is where a router does something a per-model contract cannot. Putting a cheap tier on the bulk of your traffic and escalating only the hard cases to the expensive one is the same decision as this page, expressed as configuration rather than as a rewrite — and that composition is exactly what the routing layer exists for. It also matters for the price itself: we pass provider list price through at 0% markup, so if Xiaomi cuts the rate on either checkpoint, the number on our side moves the same day rather than at the next contract renewal. That applies to the models we carry. For the MiMo-V2.6 pair specifically, today, you are still buying directly from Xiaomi.
What would settle it
Two things would turn this from a judgment call into an arithmetic one. An Artificial Analysis page for MiMo-V2.6-Flash would put both checkpoints on the same cost-per-task axis that currently places Pro at $0.133 per Intelligence Index task, and the comparison would resolve itself. A third-party replication of the security cluster — ExploitGym, ExploitBench, SEC Bench Pro — would confirm whether the 3x gap on those rows is a property of the model or a property of the grader.
Until then, the defensible position is that Xiaomi MiMo-V2.6-Flash is the better buy for most teams most of the time, and that Xiaomi MiMo-V2.6-Pro is the better buy for the narrow class of work where failure is the expensive outcome. The one row where Flash wins does not overturn that — but it is a useful reminder that the flagship's premium is a claim about a workload, not about a model.

