OrcaRouter model radar hero card headed 'MiMo-V2.6-Flash vs MiMo-V2.6', subtitled 'One series, two checkpoints, and a name that does double duty', with a left panel for MiMo-V2.6-Flash reading 309B total / 15B active, 172.9 GB of FP8 weights and Efficiency checkpoint, a right panel for MiMo-V2.6-Pro reading 1.02T total / 42B active, 'the name people write for the flagship' and Same MIT licence, and three badges below reading Published 21 Sept 2026, 1M-token context claim and No Xiaomi endpoint to call.
Guides & Insights

MiMo-V2.6-Flash vs MiMo-V2.6: One Series, Two Checkpoints, and a Name That Does Double Duty

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ask two people what MiMo-V2.6 is and you can get two different models. Xiaomi's September 2026 release is a series with two published checkpoints — MiMo-V2.6-Flash at 309 billion parameters and MiMo-V2.6-Pro at 1.02 trillion — and the bare name MiMo-V2.6 gets used for both the family and, in most coverage, the generation's top tier. So a page headed "MiMo-V2.6-Flash vs MiMo-V2.6" is really the efficiency checkpoint against the name people write when they mean the flagship. Both checkpoints landed on Hugging Face eighteen seconds apart on 21 September 2026, both carry the MIT licence, and both are downloads first — there is no Xiaomi endpoint to call. Almost nothing else about them is the same.

That matters more than a naming quibble usually would, because the two are not a size ladder you can walk up later. They are separate training runs with separate benchmark tables, separate memory footprints, and — as the released numbers show — a couple of places where the smaller one wins outright.

Two checkpoints, one published minute

The repositories are XiaomiMiMo/MiMo-V2.6-Flash-RL, created at 15:39:51 UTC, and XiaomiMiMo/MiMo-V2.6-Pro-RL, created at 15:39:33 UTC. Neither is gated. Both attach a technical report and a model card, and both are described by Xiaomi as the product of a single mixed reinforcement-learning run that folded coding, general-agent, visual and cybersecurity tasks into one training pass rather than separate per-domain runs.

• Parameters — MiMo-V2.6-Flash: 309B total, 15B activated per token. MiMo-V2.6-Pro: 1.02T total, 42B activated. Both sparse mixture-of-experts.

• Backbone — Flash is 48 layers, 39 sliding-window and 9 full attention, hidden size 4096, 256 routed experts with 8 activated, a 128-token sliding window and no shared experts. Pro runs the same hybrid attention idea at a much larger width.

• Modality — both are native omnimodal, taking text, image, video and audio into one model rather than three bolted-together pipelines. Flash carries a 681M-parameter vision transformer, a 308M audio tokenizer and a 127M audio patch encoder.

• Context — 1M tokens, claimed for both, with the config backing a 1,048,576-token maximum position embedding.

• Licence — MIT on both, with no revenue threshold, no acceptable-use gate beyond the usual, and no research clause. Commercial deployment, modification and redistribution are all permitted.

• Weight files — Flash publishes 172.9 GB of FP8 weight data across 65 shards. Pro is roughly three times the parameter count, so its download lands in the half-terabyte range before any quantisation you apply yourself.

• Where they are served — Hugging Face, Xiaomi's own API platform, AI Studio, MiMo Code and the MiMo desktop app. The card notes that neither checkpoint is currently deployed by a third-party inference provider on the Hub itself.

The collision that costs people hardware

The confusion in this pairing is not really between Flash and Pro. It is between MiMo-V2.6-Flash and the model that came before it.

MiMo-V2-Flash shipped in December 2025, and Xiaomi's own blog post announcing it describes a mixture-of-experts model with 309 billion total parameters, 15 billion active, a hybrid attention scheme interleaving sliding-window and full attention, and an aggressive 128-token sliding window. Read that against the bullet list above. The 2026 checkpoint and the 2025 checkpoint have the same headline shape, the same sliding-window size, and the same activation budget — nine months apart, from different training runs, with different capabilities and a different benchmark table.

So a search for "MiMo V2.6 Flash" will happily hand you pages about MiMo-V2-Flash, complete with a 256K context figure and a price of a tenth of a dollar per million input tokens that belongs to the older model. If you are sizing a node, that is not a cosmetic error. The older checkpoint and the new one are different downloads with different serving recipes, and the new one claims four times the context.

There is a second, quieter trap in the naming. Both repositories end in -RL, which reads like a reinforcement-learning adapter sitting on top of some other base model. They are not adapters. The suffix marks the post-training lineage: these are the full checkpoints that came out of the streamed RL run. The weight index confirms it — tens of thousands of tensors covering the whole backbone, the vision encoder, the audio stack and the speculative drafter.

What the vendor's own table says

Every figure in this section comes from Xiaomi's evaluation tables in the two model cards. Xiaomi ran the harnesses on its own checkpoints. None of it has been reproduced outside the lab, none of it appears on a public leaderboard, and the comparison columns are the vendor's own runs of the same harnesses against other companies' models. Treat the ranking as informative and the decimal places as decoration.

• Code agent — DeepSWE v1.1: Flash 67.9, Pro 71.9. Terminal Bench 2.1: Flash 87.6, Pro 89.9. ProgramBench: Flash 26.0. MiMo Code Bench: Flash 61.2.

• General agent — AutomationBench v1.0.6: Flash 52.3, Pro 53.1. Toolathlon-Verified: Flash 73.6, Pro 76.9. OSWorld-Verified: Flash 80.8, Pro 82.0. Agents' Last Exam: Flash 27.6. Terminal Bench 4.0: Flash 28.8.

• Cybersecurity — the one column where the smaller model leads. CyberGym: Flash 95.1 against Pro's 94.0. MiMo Cyber Bench: Flash 77.2. Then the floor drops out: ExploitGym 6.0, ExploitBench 25.3, SEC Bench Pro 47.5.

• Visual agent — MiMo VisualCoding: Flash 71.5, Pro 72.3.

Read down those bullets and the honest summary is that Pro is ahead almost everywhere, by small margins, and that the margins are small enough to be inside the noise of a self-run harness. On DeepSWE the two are four points apart on a scale where the same checkpoint moved two points between two of Xiaomi's own evaluations.

That last point deserves its own paragraph, because it is the most useful thing in either card. Xiaomi's public RL dashboard, which streamed both training runs through September, published Flash at 65.68 on DeepSWE v1.1 using a mini-swe-agent harness at average-of-three. The finished card prints 67.9 for the released weights. Meanwhile Pro's dashboard figure was 72.57 and its card prints 71.9 — it went down. Two checkpoints from the same run, evaluated twice, and the gap between the two evaluations is the same size as the gap between the two models. If you have been quoting dashboard numbers since last week, they describe training snapshots, not the artifacts you would download today.

Scoreboard card headed 'MiMo-V2.6-Flash vs MiMo-V2.6-Pro', with paired rows for parameters (309B total / 15B active against 1.02T total / 42B active), weights on disk (172.9 GB FP8 across 65 shards against about three times Flash), DeepSWE v1.1 (67.9 against 71.9, with the RL dashboard's earlier 65.68 and 72.57 noted), CyberGym (95.1 against 94.0), independent score (None published against an Artificial Analysis Index of 46 at 134.3 output tokens per second) and price, plus a sourcing footer stating that every benchmark is Xiaomi's own run except the Pro index score.

Neither one is on a router

Here is the part that decides most real evaluations. Xiaomi sells neither checkpoint: there is no per-token price published for the MiMo-V2.6 generation on its own site, and no callable model identifier it has announced for either one. What exists instead is a download, plus — for Pro only — a single API provider that Artificial Analysis counts as serving it at $0.435 per million input tokens and $0.87 per million output. That is a provider's listing rather than a vendor rate card, and there is nothing comparable for Flash at all. Figures circulating on third-party catalogue pages should be read the same way.

So the choice between Flash and Pro is not a choice between two endpoints with different meters. It is a choice between two GPU bills, and the smaller one is roughly a third of the larger. That is the entire commercial argument for Flash, and it is a stronger one than the benchmark table suggests, because the two models are within a few points of each other on the tasks most people are buying for.

What that leaves is the familiar three-way split. Rent the frontier, own a small model, or own a large one. The frontier column is the one Xiaomi benchmarks against in its own tables — Claude Opus 5, GPT-5.6 Sol and Claude Fable 5 all sit a few points above Pro on DeepSWE in the vendor's own comparison, and all three are callable today through one API key on OrcaRouter at provider list price with 0% markup passed through. That matters for the specific decision in front of you: a vendor price change lands on your side the same day rather than at the next contract renewal, so the rent-versus-own arithmetic stays honest as the hosted prices move. Automatic failover also means a model you are still evaluating never has to sit on a production path alone while you decide.

What you cannot do on any router today — ours included — is call MiMo-V2.6-Flash or MiMo-V2.6-Pro. We route no Xiaomi model, and the availability line in Xiaomi's own card points at Xiaomi's own channels: the MiMo API platform, AI Studio, MiMo Code and the desktop app.

Which checkpoint, and when

Take MiMo-V2.6-Flash if you want the MiMo-V2.6 generation and the hardware is the binding constraint. You give up roughly four DeepSWE points and a point or two on most agent benchmarks against the flagship. You gain a checkpoint that is about a third of the memory, that leads Pro on CyberGym in Xiaomi's own table, and that shares the same licence, the same 1M-token context claim and the same multimodality. For a team doing cybersecurity evaluation work, that CyberGym column is not a rounding error.

Take MiMo-V2.6-Pro if the workload is coding or long-horizon agentic work and the node is already paid for. It is the better model on nearly every column, it is the one Artificial Analysis has actually measured — an Intelligence Index of 46 on the recalibrated v4.3 scale, first of the 114 models in its open-weights class, 134.3 output tokens per second, and listed at $0.435 per million input tokens and $0.87 per million output — and independent measurement is worth more than a vendor table when you are planning production.

Take neither until you have read the config files rather than the summaries. The Flash card describes a five-layer speculative drafter predicting seven tokens per pass; the config.json in the same repository sets num_nextn_predict_layers to 3. The card's summary is not what the runtime reads. And the Safetensors block on the repository page reports 159B parameters for Flash and 524B for Pro, neither of which matches the total or the active count in either model card. Xiaomi has not explained the discrepancy. Size from the published weight data instead, because that is the number you pay for in VRAM regardless of how the parameters are counted.

Screenshot of the Hugging Face model page for XiaomiMiMo/MiMo-V2.6-Flash-RL, showing the MIT licence and multimodal tags, the model card heading 'Scaling Reinforcement Learning Toward Self-Improvement', a Safetensors block reporting a 159B model size, and an inference-providers panel stating that the model is not deployed by any inference provider.

What would settle it

Three things, in rough order of usefulness. A third-party run on the same harnesses — DeepSWE or Terminal Bench, submitted rather than self-reported — would tell you whether the four-point gap between Flash and Pro is a real position or a favourable configuration. A published price list would turn the whole comparison from a hardware project into a line item you can put beside the hosted frontier. And the RL environments, which Xiaomi has said it will open-source, are the artifact most teams would actually want, because they are what you would need to reproduce the loop on your own traces.

Until then the practical reading is narrow. MiMo-V2.6 is two checkpoints, not one, and the bare name usually means the expensive one. If you are choosing inside the series today, the honest default is Flash unless a benchmark column you specifically care about says otherwise — and if you are not ready to buy a node, neither is the right answer yet.

Screenshot of the Artificial Analysis model page for MiMo-V2.6-Pro, showing an Intelligence Index of 46 on the updated scale, ranked first of 114, output speed of 134.3 tokens per second, a price of $0.435 per million input tokens and $0.87 per million output with a 99% cache discount and $0.13 cost per Intelligence Index task, a 1M-token context window, and 1.02T total with 42B active parameters under the MIT licence with weights on Hugging Face.