Hero title card for the article 'Xiaomi MiMo-V2.6-Pro: 72.57 on DeepSWE', headed by the kicker 'OrcaRouter · model radar — unreleased', with the subtitle 'A run that stopped, a model that has not shipped — what is confirmed, and what is not.' Below it two cards: left, 'On Xiaomi's own dashboard', reading 'DeepSWE v1.1: 72.57 — run stopped at step 30, Sep 20'; right, 'Still unpublished', reading 'No model card, no API id, no price, no weights'. Four chips read 'Flash sibling: 65.68', 'Total spend: $3,474,715', 'Top of board: 74%' and 'Vendor-run eval, not submitted'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Xiaomi MiMo-V2.6-Pro: A 72.57 DeepSWE Score, a Run That Stopped, and Still No Release Date

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Xiaomi MiMo-V2.6-Pro stopped its publicly streamed reinforcement-learning run on September 20 with a DeepSWE v1.1 score of 72.57 sitting on Xiaomi's own dashboard — 1.4 points under the 74% that GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5 share at the top of the same leaderboard. Xiaomi MiMo-V2.6-Flash, the smaller model trained alongside it, stopped on September 19 at 65.68. Both runs ended at step 30, and the combined bill Xiaomi published alongside them read $3,474,715 when we captured the page this morning.

That is the whole of the good news, and it is genuinely good. The rest of the story is a gap between a number and a product. Neither Xiaomi MiMo-V2.6-Pro nor Xiaomi MiMo-V2.6-Flash has a release date, a model card, an API identifier, a published price, or a set of downloadable weights. There is no endpoint to call. What exists is a training dashboard that anyone can watch, a benchmark figure Xiaomi's own harness produced, and a community that has spent six days reading a telemetry page the way people used to read tea leaves.

So this is a what-we-know-so-far piece, and the honest version of it separates three things that the last week of coverage has been quietly blending: what Xiaomi has put on the record, what a single observer reported about it, and what any of it means for someone choosing a coding model today.

What Xiaomi actually put online

On September 16 the MiMo team, led by Luo Fuli, opened a live page at Xiaomi's own domain showing the post-training of two unreleased models in real time: step counts, reward curves, token throughput, sandbox counts, GPU fault notices, and a cost counter that ticks up as you watch. Xiaomi's framing, per the coverage, is that it spent roughly six months working on one question — how far reinforcement learning can be scaled — and decided to run that experiment in public rather than announce a result.

The dashboard is unusually candid for a vendor artifact. It lists its own failures in a notices feed: a Pro restart at step 17 caused by a GPU out-of-memory fault from expert load imbalance, which the team says it fixed by adjusting its training parallelism strategy; a network connectivity fault between the Pro training cluster and the grading deployment that forced a restart and a decision to drop the cyber dataset from the upcoming Pro run after "bad patterns in the rollout logs"; a Flash restart from step 15 after a class of infrastructure error went undetected for about three hours; and a Pro restart from a VRAM fault on a single node. It also notes that tasks judged "relatively easy for the current pro model" were filtered out — an admission most post-training reports leave unsaid.

Screenshot of Xiaomi's public MiMo-V2.6 reinforcement-learning dashboard captured September 21, 2026. A total-cost counter reads $3,474,715, and a notices feed lists a Pro restart at step 17 from a GPU out-of-memory fault caused by expert load imbalance, a network connectivity fault between the Pro training cluster and the grader deployment, a Flash restart from step 15 after an infrastructure error went undetected for about three hours, a Pro restart from a node VRAM fault, and a note that tasks relatively easy for the current Pro model were filtered out. Two run cards show mimo-v2.6-pro stopped at step 30 — started 2026-09-15 10:32 UTC, stopped 2026-09-20, $2,620,670 spent, 75B tokens, 753k samples, batch 1,568 x 16 — and mimo-v2.6-flash stopped at step 30 — started 2026-09-15 15:16 UTC, stopped 2026-09-19, $854,044 spent, 81.4B tokens. Three benchmark panels read DeepSWE v1.1 mini-swe-agent avg@3: mimo-v2.6-pro 72.57, mimo-v2.6-flash 65.68; an in-house coding bench at 65.43 and 62.87; and AutomationBench v1.0.6 at 53.10 and 52.70.

The two run cards are the part that matters for anyone tracking timing. Both models are labelled stopped: MiMo-V2.6-Pro after 5 days and 7 hours of wall-clock, MiMo-V2.6-Flash after 3 days 11 hours, each at step 30, on September 20 and September 19 respectively. Pro consumed 75B tokens and 753k samples for its $2.62M; Flash consumed 81.4B tokens for $854k. Each step draws a batch of 1,568 prompts with 16 rollouts per prompt — 25,088 trajectories per step, run fully asynchronously.

Worth stating plainly: Xiaomi has not said whether "stopped" means the run is finished, paused, or checkpointed. A stopped card is not a completion notice, and nobody outside the team knows which one it is.

What is confirmed, and what is not

The 72.57 is not a rumour. It is printed on Xiaomi's dashboard, in a chart labelled DeepSWE v1.1, mini-swe-agent, avg@3, next to the Pro run's orange curve. The Flash model's 65.68 sits on the same chart. The figure also appears — as 62.24 and later 65.97 — in earlier snapshots of the same panel, which is what gives the curve its shape: a steady climb across 30 steps rather than a single lucky evaluation.

What is report-only is the starting point. The claim circulating on X on September 21, from the account @nrehiew_, is that the Pro model went "from 58.41 to 72.57" on DeepSWE and that the runs have completed. The endpoint matches the vendor's dashboard exactly. The 58.41 baseline is that account's reading of where the curve begins — the chart's leftmost Pro point does sit in the high 50s — and Xiaomi has not published a step-1 figure. The completion claim is likewise an interpretation of two cards marked stopped, which is a reasonable reading and not a Xiaomi statement. Treat the 14-point improvement as directionally solid and numerically approximate.

A two-column scoreboard titled 'MiMo-V2.6-Pro vs the DeepSWE top tier', subtitled 'What Xiaomi published about its own checkpoint, against what the public leaderboard lists — September 21, 2026'. The left column, badged VENDOR-REPORTED and headed 'MiMo-V2.6-Pro', reads: DeepSWE v1.1 — 72.57; Basis — Xiaomi's own offline eval, mini-swe-agent avg@3; Confidence — not published; Cost per task — not published; Release — not announced; Independent runs — none. The right column, badged PUBLIC LEADERBOARD and headed 'Top tier — three-way tie', reads: DeepSWE v1.1 — 74%; Basis — submitted configs, Pass@1, GPT-6 Astra · Gemini 3.8 Flash · Claude Opus 5; Confidence — plus or minus 1 to 4 points; Cost per task — $2.36 to $11.84; Release — shipping today; Independent runs — on the public board. A footer reads 'MiMo-V2.6-Pro is Xiaomi's own offline evaluation, read from the vendor's public RL dashboard on 2026-09-21 and not submitted to the leaderboard. Top-tier figures per Datacurve's DeepSWE v1.1 leaderboard, updated 2026-09-03.' The OrcaRouter logo is composited in the bottom-right corner.

There is one more distinction the coverage has skipped, and it is the one that decides how much the number is worth. The dashboard's benchmark panels are Xiaomi's own evaluations: the company ran the harness on its own checkpoints and published the results. The harness is the same one the public leaderboard uses — mini-swe-agent, DeepSWE v1.1 — which makes the figure more comparable than a vendor's homegrown benchmark would be. But a self-run evaluation is not a leaderboard entry, and 72.57 does not appear on Datacurve's board, because nobody has submitted it.

The 74% ceiling is tighter than the headline suggests

Which brings the comparison into focus. Datacurve's DeepSWE v1.1 leaderboard, last updated September 3, 2026, lists 113 tasks across 91 repositories in five languages, and its top three configurations are tied at 74% Pass@1: GPT-6 Astra at xhigh reasoning effort, Gemini 3.8 Flash at high, and Claude Opus 5 at max. The numbers that sit next to those scores are the interesting ones.

• Confidence — GPT-6 Astra 74% ±3, Gemini 3.8 Flash 74% ±1, Claude Opus 5 74% ±4. On the same board, GPT-5.6 Sol sits at 73% ±3 and GLM-5.3 and Kimi K3 at 69%. The intervals overlap well past the second decimal place.

• Cost per task — Gemini 3.8 Flash averages $2.36, GPT-6 Astra $6.52, Claude Opus 5 $11.84. The leaderboard's own chart flags Gemini 3.8 Flash as the most efficient configuration on the board, and the spread between those three is roughly five to one for a pass rate the benchmark cannot separate.

• Steps — Gemini 3.8 Flash needed 166 agent steps per task on average, Claude Opus 5 99, GPT-6 Astra 29. The tied score hides three very different working styles, and the cheap one is the one that works hardest.

Screenshot of Datacurve's DeepSWE v1.1 leaderboard captured September 21, 2026, showing 113 tasks, 91 repositories, five languages and 28 models, with the v1.1 and Best tabs selected and the board marked 'updated September 3, 2026'. The Pass@1 table lists gpt-6-astra (xhigh) at 74 percent plus or minus 3 with an average cost of $6.52, 30k output tokens and 29 steps; gemini-3.8-flash (high) at 74 percent plus or minus 1 at $2.36, 143k tokens and 166 steps; claude-opus-5 (max) at 74 percent plus or minus 4 at $11.84, 118k tokens and 99 steps; gpt-5.6-sol at 73 percent; claude-fable-5 at 70 percent; glm-5.3 and kimi-k3 at 69 percent; and grok-4.6 and gpt-5.6-luna at 67 percent. The score-versus-average-cost chart above the table labels gemini-3.8-flash as the most efficient configuration. No Xiaomi MiMo model appears on the board.

So when the reported 72.57 is described as "approaching the top of the board," the accurate version is narrower and less dramatic. It is within about a point and a half of a three-way tie whose members cannot be separated from each other by the benchmark's own error bars. It is a strong result for a model still in post-training, produced by the vendor rather than the benchmark's maintainers, on a board the model has not been submitted to. The leaderboard's last update predates Xiaomi's run entirely, which is why nothing MiMo appears on it yet.

Why Xiaomi is training in public

The transparency is not incidental. Xiaomi's last open-weight release was Xiaomi MiMo-V2.5 and Xiaomi MiMo-V2.5-Pro at the end of April, both under the MIT licence — commercially usable, fine-tunable, redistributable. MiMo-V2.5-Pro is a 1.02-trillion-parameter mixture-of-experts model with 42B active parameters, a hybrid-attention architecture and a 1M-token context window, and its launch material made the agentic claims explicit: a compiler written from scratch in Rust across hundreds of tool calls, an eight-thousand-line desktop application built over eleven hours of agent work.

Streaming the next generation's post-training is a different kind of claim. A lab that publishes its reward curves, its sandbox counts and its GPU faults in real time is asking to be judged on process rather than on a launch deck, and it is signalling a timeline. Luo Fuli has said the technical details of the RL work will be open-sourced piece by piece over the coming weeks. That is the closest thing to a schedule anyone has offered: methodology first, model later.

It also fits a pattern Xiaomi has used before, where a model appears in public under another name before the company claims it. The company's habits are to train quietly, test in the open, and then release with weights.

What you can run today

Nothing in the MiMo-V2.6 family. If you want a Xiaomi MiMo model right now, the V2.5 generation is what exists — open weights through Xiaomi's own channels and several third-party platforms, with published cards and pricing. There is no MiMo-V2.6 API, no model identifier, and no price list, and any page quoting a MiMo-V2.6 identifier or per-token rate is filling in a blank Xiaomi has left empty. The `mimo-v2.6-pro` and `mimo-v2.6-flash` strings on the dashboard are training job names, not published endpoints.

The other practical consequence is what to do in the meantime. If the reason MiMo-V2.6-Pro is interesting to you is that it might be a cheaper way to hit frontier-level coding performance, the configurations it is being measured against are already callable, and the leaderboard says the cheap one is competitive: Gemini 3.8 Flash ties the top pass rate at $2.36 a task and is flagged the board's most efficient configuration. Gemini 3.8 Flash, Claude Opus 5 and GPT-6 Astra are all reachable through one OrcaRouter API key at provider list price with 0% markup passed through — so a vendor price change is live on our side the same day rather than at the next billing cycle — with failover configured per model instead of per vendor. And when a lab does open a model like this one, routing it behind the same key is the low-commitment way to evaluate it: you can put it in front of real traffic without betting a production path on a checkpoint nobody has characterised.

What would actually change this story

Four signals, roughly in order of how much they would settle:

• Weights on Hugging Face under a stated licence, which is how the V2.5 generation arrived and the strongest evidence that the RL run produced a shippable model rather than an experiment that ended.

• A model card with parameter counts, context length and architecture — none of the V2.6 numbers are public, and the dashboard does not show them. Assuming V2.6-Pro inherits V2.5-Pro's 1.02T/42B shape would be a guess.

• An API identifier and a price sheet, which is the moment the DeepSWE figure becomes a purchasing decision rather than a spectator sport.

• A submission to Datacurve's DeepSWE board, which would put MiMo-V2.6-Pro on the same table and under the same error bars as the models it is being compared to. A vendor-run eval and a leaderboard submission are not the same claim, and the gap between them is exactly one row of that table.

Until then, the honest summary is that Xiaomi has shown the most transparent post-training run any major lab has published, on two models it has not shipped, with one self-reported benchmark figure that lands close to — but does not join — the top of the board. The methodology write-up is promised in the coming weeks. The release date is not promised at all.