
Xiaomi MiMo-V2.6-Pro: A 72.57 DeepSWE Score, a Run That Stopped, and Still No Release Date
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Xiaomi MiMo-V2.6-Pro stopped its publicly streamed reinforcement-learning run on September 20 with a DeepSWE v1.1 score of 72.57 sitting on Xiaomi's own dashboard — 1.4 points under the 74% that GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5 share at the top of the same leaderboard. Xiaomi MiMo-V2.6-Flash, the smaller model trained alongside it, stopped on September 19 at 65.68. Both runs ended at step 30, and the combined bill Xiaomi published alongside them read $3,474,715 when we captured the page this morning.
That is the whole of the good news, and it is genuinely good. The rest of the story is a gap between a number and a product. Neither Xiaomi MiMo-V2.6-Pro nor Xiaomi MiMo-V2.6-Flash has a release date, a model card, an API identifier, a published price, or a set of downloadable weights. There is no endpoint to call. What exists is a training dashboard that anyone can watch, a benchmark figure Xiaomi's own harness produced, and a community that has spent six days reading a telemetry page the way people used to read tea leaves.
So this is a what-we-know-so-far piece, and the honest version of it separates three things that the last week of coverage has been quietly blending: what Xiaomi has put on the record, what a single observer reported about it, and what any of it means for someone choosing a coding model today.
What Xiaomi actually put online
On September 16 the MiMo team, led by Luo Fuli, opened a live page at Xiaomi's own domain showing the post-training of two unreleased models in real time: step counts, reward curves, token throughput, sandbox counts, GPU fault notices, and a cost counter that ticks up as you watch. Xiaomi's framing, per the coverage, is that it spent roughly six months working on one question — how far reinforcement learning can be scaled — and decided to run that experiment in public rather than announce a result.
The dashboard is unusually candid for a vendor artifact. It lists its own failures in a notices feed: a Pro restart at step 17 caused by a GPU out-of-memory fault from expert load imbalance, which the team says it fixed by adjusting its training parallelism strategy; a network connectivity fault between the Pro training cluster and the grading deployment that forced a restart and a decision to drop the cyber dataset from the upcoming Pro run after "bad patterns in the rollout logs"; a Flash restart from step 15 after a class of infrastructure error went undetected for about three hours; and a Pro restart from a VRAM fault on a single node. It also notes that tasks judged "relatively easy for the current pro model" were filtered out — an admission most post-training reports leave unsaid.

The two run cards are the part that matters for anyone tracking timing. Both models are labelled stopped: MiMo-V2.6-Pro after 5 days and 7 hours of wall-clock, MiMo-V2.6-Flash after 3 days 11 hours, each at step 30, on September 20 and September 19 respectively. Pro consumed 75B tokens and 753k samples for its $2.62M; Flash consumed 81.4B tokens for $854k. Each step draws a batch of 1,568 prompts with 16 rollouts per prompt — 25,088 trajectories per step, run fully asynchronously.
Worth stating plainly: Xiaomi has not said whether "stopped" means the run is finished, paused, or checkpointed. A stopped card is not a completion notice, and nobody outside the team knows which one it is.
What is confirmed, and what is not
The 72.57 is not a rumour. It is printed on Xiaomi's dashboard, in a chart labelled DeepSWE v1.1, mini-swe-agent, avg@3, next to the Pro run's orange curve. The Flash model's 65.68 sits on the same chart. The figure also appears — as 62.24 and later 65.97 — in earlier snapshots of the same panel, which is what gives the curve its shape: a steady climb across 30 steps rather than a single lucky evaluation.
What is report-only is the starting point. The claim circulating on X on September 21, from the account @nrehiew_, is that the Pro model went "from 58.41 to 72.57" on DeepSWE and that the runs have completed. The endpoint matches the vendor's dashboard exactly. The 58.41 baseline is that account's reading of where the curve begins — the chart's leftmost Pro point does sit in the high 50s — and Xiaomi has not published a step-1 figure. The completion claim is likewise an interpretation of two cards marked stopped, which is a reasonable reading and not a Xiaomi statement. Treat the 14-point improvement as directionally solid and numerically approximate.

There is one more distinction the coverage has skipped, and it is the one that decides how much the number is worth. The dashboard's benchmark panels are Xiaomi's own evaluations: the company ran the harness on its own checkpoints and published the results. The harness is the same one the public leaderboard uses — mini-swe-agent, DeepSWE v1.1 — which makes the figure more comparable than a vendor's homegrown benchmark would be. But a self-run evaluation is not a leaderboard entry, and 72.57 does not appear on Datacurve's board, because nobody has submitted it.
The 74% ceiling is tighter than the headline suggests
Which brings the comparison into focus. Datacurve's DeepSWE v1.1 leaderboard, last updated September 3, 2026, lists 113 tasks across 91 repositories in five languages, and its top three configurations are tied at 74% Pass@1: GPT-6 Astra at xhigh reasoning effort, Gemini 3.8 Flash at high, and Claude Opus 5 at max. The numbers that sit next to those scores are the interesting ones.
• Confidence — GPT-6 Astra 74% ±3, Gemini 3.8 Flash 74% ±1, Claude Opus 5 74% ±4. On the same board, GPT-5.6 Sol sits at 73% ±3 and GLM-5.3 and Kimi K3 at 69%. The intervals overlap well past the second decimal place.
• Cost per task — Gemini 3.8 Flash averages $2.36, GPT-6 Astra $6.52, Claude Opus 5 $11.84. The leaderboard's own chart flags Gemini 3.8 Flash as the most efficient configuration on the board, and the spread between those three is roughly five to one for a pass rate the benchmark cannot separate.
• Steps — Gemini 3.8 Flash needed 166 agent steps per task on average, Claude Opus 5 99, GPT-6 Astra 29. The tied score hides three very different working styles, and the cheap one is the one that works hardest.

So when the reported 72.57 is described as "approaching the top of the board," the accurate version is narrower and less dramatic. It is within about a point and a half of a three-way tie whose members cannot be separated from each other by the benchmark's own error bars. It is a strong result for a model still in post-training, produced by the vendor rather than the benchmark's maintainers, on a board the model has not been submitted to. The leaderboard's last update predates Xiaomi's run entirely, which is why nothing MiMo appears on it yet.
Why Xiaomi is training in public
The transparency is not incidental. Xiaomi's last open-weight release was Xiaomi MiMo-V2.5 and Xiaomi MiMo-V2.5-Pro at the end of April, both under the MIT licence — commercially usable, fine-tunable, redistributable. MiMo-V2.5-Pro is a 1.02-trillion-parameter mixture-of-experts model with 42B active parameters, a hybrid-attention architecture and a 1M-token context window, and its launch material made the agentic claims explicit: a compiler written from scratch in Rust across hundreds of tool calls, an eight-thousand-line desktop application built over eleven hours of agent work.
Streaming the next generation's post-training is a different kind of claim. A lab that publishes its reward curves, its sandbox counts and its GPU faults in real time is asking to be judged on process rather than on a launch deck, and it is signalling a timeline. Luo Fuli has said the technical details of the RL work will be open-sourced piece by piece over the coming weeks. That is the closest thing to a schedule anyone has offered: methodology first, model later.
It also fits a pattern Xiaomi has used before, where a model appears in public under another name before the company claims it. The company's habits are to train quietly, test in the open, and then release with weights.
What you can run today
Nothing in the MiMo-V2.6 family. If you want a Xiaomi MiMo model right now, the V2.5 generation is what exists — open weights through Xiaomi's own channels and several third-party platforms, with published cards and pricing. There is no MiMo-V2.6 API, no model identifier, and no price list, and any page quoting a MiMo-V2.6 identifier or per-token rate is filling in a blank Xiaomi has left empty. The `mimo-v2.6-pro` and `mimo-v2.6-flash` strings on the dashboard are training job names, not published endpoints.
The other practical consequence is what to do in the meantime. If the reason MiMo-V2.6-Pro is interesting to you is that it might be a cheaper way to hit frontier-level coding performance, the configurations it is being measured against are already callable, and the leaderboard says the cheap one is competitive: Gemini 3.8 Flash ties the top pass rate at $2.36 a task and is flagged the board's most efficient configuration. Gemini 3.8 Flash, Claude Opus 5 and GPT-6 Astra are all reachable through one OrcaRouter API key at provider list price with 0% markup passed through — so a vendor price change is live on our side the same day rather than at the next billing cycle — with failover configured per model instead of per vendor. And when a lab does open a model like this one, routing it behind the same key is the low-commitment way to evaluate it: you can put it in front of real traffic without betting a production path on a checkpoint nobody has characterised.
What would actually change this story
Four signals, roughly in order of how much they would settle:
• Weights on Hugging Face under a stated licence, which is how the V2.5 generation arrived and the strongest evidence that the RL run produced a shippable model rather than an experiment that ended.
• A model card with parameter counts, context length and architecture — none of the V2.6 numbers are public, and the dashboard does not show them. Assuming V2.6-Pro inherits V2.5-Pro's 1.02T/42B shape would be a guess.
• An API identifier and a price sheet, which is the moment the DeepSWE figure becomes a purchasing decision rather than a spectator sport.
• A submission to Datacurve's DeepSWE board, which would put MiMo-V2.6-Pro on the same table and under the same error bars as the models it is being compared to. A vendor-run eval and a leaderboard submission are not the same claim, and the gap between them is exactly one row of that table.
Until then, the honest summary is that Xiaomi has shown the most transparent post-training run any major lab has published, on two models it has not shipped, with one self-reported benchmark figure that lands close to — but does not join — the top of the board. The methodology write-up is promised in the coming weeks. The release date is not promised at all.
