
Xiaomi MiMo-V2.6-Pro-UltraSpeed vs Xiaomi MiMo-V2.6: One Checkpoint, Ten Times the Price
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Xiaomi MiMo-V2.6-Pro-UltraSpeed and Xiaomi MiMo-V2.6-Pro are not two models. They are one model sold two ways. Both are served from the same one-trillion-parameter MiMo-V2.6 checkpoint that Xiaomi published in September 2026 — the Pro tier of the Xiaomi MiMo-V2.6 series, whose smaller sibling is Xiaomi MiMo-V2.6-Flash. The difference between the two Pro offerings is entirely in how fast the tokens come out and how much you pay for them. The standard tier asks $0.435 per million input tokens and $0.87 per million output tokens. The UltraSpeed tier asks $4.35 and $8.70. That is a clean ten-to-one on both sides of the meter, and it buys a claim of roughly ten times the output speed — a claim Xiaomi makes about its own serving stack, and one that no independent benchmark has published a number for yet.
So the interesting question is not which model is better. It is whether a ten-times multiple is the right price for ten times the speed, and that turns out to have a precedent worth reading before you commit a production path to it.
The fast tier is a serving decision, not a model decision
Xiaomi's own framing for the UltraSpeed line is that it matches the standard model in quality and differs only in throughput. That matters more than it sounds, because it removes the usual reason to hesitate over a new tier. You are not betting on a different checkpoint's personality, its refusal behaviour, or its formatting quirks. You are betting that the same weights, run through a more aggressive serving configuration, will behave the same way on your prompts. If that holds, switching between the two tiers is a config change, not a migration.
What is inside that serving configuration has not been documented for this generation. The previous UltraSpeed tier, built on the MiMo-V2.5-Pro checkpoint in June 2026, was described by Xiaomi as using FP4 quantization on the expert layers only, a block-parallel speculative decoding scheme called DFlash, and a runtime called TileRT, and it ran on a single eight-GPU node. Xiaomi has not published the equivalent detail for the V2.6 tier. Treat the earlier stack as context for how the speed is obtained, not as a description of what is running now.
The line-up it sits in is short. Alongside the two Pro tiers there is MiMo-V2.6-Flash, the smaller sibling at 309 billion total parameters and 15 billion active. Pro is the sparse mixture-of-experts flagship: Artificial Analysis lists it at 1.0 trillion total parameters with 42 billion active per token, a 1-million-token context window, and an MIT licence with weights on Hugging Face. Those are the numbers to hold onto, because the whole comparison rests on them.
What is actually verified, and what is not

The asymmetry between the two columns above is the single most important thing in this article. Almost everything measurable about this pairing has been measured on the slow side.
Artificial Analysis has benchmarked MiMo-V2.6-Pro and publishes an Intelligence Index of 46 for it — the top score in the 114-model class it groups the model into. It measures 134.3 output tokens per second through Xiaomi's API, against a class median of 77.9, and a time to first chunk of 2.15 seconds against a median of 2.34. It lists the cost per Intelligence Index task at $0.13, a 99% cache discount on the input price, and output verbosity at 140 million tokens. Those are independent figures on a named configuration, and they are the only hard throughput anchor this comparison has.
For UltraSpeed, the verified list is much shorter:
• Output speed — about ten times the standard tier, stated by Xiaomi and repeated by the platforms listing it. No third party has published a tokens-per-second measurement for this specific tier.
• Price — $4.35 per million input tokens and $8.70 per million output tokens, a ten-times multiple of the standard tier on both. No cache tier is listed alongside those two rates.
• Quality — "matches the original", again a vendor statement. No independent evaluation of the UltraSpeed endpoint has been published, so the claim that quality is unchanged is untested rather than disproven.
• Context and modality — 1,048,576 tokens and native text, image, video and audio input, inherited from the base checkpoint.
There is also a vendor benchmark worth naming precisely because of how it has travelled. Xiaomi's public reinforcement-learning dashboard, which streamed the V2.6 post-training runs in September 2026, reports DeepSWE v1.1 scores of 72.57 for the Pro run and 65.68 for Flash. Those come from Xiaomi's own offline evaluation using a mini-swe-agent harness at average-of-three, they were never submitted to the public leaderboard, and they have not been reproduced outside the lab. If you have seen 72.57 quoted as a leaderboard score, that is a vendor number wearing a leaderboard's clothes.

The precedent: last time Xiaomi charged three times, not ten
This is not Xiaomi's first speed tier, and the previous one is the most useful comparison in the whole story — because the multiplier changed.
MiMo-V2.5-Pro-UltraSpeed arrived in June 2026, built on the V2.5-Pro checkpoint, and it was priced at roughly three times the standard model's output rate. Xiaomi's pitch then was the same pitch it is making now: about ten times the output speed. Coverage at the time reported sustained throughput around 1,000 tokens per second with peaks near 1,200, on a single eight-GPU node, though the model page itself listed a wider 500-to-1,000 range — and none of it was independently verified. Access was gated behind an application form rather than open sign-up.
Set the two side by side and the shift is the story. The speed multiple stayed at roughly ten. The price multiple moved from three to ten. That is a very different deal for the same kind of upgrade, and Xiaomi has not published an explanation for the change. It could reflect genuinely more expensive serving — a larger checkpoint, a more aggressive decoding configuration, or tighter capacity at launch. It could reflect launch-window pricing on a tier that has been available for a single day. What it does not yet have is a public justification you can check.
One practical difference is worth flagging for anyone who used the V2.5 tier: that one was a gated trial. The V2.6 tier is listed as generally available through Xiaomi's own API and several third-party platforms, at the published rates, without an application step. Whatever else changed, the friction to try it went down.
What ten times costs on a real workload
Percentages hide the shape of this, so here is the arithmetic on a concrete agent run — one that reads 20 million input tokens and emits 2 million output tokens over its lifetime. That is a heavy but not exotic agentic workload, the kind where a model loops over tool calls and re-reads its own context.
• Standard tier, input — 20M tokens at $0.435 per million: $8.70.
• Standard tier, output — 2M tokens at $0.87 per million: $1.74.
• Standard tier, total — $10.44 per run.
• UltraSpeed tier, input — 20M tokens at $4.35 per million: $87.00.
• UltraSpeed tier, output — 2M tokens at $8.70 per million: $17.40.
• UltraSpeed tier, total — $104.40 per run.
The difference is $93.96, and it is exactly ten times the bill rather than ten times one line of it, because both rates scale together. Now the other side of the ledger. At the 134.3 output tokens per second Artificial Analysis measures for the standard tier, generating 2 million output tokens takes a little over four hours of pure generation time. At ten times that rate it takes about twenty-five minutes. So the extra $94 buys back somewhere in the region of three and a half hours of wall clock on this run — roughly $27 for every hour saved, if you want to put it that way.
Whether that trade is good depends entirely on what the wall clock is worth to you, and there is one wrinkle that cuts against a naive read. The standard tier's input price carries a 99% cache discount on Artificial Analysis's page, which means the input half of the bill can collapse to near nothing on workloads that re-read the same context. UltraSpeed's listing shows the two headline rates and no cache tier. If your workload is cache-heavy, the effective multiple is not ten — it is much worse, because you lose the discount on the side where you were paying almost nothing. If your workload is output-heavy and cache-light, ten times is close to the real number. Measure your own mix before you trust either figure.
The open-weights asymmetry
There is a structural difference between the two tiers that has nothing to do with price or speed, and it may matter more than either.
MiMo-V2.6-Pro ships under an MIT licence with weights published on Hugging Face. That permits commercial deployment, modification and further training, and it means the standard model has a floor under its price that no vendor can lift: if the hosted rate stops making sense, you can serve the checkpoint yourself. You will not match Xiaomi's throughput on your own hardware, and you will pay for the GPUs, but the option exists and it caps what the hosted tier can charge.
UltraSpeed has no such floor. There are no UltraSpeed weights, because the tier is the serving stack rather than the model — the checkpoint is the same one that is already public, and what you are renting is Xiaomi's ability to run it fast. That is a legitimate thing to sell, and it is the same shape as every other speed tier on the market. It does mean the ten-times rate has no competitive check on it except other providers running the same open weights, and no self-hosting escape hatch if the price moves again.

Read that against the availability picture. The standard checkpoint is reachable through Xiaomi's own channels and several third-party platforms, and it is also a download. The UltraSpeed tier is reachable through Xiaomi's own API and several third-party platforms, and that is the only way to get it. For anyone weighing a long-lived dependency, that difference in optionality is worth pricing alongside the per-token rate.
Who should take which
The decision splits cleanly along how much of your cost is output tokens and how much of your latency budget is generation rather than queueing.
• Take UltraSpeed when the workload is output-heavy, cache-light, and interactive — long-form generation, agent loops where the model writes a lot of reasoning between tool calls, or anything where a human is waiting on the other end of the request.
• Take the standard tier when the workload is cache-heavy, batch-tolerant, or cost-sensitive at the margin — bulk classification, retrieval-augmented answering over a fixed corpus, overnight evaluation runs, anything where four hours of generation is fine because nobody is watching.
• Take the standard tier when you want the option to self-host. If your compliance posture requires weights you can hold, UltraSpeed is not available to you at any price.
• Take neither until you have measured. The ten-times claim is the load-bearing number here and it is currently vendor-stated. A single afternoon of running your own prompts through both tiers tells you more than any of the published figures, because the only throughput number anyone has independently verified belongs to the slow side of the comparison.
On that last point, the mechanics of testing a serving tier are the same as testing a model, and they are where a routing layer earns its place. OrcaRouter carries more than 200 models behind a single API key at provider list price with a 0% markup passed through, so a vendor price change lands on your side the same day rather than at the next contract renewal. Automatic failover means a tier you are still evaluating never sits on a production path by itself, and the routing DSL lets you send one class of request to the fast endpoint and everything else to the standard one without maintaining two integrations. None of that requires Xiaomi's models specifically — the point is that tier-level decisions like this one are cheap to reverse when the tiers sit behind one key.
What would settle this
The honest state of the MiMo-V2.6-Pro-UltraSpeed versus MiMo-V2.6-Pro question is that half of it is measured and half of it is asserted. The standard tier has an independent Intelligence Index, an independent throughput figure, and published open weights. The fast tier has a price, a context window, and a vendor promise that it is the same model going faster.
Three things would close the gap. An independent throughput measurement on the UltraSpeed endpoint — Artificial Analysis measured the standard tier through Xiaomi's API, so the same treatment is possible and simply has not happened yet. A cache policy for the fast tier, since the absence of one is what turns a ten-times bill into something worse on cache-heavy work. And any published detail on the V2.6 serving stack, in the way Xiaomi documented FP4 experts, DFlash and TileRT for the V2.5 generation.
Until then, the ten-times price is real and the ten-times speed is a claim. If your workload is output-heavy and someone is waiting, that claim is worth testing on your own prompts — and worth testing behind a layer that lets you switch back the same afternoon if it does not hold.
