
GPT-6.1 Ultrafast vs Xiaomi MiMo v2.6 Pro Ultraspeed: Two Fast Lanes, One Says What It Cost
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 111 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 55 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 377 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The two fastest ways to buy a frontier model went on sale sixteen days apart, and only one of them told you how the speed was produced. GPT-6.1 Ultrafast — the serving tier over GPT-6.1 Sol, which OpenAI released on September 29, 2026 and which gained this tier on October 8, 2026 — is a closed flagship you reach through an API, priced at six times Standard with no published account of what happens between the weights and your request. Xiaomi MiMo v2.6 Pro Ultraspeed, released September 22, 2026, is a speed tier over MiMo-V2.6-Pro, an MIT-licensed checkpoint whose reinforcement-learning environment Xiaomi also published. Both sell latency at a multiple. Only Xiaomi's multiple sits on top of something a customer can inspect, and that difference is worth more to the decision than either speed claim.
The price of the two lanes is the easy part and it is covered below. The hard part is that a speed tier is a promise about serving, and the two vendors have made that promise at very different levels of disclosure — one borrowed its headline number from a sibling model, the other shipped the training recipe alongside the price list. If you are about to commit a production workload to a fast lane, the disclosure gap is the risk you are actually taking on.
The prices, with both sides on the same line
Both tiers are sold as a multiple of a standard lane, and the multiples are close enough that price is not what separates them.
• Standard lane — GPT-6.1 Sol at $2.00 input / $10.00 output per million tokens, against MiMo-V2.6-Pro at $0.44 / $0.87.
• Fast lane — GPT-6.1 Ultrafast at $12.00 / $60.00, against MiMo-V2.6-Pro-UltraSpeed at $4.35 / $8.70.
• The multiple — exactly 6x on every line for the OpenAI tier, roughly 9.9x on input and exactly 10x on output for Xiaomi's.
• Cached input — GPT-6.1 Sol reads $0.10 per million on Standard and $0.60 on Ultrafast; the MiMo rate card we can read does not publish a separate cached-input line for either lane.
• Long context — GPT-6.1 Sol reprices the whole request at 2x input and cache rates and 1.5x output above 272,000 input tokens, putting Ultrafast at $24.00 / $1.20 / $30.00 / $90.00. MiMo-V2.6-Pro carries a 1M-token context window with no equivalent published cliff.
• Absolute spread — the two fast lanes are 2.8x apart on input and 6.9x apart on output, which is the gap between a frontier closed model and an open-weight one showing up inside the speed tier rather than at the standard rate.
The one-line summary of that block: Xiaomi charges a bigger multiple on a much cheaper model, and OpenAI charges a smaller multiple on a much more expensive one. On output tokens you pay $60.00 per million for the OpenAI lane and $8.70 for Xiaomi's. Neither vendor is offering a discount for speed; both are pricing it as a separate product, which is the first sign that in both cases the standard lane is still the default and the fast lane is an exception you justify per workload.
What each vendor will tell you about the speed
This is where the two rollouts stop being symmetrical, and it is the most useful thing on the page.
OpenAI's published number for Ultrafast is that GPT-6 Astra Ultrafast generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex. That is a measurement of a different model, in a different client, phrased as a ceiling. The documentation that covers Ultrafast for GPT-6.1 Sol says the tier reduces the time between generated output tokens and points at a rate card; it does not put a number on the Sol tier at all. So the headline number attached to this rollout is borrowed from the sibling model that has been on Ultrafast since September 29, and no independent party has published a tokens-per-second measurement for either. It may well hold — same serving stack, same mechanism — but it is a vendor ceiling that OpenAI itself did not measure for this model.
Xiaomi's disclosure is different in kind, and the difference is not that it is more precise. Its own material claims UltraSpeed reaches up to 20 times the output speed of the standard Pro service at the same quality. Third-party catalogue listings describing the same model on the same day say roughly 10 times. Both numbers are in circulation, they are not the same number, and nobody has published a measurement that reconciles them. Reading that honestly: "up to 20x" is a vendor ceiling under conditions Xiaomi has not specified, "roughly 10x" is what a catalogue was willing to assert as typical, and the real multiplier sits somewhere in a wide band until someone publishes per-request latency distributions at a stated concurrency level.
So Xiaomi gives you two unreconciled numbers, and OpenAI gives you one number that belongs to a different model. Neither is a fact you can budget against, and the practical consequence is the same for both: measure the tier on your own prompts before you commit a production path to it.
The asymmetry that does matter: what the extra digits buy
Speed claims aside, the two lanes are attached to very different kinds of product, and this is where the choice stops being a price comparison.
MiMo-V2.6-Pro is a sparse mixture-of-experts design that Xiaomi has described in public: 1.02 trillion total parameters with 42 billion activated per token, a 1M-token context window, native multimodality across text, image, video and audio, and an MIT licence tag on the released checkpoints. The release bundle includes a technical report and deployment notes, and Xiaomi says it is open-sourcing the reinforcement-learning training environment and its code so the post-training can be examined and reproduced. The RL run is documented too — 30 steps and roughly 750,000 trajectories per model, under six days, at a combined published spend of about $3,474,715 across the Pro and Flash runs. On DeepSWE v1.1, Xiaomi's own harness with mini-swe-agent at avg@3, the company reports MiMo-V2.6-Flash moving from 48.8 to 65.68 and MiMo-V2.6-Pro from 58.4 to 72.57 across that run. Those are vendor-reported and unreproduced; the distinction between "Xiaomi's harness says 72.57" and "72.57 is the score" is the difference between a claim and a fact.
GPT-6.1 Sol publishes none of that. There is no parameter count, no architecture note, no training-spend figure, no checkpoint, and no licence to read — there is a model card, a context window, a knowledge cutoff of April 30, 2026, and an API. That is the normal arrangement for a frontier lab and it is not a criticism. But it has a consequence for a buying decision that launch coverage tends to skip: with the open-weight lane, a disappointing fast tier is recoverable. If UltraSpeed turns out not to be worth 10x your traffic, the same checkpoint is downloadable and you can serve it yourself or through another provider, which caps your downside at the cost of the migration. With the closed lane, a disappointing fast tier is a bill and a rate-limit budget you adjust back down; you cannot route around it by running the weights, because there are no weights to run.
One caution about that MIT tag, because it gets flattened in launch coverage. The licence applies to the checkpoints Xiaomi published. UltraSpeed is a hosted service, and hosted services are governed by terms of service rather than by a weight licence. Holding the right to run the checkpoint on your own hardware and holding the right to resell somebody's accelerated serving of it are two different rights, and only the first one comes from the licence file.


Two naming traps before you copy either identifier
Neither name in the title of this page is a checkpoint, and both vendors' documentation makes the mistake easy to repeat.
• GPT-6.1 Ultrafast — not a model. The identifier in the request stays the same as the standard model's; the tier is set by a request field, and the result is the same checkpoint scheduled differently.
• MiMo v2.6 Pro Ultraspeed — also not a distinct set of weights. It is the MiMo-V2.6-Pro checkpoint served on a faster path and priced as its own line item.
• What that means for benchmarking — if you compare the two lanes of either model on identical prompts, the expected result is the same answers at different rates. Divergent content on the same input is a defect to report, not a capability you bought.
• What that means for the licence check — for Xiaomi, read the model card. For OpenAI, there is nothing to read, because the weights are not distributed.
There is also a softer asymmetry that shows up in how each tier is sold. GPT-6.1 Sol supports US and EU data residency and global processing, including under Ultrafast, so a workload pinned to European processing has a documented answer. The MiMo rate card we can read does not state residency at all — Xiaomi is selling a model and a serving path, not a compliance posture. For a regulated buyer that single line may decide the comparison before price is considered.
Where each lane is actually callable
Both fast lanes are reachable through their vendor's own API, and both appear on third-party platforms. Neither is on OrcaRouter, and it is worth saying that plainly rather than letting a comparison page imply otherwise.
What is on OrcaRouter is the standard lane of the OpenAI model: openai/gpt-6.1-sol at OpenAI's own list rates of $2.00 per million input tokens and $10.00 per million output tokens, with 0% markup and the provider's price passed straight through, so a vendor repricing lands on our side the same day. Ultrafast is a service-tier flag billed on your own OpenAI account, and no Xiaomi model is hosted here at all. What one key does buy is the standard lane alongside more than 200 other models behind one OpenAI-compatible endpoint, with automatic failover across providers — which is the useful posture when the workload you are protecting is the one that would notice a rate-limit ceiling at three in the morning.

Which lane to test, in order
If your workload is interactive and the generation time sits on a person's clock, both tiers are worth a trial and the deciding factor will not be price — it will be the expiry date on your trust. The open-weight lane lets you verify and then leave; the closed lane asks you to verify with the one vendor who can serve it and stay. That asymmetry is real, but it is also not free: MiMo-V2.6-Pro-UltraSpeed is a hosted service, so the exit is a migration rather than a config change, and the checkpoint you would migrate to may not match the tier's serving optimisations.
If your workload is batch or scheduled, both are the wrong purchase, at 6x and at 10x, and the honest move is to run it on the standard lane and spend the difference on evaluation instead.
What would change this page: an independent measurement of either tier's output speed at a stated concurrency level, a published residency statement from Xiaomi, or a rate-limit number from OpenAI for the Sol tier that a buyer can plan against instead of requesting. Until those exist, the comparison is a price list against a price list, with one of the two vendors having published considerably more about the object being priced.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
