
GPT-5.6 Sol Ultrafast vs Xiaomi MiMo v2.6 Pro Ultraspeed: Two Speed Tiers, One Missing Price
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 661 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 188 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1306 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 225 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
Two vendors shipped the same product idea three weeks apart, and the difference between them is the best way to see what each is really selling. The Xiaomi side landed on September 22, 2026: Xiaomi MiMo v2.6 Pro Ultraspeed is the same 1.02-trillion-parameter checkpoint as MiMo-V2.6-Pro, served faster, at $4.35 per million input tokens and $8.70 per million output against the standard Pro lane's $0.44 and $0.87. The vendor's GPT-5.6 Sol Ultrafast is the same manoeuvre applied to GPT-5.6 Sol: identical weights, up to 750 output tokens per second, up to 14× the standard tier, running on Cerebras wafer-scale hardware — and no published price at all, six weeks after the mode was announced on August 13.
Put the two side by side and the anomaly is immediate. Xiaomi published a rate card before it published most of the benchmarks. OpenAI published the benchmarks, the hardware partner, and the customer names — and still has not published a number you can budget against. One of these tiers can be purchased today. The other is a value in an API schema. Comparing them is therefore less about which is faster, and more about what a speed tier is worth when one vendor has told you the price and the other has not.
What each one actually is
Neither is a new model. That is the shared premise and it is worth stating flatly, because both launches attracted coverage that treated a serving change as a capability release.
MiMo v2.6 Pro Ultraspeed is a serving configuration of Xiaomi's flagship: same 1.02T-parameter sparse mixture-of-experts checkpoint with 42B activated per token, same 1M-token context, same native multimodality across text, image, video and audio, same MIT-licensed weights on the public repositories. Xiaomi's own material claims up to 20 times the output speed of the standard Pro service at the same quality; third-party catalogue listings describing the same model on the same day say roughly 10 times. Nobody has published a measurement reconciling the two.
GPT-5.6 Sol Ultrafast is a serving configuration of OpenAI's flagship: same checkpoint, same reasoning behaviour, same 1,050,000-token context window, same 128K maximum output, same answers. The acceleration is hardware rather than scheduling — the model's weights sit in on-chip SRAM on Cerebras wafer-scale silicon instead of being shuttled between GPU memory and compute — and it is the first product of the January 2026 OpenAI–Cerebras compute partnership, reported at roughly $10 billion over three years.
So the honest axis of comparison is not intelligence. It is what the vendor chose to change, what it charged for the change, and how much of the change it let you verify.

The rate cards, including the blank one
• MiMo v2.6 Pro Ultraspeed — $4.35 input / $8.70 output per million tokens, published by Xiaomi.
• MiMo-V2.6-Pro standard — $0.44 / $0.87, published by Xiaomi's commercial catalogues. Roughly 9.9× on input and exactly 10× on output.
• GPT-5.6 Sol Ultrafast — no published rate card as of September 27, 2026. OpenAI's pricing page still carries four tabs — Standard, Batch, Flex, Fast — and the tier is not among them.
• GPT-5.6 Sol standard — $4.00 / $20.00 per million, OpenAI's current promotional rate, with cached input at $0.40 and a long-context step to $8.00 / $30.00 past roughly 272K tokens of input.
• The only comparable premium OpenAI has ever priced — Fast mode, at $8.00 / $40.00, exactly double the standard rate for up to about 2.5× output speed.
That last row is the whole reason the blank matters. Xiaomi's speed tier costs ten times the standard lane and claims ten to twenty times the speed, which is at least internally coherent as a trade: you pay proportionally more for proportionally less waiting, and the bet is that wall-clock time is the product. OpenAI's speed tier is a bigger claim on scarcer silicon than Fast mode, so whatever it eventually costs, the precedent says a premium rather than a discount. But "the precedent says a premium" is not a number, and until one exists the tier cannot be compared to UltraSpeed on price at all — only on the fact that Xiaomi was willing to say its number and OpenAI was not.
Where the numbers came from
The two tiers have opposite evidentiary problems, and it is worth naming both rather than splitting the difference.
Xiaomi's speed claim is vague in magnitude but anchored to a published price. The 20× is a vendor ceiling under conditions Xiaomi has not specified, the 10× is what a catalogue listing was willing to assert as typical, and the real multiplier is somewhere in a wide band that nobody has measured in public at a stated concurrency level. Against that, the price is exact, the weights are downloadable under MIT, and the technical report and reinforcement-learning training environment are public — you can inspect what Xiaomi did, you simply cannot yet verify how fast it runs.
OpenAI's speed claim is specific in magnitude but anchored to nothing purchasable. Up to 750 output tokens per second, up to 14× standard, stated plainly. It is also a throughput figure for output tokens, not a flat end-to-end multiplier: input processing and the model's own reasoning do not compress at that ratio, which is why the company labels it a maximum. The supporting benchmarks are vendor-reported and in one case cross-vendor — a 2,500-question Humanity's Last Exam run reported as finishing in 11 hours 11 minutes against 78 hours 27 minutes for Claude Fable 5, with comparable accuracy. Independent coverage of the launch quoted a narrower roughly 11× generation-speed comparison on the same run, while Cerebras's own figures imply about 7× for total test time. Three parties, three fractions of one workload, no third-party reproduction of any of them. Cerebras separately reports 5.6× end-to-end on GDP-Val at no measurable quality loss, also unreproduced.
Both sets of numbers are vendor-stated. The difference is that Xiaomi's vagueness is about speed and OpenAI's is about cost, and for anyone writing a budget, cost is the harder one to guess.

What happened this week, and why it does not close the gap
The reason this comparison has a news peg at all is a commit rather than an announcement. The ultrafast value has been in OpenAI's public openai-openapi schema since August 13, 2026 — the same day the mode was announced — carrying the description that scopes it to gpt-5.6-sol and calls it access-controlled. The commit that moved this week is later and smaller: on September 25, 2026, fe4f7a1 added ultrafast to two further enumerations, the request-level service-tier parameter and the agent service-tier policy field. Before that commit those two lists read auto, default, flex, priority, fast. After it, there is an entry described as "Uses the ultrafast service tier."ultrafast to the service_tier enumerations in its public openai-openapi specification — the field the schema describes as "the service tier used for model requests," and the policy field attached to agents. Before the change that list read default, flex, priority, fast. After it there is an entry described as "Uses the ultrafast service tier."
The distinction matters for anyone reading the coverage. This is not the tier being created; it is the tier being wired into the field an agent is configured with, which means a client can eventually be written against it without any new endpoint. Reporting the following day describes a Fast / Standard / Ultrafast selector coming to the Responses API Playground, with wider access expected after OpenAI's DevDay conference. We have not seen the selector and OpenAI has not published a rollout note, so that part is single-source. What is checkable is the added schema entries and the continued absence of a rate card six weeks after the announcement.
Meanwhile Xiaomi's tier has been purchasable for five days and is still the only one of the two you can put on an invoice. Its own open question is different: whether a 10× premium holds, or whether it behaves like most launch pricing on a new premium tier and drifts toward the standard lane once the launch coverage stops. Standard Pro at $0.44 / $0.87 is the anchor, and a tier that opens at ten times its anchor rarely finishes there.
Choosing between them, given that you probably cannot
The uncomfortable practical fact is that for most readers this is a comparison between one tier you cannot buy and one you probably should not buy at list. GPT-5.6 Sol Ultrafast is a waitlisted preview for select customers, expanding as capacity grows, with no general-availability date and no price; MiMo v2.6 Pro Ultraspeed is available, expensive, and built for a specific kind of buyer.
The workload test is the same for both, and it is about the shape of your request graph rather than the size of your prompt. A single long generation gains far less than the headline multiplier, because input processing and reasoning do not speed up at the same ratio. A forty-step agent loop behind one user-visible action gains something much closer to the full number, because every round trip is latency a person is sitting through. If your traffic looks like the second shape, both tiers are aimed at you; if it looks like the first, the standard lane on either model was always the correct call and the standard MiMo-V2.6-Pro lane at $0.44 / $0.87 is one of the cheapest ways to call a trillion-parameter-class model anywhere.
What neither tier changes is that you are buying latency, not quality. If you benchmark UltraSpeed or Ultrafast against its own standard lane on identical prompts and get different answers, that is a finding worth reporting, not a feature — nothing in either vendor's description claims the weights changed.
This is also the case where a router earns its place, and specifically for the tier you cannot get. GPT-5.6 Sol is live on OrcaRouter as openai/gpt-5.6-sol at the provider's own rate with zero markup on tokens over the OpenAI-compatible endpoint at api.orcarouter.ai/v1, so the standard lane you would fall back to is one base-URL change away, and the same key carries 200+ models. That matters because the honest answer to "should I use Ultrafast" today is "you cannot, so decide what carries the interactive traffic in the meantime" — and the way to buy latency without a waitlist is to route the interactive calls to whichever model on the catalogue clears your quality bar with the lowest wall-clock time, and leave the overnight work on the cheapest lane that passes. When the tier opens, that is the switch you flip. Note plainly what we do not host: no Xiaomi model is on OrcaRouter, so MiMo v2.6 Pro Ultraspeed comes from Xiaomi's own API and several third-party platforms, and this page is not a route to it.

What would settle it
Three measurements, none of which exists yet. A published Ultrafast rate card would make this the comparison it deserves to be rather than a comparison with a hole in it. An independent latency distribution at a stated concurrency level would tell you whether Xiaomi's real multiplier sits nearer 10× or 20×, and whether it holds under load rather than in a single-stream demonstration. And an independent reproduction of either vendor's benchmark suite would move both sets of figures out of the claims column.
Until then the defensible reading is narrow. Xiaomi shipped a speed tier with a price and vague speed; OpenAI shipped a speed tier with specific speed and no price; both are the same model as the thing they are priced against, and neither has been independently measured. If you need the latency now and the workload is interactive, UltraSpeed is real, purchasable, and priced for exactly the buyer it wants. If you need it now and the workload is not, the standard lanes on both models are cheaper, proven, and available today — and the tier everyone is arguing about will still be there when someone publishes a number.
