A generated hero card headed 'Two fast lanes, sixteen days apart' with the subtitle 'GPT-6.1 Ultrafast vs Xiaomi MiMo v2.6 Pro Ultraspeed', a left panel reading '6x standard', 'Closed API' and 'Speed multiple borrowed from GPT-6 Astra', and a right panel reading 'Roughly 10x standard', 'MIT weights' and '1.02T total / 42B active', over a footer reading 'Vendor and catalogue figures, September-October 2026; no independent measurements of either tier.'
Engineering & Research

GPT-6.1 Ultrafast vs Xiaomi MiMo v2.6 Pro Ultraspeed: Two Fast Lanes, One Says What It Cost

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The two fastest ways to buy a frontier model went on sale sixteen days apart, and only one of them told you how the speed was produced. GPT-6.1 Ultrafast — the serving tier over GPT-6.1 Sol, which O​penAI released on September 29, 2026 and which gained this tier on October 8, 2026 — is a closed flagship you reach through an API, priced at six times Standard with no published account of what happens between the weights and your request. X​iaomi MiMo v2.6 Pro Ultraspeed, released September 22, 2026, is a speed tier over MiMo-V2.6-Pro, an MIT-licensed checkpoint whose reinforcement-learning environment X​iaomi also published. Both sell latency at a multiple. Only X​iaomi's multiple sits on top of something a customer can inspect, and that difference is worth more to the decision than either speed claim.

The price of the two lanes is the easy part and it is covered below. The hard part is that a speed tier is a promise about serving, and the two vendors have made that promise at very different levels of disclosure — one borrowed its headline number from a sibling model, the other shipped the training recipe alongside the price list. If you are about to commit a production workload to a fast lane, the disclosure gap is the risk you are actually taking on.

The prices, with both sides on the same line

Both tiers are sold as a multiple of a standard lane, and the multiples are close enough that price is not what separates them.

• Standard lane — GPT-6.1 Sol at $2.00 input / $10.00 output per million tokens, against MiMo-V2.6-Pro at $0.44 / $0.87.

• Fast lane — GPT-6.1 Ultrafast at $12.00 / $60.00, against MiMo-V2.6-Pro-UltraSpeed at $4.35 / $8.70.

• The multiple — exactly 6x on every line for the O​penAI tier, roughly 9.9x on input and exactly 10x on output for X​iaomi's.

• Cached input — GPT-6.1 Sol reads $0.10 per million on Standard and $0.60 on Ultrafast; the MiMo rate card we can read does not publish a separate cached-input line for either lane.

• Long context — GPT-6.1 Sol reprices the whole request at 2x input and cache rates and 1.5x output above 272,000 input tokens, putting Ultrafast at $24.00 / $1.20 / $30.00 / $90.00. MiMo-V2.6-Pro carries a 1M-token context window with no equivalent published cliff.

• Absolute spread — the two fast lanes are 2.8x apart on input and 6.9x apart on output, which is the gap between a frontier closed model and an open-weight one showing up inside the speed tier rather than at the standard rate.

The one-line summary of that block: X​iaomi charges a bigger multiple on a much cheaper model, and O​penAI charges a smaller multiple on a much more expensive one. On output tokens you pay $60.00 per million for the O​penAI lane and $8.70 for X​iaomi's. Neither vendor is offering a discount for speed; both are pricing it as a separate product, which is the first sign that in both cases the standard lane is still the default and the fast lane is an exception you justify per workload.

What each vendor will tell you about the speed

This is where the two rollouts stop being symmetrical, and it is the most useful thing on the page.

O​penAI's published number for Ultrafast is that GPT-6 Astra Ultrafast generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex. That is a measurement of a different model, in a different client, phrased as a ceiling. The documentation that covers Ultrafast for GPT-6.1 Sol says the tier reduces the time between generated output tokens and points at a rate card; it does not put a number on the Sol tier at all. So the headline number attached to this rollout is borrowed from the sibling model that has been on Ultrafast since September 29, and no independent party has published a tokens-per-second measurement for either. It may well hold — same serving stack, same mechanism — but it is a vendor ceiling that O​penAI itself did not measure for this model.

X​iaomi's disclosure is different in kind, and the difference is not that it is more precise. Its own material claims UltraSpeed reaches up to 20 times the output speed of the standard Pro service at the same quality. Third-party catalogue listings describing the same model on the same day say roughly 10 times. Both numbers are in circulation, they are not the same number, and nobody has published a measurement that reconciles them. Reading that honestly: "up to 20x" is a vendor ceiling under conditions X​iaomi has not specified, "roughly 10x" is what a catalogue was willing to assert as typical, and the real multiplier sits somewhere in a wide band until someone publishes per-request latency distributions at a stated concurrency level.

So X​iaomi gives you two unreconciled numbers, and O​penAI gives you one number that belongs to a different model. Neither is a fact you can budget against, and the practical consequence is the same for both: measure the tier on your own prompts before you commit a production path to it.

The asymmetry that does matter: what the extra digits buy

Speed claims aside, the two lanes are attached to very different kinds of product, and this is where the choice stops being a price comparison.

MiMo-V2.6-Pro is a sparse mixture-of-experts design that X​iaomi has described in public: 1.02 trillion total parameters with 42 billion activated per token, a 1M-token context window, native multimodality across text, image, video and audio, and an MIT licence tag on the released checkpoints. The release bundle includes a technical report and deployment notes, and X​iaomi says it is open-sourcing the reinforcement-learning training environment and its code so the post-training can be examined and reproduced. The RL run is documented too — 30 steps and roughly 750,000 trajectories per model, under six days, at a combined published spend of about $3,474,715 across the Pro and Flash runs. On DeepSWE v1.1, X​iaomi's own harness with mini-swe-agent at avg@3, the company reports MiMo-V2.6-Flash moving from 48.8 to 65.68 and MiMo-V2.6-Pro from 58.4 to 72.57 across that run. Those are vendor-reported and unreproduced; the distinction between "X​iaomi's harness says 72.57" and "72.57 is the score" is the difference between a claim and a fact.

GPT-6.1 Sol publishes none of that. There is no parameter count, no architecture note, no training-spend figure, no checkpoint, and no licence to read — there is a model card, a context window, a knowledge cutoff of April 30, 2026, and an API. That is the normal arrangement for a frontier lab and it is not a criticism. But it has a consequence for a buying decision that launch coverage tends to skip: with the open-weight lane, a disappointing fast tier is recoverable. If UltraSpeed turns out not to be worth 10x your traffic, the same checkpoint is downloadable and you can serve it yourself or through another provider, which caps your downside at the cost of the migration. With the closed lane, a disappointing fast tier is a bill and a rate-limit budget you adjust back down; you cannot route around it by running the weights, because there are no weights to run.

One caution about that MIT tag, because it gets flattened in launch coverage. The licence applies to the checkpoints X​iaomi published. UltraSpeed is a hosted service, and hosted services are governed by terms of service rather than by a weight licence. Holding the right to run the checkpoint on your own hardware and holding the right to resell somebody's accelerated serving of it are two different rights, and only the first one comes from the licence file.

A screenshot of the XiaomiMiMo MiMo-V2.6-Pro-RL model card on Hugging Face, showing the released reinforcement-learning checkpoints, the parameter and context figures, the licence tag and the model-card sections describing the MiMo-V2.6 family.A generated disclosure card headed 'What each vendor published', with an 'OpenAI — GPT-6.1 Ultrafast' column reading 'No speed multiple for this model', 'Headline 8x figure is GPT-6 Astra's', 'No checkpoint, no parameter count' and 'No licence to read', and a 'Xiaomi — MiMo-V2.6-Pro-UltraSpeed' column reading '1.02T total / 42B active', '1M-token context', 'MIT-licensed checkpoints', 'RL environment and spend published' and 'Speed: up to 20x or roughly 10x', footnoted that all figures are vendor-reported and unreproduced.

Two naming traps before you copy either identifier

Neither name in the title of this page is a checkpoint, and both vendors' documentation makes the mistake easy to repeat.

• GPT-6.1 Ultrafast — not a model. The identifier in the request stays the same as the standard model's; the tier is set by a request field, and the result is the same checkpoint scheduled differently.

• MiMo v2.6 Pro Ultraspeed — also not a distinct set of weights. It is the MiMo-V2.6-Pro checkpoint served on a faster path and priced as its own line item.

• What that means for benchmarking — if you compare the two lanes of either model on identical prompts, the expected result is the same answers at different rates. Divergent content on the same input is a defect to report, not a capability you bought.

• What that means for the licence check — for X​iaomi, read the model card. For O​penAI, there is nothing to read, because the weights are not distributed.

There is also a softer asymmetry that shows up in how each tier is sold. GPT-6.1 Sol supports US and EU data residency and global processing, including under Ultrafast, so a workload pinned to European processing has a documented answer. The MiMo rate card we can read does not state residency at all — X​iaomi is selling a model and a serving path, not a compliance posture. For a regulated buyer that single line may decide the comparison before price is considered.

Where each lane is actually callable

Both fast lanes are reachable through their vendor's own API, and both appear on third-party platforms. Neither is on OrcaRouter, and it is worth saying that plainly rather than letting a comparison page imply otherwise.

What is on OrcaRouter is the standard lane of the O​penAI model: openai/gpt-6.1-sol at O​penAI's own list rates of $2.00 per million input tokens and $10.00 per million output tokens, with 0% markup and the provider's price passed straight through, so a vendor repricing lands on our side the same day. Ultrafast is a service-tier flag billed on your own O​penAI account, and no X​iaomi model is hosted here at all. What one key does buy is the standard lane alongside more than 200 other models behind one O​penAI-compatible endpoint, with automatic failover across providers — which is the useful posture when the workload you are protecting is the one that would notice a rate-limit ceiling at three in the morning.

A screenshot of the OrcaRouter model page for GPT-6.1 Sol, model id openai/gpt-6.1-sol, showing a 1,050,000-token context window, 128,000 maximum output tokens, text and image input with text output, input price $2.00 and output price $10.00 per 1M tokens, with the EN language toggle visible in the page header.

Which lane to test, in order

If your workload is interactive and the generation time sits on a person's clock, both tiers are worth a trial and the deciding factor will not be price — it will be the expiry date on your trust. The open-weight lane lets you verify and then leave; the closed lane asks you to verify with the one vendor who can serve it and stay. That asymmetry is real, but it is also not free: MiMo-V2.6-Pro-UltraSpeed is a hosted service, so the exit is a migration rather than a config change, and the checkpoint you would migrate to may not match the tier's serving optimisations.

If your workload is batch or scheduled, both are the wrong purchase, at 6x and at 10x, and the honest move is to run it on the standard lane and spend the difference on evaluation instead.

What would change this page: an independent measurement of either tier's output speed at a stated concurrency level, a published residency statement from X​iaomi, or a rate-limit number from O​penAI for the Sol tier that a buyer can plan against instead of requesting. Until those exist, the comparison is a price list against a price list, with one of the two vendors having published considerably more about the object being priced.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily