A generated title card with the headline 'GPT-6 Astra Ultrafast vs MiMo v2.6 Pro Ultraspeed', the subhead 'Same manoeuvre, two different deals', and two cards reading 'Price multiple: 6x vs 10x' and 'Speed claim: 8x vs 20x', with a footer reading 'Vendor-reported figures, September-October 2026'.
Guides & Insights

GPT-6 Astra Ultrafast vs Xiaomi MiMo v2.6 Pro Ultraspeed: Same Manoeuvre, Two Different Deals

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Two labs performed the same manoeuvre in the same fortnight, and the interesting part is that they performed it on opposite assumptions about what a customer is buying. GPT-6 Astra Ultrafast is the vendor's flagship sold through the Ultrafast serving tier — the model released September 3, 2026, the tier broadly available since September 29, 2026, and named on NVIDIA's October 1 post as running on Blackwell GPUs. Xiaomi MiMo v2.6 Pro Ultraspeed is Xiaomi's version, released September 22, 2026: the same 1.02-trillion-parameter checkpoint as MiMo-V2.6-Pro, served faster, at roughly ten times the standard rate.

One of them is the fastest way to call a frontier model. The other is the cheapest fast tier in existence. They are not competitors in the ordinary sense — you would have to write a very strange application to be choosing between a $300-per-million-token closed flagship and a $8.70-per-million-token open-weight model. What makes the pairing worth a page is that both are speed tiers rather than models, so comparing them isolates the one question a speed tier poses: what exactly are you paying the multiple for?

Both tiers sell the same non-thing

Neither name is a checkpoint. This is the part that launch coverage tends to blur, so it is worth being mechanical about it.

• What the name points at — GPT-6 Astra Ultrafast is GPT-6 Astra with the request field service_tier: "ultrafast" set; the model id in the request is still gpt-6-astra. Xiaomi MiMo v2.6 Pro Ultraspeed is the MiMo-V2.6-Pro checkpoint served on a faster path, priced as its own line item.

• What changes — batching, speculative decoding, hardware allocation, concurrency limits, connection handling. Everything between the weights and your HTTP request, and nothing inside the weights.

• What does not change — the answers. Same checkpoint at the same settings should produce the same content at a different rate. If two lanes of the same model diverge on identical input, that is a defect to report, not a capability you purchased.

• Why that matters for this comparison — because neither tier claims a benchmark improvement over its own standard lane, the entire purchase decision collapses into price per token against seconds saved. Which means the two deals are decided by arithmetic, not by evaluation.

There is one asymmetry in the framing, though, and it is not in the tiers. Xiaomi's UltraSpeed sits over an openly licensed model — the MiMo-V2.6 weights carry an MIT licence, per Xiaomi's own model cards, and the family was released with its RL training environment. OpenAI's Ultrafast sits over a closed flagship you can only reach through an API. Paying a speed multiple for open weights is a reversible decision; you can always serve the checkpoint yourself. Paying it for a closed flagship is not.

A screenshot of the XiaomiMiMo/MiMo-V2.6-Pro-RL model card on Hugging Face, showing 64 likes, the TextGeneration, Transformers, Safetensors, English, Chinese, conversational, custom_code, 8-bit precision and fp8 tags, an MIT licence tag, a Safetensors model size of 524B params and the model card's opening description of MiMo-V2.6-Pro-RL as the flagship checkpoint of the MiMo-V2.6 series, covering native omnimodal input and 1M-token context.

The two price multiples, and what they buy

Here is where the pair stops being symmetric, and the numbers are unusually clean on both sides because each vendor priced a multiple rather than an absolute.

• GPT-6 Astra Ultrafast — $60.00 per million input tokens, $6.00 cached input, $75.00 cache write and $300.00 output, against the standard tier's $10.00 and $50.00. Uniformly 6.00x across every column, stepping to $120.00 and $450.00 above 272,000 input tokens. The standard tier is live in our catalogue at OpenAI's list price, which is the cheapest way to reach the flagship.

• Xiaomi MiMo v2.6 Pro Ultraspeed — $4.35 input and $8.70 output per million tokens, against the standard Pro lane at $0.44 and $0.87. That is roughly 9.9x on input and exactly 10x on output.

• The speed claims attached to those multiples — OpenAI and NVIDIA both cap Astra Ultrafast at "up to 8x" faster token generation than the Astra standard mode. Xiaomi's own material claims up to 20x output speed over the standard Pro service; third-party catalogue listings describing the same model on the same day say roughly 10x. No measurement reconciling those two figures has been published.

• The ratio of multiple to speed — OpenAI's 6x price buys a claimed ceiling of 8x, so the tier is priced below its own best case. Depending on whose number you believe, Xiaomi's 10x price buys a claimed ceiling of 10x to 20x, so at the vendor's ceiling the trade is favourable and at the lower figure it is a straight wash.

That is the single most decision-relevant difference between the two, and it is narrower than the headline prices suggest. Per unit of latency removed at the vendors' own claims, Astra Ultrafast is the better of the two deals. Per token of work, Xiaomi UltraSpeed is roughly fourteen times cheaper on input and thirty-four times cheaper on output against the Ultrafast rates, and between two and six times cheaper than Astra's standard lane — but it is a different model, and a considerably less capable one, which is the trade you are actually making. Xiaomi's own model cards for the family publish architecture and training facts rather than an independent aggregate score, and no reading from the index OpenAI's flagship is measured on has been published for it at all.

A screenshot of OpenAI's API pricing page with the Ultrafast tab selected, showing the gpt-6-astra row at $60.00 input, $6.00 cached input, $75.00 cache writes and $300.00 output per 1M tokens for short-context requests, stepping to $120.00 / $12.00 / $150.00 / $450.00 above 272,000 input tokens, alongside gpt-5.6-sol at $4.00 / $0.40 / $5.00 / $20.00 short-context and $8.00 / $0.80 / $10.00 long-context, with the page's notes on the 10% regional-processing uplift and on GPT-5.6 Sol's promotional pricing running at least through November 21, 2026.

What the two vendors chose to disclose

Speed tiers are a disclosure test more than a performance test, because the thing you would need to verify the claim — a latency distribution at a stated concurrency — is entirely in the vendor's hands.

NVIDIA's October 1 post is the most specific public statement behind the Astra tier, and what it contributes is attribution rather than measurement: Blackwell GPUs, availability in the OpenAI API and to eligible ChatGPT Work and Codex users, and two OpenAI engineers describing the loop. Philippe Tillet, OpenAI's inference lead, says Astra can "turn that knowledge into high-performance kernels that make NVIDIA hardware compelling"; Uday Ruddarraju, OpenAI's chief technology officer of compute, says "we used our internal models to optimize inference on NVIDIA GPUs." Neither sentence comes with a throughput figure for the tier. The 8x stands on its own, unqualified beyond "up to."

Xiaomi's disclosure went the other direction — it published the training economics, including the reinforcement-learning environment and code, and its model cards carry architecture facts rather than serving facts. The UltraSpeed tier itself came with a 20x claim and no published latency curve. Which means that on the question a speed tier exists to answer, both vendors are equally unhelpful, and the tie there is the honest finding.

Choosing, if you genuinely have to

The two tiers do not overlap much, so the decision is less about picking a winner than about recognising which of the two purchases yours is.

Pick Astra Ultrafast if the work is agentic and the model choice is already made. A 40,000-output-token job moves from $2.00 to $12.00, and the tier is the only way to get frontier-model quality at a faster rate. OpenAI's own documentation is explicit about the operational precondition: it recommends WebSockets "especially for agentic applications that make many tool calls in quick succession," warning that "without a persistent connection, network overhead can reduce the latency gains." Turn the tier on and leave per-request HTTP underneath it and you have paid 6x for a fraction of the speedup. Also worth knowing before you wire it in: the tier supports US data residency and global processing only, with no EU or other non-US regional processing endpoints, and it launched at low initial rate limits even for accounts that would otherwise qualify for more.

Pick MiMo v2.6 Pro Ultraspeed if the model is a commodity input to a pipeline you could host yourself. The MIT licence means the standard Pro lane is not your only fallback — self-hosting is. At $4.35 and $8.70 it is cheap enough that a faster tier is worth enabling speculatively rather than justified per-task, and its 1M-token context matches the flagship's. Understand what you are buying, though: a roughly ten-times rate for a model that trails the frontier, which is the correct trade for high-volume extraction and the wrong one for anything where the answer quality gates the workflow.

Neither fast tier is on OrcaRouter, and that is worth stating rather than implying around. What is on OrcaRouter is openai/gpt-6-astra — the standard-tier flagship, at $10.00 and $50.00 per million tokens with the long-context step at $20.00 and $75.00, a 1,050,000-token context and 128,000-token maximum output, over an OpenAI-compatible endpoint. No Xiaomi model is hosted here at all, so MiMo-V2.6-Pro and its UltraSpeed lane are reachable only through Xiaomi's own API and several third-party platforms. The practical use of a 200-plus-model endpoint in this comparison is not access to the fast tiers. It is that when you want to know whether the expensive lane is earning its multiple, you can send the same prompt to a cheaper model on the same credential and the same key, in the next request, and find out. That is the cheapest experiment available and it is the one that answers the question — because no vendor has published the latency curve that would let you answer it without measuring.

A screenshot of the OrcaRouter model page for GPT-6 Astra, model id openai/gpt-6-astra, showing a 1,050,000-token context window, 128k maximum output, text, image and file input, public benchmarks attributed to OpenAI dated 2026-09-04, input price $10.00 and output price $50.00 per 1M tokens, p50 time to first token 4.69 s, a 10.00 s figure and 155.1M tokens of traffic, with an OpenAI-compatible Python code sample and the EN language toggle in the header.

The honest reading

Both vendors shipped a rate card for latency in the last three weeks, and neither shipped a measurement. That symmetry is the story, and it has a practical consequence: the only figures you can act on today are the prices, and those are unusually informative because both sides priced a clean multiple rather than a bespoke number.

OpenAI's fast tier costs six times its own standard rate and claims a ceiling of eight. Xiaomi's costs ten times its own standard rate and claims a ceiling somewhere between ten and twenty, depending on whose number you believe. Read as arithmetic on the vendors' own figures, the two are close to a wash: OpenAI's tier buys a claimed ceiling worth 1.33 units of speed per unit of price, Xiaomi's buys between 1.0 and 2.0 by the same reckoning — and the reason they can both be defensible is that the two tiers are selling to different buyers who will never seriously consider each other's option. The prices are also the only part of either deal that is auditable today, since a rate card is a published fact and a latency claim is not.