Een gegenereerde herokaart met als kop 'Twee snelle rijstroken, zestien dagen uit elkaar' en als ondertitel 'GPT-6.1 Ultrafast vs Xiaomi MiMo v2.6 Pro Ultraspeed', een linkerpaneel met '6x standaard', 'Gesloten API' en 'Snelheidsfactor ontleend aan GPT-6 Astra', en een rechterpaneel met 'Ruwweg 10x standaard', 'MIT-gewichten' en '1,02T totaal / 42B actief', boven een voettekst met 'Leveranciers- en cataloguscijfers, september-oktober 2026; geen onafhankelijke metingen van geen van beide tiers.'
Engineering & Research

GPT-6.1 Ultrafast vs Xiaomi MiMo v2.6 Pro Ultraspeed: Twee snelle rijstroken, één zegt wat het kost

Auteur

Rowan Sterling

Publicatiedatum

Nieuwste modellen · 20Bekijk alle modellen →
Benchmarks: Artificial Analysis · dagelijks bijgewerkt
Terug naar alle berichten

The two fastest ways to buy a frontier model went on sale sixteen days apart, and only one of them told you how the speed was produced. GPT-6.1 Ultrafast — the serving tier over GPT-6.1 Sol, which O​penAI released on September 29, 2026 and which gained this tier on October 8, 2026 — is a closed flagship you reach through an API, priced at six times Standard with no published account of what happens between the weights and your request. X​iaomi MiMo v2.6 Pro Ultraspeed, released September 22, 2026, is a speed tier over MiMo-V2.6-Pro, an MIT-licensed checkpoint whose reinforcement-learning environment X​iaomi also published. Both sell latency at a multiple. Only X​iaomi's multiple sits on top of something a customer can inspect, and that difference is worth more to the decision than either speed claim.

De prijs van de twee lanes is het gemakkelijkste deel, en die komt hieronder aan bod. Het moeilijke deel is dat een snelheidscategorie een belofte over serving is, en de twee leveranciers hebben die belofte op zeer uiteenlopende niveaus van informatieverstrekking gedaan — de een leende zijn belangrijkste cijfer van een zustermodel, de ander leverde samen met de prijslijst ook het trainingsrecept mee. Als je op het punt staat een productieworkload aan een snelle lane toe te vertrouwen, dan is het verschil in informatieverstrekking het risico dat je feitelijk op je neemt.

De prijzen, met beide zijden op dezelfde regel

Beide niveaus worden verkocht als een veelvoud van een standaardbaan, en de veelvouden liggen dicht genoeg bij elkaar dat de prijs niet is wat ze onderscheidt.

• Standaardroute — GPT-6.1 Sol voor $2,00 input / $10,00 output per miljoen tokens, tegenover MiMo-V2.6-Pro voor $0,44 / $0,87.

• Snelle rijstrook — GPT-6.1 Ultrafast voor $12,00 / $60,00, vergeleken met MiMo-V2.6-Pro-UltraSpeed voor $4,35 / $8,70.

• The multiple — exactly 6x on every line for the O​penAI tier, roughly 9.9x on input and exactly 10x on output for X​iaomi's.

• Gecachte invoer — GPT-6.1 Sol rekent $0,10 per miljoen op Standard en $0,60 op Ultrafast; de MiMo-tariefkaart die wij kunnen raadplegen, publiceert voor geen van beide varianten een aparte regel voor gecachte invoer.

• Lange context — GPT-6.1 Sol herprijst het hele verzoek tegen 2x de invoer- en cachetarieven en 1,5x de uitvoertarieven boven 272.000 invoertokens, waardoor Ultrafast op $24,00 / $1,20 / $30,00 / $90,00 uitkomt. MiMo-V2.6-Pro heeft een contextvenster van 1M tokens zonder een vergelijkbare gepubliceerde prijssprong.

• Absolute spreiding — de twee snelle lanes liggen 2,8x uit elkaar op input en 6,9x uit elkaar op output, wat het verschil is tussen een gesloten frontiermodel en een open-weightmodel dat zich binnen de snelheidstier voordoet in plaats van bij het standaardtarief.

The one-line summary of that block: X​iaomi charges a bigger multiple on a much cheaper model, and O​penAI charges a smaller multiple on a much more expensive one. On output tokens you pay $60.00 per million for the O​penAI lane and $8.70 for X​iaomi's. Neither vendor is offering a discount for speed; both are pricing it as a separate product, which is the first sign that in both cases the standard lane is still the default and the fast lane is an exception you justify per workload.

Wat elke leverancier u zal vertellen over de snelheid

Dit is het punt waarop de twee uitrollen niet langer symmetrisch zijn, en het is het nuttigste ding op de pagina.

O​penAI's published number for Ultrafast is that GPT-6 Astra Ultrafast generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex. That is a measurement of a different model, in a different client, phrased as a ceiling. The documentation that covers Ultrafast for GPT-6.1 Sol says the tier reduces the time between generated output tokens and points at a rate card; it does not put a number on the Sol tier at all. So the headline number attached to this rollout is borrowed from the sibling model that has been on Ultrafast since September 29, and no independent party has published a tokens-per-second measurement for either. It may well hold — same serving stack, same mechanism — but it is a vendor ceiling that O​penAI itself did not measure for this model.

X​iaomi's disclosure is different in kind, and the difference is not that it is more precise. Its own material claims UltraSpeed reaches up to 20 times the output speed of the standard Pro service at the same quality. Third-party catalogue listings describing the same model on the same day say roughly 10 times. Both numbers are in circulation, they are not the same number, and nobody has published a measurement that reconciles them. Reading that honestly: "up to 20x" is a vendor ceiling under conditions X​iaomi has not specified, "roughly 10x" is what a catalogue was willing to assert as typical, and the real multiplier sits somewhere in a wide band until someone publishes per-request latency distributions at a stated concurrency level.

So X​iaomi gives you two unreconciled numbers, and O​penAI gives you one number that belongs to a different model. Neither is a fact you can budget against, and the practical consequence is the same for both: measure the tier on your own prompts before you commit a production path to it.

De asymmetrie die er wel toe doet: wat de extra cijfers opleveren

Los van de snelheidsclaims: de twee sporen zijn gekoppeld aan heel verschillende soorten producten, en dit is het punt waarop de keuze niet langer een prijsvergelijking is.

MiMo-V2.6-Pro is a sparse mixture-of-experts design that X​iaomi has described in public: 1.02 trillion total parameters with 42 billion activated per token, a 1M-token context window, native multimodality across text, image, video and audio, and an MIT licence tag on the released checkpoints. The release bundle includes a technical report and deployment notes, and X​iaomi says it is open-sourcing the reinforcement-learning training environment and its code so the post-training can be examined and reproduced. The RL run is documented too — 30 steps and roughly 750,000 trajectories per model, under six days, at a combined published spend of about $3,474,715 across the Pro and Flash runs. On DeepSWE v1.1, X​iaomi's own harness with mini-swe-agent at avg@3, the company reports MiMo-V2.6-Flash moving from 48.8 to 65.68 and MiMo-V2.6-Pro from 58.4 to 72.57 across that run. Those are vendor-reported and unreproduced; the distinction between "X​iaomi's harness says 72.57" and "72.57 is the score" is the difference between a claim and a fact.

GPT-6.1 Sol publiceert niets daarvan. Er is geen aantal parameters, geen architectuurnotitie, geen cijfer over trainingsuitgaven, geen checkpoint en geen licentie om te lezen — er is een modelkaart, een contextvenster, een kennisafsluitdatum van 30 april 2026 en een API. Dat is de normale gang van zaken voor een frontierlab, en het is geen kritiek. Maar het heeft een gevolg voor een aankoopbeslissing, en dat gevolg wordt in de berichtgeving rond de lancering vaak overgeslagen: met het open-weight-spoor is een teleurstellende snelle tier herstelbaar. Als UltraSpeed niet 10x je verkeer waard blijkt te zijn, is hetzelfde checkpoint downloadbaar en kun je het zelf serveren of via een andere provider, wat je neerwaartse risico beperkt, tegen de kosten van de migratie. Met het gesloten spoor is een teleurstellende snelle tier een rekening en een rate-limitbudget dat je weer naar beneden bijstelt; je kunt er niet omheen routeren door de gewichten te draaien, want er zijn geen gewichten om te draaien.

One caution about that MIT tag, because it gets flattened in launch coverage. The licence applies to the checkpoints X​iaomi published. UltraSpeed is a hosted service, and hosted services are governed by terms of service rather than by a weight licence. Holding the right to run the checkpoint on your own hardware and holding the right to resell somebody's accelerated serving of it are two different rights, and only the first one comes from the licence file.

A screenshot of the XiaomiMiMo MiMo-V2.6-Pro-RL model card on Hugging Face, showing the released reinforcement-learning checkpoints, the parameter and context figures, the licence tag and the model-card sections describing the MiMo-V2.6 family.A generated disclosure card headed 'What each vendor published', with an 'OpenAI — GPT-6.1 Ultrafast' column reading 'No speed multiple for this model', 'Headline 8x figure is GPT-6 Astra's', 'No checkpoint, no parameter count' and 'No licence to read', and a 'Xiaomi — MiMo-V2.6-Pro-UltraSpeed' column reading '1.02T total / 42B active', '1M-token context', 'MIT-licensed checkpoints', 'RL environment and spend published' and 'Speed: up to 20x or roughly 10x', footnoted that all figures are vendor-reported and unreproduced.

Twee naamgevingsvalkuilen voordat je een van beide identifiers kopieert

Geen van beide namen in de titel van deze pagina is een checkpoint, en de documentatie van beide leveranciers maakt het gemakkelijk om de fout te herhalen.

• GPT-6.1 Ultrafast — geen model. De identifier in het verzoek blijft hetzelfde als die van het standaardmodel; de tier wordt ingesteld via een veld in het verzoek, en het resultaat is dezelfde checkpoint die anders wordt ingepland.

• MiMo v2.6 Pro Ultraspeed — ook geen afzonderlijke set gewichten. Het is de MiMo-V2.6-Pro-checkpoint die via een sneller pad wordt aangeboden en als een eigen regelitem wordt geprijsd.

• Wat dat betekent voor benchmarking — als je de twee lanes van elk model op identieke prompts vergelijkt, is het verwachte resultaat dezelfde antwoorden met verschillende snelheden. Afwijkende inhoud bij dezelfde input is een defect om te melden, niet een capaciteit die je hebt gekocht.

• What that means for the licence check — for X​iaomi, read the model card. For O​penAI, there is nothing to read, because the weights are not distributed.

There is also a softer asymmetry that shows up in how each tier is sold. GPT-6.1 Sol supports US and EU data residency and global processing, including under Ultrafast, so a workload pinned to European processing has a documented answer. The MiMo rate card we can read does not state residency at all — X​iaomi is selling a model and a serving path, not a compliance posture. For a regulated buyer that single line may decide the comparison before price is considered.

Waar elke lane daadwerkelijk aanroepbaar is

Beide snelle routes zijn bereikbaar via de eigen API van hun leverancier, en beide komen voor op platforms van derden. Geen van beide staat op OrcaRouter, en het is de moeite waard om dat ronduit te zeggen in plaats van een vergelijkingspagina de indruk te laten wekken dat het anders is.

What is on OrcaRouter is the standard lane of the O​penAI model: openai/gpt-6.1-sol at O​penAI's own list rates of $2.00 per million input tokens and $10.00 per million output tokens, with 0% markup and the provider's price passed straight through, so a vendor repricing lands on our side the same day. Ultrafast is a service-tier flag billed on your own O​penAI account, and no X​iaomi model is hosted here at all. What one key does buy is the standard lane alongside more than 200 other models behind one O​penAI-compatible endpoint, with automatic failover across providers — which is the useful posture when the workload you are protecting is the one that would notice a rate-limit ceiling at three in the morning.

A screenshot of the OrcaRouter model page for GPT-6.1 Sol, model id openai/gpt-6.1-sol, showing a 1,050,000-token context window, 128,000 maximum output tokens, text and image input with text output, input price $2.00 and output price $10.00 per 1M tokens, with the EN language toggle visible in the page header.

Welke baan te testen, in volgorde

Als je workload interactief is en de generatietijd op iemands klok drukt, zijn beide tiers een proef waard, en de beslissende factor zal niet de prijs zijn — het is de houdbaarheidsdatum van je vertrouwen. De open-weight-route laat je verifiëren en dan vertrekken; de gesloten route vraagt je te verifiëren bij de enige leverancier die het kan leveren, en te blijven. Die asymmetrie is reëel, maar ze is ook niet gratis: MiMo-V2.6-Pro-UltraSpeed is een gehoste dienst, dus de exit is een migratie in plaats van een configuratiewijziging, en de checkpoint waarnaar je zou migreren komt mogelijk niet overeen met de serving-optimalisaties van de tier.

Als je workload batch of gepland is, zijn beide de verkeerde aankoop, tegen 6x en tegen 10x, en de eerlijke zet is om het op de standaardbaan te draaien en het verschil in plaats daarvan aan evaluatie te besteden.

What would change this page: an independent measurement of either tier's output speed at a stated concurrency level, a published residency statement from X​iaomi, or a rate-limit number from O​penAI for the Sol tier that a buyer can plan against instead of requesting. Until those exist, the comparison is a price list against a price list, with one of the two vendors having published considerably more about the object being priced.

Vergeleken in dit artikel1

Herkend uit dit artikel · Benchmarks: Artificial Analysis · dagelijks bijgewerkt