
GPT-6.1 Ultrafast Is Pro $500 Only, and That Detail Was Three Posts Deep
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 111 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 55 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 377 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
GPT-6.1 Sol Ultrafast arrived on October 8, 2026 with the rarest thing a speed tier can ship with: a published price. $12.00 per million input tokens, $60.00 per million output tokens, exactly six times what GPT-6.1 Sol — the model OpenAI released on September 29, 2026 at its DevDay 2026 keynote — costs on the standard lane. What the announcement thread did not lead with is where a ChatGPT subscriber can actually switch it on, and the answer, three posts further down, is the $500 Pro plan. That ordering is worth more attention than it got, because the API half and the subscription half of this rollout are gated in opposite ways, and the difference decides whether Ultrafast is something you can try this week or something you have to buy a plan for.
Nothing about the model changed. GPT-6.1 Sol Ultrafast is the same checkpoint, the same 1,050,000-token context window, the same 128,000-token output ceiling, the same April 30, 2026 knowledge cutoff and the same answers as GPT-6.1 Sol. Ultrafast is a service tier — a scheduling priority you request — not a second model. The whole product is latency, sold at a multiple, and the only two questions that matter are how much of it you get and who is allowed to buy it.
What the record actually says, in the order it says it
The grievance version of this story is that OpenAI buried the gate. The accurate version is more specific, and it is checkable line by line against OpenAI's own documentation.
The developer changelog carries the entry dated October 8: adding Ultrafast mode for GPT-6.1 Sol in the Responses API, requested by sending model: "gpt-6.1-sol" with service_tier: "ultrafast", described as reducing the time between generated output tokens, and stated as available to all API users subject to rate limits, with global processing and US and EU data residency. Read on its own, that is an unrestricted launch at a published rate.
The tier documentation says Ultrafast is the fastest service tier in the API, broadly available for GPT-6 Astra and GPT-6.1 Sol with preview access for GPT-5.6 Sol, and recommends WebSockets for it. Still no mention of a plan.

The subscription detail lives on the ChatGPT-side speed page, and it is three limitations stacked in one paragraph: Ultrafast is available in Codex and ChatGPT Work, on Pro $500, and on eligible Enterprise and Edu plans. Enterprise administrators get it switched off by default and must enable it per user or per workspace. And the consumer Chat surface is not offered the tier at all — GPT-6.1 Sol itself lives in Work and Codex, so Chat was never in scope.
Put the three pages together and the shape is clear. An API key is a ticket. A ChatGPT subscription is a ticket only if it is the most expensive one OpenAI sells, or if an administrator has enabled it for you. Nobody reading the first post of the announcement would guess that.
Why the gate exists, and it is not arbitrary
The obvious reading is that OpenAI is using the fast lane to sell $500 plans. The mechanism underneath is less cynical and more useful: Ultrafast is an expensive meter, and a $500 plan is the only consumer plan where an 8x meter fits inside the included usage.
On a subscription, included usage is consumed at eight times the standard rate when you use Ultrafast, so a single Ultrafast turn draws down the same allowance as eight standard turns. That is a consequence of what the tier does — it buys wall-clock by spending serving capacity, and a subscription is a fixed-capacity product. Charge the same allowance for eight times the capacity consumption and the plan economics break at $20 a month. Charge eight times the allowance and the tier simply becomes unattractive on a small plan. Either way, the tier lands where the allowance is large enough to absorb it.
The API has no such constraint, because it is metered. There, Ultrafast gets its own rate-limit budget, separate from the Standard and Fast budgets, and OpenAI describes those limits as set per organization rather than published on the tier page. That is a second thing to watch that the announcement did not headline: switching a workload to Ultrafast does not just multiply the bill, it moves the workload onto a different ceiling, and if you raise traffic beyond what the organization was granted, the tier's own limits become the failure mode.
The arithmetic of buying speed two different ways
Here is the decision in numbers, and it is the part worth doing on your own token counts rather than mine. Take an agent session that sends 30,000 input tokens and receives 1,500 output tokens per turn, over 40 turns.
• Total tokens per run — 1,200,000 input, 60,000 output, spread over 40 requests that each sit comfortably under the 272,000-token threshold where OpenAI reprices the whole request.
• On standard GPT-6.1 Sol — 1.2 million tokens at $2.00 per million is $2.40, plus 60,000 at $10.00 per million is $0.60. Three dollars a run.
• On GPT-6.1 Sol Ultrafast — 1.2 million at $12.00 is $14.40, plus 60,000 at $60.00 is $3.60. Eighteen dollars a run.
• The difference — $15.00 per run to remove most of the generation time on a job that does not otherwise change.
That $15.00 per run is the number to hold against the subscription. If the only reason you are considering the Pro plan is Ultrafast, roughly 33 runs a month at these token counts is where the plan's $500 and the API's per-token meter cross over — below that the meter is the cheaper way to buy the same tier, and above it the plan starts to win. Your token counts will move that crossover, sometimes a lot, so run your own numbers before treating 33 as a threshold. It is our arithmetic on a stated workload, not a published break-even.

p>Two structural notes that survive any token count.
The speed claim is borrowed, and it is the one thing to read carefully
Ultrafast is sold on "up to 8x faster", and the only model-specific measurement OpenAI publishes behind that phrase belongs to a different model. The sentence in its documentation reads that GPT-6 Astra Ultrafast generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex. That is an Astra figure, measured in Codex. The documentation for GPT-6.1 Sol's tier describes it as reducing inter-token time and points at a rate card; it does not put a number on it, and no independent party has published a tokens-per-second measurement for the Sol variant either.
It may well hold. Both models run on the same serving stack and the tier mechanism is the same, so an Astra multiple is a reasonable prior for Sol. But "up to 8x" is a vendor ceiling, unreproduced, taken from a sibling model — and the "up to" is doing real work in that sentence. Treat the multiple as a hypothesis you test on your own prompts, because the tier's own documentation never tests it for you.
The same caveat is cheaper to state for the older tiers. Ultrafast over GPT-5.6 Sol was announced on August 13, 2026 at "up to 14x faster than Standard processing" in limited preview, and three months later the documentation still says preview access while the Ultrafast rate table carries exactly two rows — GPT-6 Astra and GPT-6.1 Sol. A 14x number with no published price and no general availability is a claim rather than a product, and it is a useful reminder of how much of this tier ladder is still aspirational.
If the $500 plan is not happening
For anyone who is not going to buy Pro $500 to test a latency tier, the useful move is to separate the two things the announcement bundled. The tier is available on the API to all API users, so the fast lane is purchasable without a subscription — at $12.00 and $60.00 per million tokens, with its own rate-limit budget, on a workload small enough to measure. Test it there first, on your own prompts, against the standard lane.
The standard lane is the one that exists without any of this. OrcaRouter serves GPT-6.1 Sol as openai/gpt-6.1-sol at OpenAI's own list rates — $2.00 per million input tokens and $10.00 per million output tokens — with 0% markup and the provider's price passed straight through, so a vendor repricing lands on our side the same day. Ultrafast is not something we sell: it is a service-tier flag billed on your own OpenAI account, and we would rather say that plainly than let a model page imply otherwise. What one key does buy is the ability to route the standard lane and the rest of the catalogue — more than 200 models behind one OpenAI-compatible endpoint — and to fail over automatically across providers when one is degraded, which is the cheapest insurance available if you are about to put a latency tier in front of a production path.

And if you are on Plus or Team and the answer you were looking for was "can I just turn this on", the answer is no, with a specific shape. Codex and ChatGPT Work on Pro $500 or an eligible Enterprise or Edu plan, enabled by an administrator if it is the enterprise kind. Everything else reaches Ultrafast through the API and pays per token.
What to watch next
Three things would turn this from a gated rollout into an ordinary one, and all three are observable. A published Ultrafast rate limit for GPT-6.1 Sol would end the guesswork about whether the tier's own ceiling binds before the bill does. A model-specific speed measurement — or an independent one — would replace the borrowed Astra number with something the Sol tier can be judged on. And Ultrafast reaching ChatGPT Plus would mean the capacity math changed, which is the signal that the gate was launch scarcity rather than the permanent shape of the tier.
Until then the honest summary is narrow and checkable. GPT-6.1 Sol shipped on September 29, 2026 and is unchanged. On October 8, 2026 it got a fast lane with a real price, open to API users and gated in ChatGPT behind the $500 plan — a detail that deserved the first post rather than the third, and one that decides for most readers whether this is a line item or a purchase order.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
