
GPT-6.1 Sol Ultrafast Rolls Out to the API, Codex and ChatGPT Work
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 111 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 55 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 377 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
OpenAI is rolling out Ultrafast mode for GPT-6.1 Sol — the fastest service tier it sells, and the first time the near-Astra Sol tier has had one. The move puts GPT-6.1 Sol in the same bracket as GPT-6 Astra Ultrafast in the API, in Codex, and in ChatGPT Work, and it arrives with a published rate card rather than a waitlist: $12.00 per million input tokens and $60.00 per million output tokens, exactly six times what standard GPT-6.1 Sol costs. The speed claim on the tin is "up to 8x". The interesting part is what that phrase is measuring, and the answer is not the model you are about to pay for.
The developer changelog carries the entry dated October 8: "Added Ultrafast mode for GPT-6.1 Sol in the Responses API. Use gpt-6.1-sol with service_tier: "ultrafast" to reduce the time between generated output tokens. It is available to all API users, subject to rate limits, with global processing and US and EU data residency." The Ultrafast documentation page now reads "broadly available for GPT-6 Astra and GPT-6.1 Sol, with preview access for GPT-5.6 Sol", and the ChatGPT-side speed page confirms the subscription half: Ultrafast is available in Codex and ChatGPT Work on Pro $500 and on eligible Enterprise and Edu plans.
Ultrafast is a service tier, not a second model
Nothing about the weights changes. Same checkpoint, same 1,050,000-token context window, same 922,000-token maximum input, same 128,000-token output ceiling, same April 30, 2026 knowledge cutoff, same answers. What you are buying is a scheduling priority: the model generates tokens faster, and OpenAI's own framing is deliberately narrow — the tier "reduces the time between generated output tokens". It does not make the model smarter, and it does not make the thinking phase shorter.
That distinction is the whole decision. For an agent loop that makes forty tool calls in a row, the win compounds because every turn's output arrives sooner. For a single hard reasoning request that spends most of its wall-clock deliberating, you can pay six times as much and watch the same spinner.
The tier ladder, once, in plain terms:
• Batch and Flex — half of Standard, the cheap lane for work nobody is waiting on
• Standard — $2.00 input / $0.10 cached / $2.50 cache write / $10.00 output per million tokens
• Fast — twice Standard, or $4.00 / $0.20 / $5.00 / $20.00; OpenAI renamed Priority processing to Fast mode on July 30, 2026
• Ultrafast — six times Standard, or $12.00 / $0.60 / $15.00 / $60.00
• Above 272K input tokens — the whole request is repriced at 2x input and cache rates and 1.5x output; Ultrafast long-context is $24.00 / $1.20 / $30.00 / $90.00
The arithmetic nobody puts on the marketing page
Six times the price for eight times the speed is not a loss-making trade, and it is not a free lunch either. Ultrafast is billed per token, so a job that costs $1.00 at Standard costs $6.00 at Ultrafast — and finishes in roughly one-eighth of the time. The saving is not in the bill, it is in the wall-clock. You are paying five extra dollars to remove seven-eighths of the queue time on that job, and whether that is cheap or absurd depends entirely on what the person or the pipeline waiting on the other end is worth per minute.
Two implications follow, and they are the ones that get missed:
• Interactive work is the easy yes. A developer watching a diff land, an agent that cannot proceed until it has read the result — those are the cases where the latency is the product. Eight times faster output on a forty-turn loop is not a marginal improvement to a human perception; it is the difference between a tool you use and a tool you background.
• Scheduled and batch work is the easy no. If a job has until morning anyway, Ultrafast is a 6x bill for nothing, and the half-price Batch lane is sitting right there.
Cached input is worth a second look because it moves too: $0.60 per million on Ultrafast against $0.10 on Standard, still 5% of the uncached input rate. Prompt-cached agents that resend a large system prompt every turn will find the cache discount scales with the tier rather than shielding them from it.

Where the "up to 8x" number comes from
This is the part to read carefully, because it is where the rollout's headline and the rollout's documentation quietly describe different things.
The only model-specific speed measurement OpenAI publishes is this sentence: "GPT-6 Astra Ultrafast generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex." That is an Astra figure, measured in Codex. There is no equivalent published multiple for GPT-6.1 Sol. The document that covers Ultrafast for this model says it reduces the time between generated output tokens and points at a pricing table; it does not put a number on it. The ChatGPT-side speed page repeats the Astra sentence and stops.
So the honest reading of "up to 8x faster" is: it is the tier's headline, sourced from the sibling model that has been on Ultrafast since September 29, and it is a vendor claim that no independent party has reproduced for either model. It may well hold for GPT-6.1 Sol — the two run on the same serving stack and the tier's mechanism is the same — but nobody outside OpenAI has published a token-per-second measurement for the Sol variant, and the "up to" is doing real work in that sentence.
It is also worth keeping the tiers separated by model, because the older one is still unfinished business. Ultrafast was announced on August 13, 2026 for GPT-5.6 Sol at "up to 14x faster than Standard processing", in limited preview to select customers. Three months later, the docs still say GPT-5.6 Sol has "preview access" and the Ultrafast pricing table contains exactly two rows — GPT-6 Astra and GPT-6.1 Sol. A 14x number that has never reached general availability and has no published rate is a claim, not a product.

Who can actually switch it on
The API side is broad; the subscription side is narrow, and the difference catches people out.
• API — available to all API users on gpt-6.1-sol, subject to separate rate limits from Standard and Fast. That means a second budget to watch, not just a second price.
• Ultrafast rate limits for GPT-6.1 Sol — 1,000,000 tokens per minute on Build, 4,000,000 on Launch, 40,000,000 on Grow. OpenAI simplified its usage tiers from five to three on October 6, so those three names are the current ladder, not a legacy one.
• Data residency — GPT-6.1 Sol supports US and EU data residency and global processing, including under Fast and Ultrafast. Astra's Ultrafast does not get EU residency, so if your workload is EU-pinned this is the first Ultrafast configuration you can buy at all.
• Codex and ChatGPT Work — Ultrafast is available on the Pro $500 plan and on eligible Enterprise and Edu plans. Enterprise administrators get it switched off by default and have to enable it per user or per workspace.
• Chat — not included. GPT-6.1 Sol itself lives in Work and Codex; the consumer Chat surface is not part of this.
• How it is billed on a subscription — included usage is consumed at 8x the Standard rate, and purchased credits or Enterprise pay-as-you-go are billed at 6x the Standard rate. Both numbers are worth knowing before you leave it toggled on.
One implementation detail decides whether you see any of the speed at all. OpenAI's guidance is blunt about it: use WebSockets, "especially for agentic applications that make many tool calls in quick succession", because "without a persistent connection, network overhead can reduce the latency gains". If your client opens a fresh HTTP connection per call, per-request overhead can eat the entire advantage of the tier you just paid 6x for. HTTP still works. It just may not be worth the tier.
What we route, and what we do not
Standard-tier GPT-6.1 Sol is on OrcaRouter as openai/gpt-6.1-sol at the provider's own $2.00 input and $10.00 output per million tokens, served at 0% markup — the vendor's list price passed through, so when OpenAI moves a rate the change lands in your bill the same day it lands on the rate card rather than at your next contract renewal. The same key and endpoint carry more than 200 models, so a GPT-6.1 Sol call and a Claude or Qwen call sit behind one integration with automatic failover and no second contract, and the routing DSL will compose models into a single call when a task needs more than one opinion.
Ultrafast is not one of those models, and it would be misleading to imply otherwise. It is a service-tier flag on an OpenAI account — you send service_tier: "ultrafast" to gpt-6.1-sol and OpenAI bills your own account at the rates above — so it lives inside your direct OpenAI integration. If you turn it on, keep the tiered traffic on your vendor key and route the standard-volume traffic through us. That is the honest split, and it is also the cheaper one for anything that is not waiting on a human.

What to do with it today
If you have a Codex or ChatGPT Work seat on Pro $500, the toggle is already there and the trial costs you nothing but the usage multiplier — run the same task both ways and see whether the loop you actually run has enough turns for the eight-fold to show up. If you are on the API, the experiment worth running is narrower than the announcement suggests: instrument how much of your latency is token generation and how much is the model thinking, then point Ultrafast at the part it can actually compress.
And if you have a choice between Fast at 2x and Ultrafast at 6x for the same request, Fast is the tier that got here first and the one whose behaviour you can reason about from a rate card alone. Ultrafast is new for this model, unpriced by anyone's independent measurement, and worth exactly as much as your own stopwatch says it is.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
