
GPT-6.1 Ultrafast vs GPT-6.1 Sol: Three Jobs, One Verdict Each
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 111 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 55 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 377 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Ask whether GPT-6.1 Ultrafast is worth six times the price of GPT-6.1 Sol and the honest answer is that two of your three workloads should stay exactly where they are. The interactive coding agent should move. The overnight evaluation sweep should not, and neither should the batch summariser, and the reason is not that the tier is overpriced — it is that they were never buying the thing it sells. GPT-6.1 Sol shipped on September 29, 2026 and Ultrafast arrived as a mode over it on October 8, 2026: same checkpoint, same 1,050,000-token context window, same 128,000-token output ceiling, same April 30, 2026 knowledge cutoff, same answers, at $12.00 per million input tokens and $60.00 per million output against $2.00 and $10.00. A speed tier is a purchase of wall-clock, and wall-clock is only worth money to a job that someone is waiting on.
That reframing is the whole article. The rest is the arithmetic that tells you which of your jobs is the waiting kind, and the two costs that arrive alongside the multiple whether you planned for them or not.
What actually differs: one request field and one bill
There is no Ultrafast checkpoint, no separate context limit, no alternate knowledge cutoff, and nothing to pin. You send the same model identifier with a service-tier field set, and OpenAI schedules the request differently; send it without the field and you are on Standard. Everything a developer codes against is identical, and it is worth being mechanical about the list, because the length of the list is the argument.
• Model id — the same identifier on both, one default snapshot, nothing dated to pin.
• Context — a 1,050,000-token window, 922,000-token maximum input and 128,000-token maximum output on both.
• Reasoning ladder — low, medium (default), high, xhigh and max on both, with none and minimal unsupported on both.
• Tool surface — web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search, over the Responses API, on both.
• Price — $2.00 input / $0.10 cached / $2.50 cache write / $10.00 output per million tokens on Standard, against $12.00 / $0.60 / $15.00 / $60.00 on Ultrafast.
• Above 272,000 input tokens — the whole request reprices at 2x input and cache rates and 1.5x output on both, which puts Ultrafast long-context at $24.00 / $1.20 / $30.00 / $90.00.
• Speed — Standard is the baseline by definition; Ultrafast is the top rung, and the only model-specific multiple OpenAI publishes belongs to GPT-6 Astra, not to this model.
That last line is the caveat that runs under everything else, and it gets its own section further down. First, the jobs.
Job one: the interactive agent. This is the one that moves
An agent loop that makes forty tool calls in a row is the buyer this tier was built for, because every turn's generation time gates the next turn, and the user is watching. Take a session that sends 30,000 input tokens and receives 1,500 output tokens per turn across 40 turns — 1,200,000 input tokens and 60,000 output tokens in total, every request well under the 272,000-token repricing threshold.
• Standard — 1.2 million input tokens at $2.00 is $2.40; 60,000 output tokens at $10.00 is $0.60. Three dollars for the run.
• Ultrafast — 1.2 million input tokens at $12.00 is $14.40; 60,000 output tokens at $60.00 is $3.60. Eighteen dollars for the run.
• The delta — $15.00 to remove most of the generation latency from a job whose output is otherwise identical.
Whether $15.00 is cheap is a question about minutes, not tokens. If the session takes twenty minutes on Standard and four minutes on Ultrafast, you have bought sixteen minutes for fifteen dollars — about $0.94 a minute — and the comparison that matters is against what those sixteen minutes cost you. A developer at a fully loaded rate is worth more than a dollar a minute, so for the human-in-the-loop case the answer is not close. An agent waiting on a human review queue is worth nothing per minute, and there the same sixteen minutes are free sixteen minutes and the tier is a waste.
That is the test to apply, and it has nothing to do with the model. Whose clock does the generation time sit on, and what is that clock worth? If the answer is "a person's, and a lot", Ultrafast is the cheapest thing on your invoice. If the answer is "a scheduler's, and nothing", it is the most expensive.

Job two: the overnight evaluation sweep. This is a no
An eval sweep runs a few thousand prompts through the model, writes results to object storage, and a human reads the table in the morning. Its latency budget is not minutes; it is a night. Nothing in the pipeline is waiting on the model except the next request in the queue.
Ultrafast does not remove the sweep's wall-clock cost, because the sweep's wall-clock cost is a scheduling decision you made, not a latency problem you have. What it does is multiply the bill by six and move the job onto a budget with its own per-organization rate limits — which is a real risk in an unattended batch, because a rate-limit ceiling is exactly the failure mode a big sweep hits and the tier page does not publish the number.
And there is a lane on the same rate card that is priced for this job and priced in the opposite direction. Batch and Flex run at half of Standard, and the API's Batch processing is designed for exactly this shape: large volumes, no interactive deadline, results returned asynchronously. On the numbers above, the same 40-turn token total costs $1.50 on Batch against $18.00 on Ultrafast. That is a 12x spread for a job that cannot tell the difference.
The mistake to avoid is treating Ultrafast as the general-purpose upgrade. It is the top of a ladder — Batch and Flex at half, Standard at one, Fast at two, Ultrafast at six — and a ladder is not a menu of better versions. Picking the wrong rung is worth more money than picking the wrong model.
Job three: the batch summariser. Also a no, for a different reason
Suppose the workload is a nightly pass over a document store: long inputs, short outputs, no human in the loop, and an SLA measured in hours. This is where the token mix turns against Ultrafast rather than for it.
The 272,000-token threshold is the reason. On either tier, a single request that crosses it reprices the whole request — every input token, every cached read, every output token — at twice the input and cache rates and 1.5x the output. Long-context Ultrafast therefore reads $24.00 per million input tokens and $90.00 per million output tokens, and the repricing is triggered by the request, not by the portion that exceeds the line. A document summariser that occasionally ships a request at 300,000 input tokens pays the long-context rate on all 300,000 of them.
Cache behaviour compounds it. Cached reads are the least expensive tokens on this model and the most effective cost lever, and they scale with the tier rather than absorbing it — $0.10 per million cached input on Standard, $0.60 on Ultrafast, both at 5% of the uncached input rate. There is no mix of cached and fresh tokens that softens the multiple, so a sharply optimised cached pipeline does not earn a discount from the fast lane. It just pays six times a smaller number.
Put the two together and the summariser is the case where Ultrafast's premium is largest in absolute terms and its benefit is smallest. If the documents are genuinely long and the deadline is genuinely hours, the correct configuration is Batch or standard Standard, and the correct use of Ultrafast is the mid-development loop where you are iterating on the prompt and a person is waiting for each revision.
The two costs that arrive with the multiple
The six-times rate is the visible part of the price. Two invisible parts matter more in production.
The first is the rate-limit budget. Ultrafast runs on its own limits, separate from the Standard and Fast budgets, and OpenAI sets them per organization rather than publishing them on the tier page; the guidance is to check your organization's limits before raising traffic and to contact an account team if they need lifting. So switching a workload to Ultrafast does two things at once: it multiplies the bill and it moves the workload onto a ceiling you may not be able to read. For an unattended agent, the ceiling binds first.
The second is the shape of the savings. Ultrafast reduces inter-token time, not time-to-first-token and not the deliberation phase. A request that spends most of its wall-clock thinking before it emits anything can be switched to Ultrafast and still feel slow, because the tier accelerates the part of the request that was never the bottleneck. This is the failure mode that produces "we paid six times and it's the same speed" reports: it is measuring an application whose latency lives somewhere the tier does not reach. Before committing, measure where the seconds actually go — time to first token against inter-token time — because the tier only owns one of those.
Why "faster" is not a number you can hold OpenAI to yet
Ultrafast is sold on "up to 8x", and the measurement behind that phrase belongs to a different model. The published sentence is about GPT-6 Astra Ultrafast generating tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex. There is no equivalent published multiple for GPT-6.1 Sol, and no independent party has published a tokens-per-second figure for the Sol variant either. The documentation for this tier describes it as reducing the time between generated output tokens and points at a rate card.
It is a reasonable prior that the multiple carries — both models sit on the same serving stack and the tier's mechanism is the same — but "up to" is doing real work in that sentence, and a vendor ceiling measured on a sibling model in a different client is not the number your workload will see. The same is true of the older rung: Ultrafast over GPT-5.6 Sol was announced in August 2026 at "up to 14x faster than Standard processing" in limited preview, and the documentation still says preview access with the Ultrafast rate table holding exactly two rows.
What you can hold OpenAI to is the price, because the price is published and applies to every token. Which means the decision this page is about is a decision about your own latency budget, not about the vendor's speed claim.

Testing it without committing, and switching back
The tier is available to all API users, so the cheap way to answer the wall-clock question is to run the same prompt set twice with the field on and off and compare the token counts and the timings. Two things to watch in that test: whether your requests cross the 272,000-token line, and whether time-to-first-token dominates inter-token time. Either one can make the trial look like a null result when the tier is working exactly as advertised.
Switching back is a request field, not a migration, and the standard lane is served at the vendor's own list rate through OrcaRouter as openai/gpt-6.1-sol — $2.00 per million input tokens and $10.00 per million output tokens, 0% markup with the provider's price passed straight through, so a vendor repricing is live on our side the same day. Ultrafast itself is not something we sell; it is a service-tier flag billed on your own OpenAI account, and saying so is more useful than implying otherwise. What a single key does buy is the standard lane plus the rest of the catalogue behind one OpenAI-compatible endpoint, and automatic failover across providers, which is worth having if you are about to put an expensive tier in front of a production agent and want the standard lane to be a routing decision rather than a code change.

One practical note on the fast lane that does not apply to the cheap one: WebSockets. OpenAI recommends a persistent WebSocket connection for Ultrafast, which is the right shape for an agent that makes many sequential calls and the wrong shape for a request-per-process batch script. If your client is the second kind, the tier's own recommended transport is another reason the night job belongs elsewhere.
When the verdict changes
Three developments would move jobs off the "stay" list. A published Ultrafast rate limit for GPT-6.1 Sol would remove the ceiling risk from the batch case. A published or independently reproduced speed measurement for the Sol tier would let you budget the wall-clock saving instead of assuming it. And a discount rung on the fast lane — an Ultrafast equivalent of Batch, where the tier is still fast but not six times the price — would change the economics of every job in the middle of the ladder.
None of those exists today. What exists is a tier that is exactly what it says it is: the same GPT-6.1 Sol, scheduled differently, at six times the price on every line. Move the interactive agent, leave the other two alone, and measure where your seconds actually go before you decide which one yours is.
