
GPT-5.6 Sol Ultrafast: Ultrafast Shipped at 6× Standard Price, but This Model Is Still Preview-Only
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 127 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 68 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 362 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 233 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The wait is over for the tier, but not for this model. At DevDay on September 29, 2026, the vendor turned Ultrafast into something you can buy: the new ChatGPT Pro 500 plan at $500 a month bundles it, the API now publishes an Ultrafast rate card at six times the standard rate — $60.00 input / $300.00 output per million tokens on GPT-6 Astra — and the vendor's Codex documentation confirms that the Pro tiers carry no five-hour usage limit at all, a window that applies to Plus. All of that is vendor-stated, published on the vendor's own DevDay recap, help pages and developer documentation. What has not moved is the model this page is about. GPT-5.6 Sol Ultrafast, announced on August 13, 2026 as the Cerebras-accelerated serving tier for GPT-5.6 Sol at up to 14× standard throughput and 750 output tokens per second, is still preview access only — arranged through the vendor's account team — and still has no published price of its own. The tier shipped. The Sol variant is still waiting.
Two price movements landed since we first wrote this up, and they point in opposite directions. The model underneath got cheaper: standard GPT-5.6 Sol now bills at $4.00 per million input tokens and $20.00 per million output — OpenAI's promotional rate, which the vendor says is available at least through November 21, 2026 — not the $5.00 / $30.00 that was current in August. The speed tier on top of it got a price for the first time, and it is not cheap: Ultrafast on GPT-6 Astra bills at $60.00 / $300.00 per million. Every figure in this piece was re-checked against OpenAI's published rate card, its API reference and its Codex pricing documentation on 2026-09-30.
What Ultrafast mode actually is
Ultrafast is a serving tier, not a new model. OpenAI announced it on August 13, 2026 as the first product of its January 2026 compute partnership with chipmaker Cerebras — a deal reported at roughly $10 billion over three years. Where standard GPT-5.6 Sol inference runs on GPU clusters and spends much of its time moving weights between memory and compute, Ultrafast loads the model's weights into 44 GB of on-chip SRAM on Cerebras's wafer-scale chips; a model Sol's size spans several wafers in layers, so the weights sit where the compute is. The result is a throughput ceiling set by silicon rather than by memory bandwidth.
The hardware story moved on 2026-08-18. At its Supernova event, Cerebras unveiled the CS-4, the next-generation rack-scale system built on its Nexus platform — three Wafer-Scale Engine chips per rack, up to 2× the token-generation speed of the prior CS-3, and, per Cerebras, more than 1,000 tokens per second on models above 10 trillion parameters. Those are Cerebras's claims for a system in early access, not yet generally available. Since the unveiling, a single post on X claims CS-4 already serves GPT-5.6 Sol at roughly 1,300 tokens per second — consistent in direction with Cerebras's own numbers, but confirmed by nobody so far. Treat it as an early signal: plausible, unverified, and worth re-checking before you plan capacity around it.
The part that matters for your code: the model id, the weights, the reasoning behavior, and the output are identical to standard GPT-5.6 Sol. OpenAI frames Ultrafast as "more useful work per second" rather than a quality tier, and the preview runs the full flagship — speed does not come from swapping in a cheaper model. What changes is how fast tokens come out.
Don't confuse Ultrafast with fast mode. Fast mode is a separate, purchasable tier on standard hardware: up to roughly 2.5× output speed at a 2× per-token premium, with no quality change — it was renamed from Priority Processing on July 30, 2026, and on GPT-5.6 Sol it is $8.00 input / $40.00 output per 1M. Ultrafast is a different thing — Cerebras hardware, up to 14× on GPT-5.6 Sol and up to 8× on GPT-6 Astra, costing 6× the standard rate on Astra and still unpriced on Sol. The two names now sit next to each other in the same model picker, which remains the single most confusing thing about this release.
How fast — the numbers that matter

The headline figure is 14×, and it needs a definition. OpenAI says Ultrafast generates up to 750 output tokens per second versus standard GPT-5.6 Sol processing. That is a throughput number for output tokens, not a flat "everything runs 14× faster" claim — end-to-end request time also includes input processing and the model's own reasoning, which is why 14× is labeled a maximum.
• Max output throughput — up to 750 output tokens per second, per OpenAI's announcement (2026-08-13).
• Vs standard processing — up to 14× faster, per the same announcement; actual gains vary with input length and task type.
• Humanity's Last Exam, 2,500 questions — OpenAI/Cerebras report GPT-5.6 Sol Ultrafast finishing in 11 hours 11 minutes versus 78 hours 27 minutes for Claude Fable 5, with comparable accuracy. Vendor-reported, and the comparison spans two vendors' hardware stacks — read it as directional, not a verdict.
• GDP-Val, end-to-end — Cerebras reports 5.6× faster overall with no measurable quality loss. Also vendor-reported.
• Fast mode and Ultrafast compared — fast mode tops out at roughly 2.5× output speed for a 2× price premium, capped at the same ~272K-token boundary as standard pricing. Ultrafast is the bigger jump: 14× on GPT-5.6 Sol per the August announcement, and up to 8× on GPT-6 Astra in Codex per OpenAI's current docs, with launch coverage putting that at up to 300 output tokens per second.
• CS-4, the next hardware — Cerebras claims the CS-4 rack delivers more than 1,000 tokens per second on models above 10 trillion parameters, up to 2× CS-3's token generation. An unverified community report puts GPT-5.6 Sol on CS-4 at ~1,300 tokens per second — same direction, not yet confirmed by OpenAI or Cerebras.
One caution before you quote the 14×: it is an output-throughput ceiling on one vendor's hardware, and independent comparisons of the August launch ranged from roughly 7× to 11× depending on which part of a workload was measured. The same discipline applies to the CS-4 speed signal — the ~1,300-token/s figure for GPT-5.6 Sol on CS-4 started as a single X post, and neither OpenAI nor Cerebras has confirmed it. Treat all of these as vendor and community claims, not benchmark verdicts.
What Ultrafast costs now — and why that still does not price Sol

Here is the state of the price question on 2026-09-30, stated plainly. Ultrafast has a published rate — for GPT-6 Astra, and only for GPT-6 Astra. OpenAI's pricing page now carries an Ultrafast column alongside Standard, Batch, Flex and Fast: $60.00 input / $6.00 cached input / $75.00 cache writes / $300.00 output per million tokens in short context, doubling to $120.00 / $12.00 / $150.00 / $450.00 past roughly 272K tokens. That is six times the standard Astra rate, where fast mode is two times — and it is the first hard number OpenAI has put on the tier. The card above is our own 2026-08-18 snapshot, and its dollar figures are superseded twice over, by the standard-rate cut and now by this Ultrafast row. What is still missing from that table is any Ultrafast rate for GPT-5.6 Sol: the Ultrafast column lists exactly one model, gpt-6-astra. The tier is priced; this model is not.
What you can price today, model by model:
• Standard GPT-5.6 Sol — $4.00 input / $20.00 output per 1M tokens for prompts up to roughly 272K tokens. OpenAI lists this as promotional pricing available at least through November 21, 2026. This is the baseline you can call right now.
• Fast mode — $8.00 / $40.00 short-context on GPT-5.6 Sol, doubling to $16.00 / $60.00 past roughly 272K tokens. Twice the standard rate for up to about 2.5× the output speed, and on Sol it is still the fastest thing you can actually buy.
• Prompt cache — cache reads bill at $0.40 per 1M (90% off fresh input); cache writes bill at $5.00 per 1M.
• Long-context tier — inputs past roughly 272K tokens bill at $8.00 / $30.00 on standard processing, within the model's ~1.05M-token window and 128K-token maximum output.
• Batch API — $2.00 / $10.00 short-context and $4.00 / $15.00 long-context, a flat 50% off with a roughly 24-hour turnaround.
For GPT-5.6 Sol, the Astra number is the best available clue and nothing more. OpenAI prices Ultrafast at 6× the underlying model's standard rate; apply that multiplier to Sol's $4 / $20 and you land at $24 / $120 per million. That is arithmetic, not a published price — OpenAI has said nothing about a Sol Ultrafast rate, and the August preview cohorts were not billed at a tier rate at all. Plan against standard Sol at $4 / $20 or fast mode at $8 / $40 until the vendor publishes the row.
How to use it — broadly available on Astra, still a request for Sol
Access split in two on September 29, and the split runs against the model this page covers. Ultrafast is now broadly available on GPT-6 Astra: OpenAI's API documentation says it is open to all API users at low rate limits — 500,000 tokens per minute on usage tiers 1–3, 1 million on tier 4, 5 million on tier 5 — and it ships inside ChatGPT Work and Codex on the new Pro 500 plan and on eligible Enterprise and Edu plans, where a workspace owner has to switch it on, and where it is unsupported for workspaces that require inference outside the United States. GPT-5.6 Sol Ultrafast did not move with it. The same documentation still describes it as preview access through an OpenAI account team, and the parameter's own description in the API schema still reads "currently available for gpt-5.6-sol" — a line that now lags the rest of the docs rather than leading them. If you want the tier today, you want it on Astra.
The interface question is settled too. TestingCatalog reported on 2026-09-26 that a speed selector offering Standard, Fast and Ultrafast was sitting unpublished in the Responses API Playground; three days later OpenAI shipped the consumer version of exactly that, with Ultrafast now an option in the model picker in ChatGPT Work and Codex on Pro 500. We flagged that sighting as one publication's read of unreleased UI and said OpenAI had confirmed none of it. It was directionally right, and the rollout it pointed at arrived on DevDay rather than after it.
The preview cohorts OpenAI named in August were coding, commerce, financial research, support, and other interactive workloads — teams whose agents make many calls per request and where response latency is the bottleneck. Jane Street, Rogo, and Podium were among the named early testers. If your workload is a single-turn chat, the speed tier matters far less than it does for a 40-call agent loop. The practical answer now depends on which model you are asking about: on GPT-6 Astra you can set service_tier: "ultrafast" this afternoon, and on GPT-5.6 Sol you still cannot.
What you can do today — the practical answer

While GPT-5.6 Sol Ultrafast stays behind the preview door, both flagship models are live on OrcaRouter at the provider rate with $0 per-token markup — model ID openai/gpt-5.6-sol through the OpenAI-compatible endpoint at api.orcarouter.ai/v1, and GPT-6 Astra beside it at the standard $10.00 / $50.00, which is the model OpenAI has now put an Ultrafast tier on top of and which we do not route at that tier. The capture below is from August 2026 and still shows the then-current $5.00 / $30.00 line; today's Sol rate is $4.00 / $20.00, which is what we pass through. That pass-through is why vendor price moves reach our side the same day: $4.00 input / $20.00 output, the same figures OpenAI publishes, with no per-token markup added on top. If you already have an OpenAI client, migration is a base-URL and model-id change; nothing else. The same key and endpoint carry 200+ models, so Sol sits next to its own siblings — GPT-5.6 Terra at $2.00 / $12.00 and GPT-5.6 Luna at $0.20 / $1.20 — and the open-weight models you would compare it against, and you can route by difficulty instead of re-integrating each one.
Two OrcaRouter specifics are worth naming on a speed page. BYOK: bring your own OpenAI key and OpenAI bills you directly at its own rates — OrcaRouter adds $0 per token and keeps your provider rate limits and credits intact. Guardrails: the PII Shield and content policy run before billing, so a blocked request returns a clean 400 and is never charged, and the Agent Firewall grades tool and MCP calls before they execute. Ultrafast itself is broadly available on GPT-6 Astra but not on routable terms and not for Sol — it is an access-controlled tier billed at 6× the standard rate, and we do not serve it. If OpenAI ever offers it on standard terms, plugging it into the same endpoint is a model-id change; the setup you build today carries over.
The honest boundary — when Ultrafast isn't the answer
Five places the speed story still does not hold, now that half of it has settled:
Available is not the same as available to you. Ultrafast opened to all API users on GPT-6 Astra, but it opened at low rate limits, it is still described as access-controlled in OpenAI's schema, it does not support EU or other non-US regional processing, and in Enterprise workspaces it is off until an owner enables it. On GPT-5.6 Sol nothing opened at all. Check your own access before you architect on either.
14× is a maximum, not a guarantee. It is an output-throughput ceiling on Cerebras hardware; end-to-end latency still includes input processing, the model's reasoning time, and your own network. A reasoning-heavy prompt can spend most of its wall-clock thinking, where the headline speedup shrinks.
It has a price on one model and no price on the other. The Astra Ultrafast rate is published — 6× standard — and there is no Ultrafast row for GPT-5.6 Sol anywhere. Budget against standard Sol at $4 / $20 or fast mode at $8 / $40, and treat any "GPT-5.6 Sol Ultrafast cost" figure you see, including the $24 / $120 extrapolation above, as arithmetic rather than a rate card.
The benchmarks are vendor-reported. The 11-hour HLE run and the 5.6× GDP-Val figure are OpenAI/Cerebras numbers from the August announcement; independent estimates of the speedup range from roughly 7× to 11× depending on what is measured. Add the CS-4 signal to that list: the ~1,300-token/s figure for GPT-5.6 Sol on CS-4 comes from a single X post and has no independent confirmation. The new Astra Ultrafast numbers — 8× in Codex, 6× the price — are also OpenAI's own.
And the one that costs money rather than time: it won't fix a bottleneck you don't have. If your agents are slow because of your own orchestration, tool latency, or input-heavy prompts, a faster output path moves the tail only slightly. The tier-matching decision — GPT-5.6 Sol at $4 / $20 versus GPT-5.6 Terra at $2 / $12 versus GPT-5.6 Luna at $0.20 / $1.20 — still moves your bill more than any speed tier does.
Bottom line. Ultrafast is no longer a waitlist tier you can only read about. As of September 29, 2026 OpenAI sells it: included with the $500 ChatGPT Pro 500 plan in Work and Codex, open to all API users on GPT-6 Astra at low rate limits, and priced at six times the standard rate — $60 / $300 per million short-context. GPT-5.6 Sol Ultrafast, the August preview this page was built around, has not moved: still access-controlled, still limited to gpt-5.6-sol in the API schema, still unpriced, and the ~1,300-token/s figure reported for Sol on Cerebras's CS-4 is still a single unverified X post. Standard GPT-5.6 Sol at $4 / $20 is what you can call today on OrcaRouter at $0 markup; GPT-6 Astra at $10 / $50 is what you can call if you want the model that got the Ultrafast row.
Ultrafast now costs real money on GPT-6 Astra — 6× standard, or bundled into the $500 Pro 500 plan — while GPT-5.6 Sol Ultrafast stays a preview. The full GPT-5.6 Sol is live on OrcaRouter at $4 / $20 per million tokens with zero markup, ready for your own key.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
