
GPT-6 Astra Ultrafast at 6x: What the Published Rate Card Actually Costs You
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 383 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 209 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 680 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 50 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 102 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
GPT-6 Astra Ultrafast stopped being a waitlist on September 29, 2026. It is now a purchasable service tier with a published price, and that price is exactly six times standard on every single line. The model underneath — GPT-6 Astra, the company's flagship, shipped on September 3, 2026 — is unchanged; what changed at DevDay is that the fast lane over it became generally available, and the company's API documentation now reads plainly that Ultrafast "is broadly available for GPT-6 Astra, with preview access for GPT-5.6 Sol." For the first time you can put this tier on a budget line, and the number to put there is $60.00 per million input tokens and $300.00 per million output tokens, read from the company's own pricing page on September 30, 2026.
That makes this a different question from the one the launch coverage answered. The interesting fact this week is not that the tier exists — the preview was announced on August 13, 2026, and covered to death. The interesting fact is that a tier with no price now has one, and the arithmetic that follows from it is unflattering in a specific, quantifiable way. Six times is not a premium you absorb with a shrug; it is a premium you have to earn back in fewer turns, and whether you can is a property of your workload rather than of the model.
The multiple holds on every line, and that is the first thing to understand
Reading OpenAI's pricing page on September 30, 2026, the Ultrafast column for GPT-6 Astra is a flat 6x of the Standard column rather than a markup on some lines and not others. That matters, because it means you cannot optimise your way out of it by shifting the shape of your requests.
• GPT-6 Astra, Standard service, requests up to 272,000 input tokens — $10.00 input, $1.00 cached input, $12.50 cache writes, $50.00 output, per 1M tokens.
• GPT-6 Astra Ultrafast, same threshold — $60.00 input, $6.00 cached input, $75.00 cache writes, $300.00 output, per 1M tokens. Exactly 6x on all four lines.
• Above 272,000 input tokens the whole request reprices — Standard moves to $20.00 / $2.00 / $25.00 / $75.00, Ultrafast to $120.00 / $12.00 / $150.00 / $450.00. Still 6x. The long-context rule OpenAI documents for GPT-6 Astra is 2x input and cache rates with 1.5x output on the full request once you cross the threshold, and it applies identically to both tiers.
• The other tiers on the same card — Batch and Flex are 50% of Standard, and Fast mode is 2x Standard. So the multiplier ladder for this one model runs 0.5x, 1x, 2x, 6x, with no rung in between the last two.
There is no Ultrafast equivalent of Batch. Batch and Flex are discounts you buy with looser scheduling on the standard lane; there is no slow-cheap version of the fast lane. If you want the speed, you pay the 6x on every token you send and every token you get back, and there is no mix of cached and fresh input that improves the ratio.

What one call actually costs
The abstraction gets easier to reason about with two concrete requests. Both use OpenAI's published per-million rates; the arithmetic is ours.
A single-shot call with 20,000 input tokens and a 2,000-token answer bills $0.30 on Standard — 20,000 tokens at $10 per million is $0.20, plus 2,000 at $50 per million is $0.10. The same call on Ultrafast bills $1.80. Exactly six times, as advertised.
Now the case that actually describes an agent loop: a 200,000-token conversation where 190,000 tokens are a cached prefix and only 10,000 are fresh, producing a 1,000-token step. Standard bills $0.19 for the cached read, $0.10 for the fresh input and $0.05 for the output — $0.34 per step. Ultrafast bills $1.14, $0.60 and $0.30 — $2.04 per step. Again 6x, but look at what the mix did: caching is the single most effective cost lever on this model and it is still exactly 6x cheaper on the standard lane, so a heavily cached agent gains nothing proportionally from the fast tier and it does not soften the premium either.
The consequence is the part worth internalising. Because the multiple is flat, the break-even is not about your token mix at all. It is entirely about turn count. Ultrafast pays for itself in a workload only when running it cuts the number of billed turns by roughly a factor of six — either by collapsing a retry loop into a single pass, or by replacing a sequence of short clarifying calls with one faster interactive session. If your workload issues the same number of turns as before, only sooner, you are buying latency at exactly 6x and nothing else.
Rate limits and geography are the unadvertised ceilings
Two constraints sit outside the price and are easy to miss when you are reading a rate card.
The first is throughput. Ultrafast for GPT-6 Astra is available to all API users, but initially at low rate limits: OpenAI documents 500,000 tokens per minute for tiers 1 through 3, 1,000,000 for tier 4 and 5,000,000 for tier 5. For a tier sold on speed, that is a real ceiling — a 200,000-token agent context at half a million tokens per minute is two and a half requests per minute on a low tier. OpenAI's own guidance points at the same problem from the other side, strongly recommending WebSockets precisely because per-request network overhead can eat the latency the tier is paying for.
The second is residency. Ultrafast supports US data residency and global processing only. It does not support EU or other non-US regional processing endpoints. If your deployment depends on an EU residency endpoint for GPT-6 Astra, this tier is not available to you at any price, and that is a hard boundary rather than a configuration choice. Note also that residency endpoints carry a 10% uplift for models released on or after March 5, 2026, which GPT-6 Astra is.
The speed headline came down from 14x to 8x
This is worth stating carefully, because the number circulating in most coverage is stale. OpenAI's Ultrafast documentation page, as published today, carries the line "our fastest API service tier, with up to 8x faster speeds than Standard mode." The August 13, 2026 preview announcement for GPT-5.6 Sol promised "up to 14x faster than Standard processing" and up to 750 output tokens per second on Cerebras hardware.
Those are not the same claim, and the difference is not a retraction so much as a change of subject. The 14x figure was measured on GPT-5.6 Sol in a limited preview; the 8x figure is what OpenAI publishes for the tier as it stands now, with GPT-6 Astra as its broadly available model. Anyone budgeting against 14x on Astra is budgeting against a number OpenAI has not published for Astra. The honest reading is that both numbers are vendor-reported ceilings on output throughput, that neither is a latency guarantee for your workload, and that end-to-end time still includes input processing, tool calls and your own orchestration. Treat 8x as the planning figure and treat anything better as a bonus you verify on your own traffic rather than assume.
What to do with this on Monday
The decision rule that falls out of the arithmetic is narrower than the marketing implies, and it is worth writing down before someone proposes enabling Ultrafast across an estate.
Turn it on where the workload is latency-bound and interactive, and where a human is waiting. Incident triage, a live support escalation, a research loop where an analyst is iterating inside one sitting — these are the cases where the value of the answer arriving now rather than in four minutes is not proportional to token price, and where a 6x bill on a small number of turns is trivially smaller than the cost of the person waiting. OpenAI's own examples in the preview announcement were exactly this shape: reading logs while an outage unfolds, tightening an overnight research loop into a working session.
Leave it off where the workload is throughput-bound or batch-shaped. A nightly re-index, a bulk classification pass, a scheduled report, a backfill — none of these care about wall-clock latency, all of them are dominated by token count, and all of them have a 0.5x option sitting right there on the same rate card in Batch and Flex. Paying 12x the batch rate to finish sooner is the clearest way to waste money in this table.
The awkward middle is the agentic coding loop, and that is where the turn-count arithmetic decides it. If the speed genuinely removes retries — if the model's first pass at a multi-file change lands more often because it is not fighting a timeout — then Ultrafast can be cheaper per completed task despite being 6x per token. That is a measurable claim about your own traces, not a general one, and the measurement is cheap: count billed turns per completed task on both tiers and multiply by the rate. If the ratio is not near six, the fast tier is a latency purchase and should be justified as one.
The tier is priced; the routes around it are not the same question
One practical note on how to hold this. GPT-6 Astra at standard service is live on OrcaRouter as openai/gpt-6-astra at the provider's own rate — $10.00 input and $50.00 output per million tokens, with the 272K long-context step to $20.00 / $75.00 — served at 0% markup, meaning the provider's list price is passed through and any vendor move lands in your bill the same day it lands on OpenAI's rate card instead of at your next contract renewal. The same key reaches more than 200 models, so the standard tier can sit behind the same client as the cheap tier or a competitor's model without a second integration.
What we do not do is serve the Ultrafast tier. It is a service-tier flag on an OpenAI account rather than a separately routable model — you set service_tier: "ultrafast" on a request to gpt-6-astra — and it is billed by OpenAI to your own account at the rates above. If you enable it, it lives in your direct OpenAI integration, and OrcaRouter is the right place for the standard-tier traffic you route alongside it. Saying that plainly is more useful than implying a capability we do not have.

The decision rule, in one paragraph
Six times is the price of a tier, not the price of a model: the weights, the 1,050,000-token context window, the 128,000-token output ceiling and the answers are all identical to standard GPT-6 Astra. What you are buying is up to 8x faster output throughput, on a rate card that is 6x on every line, subject to low initial rate limits and a US-residency restriction that excludes the EU entirely. That trade is obviously correct for a human waiting on an answer, obviously wrong for a scheduled batch job that has a half-price lane available, and genuinely undecided for agentic loops where the answer depends on whether speed removes turns. Measure the turn count on your own traces before you commit, and route the rest at list price.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
