A generated title card with the headline 'Ultrafast vs Ultrafast', the subtitle 'GPT-6 Astra against GPT-5.6 Sol' and two badges reading 'one has a published price' and 'the other is still a preview'.
Guides & Insights

GPT-6 Astra Ultrafast vs GPT-5.6 Sol Ultrafast: One Is on the Rate Card, One Is Not

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The two Ultrafast tiers are not twins, and the difference between them is not a benchmark. GPT-6 Astra Ultrafast runs on GPT-6 Astra, the company's flagship, which shipped on September 3, 2026 and went broadly available on the Ultrafast tier on September 29, 2026. GPT-5.6 Sol Ultrafast runs on GPT-5.6 Sol, the model that held the top of the company's line from July 9, 2026, and it has been in limited preview since August 13, 2026 — the state it is still in today. Read the company's own API documentation on September 30, 2026 and the sentence is unambiguous: Ultrafast "is broadly available for GPT-6 Astra, with preview access for GPT-5.6 Sol."

That single sentence produces the whole comparison, and it produces an asymmetry that a spec table made of two identical rows would hide. Both tiers expose the identical weights of their parent models, both are the same serving concept, and yet one of them has a price on OpenAI's rate card and the other does not. A buyer deciding between them is therefore not choosing between two prices. They are choosing between a price and a waitlist, which is a different kind of decision and needs a different page.

What is actually the same

Start with the parts that do not differ, because they are most of the substance and they are what makes the naming confusing in the first place.

• The weights — GPT-6 Astra Ultrafast is standard GPT-6 Astra on a faster serving path, and GPT-5.6 Sol Ultrafast is standard GPT-5.6 Sol on a faster serving path. Neither is a distillate, a fine-tune or a separate checkpoint. OpenAI's request schema for the tier sets model to the base model id and adds service_tier: "ultrafast"; there is no gpt-6-astra-ultrafast model string.

• The answers — the same checkpoint at the same effort setting should return the same content, modulo sampling. If your two lanes diverge on an identical prompt at identical settings, that is a finding worth investigating rather than a feature you bought.

• The call surface — both are reachable over the Responses API, and OpenAI recommends WebSockets for either, warning that per-request network overhead can erode the latency the tier is paying for.

• The residency restriction — Ultrafast supports US data residency and global processing only, and does not support EU or other non-US regional processing endpoints. That applies to the tier, not to the model, so it constrains both sides equally.

Where they genuinely diverge

Everything that separates these two lives in the columns the marketing material does not print side by side.

• Availability — GPT-6 Astra Ultrafast is available to all API users today at low initial rate limits. GPT-5.6 Sol Ultrafast is a limited preview for a select group; OpenAI's guidance today is to contact your account team to request higher rate limits or preview access for GPT-5.6 Sol.

• Price — GPT-6 Astra Ultrafast is on the published rate card at $60.00 input, $6.00 cached input, $75.00 cache writes and $300.00 output per 1M tokens for requests up to 272,000 input tokens, stepping to $120.00 / $12.00 / $150.00 / $450.00 above that line. GPT-5.6 Sol Ultrafast has no published rate card at all. OpenAI's Ultrafast pricing table has exactly one row in it, and the row is gpt-6-astra.

• Standard-tier price underneath — GPT-6 Astra bills at $10.00 / $50.00 standard, GPT-5.6 Sol at $4.00 / $20.00, with Sol's rate described on OpenAI's own page as promotional and available at least through November 21, 2026. Sol is 2.5x cheaper on both axes before any speed tier is applied.

• Throughput ceiling — OpenAI's Ultrafast documentation states up to 8x faster than Standard mode. The number attached to Sol's Ultrafast preview is up to 14x, with up to 750 output tokens per second, and that is the figure the August 13, 2026 announcement carried. Those are two different claims about two different models, not a revision of one claim.

• Hardware — the preview announcement names Cerebras as the partner behind Ultrafast, with the model's weights close to the silicon. Nothing in the current documentation attaches a hardware claim to the Astra tier specifically; the Cerebras story is documented for the Sol preview.

• Rate limits — Ultrafast for GPT-6 Astra is documented at 500,000 tokens per minute for API tiers 1 through 3, 1,000,000 for tier 4 and 5,000,000 for tier 5. No equivalent published table exists for Sol on Ultrafast, because it is not generally available.

One adjacent fact worth knowing before you assume the newest model inherits the fast lane: OpenAI's pricing page also carries gpt-6.1-sol, and the Ultrafast section has no row for it either. The fast tier is, today, an Astra capability with a Sol preview attached.

A screenshot of OpenAI's Ultrafast documentation page, showing the sentence 'Our fastest API service tier, with up to 8x faster speeds than Standard mode.', the sentence 'It is broadly available for GPT-6 Astra, with preview access for GPT-5.6 Sol.', the recommendation to use WebSockets because per-request network overhead can reduce the latency gains, the note that Ultrafast for GPT-6 Astra is available to all API users at low rate limits with higher limits or Sol preview access available through an OpenAI account team, and a Python code sample setting service_tier to 'ultrafast' with model 'gpt-6-astra'.

Reading the 8x and the 14x together

The two speed numbers are the most misquoted part of this comparison, so it is worth being precise about what each one is.

The 14x figure and the 750-tokens-per-second figure come from OpenAI's August 13, 2026 preview of Ultrafast on GPT-5.6 Sol. As written on that page, Ultrafast "runs GPT-5.6 Sol up to 14x faster than Standard processing" and "generates up to 750 output tokens per second." Both are ceilings on output throughput, measured by the vendor, on a limited preview.

The 8x figure is what OpenAI's Ultrafast documentation page states today as the description of the tier itself: "our fastest API service tier, with up to 8x faster speeds than Standard mode." That page's broadly-available model is GPT-6 Astra.

The reading that survives scrutiny is that the tier's published ceiling is 8x and the Sol preview's published ceiling is 14x. Whether Astra on Ultrafast can hit the higher number is not something OpenAI has published, and neither figure is a latency promise for your traffic — end-to-end time still contains input processing, tool calls and your own orchestration. Plan on 8x for Astra, treat 14x as a Sol-preview figure that may or may not transfer, and verify against your own traces rather than a slide.

A screenshot of OpenAI's API pricing page with the Ultrafast tab selected, showing a single Ultrafast row for gpt-6-astra at $60.00 input, $6.00 cached input, $75.00 cache writes and $300.00 output per 1M tokens short-context, stepping to $120.00 / $12.00 / $150.00 / $450.00 above 272,000 input tokens, with gpt-5.6-sol shown only at its standard rates of $4.00 / $0.40 / $5.00 / $20.00 short-context and $8.00 / $0.80 / $10.00 long-context, and page notes covering the 10% data-residency uplift for models released on or after March 5, 2026, the further 10% FedRAMP uplift, the July 30, 2026 renaming of Priority processing to Fast mode, and the availability of GPT-5.6 Sol's promotional pricing at least through November 21, 2026.

For scale, the tiers below these two are worth keeping in view: Batch and Flex are priced at 50% of Standard on GPT-6 Astra, and Fast mode — renamed from Priority processing on July 30, 2026 — is priced at 2x Standard with a documented ceiling of up to 2.5x faster speeds. So the ladder runs 0.5x, 1x, 2x at up to 2.5x speed, and 6x at up to 8x speed. The step from 2x to 6x buys roughly 3.2x more speed for 3x more money, which is a roughly linear trade rather than the step change the 8x headline suggests.

Everything that matters happens at the standard tier

There is a version of this comparison in which the fast tiers do not decide anything, and it is the version most teams should be in.

Both parents share a 1,050,000-token context window and a 128,000-token output ceiling, both take text and image input and return text, and both support the same function-calling and structured-output surface. Their differences are the ordinary ones between two generations of the same line: GPT-6 Astra carries an April 30, 2026 knowledge cutoff against Sol's February 16, 2026, and exposes reasoning effort from low through max where Sol offers none through max. On the independent benchmarks OrcaRouter's catalogue carries from Artificial Analysis, as read on September 30, 2026, GPT-6 Astra posts 96.1 on GPQA Diamond and 54.7 on Humanity's Last Exam against GPT-5.6 Sol's 94.1 and 49.5, with a narrower gap on coding — 76.9 against 77.4 on the AA Coding Index — and near parity on Terminal-Bench v2.1 at 88.39 against 88.01. Those are single-run figures from one evaluator, they are taken at each model's own default configuration, and Artificial Analysis re-cuts its indices, so treat them as a snapshot rather than a standing.

The point of putting them here is that none of those numbers changes when you add Ultrafast. The fast tier buys you the same model sooner, so if your decision is quality per token, it is settled entirely at the standard tier and the fast tier never enters the argument. Which model do you want, at $10 / $50 or at $4 / $20? That question has an answer you can test today. Whether you also want to pay six times the rate to get that model's answer sooner is a separate and much narrower question, and for GPT-5.6 Sol it is a question you cannot currently buy an answer to at all.

The purchase decision, stated plainly

If you need speed today, GPT-6 Astra Ultrafast is the only one of these two you can actually have. It is generally available to any API user at documented rate limits, and its cost is knowable in advance: $60 / $300 short-context, $120 / $450 long-context, with cached reads at $6.00 per million and cache writes at $75.00. The decision rule follows the flat 6x multiple — the tier earns its price in interactive work where a person is waiting and in loops where the speed demonstrably removes billed turns, and it does not earn it in batch-shaped or throughput-bound work, where Batch and Flex sit at half of standard on the same card.

If you want speed on GPT-5.6 Sol, the honest answer is that you cannot buy it yet, and OpenAI's own page says so. The preview exists, the 14x and 750-tokens-per-second numbers are published, and access runs through an account team. Anyone telling you a price for GPT-5.6 Sol Ultrafast is extrapolating; the published rate card has one Ultrafast row and it is not Sol's. The practical position is to keep the standard tier of whichever model fits your workload in production now, and treat the Sol preview as a pipeline item to revisit when it becomes purchasable.

Where the routes sit

One operational note, since neither of these fast tiers is a separately addressable model. GPT-5.6 Sol at standard service is live on OrcaRouter as openai/gpt-5.6-sol at the provider's own rate — $4.00 input and $20.00 output per million tokens, with the long-context step to $8.00 / $30.00 above 272K — and GPT-6 Astra is live as openai/gpt-6-astra at $10.00 / $50.00 with its own long-context step. Both are served at 0% markup, meaning the provider's list price passes straight through, so a vendor price move lands in your bill the same day rather than at contract renewal, and both sit behind one key alongside more than 200 other models with automatic failover, so you can compare them on your own traffic without a second integration.

The Ultrafast tiers are not routes we serve. They are access-controlled service-tier flags on an OpenAI account — set by the caller on a request to the base model id, billed by OpenAI at the rates above — and for GPT-5.6 Sol, gated behind a preview request. Saying that precisely matters more than implying a capability that does not exist: if you are running Ultrafast, it is in your direct OpenAI integration, and the sensible place for the standard-tier traffic you route beside it is a gateway that passes list price through.

A screenshot of the OrcaRouter model page for GPT-5.6 Sol, model id openai/gpt-5.6-sol, showing a 1M-token context window, 128K maximum output, public benchmarks attributed to OpenAI dated 2026-07-09, input price $4.00 and output price $20.00 per 1M tokens, p50 time to first token 1.72 s, a 10.00 s figure, and 89.0M tokens of traffic, with the page's code sample and EN language toggle visible in the header.

What to watch

Two things would change this page, and neither is a rumour to be anticipated. The first is a rate card: the moment OpenAI publishes an Ultrafast price for GPT-5.6 Sol, the comparison becomes a straight six-times-versus-something question and the arithmetic I sketched above would need redoing against Sol's $4 / $20 base rather than Astra's $10 / $50. The second is general availability for the Sol preview, which would move the "you cannot buy this" half of the page. Until one of those lands, the honest summary is that GPT-6 Astra Ultrafast is a priced, purchasable, generally available fast lane on the more expensive and more recent flagship, and GPT-5.6 Sol Ultrafast is a documented preview on the cheaper one with a higher headline speed and no price at all.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily