
GPT-6 Astra Pricing: Why Three Different Prices Are All Correct, and What a Task Really Costs
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
The bottom line first: GPT-6 Astra costs $10.00 per million input tokens, $1.00 per million cached input, and $50.00 per million output tokens on OpenAI's standard tier. That is OpenAI's own list price, read from OpenAI's own pricing page on September 16, 2026 — and the date matters, because the model has been priced differently at different points and readers keep finding those other numbers. GPT-6 Astra itself shipped on September 3, 2026, thirteen days before this page. This is not a launch piece and not an announcement piece; it is the reference page for a model that is now generally available and, as of this week, generally deployed. GPT-5.6 Sol appears below as the price baseline Astra is sold above, not as the subject.
The reason a rate card needs a whole page is that the numbers in circulation disagree. Search for this model's price and you will find $10/$50 from OpenAI, a $5/$25 batch rate, Azure quotes at both $10/$50 and $11/$55, and at least one endpoint listed at $20/$100. None of those is a typo and none of them is the list price being quietly changed. They are different service levels and different resellers, and telling them apart is the difference between a budget and a guess.
The rate card, with every line labeled and dated
Everything in this section was read from OpenAI's own API pricing and model documentation on September 16, 2026. Where a figure comes from anyone else, the sentence carrying it says so.
• Standard, requests at or under 272,000 input tokens — $10.00 per 1M input, $1.00 per 1M cached input, $12.50 per 1M cache writes, $50.00 per 1M output (OpenAI's own pricing table, read 2026-09-16).
• The cache write line is real and people miss it — the first time you send a prefix it is a write at $12.50 per 1M, not a read at $1.00. Cache writes are billed at 1.25x the standard input rate.
• Standard, requests over 272,000 input tokens — $20.00 input, $2.00 cached input, $25.00 cache writes, $75.00 output per 1M. The whole request reprices, not just the tokens past the line (OpenAI's own pricing table, read 2026-09-16).
• Batch and Flex — half of standard: $5.00 input, $0.50 cached input, $6.25 cache writes, $25.00 output per 1M (OpenAI's own pricing table, read 2026-09-16). Over 272K input the batch rate is $10.00 / $1.00 / $12.50 / $37.50.
• Fast mode — double standard: $20.00 input, $2.00 cached input, $25.00 cache writes, $100.00 output per 1M (OpenAI's own pricing table, read 2026-09-16).
• Model specs that set the bill — context window 1,050,000 tokens, maximum input 922,000, maximum output 128,000, knowledge cutoff 2026-04-30, reasoning token support yes, reasoning effort values low / medium / high / xhigh / max (OpenAI's own model documentation, read 2026-09-16).
If you want the single sentence version: on the standard tier, a request that stays under 272K input tokens costs $10 per million in and $50 per million out, cached reads are a tenth of the input rate, and everything else in this article is either a discount for accepting slower processing or a surcharge for crossing a threshold.
Why you will see a different number, and which is which
Five distinct things produce a quote that is not $10/$50. Four of them we can tie to documentation; one we cannot, and we will say so rather than guess.
$5/$25 is the batch tier, and it is OpenAI's own price. This is not a reseller discount or a negotiated enterprise rate. OpenAI's pricing page lists batch processing at half the standard rate, and the Astra batch row reads $5.00 input and $25.00 output. If a source quotes you $5/$25 as though it were a different model, it is the same model, the same weights, on the slower queue. Batch is the correct answer for overnight evaluation runs and the wrong answer for anything interactive.
$11/$55 is exactly a 10% uplift on list, and OpenAI documents a 10% uplift for residency endpoints. OpenAI's pricing page states that residency endpoints are charged a 10% uplift for models released on or after March 5, 2026. GPT-6 Astra was released on September 3, 2026, so it qualifies. Multiply list by 1.1 and you get $11.00 and $55.00 to the cent. That is a strong explanation of where the higher Azure-style quote comes from — a residency-bound endpoint passing through a documented uplift rather than a reseller inventing a margin. We should be precise about the limit of that claim: the uplift is documented by OpenAI, and the arithmetic matches exactly, but the specific Azure listing we saw does not itself state that it is a residency endpoint. Treat it as the likely explanation, not a confirmed one.
$20/$100 is OpenAI's Fast mode rate for this model. This is the one that looks most like a bogus number and is least bogus. Fast mode doubles every line, and Astra's standard Fast row is $20.00 input and $100.00 output per 1M. An endpoint quoting $20/$100 is quoting a different service level — priority processing — not a different list price and not a different model. If you see that figure next to Astra's name, the question to ask is which processing tier, not which vendor.
Prices below list exist, and they are someone else's margin. Third-party gateways and resellers advertise Astra under list. One gateway publicly lists about 20% below OpenAI's rate. That is a commercial decision by that operator, not a price change by OpenAI — and it carries whatever data-handling and reliability terms that operator applies. We are not going to name them or count their discount as evidence about this model's price, because it is evidence about their business, not about OpenAI's rate card.
What we could not verify. We saw a router listing give the model as gpt-6-astra with a batch tier at $5/$25 and a moving alias of the ~openai/gpt-astra-latest shape. The batch figure checks out against OpenAI's own page. The alias is a gateway construct: OpenAI's model documentation lists exactly one entry under snapshots and it is the bare gpt-6-astra, with no dated form and no -latest alias. Do not put a gateway alias into a config as though the vendor published it. Separately, we found no OpenAI documentation for any separately priced "Astra Pro", "Astra-Medium" or similar tier. "Pro" appears in ChatGPT plan naming, not in the API rate card, and a router listing is not vendor documentation. If you need a tier confirmed, confirm it on OpenAI's own page.

Astra beside the rest of the OpenAI rate card
A price is only meaningful next to the alternatives on the same card. These are all OpenAI's own standard rates, all read on September 16, 2026, all per million tokens, all at or under 272K input.
• GPT-6 Astra — $10.00 input / $1.00 cached / $12.50 write / $50.00 output.
• GPT-5.6 Sol — $4.00 input / $0.40 cached / $5.00 write / $20.00 output.
• GPT-5.6 Terra — $2.00 input / $0.20 cached / $2.50 write / $12.00 output.
• GPT-5.6 Luna — $0.20 input / $0.02 cached / $0.25 write / $1.20 output.
• The long-context versions — Astra $20.00 / $2.00 / $25.00 / $75.00; Sol $8.00 / $0.80 / $10.00 / $30.00; Terra $4.00 / $0.40 / $5.00 / $18.00; Luna $0.40 / $0.04 / $0.50 / $1.80.
There is a promotion running underneath that table, and it decides the multiple everyone argues about. OpenAI's pricing page notes that GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026. Read against the promotional rate that is live today, Astra is 2.5x Sol on every single line — 2.5x on input, 2.5x on cached input, 2.5x on output. Read against Sol's non-promotional rate of $5.00 input and $30.00 output, Astra is 2x on input and about 1.67x on output. Both multiples are arithmetically correct and they describe the same two models on the same day. The difference is entirely which Sol rate you use as the denominator, and the honest answer to "is Astra 2.5x Sol?" is: yes, until November 21, 2026, or whenever OpenAI ends the promotion, whichever comes first.
We make this point at length because it is the most common error in coverage of this model, including on this blog. A per-token multiple that silently picks one Sol rate and doesn't name it is not a fact about Astra. It is a fact about when the sentence was written.

Cost per task, not cost per token
OpenAI's own framing is that comparing token prices for this model is the wrong comparison, because Astra is designed to complete long-horizon work in fewer tokens and with fewer retries. That claim is worth taking seriously enough to do the arithmetic rather than repeat it. So here is a worked example with the assumptions stated up front, so you can change them to match your own workload.
The workload: a 200-call agent run, the shape of an overnight migration. Each call carries 10,000 input tokens, of which 8,000 are cache reads and 2,000 are fresh, and produces 500 output tokens. Across the run that is 400,000 fresh input tokens, 1,600,000 cached input tokens and 100,000 output tokens, plus a single 10,000-token cache write to seed the prefix. Standard tier, under 272K per request, so no long-context surcharge applies.
• GPT-6 Astra — fresh 0.4M x $10.00 = $4.00; cached 1.6M x $1.00 = $1.60; write 0.01M x $12.50 = $0.13; output 0.1M x $50.00 = $5.00. Total $10.73.
• GPT-5.6 Sol at its promotional rate — fresh $1.60; cached $0.64; write $0.05; output $2.00. Total $4.29.
• GPT-5.6 Sol at its non-promotional rate — fresh $2.00; cached $0.80; write $0.06; output $3.00. Total $5.86.
• GPT-5.6 Terra — fresh $0.80; cached $0.32; write $0.03; output $1.20. Total $2.35.
• GPT-5.6 Luna — fresh $0.08; cached $0.03; write $0.00; output $0.12. Total $0.23.
At identical token counts, then, Astra is about 46x Luna and 4.6x Terra. Nobody buys Luna to do an Astra-sized job, and that footnote is the whole reason cost-per-token comparisons between those two models are theatre. The Sol comparison is the one that decides something, so here is the number that matters.
The break-even is 40%. Because Astra is exactly 2.5x Sol on every line at Sol's promotional rate, Astra can burn up to 40% of Sol's tokens on the same task and still land on the same bill — and that is true for any mix of fresh input, cached input and output, because the ratio is uniform. If Astra uses a third of Sol's tokens, it comes in cheaper; if it uses half, it is more expensive. Against Sol's non-promotional rate the break-even relaxes to about 55% on this cache-heavy mix, because Astra's output multiple drops to 1.67x while its input multiple drops to 2x.
That single number is the useful one, and it is where the vendor claim and the independent measurement stop agreeing. OpenAI reports Terminal-Bench 4.0 at 57.9% for Astra against 37.3% for GPT-5.6 Sol, at approximately 9% lower estimated API cost per task — vendor-reported. Artificial Analysis, in an independent evaluation published September 9, 2026, measured Astra using about one third of Sol's tokens per task on its Coding Agent Index, and still found Astra at $7.09 per task at max effort — about 15% more than Sol (max) — for a 7-point higher score. On its Intelligence Index the same evaluation put Astra between $0.82 and $3.26 per task across effort levels, about 60% more expensive than Sol at max effort.
Both of those can be true at once, and the gap between them is the honest headline. Token savings are real and large — a third is exactly the kind of number that clears a 40% break-even — but savings measured on output tokens do not shrink the cache-read and fresh-input lines that dominate a long agent session. Whether Astra is cheaper per task than Sol depends on your mix, and the direction flips between workloads. Anyone who tells you it is simply cheaper, or simply 2.5x more expensive, is quoting one line of the bill as though it were the total.
What actually drives the bill on this model
Three things dominate, and all three are specific to how Astra is meant to be used.
Reasoning tokens are output tokens. They bill at $50.00 per million, the same as visible output. OpenAI's model page lists the supported effort values as low, medium, high, xhigh and max, and there is no disable setting — the floor is higher than it was on older models. On the independent Intelligence Index measurement, the spread between low and max effort was $0.82 to $3.26 per task, a factor of four on identical work. Effort level is the largest single cost lever you control, and it is worth setting per-call rather than globally.
Retries multiply everything. An agent that fails and re-runs pays for the whole prefix again. This is the mechanism behind OpenAI's fewer-retries claim, and it is also why per-token arithmetic misleads in both directions: a model that costs 2.5x per token but completes in one pass can beat a cheaper model that needs three. If you are evaluating Astra against anything, measure completed tasks per dollar, not tokens per dollar.
Long context is the designed mode, and it is surcharged. The window is 1,050,000 tokens. Cross 272,000 input tokens in a single request and the entire request reprices at $20.00 input and $75.00 output per million — not just the overflow. A request at 280,000 tokens costs roughly double a request at 270,000 for about 4% more input. On a model built for long-horizon work this is the single easiest way to blow a budget without noticing, and it is why chunking a long-horizon task into sub-272K requests is usually cheaper than sending it whole, even though the model is capable of taking it whole.
The counterweight is caching, and it is where long agent sessions actually save money. Cached reads are $1.00 against $10.00 fresh, a 90% reduction, and a long session re-sends the same prefix dozens of times. The $12.50 cache write is paid once. On the worked example above, if every call had to re-send its 8,000-token prefix uncached, Astra's bill would rise from $10.73 to about $25.00 — the cache is doing more work for the budget than any other single line.
The record on the multiple
Several pages on this blog, including GPT-6 Astra vs GPT-5.6 Sol: is 2.5x the price worth it? and GPT-6 Astra Demand Pauses New $200 ChatGPT Pro Signups, state that Astra lists at 2.5x GPT-5.6 Sol's token price. That figure was computed against Sol's promotional rate of $4.00 input and $20.00 output, which is the rate OpenAI's pricing page carries today and describes as promotional, available at least through November 21, 2026.
So the multiple has not stopped being true — it has acquired an expiry date that the earlier pages did not carry. Against Sol's non-promotional rate of $5.00 input and $30.00 output, Astra is 2x on input and about 1.67x on output, not 2.5x. The pages that say 2.5x should be read as being about the promotional window; the direction of the comparison does not change, but the size of the gap shrinks by a fifth to a third the moment the promotion ends. Anyone budgeting past November 21, 2026 should use $5.00 and $30.00 as the Sol denominator and 2x / 1.67x as the Astra multiple. We are not restating the arguments on those pages — the benchmark and reliability reasoning there stands on its own — only correcting the denominator they were computed against.
The "free" question, answered without hedging
There is no free tier for this model, and no free API path to it. Free and Go ChatGPT users do not get GPT-6 Astra at all. Plus subscribers get it only through ChatGPT Work and Codex, not in regular chat. The API is billed from the first token at $10.00 per million input, and the cheapest legitimate way to reduce that number is the batch tier at $5.00 per million, which costs you latency rather than capability. Anyone advertising free GPT-6 Astra API access is reselling something — usually a trial credit pool on a third-party gateway with its own data terms. That is not necessarily a bad deal, but it is not free, and it is not OpenAI.
Worth knowing if you are capacity-planning around a subscription rather than the API: on September 10, 2026, OpenAI paused new sign-ups and upgrades to its $200 ChatGPT Pro tier, citing demand for Astra. The $100 tier, Plus, Go, the API and Business and Enterprise accounts stayed open. Consumers got rationed; API and enterprise access did not. If your plan depends on a consumer subscription's usage allowance, that is the risk to model.
Routing it, and what a router does not fix
Astra is one of the models available through OrcaRouter under openai/gpt-6-astra, at the same $10.00 input and $50.00 output the vendor charges, with the above-272K tiering shown on our model page. We pass provider list price through with 0% markup, which matters more on this model than on most: the numbers in this article have already moved once inside the model's lifetime, and a pass-through price moves the day the vendor moves it rather than when someone remembers to update a table.

The honest case for routing here is not price, because there is no discount to sell you — it is that the cost math above is workload-dependent and you will not know which side of the 40% break-even you land on until you run it. Putting Astra behind one endpoint alongside Sol, Terra and Luna means you can measure completed tasks per dollar on your own traffic and switch models without a second contract or a code change. The routing DSL lets you compose a call across several models rather than committing a production path to one, and automatic failover covers the case this model has already demonstrated twice — that demand for it can outrun capacity.
The tradeoffs are just as real and worth naming. A router is a hop in the request path, and on a model whose selling point is long-horizon autonomous work that hop is one more thing in the failure domain; if your tolerance for that is zero, call the vendor directly. Several things in this article only exist on a direct relationship: the Zero Data Retention option OpenAI offers to eligible API customers on supported endpoints is subject to OpenAI's approval, and enterprise enablement — which is off by default at launch and turned on by administrators under their applicable rate card — is an agreement between you and OpenAI, not something a router can grant. OpenAI also documents that Fast mode for this model is unavailable with EU data residency, so if you need residency and speed together, that constraint follows the model, not the route. And a router's 0% markup is not a discount: we are not cheaper than going direct, we are the same price with fewer integrations to maintain.
What to watch
One date and one number. The date is November 21, 2026, when Sol's promotional pricing is no longer guaranteed — if it lapses, Astra's multiple against Sol drops from 2.5x to 2x on input and 1.67x on output, and every cost-per-task comparison in circulation gets a fifth to a third more favourable to Astra without a single thing changing about Astra. The number is your own break-even: run one representative long-horizon task on both models, count completed tasks and total dollars, and you will have an answer more useful than anything published, including this page.
Putting Astra behind one endpoint alongside Sol, Terra and Luna means you can measure completed tasks per dollar on your own traffic and switch models without a second contract or a code change.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
