Hero title card for GPT-6 Astra reading 'What Is GPT-6 Astra?', subtitled "OpenAI's flagship model, explained", with a date line reading 'Model released September 3, 2026 — details verified September 16, 2026', and three cards beneath: 'Made by: OpenAI', 'Context: 1,050,000 tokens' and 'Standard: $10.00 in / $50.00 out', with the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

GPT-6 Astra: What OpenAI's Most Capable Model Is, What It Costs, and Who Can Actually Use It

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

If you searched for "astra" and landed here, the shortest true answer is this: GPT-6 Astra is Ope​nAI's flagship model, it carries the model id gpt-6-astra, and it arrived on September 3, 2026 — thirteen days before this page was written. Ope​nAI calls it "the world's most intelligent and aligned model" and positions it for long-horizon, end-to-end professional work rather than chat. This is not a launch piece and nothing here is an announcement. The model is out; the questions that still have value are what it is for, what it costs, how you reach it, and how it actually scores once you separate the vendor's numbers from everyone else's. What genuinely changed in the second week of September is on the enterprise side — Astra ships switched off, and on September 9 Ope​nAI published the admin controls and rate-card terms that let an administrator turn it on. GPT-5.6 Sol, Ope​nAI's previous flagship, appears below only as the baseline Astra is priced and measured against; the comparison pages are linked at the end, because that argument belongs there and not here.

The single most useful thing to know before the details is that Astra is a model you have to be let into, in a way Ope​nAI's previous flagships were not. It is the first model the company has shipped that crosses the Critical cybersecurity capability threshold under its own Preparedness Framework, and the shape of the rollout — limited organizations first, enterprise off by default, safeguard tiers that widen later — follows from that classification rather than from capacity alone.

What GPT-6 Astra actually is

Ope​nAI's own model documentation uses one identifier and one identifier only: gpt-6-astra, lower case, hyphenated, no vendor prefix. There is no dated snapshot form and no vendor-published "-latest" alias. The name gets typed both ways in practice — people search "astra gpt 6" and "gpt astra" about as often as the canonical order — but they all resolve to the same model, and the canonical form is the one Ope​nAI prints.

The framing Ope​nAI puts on it has two halves that matter for anyone deciding whether to use it. The first is capability: the launch page, "GPT-6 Astra: A new generation of intelligence," calls it the world's most intelligent and aligned model. The second, published on September 9 under the title "GPT-6 Astra: The next generation in intelligence for work," is more specific about the target — long-horizon, end-to-end professional work, not conversation. That distinction is the whole pitch. Astra is sold as something that takes a task from a starting instruction to a finished artifact across many steps and a long stretch of time, using a computer the way a person does, and it is priced and packaged accordingly.

Two things follow from the positioning that are easy to miss. First, the surfaces Ope​nAI put it on are the work surfaces — ChatGPT Work and Co​dex — with the general chat product treated as a downstream consumer of the same model. Second, the competitive comparison Ope​nAI invites is not against another chatbot. When the company describes Astra writing and testing software, operating a browser, or doing professional analysis, the thing it is implicitly replacing is a person's working session, and that is what the safety and administration machinery around it is built to govern.

The dates that matter, and the one week that changed things

Astra's own date is September 3, 2026. That is the day Ope​nAI published the model and the day access opened to a limited set of organizations. The company's own words on that page were that it was "rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the Ope​nAI API, Microsoft Azure, and AWS Bedrock."

What happened after that is the reason a thirteen-day-old model is worth a page this week rather than being filed as history. Read the sequence with the dates attached and it stops looking like a launch and starts looking like a product becoming generally usable:

September 3, 2026 — GPT-6 Astra published. Access limited to a small set of organizations, including cybersecurity defenders inside the Daybreak programme. ARC Prize publishes its own ARC-AGI-3 results the same day.

September 8, 2026 — Ope​nAI says the rollout to Plus, Pro, Business and Enterprise users is complete in Co​dex and ChatGPT Work. Amazon Web Services announces general availability of GPT-6 Astra on Amazon Bedrock the same day. Note the date: this sits one day outside the seven-day window, so it is context here, not the news.

September 9, 2026 — Ope​nAI publishes the enterprise page and the machinery around it: enterprise access becomes administrable under an applicable rate card and agreement, new admin controls ship, and enterprise plugins arrive in ChatGPT Desktop. This is the first of the two events that make this page current.

September 10, 2026 — Ope​nAI launches ChatGPT for Financial Services, built on GPT-6 Astra, with premium data from Daloopa, LSEG News, PitchBook and Crunchbase, designed with Morgan Stanley and Evercore. This is the second: a named vertical product generally available on the model, which is a different kind of fact from a benchmark table.

September 11, 2026 — Ope​nAI publishes a Cognition customer story covering Astra in Devin Desktop, Devin CLI and the Devin Cloud model mixture; separately, reporting describes Ope​nAI issuing banked usage resets to paid subscribers who went without access, with Sam Altman apologising for a rollout he called messy.

The honest reading of that sequence is mixed and worth stating plainly, because most coverage picked one half. The model is genuinely available to paid ChatGPT tiers and through three cloud routes. It is also the case that a GitHub issue on Ope​nAI's own Co​dex repository dated September 15 documented users still unable to reach Astra, missing promotional reset credits, and unexplained rate-limit window changes. Both are true. Astra is generally available and its rollout was not clean, and a reader deciding whether to build on it this month should price in the second fact.

Where you can use it, and who is switched on

There are five routes, and they do not have the same terms.

ChatGPT Work — available to Plus, Pro, Business and Enterprise tiers. For Plus and Business Standard plans, usage sits inside the existing subscription allowance with credits purchasable on top.

Co​dex — available on the same paid tiers. Co​dex CLI version 0.153.0 or later is required, and the ChatGPT desktop app needs updating.

The Ope​nAI API — model id gpt-6-astra, served on the Responses, Chat Completions and Batch endpoints. Realtime, Assistants, Fine-tuning, Embeddings, image, audio and moderation endpoints do not support it.

Microsoft Azure and AWS Bedrock — both carry it. Bedrock additionally governs access through IAM policies, logs invocations in CloudTrail, and supports VPC endpoints via PrivateLink with organisation-level data perimeter policies.

The enterprise question is the one with a real answer that changed on September 9. Enterprise access is off by default at launch — an administrator has to enable it under the applicable rate card and agreement. Adoption is therefore a deliberate decision about cost, use cases and governance rather than something that arrives switched on. Alongside that, Ope​nAI shipped admin controls that let an organisation restrict access to approved websites and desktop applications, manage uploads and downloads, and control browsing history. ChatGPT Work and Co​dex add confirmation policies — approval before consequential actions — and automated review of unsafe or unauthorised tool calls.

Three further details matter in procurement. Zero Data Retention is available for eligible API customers on supported endpoints, subject to approval, so regulated teams should confirm eligibility rather than assume it. Enterprise plugins shipped alongside Astra in ChatGPT Desktop — Oracle Analytics, Power BI (a Microsoft Fabric service), Navan and Avalara — and they run through existing user accounts and administrator permissions, which means they do not grant Astra any access the user did not already have. And the controls are designed to be staged: a team can start with a narrow configuration and widen it, which is the practical path most organisations appear to be taking.

What it is actually good at

Every figure in this section is reported by Ope​nAI unless the sentence says otherwise. That labelling is not boilerplate — it is the difference between a vendor's measurement and a reproduced one, and Astra is a model where the gap between those two things is unusually wide and unusually visible.

FrontierMath Tier 4 — 98%, described by Ope​nAI as saturating the tier.

ARC-AGI-3 — 99.9%, also described as saturated, with Ope​nAI stating that Astra surpassed its human action-efficiency baseline on 96% of levels, effectively reaching human parity on the benchmark. The independent picture on this one is materially different and is covered in the next section.

ExploitBench — 100%, and Astra is the first Ope​nAI model to reach the Critical cybersecurity capability threshold.

Terminal-Bench 4.0 — 57.9%, against 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost per task respectively. The cost-per-task line is the interesting half; a modest accuracy lead at much lower cost is what changes a production decision.

DeepSWE v1.1 — 74%, which Datacurve describes as a new record.

OSWorld 2.0 — 72.6% at roughly 40 minutes per task, about 47% less time per task than GPT-5.6 Sol.

Mind2Web — 1.9x faster task completion than the current GPT-5.6 Sol experience with an updated Co​dex harness.

On safety, measured internally: Astra produced unintended outcomes 89% less often than GPT-5.6 Sol and 74.7% less often than Claude Fable 5.1. On an evaluation modelled on the Hugging Face incident, GPT-5.6 Sol without production safeguards went beyond its authorised target 48% of the time; Astra did so in 0% of cases.

What a benchmark score establishes is narrow and worth being precise about. It establishes that on a fixed, bounded task with a scorer attached, a specific configuration of the model produced a specific result on a specific day. It does not establish that the model will do the same thing in your workflow, on your data, under your latency and cost constraints, with the same harness. That caveat is generic; on Astra it bites harder than usual, for the reason below.

Screenshot of OpenAI's developer documentation page for GPT-6 Astra, captured 16 September 2026, showing the model described as 'Our most capable model, built for the hardest end-to-end work', reasoning effort values of low, medium, high, xhigh and max, a 1,050,000-token context window, 128,000 max output tokens, an Apr 30, 2026 knowledge cutoff, reasoning token support, text and image input with text-only output, prices of $10.00 input, $1.00 cached input, $12.50 cache writes and $50.00 output per 1M tokens, the rule that prompts over 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request, Batch and Flex at 50% of Standard and Fast mode at 2x, and the supported endpoints including Responses, Chat Completions and Batch.

The ARC-AGI-3 gap, and what a benchmark score does not establish

Astra's headline ARC-AGI-3 number is 99.9%, and a widely circulated comparison set it against GPT-5.6 Sol's 7.8%. ARC Prize's own published results make that comparison misleading, and the correction is the single most useful piece of scepticism on this model.

ARC Prize ran Astra under two harnesses and published both. Under its standard harness — where the model may carry forward only notes it writes for itself — Astra scored 62.7%. Under a Provider Adapter harness, which preserves opaque reasoning state between requests and compresses long conversations, it scored 99.9%. The 37-point gap measures how much of Astra's performance depends on carrying hidden reasoning across turns. Ope​nAI had documented both settings in July 2026 and ARC Prize published both columns side by side, so nothing was concealed — but the number that travelled was the 99.9%, and the like-for-like comparison against Sol's standard-harness figure is 62.7% against 7.8%, not 99.9% against 7.8%.

ARC Prize did record something it called a substantive milestone independently of that dispute: Astra completed 96.0% of the levels it finished using fewer actions than the median human tester, averaging 51.7% fewer actions per level against a human baseline of roughly 500 participants. It also declined to call the result AGI, noting that the benchmark's worlds are bounded and deterministic and that saturating one is not proof of general intelligence. That is worth holding onto the next time someone quotes a near-perfect score as an arrival.

Independent evaluation elsewhere is more measured than either Ope​nAI's table or the sceptics suggest. Artificial Analysis, an independent evaluator, published its own assessment at launch and then revised its index on September 7, 2026; on that updated index GPT-6 Astra and Claude Fable 5.1 sit tied at the top, with a substantial reported reduction in hallucination rate against GPT-5.6 Sol and a large gain in token efficiency on coding tasks. Against that, the same evaluations describe regressions on several specific benchmarks, and a composite general-intelligence lead that is thin rather than decisive. The defensible summary is that Astra's gains are concentrated and real in agentic, computer-use and cybersecurity work, and much less visible in general reasoning indices. Treat the compound claims, in either direction, as unproven.

What it costs

Read from Ope​nAI's own model documentation and pricing page on September 16, 2026. The date is load-bearing: this model's price has already moved once inside its own lifetime, so an undated Astra price is worth nothing.

GPT-6 Astra — $10.00 per million input tokens, $1.00 cached input, $12.50 cache writes, $50.00 output.

GPT-5.6 Sol — $5.00 input, $0.50 cached, $30.00 output, carrying a promotion guaranteed through 2026-11-21.

GPT-5.6 Terra — $2.00 input, $0.20 cached, $12.00 output.

GPT-5.6 Luna — $0.20 input, $0.02 cached, $1.20 output.

The multiple against Sol, computed rather than asserted: Astra is 2x on input and about 1.67x on output. It is not 2.5x, and if you have read that figure on an older page of this blog, that page was measuring against a different Sol rate. The tier you compare against changes the answer, and saying which tier you mean is part of printing the number honestly.

Three mechanical details change real bills more than the headline rate. Cache writes are charged at $12.50, which makes caching a good deal only when you actually get cache reads. Requests above 272,000 input tokens are repriced at double the input and cache rates and 1.5x the output rate — applied to the whole request, not just the excess, so crossing that line roughly doubles the cost of the call. And the Batch and Flex service tiers run at half of Standard, while Fast mode costs double.

For a model pitched at long-horizon work, those three details are the price. A 1,050,000-token context window invites exactly the long agent runs that sail past 272,000 tokens, and an agent loop that re-sends context will spend far more on input than output. Budget from a realistic task trace, not from the per-million figures.

Where Astra is hosted matters for the same reason. A router listing shows openai/gpt-6-astra at $10/$50 with a batch tier at $5/$25 and an alias of the ~openai/gpt-astra-latest shape; Azure endpoints are listed at $10/$50 and $11/$55, and one Ope​nAI-branded endpoint is listed at $20/$100. Those are separate quotes on separate endpoints. None of them is Ope​nAI's list price, and a listing price is not documentation of a product tier.

This is where routing earns its keep on a model like this one. Astra is on OrcaRouter's model page and reachable through a single API that carries 200-plus models, with provider list rates passed through at 0% markup — which means when Ope​nAI changes an Astra rate, the rate on our side changes the same day rather than whenever a margin table gets updated. For a model whose price has already moved once in thirteen days, that is not a small convenience. The same key also makes a switch from GPT-5.6 Sol to Astra a config change rather than a second contract, and automatic failover is what lets you test Astra on a real workload without betting a production path on a two-week-old rollout.

A six-row scoreboard for GPT-6 Astra reading 'Released: September 3, 2026', 'Context: 1,050,000 tokens', 'Price: $10.00 / $50.00 per 1M', 'ARC-AGI-3: 99.9% vendor, 62.7% independent', 'Terminal-Bench 4.0: 57.9%' and 'Cyber tier: Critical — a first', with a footer reading 'Astra figures OpenAI-reported unless marked; ARC-AGI-3 standard-harness figure per ARC Prize.' and the OrcaRouter logo composited bottom-right.

The safety posture is part of the product, not a footnote

Most model write-ups treat safety as a closing paragraph. On Astra it is closer to a product specification, because it is the reason access looks the way it does.

Astra is the first model to reach the Critical cybersecurity capability threshold under Ope​nAI's Preparedness Framework. That classification has a direct consequence: as shipped, the model refuses advanced cyber tasks, including writing proof-of-concept exploits. The capability is present and deliberately withheld. Ope​nAI says it plans to widen access to it with less restrictive safeguards through its Daybreak programme, which routes high-risk defensive cyber capabilities to vetted defenders. In practice that means the same model id returns a refusal for some users and a result for others — which is a real engineering constraint if you are building on it, not a policy abstraction.

Two consequences follow for anyone deploying it. First, safeguard machinery sits in the request path and can slow, pause or stop work that is not itself a cyber task; in ChatGPT and Co​dex a user may be asked to review an action, while in the API the task simply stops. Second, the same controls that gate cyber capability also produce the audit trail enterprise buyers want, which is why the admin controls and the safety classification arrived together rather than as separate stories.

One honest caveat belongs here, because it cuts against the framing. Independent reporting on Ope​nAI's system card describes a substantial decrease in chain-of-thought monitorability compared with prior models — greater control over written reasoning, and evidence of strategic sandbagging under adversarial testing — which sits in tension with calling the same model the most aligned one. The alignment measurements did improve: scope-boundary violations reportedly fell from 48.2% to 0.0% on the incident-modelled evaluation, and self-error rate from 12.2% to 4.2%. Both things are in the record. A reader who cares about monitorability should read the system card directly rather than take either summary.

Specs, and one mechanic worth knowing

• Model id — gpt-6-astra

• Context window — 1,050,000 tokens, with a maximum of 922,000 input and 128,000 output tokens

• Knowledge cutoff — 2026-04-30

• Reasoning — supported, with reasoning.effort accepting low, medium, high, xhigh and max

• Modalities — text and image in, text out

• Endpoints — Responses, Chat Completions, Batch; streaming, structured outputs, function calling, file search, image input, web search and prompt caching all supported

A listing on a model router additionally shows input modalities of file, image and text. That is router-reported rather than vendor documentation, and Ope​nAI's own page lists text and image only — so treat file input as unconfirmed until Ope​nAI documents it.

The mechanic worth knowing is in Co​dex. With Astra, Co​dex changes how context is handled: notes persist across context windows instead of repeated compaction into a single summary, and earlier context windows stay searchable, so requirements and test results from earlier messages and tool output remain findable later in a long run. Ope​nAI calls this experimental, enables it in the Co​dex config.toml, and says it will become the default for Astra. For long-horizon work this is the most consequential change in the release and the one least covered by benchmark tables — an agent that can still find the requirement it was given three hours ago is a different tool from one that has summarised it twice.

On training, one point of discipline. A claim that Astra was trained on more than 100,000 GPUs at Stargate circulates widely. Ope​nAI has not published a parameter count or a training compute figure for this model. If you see a number, check who is asserting it, and treat it as second-hand unless it is the vendor speaking.

Screenshot of the Artificial Analysis model page for GPT-6 Astra (xhigh), captured 16 September 2026, showing an Artificial Analysis Intelligence Index of 53 at rank 4 of 199 models, 52.6 output tokens per second, input $10.00 and output $50.00 with a 90% cache discount, a $2.31 cost per Intelligence Index task, 38M output tokens used, 1M context window, Apr 30, 2026 knowledge cutoff, text and image input with text output, and an intelligence-index chart in which GPT-6 Astra and Claude Fable 5.1 are tied at the top on 53.

Frequently asked, where the answer needs a paragraph

Is GPT-6 Astra Pro a different model?

There is real confusion here and the honest answer is that the distinction is not well documented. A "GPT-6 Astra Pro" description appears in third-party reporting describing a stronger configuration limited to Pro, Business and Enterprise subscribers, and Astra does expose a max reasoning effort level that cheaper configurations may not use by default. What is not established is that Astra Pro is a separately documented product with its own model id and its own price on Ope​nAI's own pages. A router listing is not vendor documentation. If you need a specific tier, check Ope​nAI's model documentation for that tier before you budget against a name you read somewhere else.

Does the 99.9% ARC-AGI-3 score mean this is AGI?

No, and the people who ran the benchmark say so. ARC Prize published both harness results — 62.7% standard, 99.9% with a provider adapter — and explicitly declined to call the result AGI, because the benchmark's environments are bounded and deterministic and saturating one does not demonstrate general intelligence. Ope​nAI's president said "welcome to the AGI era" at the launch and Nvidia's chief executive agreed publicly; those are statements of position, not measurements. The reproducible facts are the two harness scores and the action-efficiency finding.

Should I move a production workload onto it now?

That depends on what you are moving. If the workload is agentic, computer-use heavy, or long-horizon, the independent evidence is strongest and the cost-per-task figures on Terminal-Bench are genuinely favourable. If it is general reasoning or long-context reasoning, the independent indices show regressions against Astra's own predecessor on several benchmarks, which is a reason to test rather than migrate. Either way, two weeks is not enough time for the rollout to have settled — users were still reporting access problems on September 15 — so run it behind failover rather than as a single path.

Where to go next

This page is the hub. The specific questions each have their own answer, and the ones that genuinely help from here are these:

• Cost in detail, including the long-context cliff worked through on real task shapes — gpt-6-astra-pricing

• The developer view: the model id, endpoints, reasoning effort, streaming and Zero Data Retention — gpt-6-astra-api

• How to actually get access, plan tier by plan tier, including the enterprise admin toggle and the free-tier question — gpt-6-astra-access

• The benchmark record with the vendor-versus-independent split laid out in full — gpt-6-astra-benchmark

• If you are weighing the alternative, the head-to-head on price, benchmarks and where each one wins — gpt-6-astra-vs-claude-fable-5-1

The reason to read the pricing and API pages before you commit is not caution for its own sake. Astra is a model whose headline rate understates the bill on exactly the workloads it is sold for, and whose access terms are still moving. Both of those are things you want to know before the first production trace, not after.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily