
GPT-6 Astra Benchmarks: Every Score, Who Measured It, and Where the Numbers Stop Meaning Anything
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
GPT-6 Astra scores 98% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3, and OpenAI itself calls both benchmarks saturated — which makes the two most-quoted numbers on the model's page the two that tell you least about whether to use it. Against GPT-5.6 Sol and Claude Fable 5.1, the rows that actually discriminate are elsewhere: a 57.9% on Terminal-Bench 4.0, a 74% on DeepSWE v1.1 that is independent and also a tie, and a time-per-task figure that decides more about usability than any percentage here. The model itself is from September 3, 2026 — thirteen days old as we publish, not launch news. This is a reference page for what GPT-6 Astra scores, who ran each measurement, and what each number does and does not establish.
The reason a thirteen-day-old model is still worth a page this week is that the things a reader can act on arrived after the launch did. Broad availability completed: OpenAI said on September 8 that Astra was fully rolled out to Plus, Pro, Business and Enterprise users in Codex and ChatGPT Work, having opened on September 3 with the line that it was "rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock." Enterprise access, off by default at launch, became something an administrator can turn on under their applicable rate card, and OpenAI's enterprise-work page for the model — published September 9 — shipped admin controls and four enterprise plugins alongside it. And the first independent evaluations of the model landed, which is what changes this from a scoreboard read off a vendor slide into something a buyer can weigh.
The full benchmark list, with the source on every row
Everything below is OpenAI-reported unless the row says otherwise. That distinction is not decoration: only one of these headline figures was produced by someone other than the company selling the model, and it is the one that turns out to be a tie.
• FrontierMath Tier 4 — 98%. Vendor-reported by OpenAI, and described by OpenAI as saturating the tier.
• ARC-AGI-3 — 99.9%. Vendor-reported, likewise described as saturated. OpenAI adds that Astra "surpassed our human action-efficiency baseline on 96% of levels, effectively reaching human parity on the benchmark" — a more specific and more interesting claim than the 99.9%, and still OpenAI's own measurement.
• ExploitBench — 100%. Vendor-reported, and a perfect score.
• Terminal-Bench 4.0 — 57.9%. Vendor-reported by OpenAI, against 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost per task respectively. The cost figures come with no published methodology in the material we read.
• DeepSWE v1.1 — 74%. Measured by Datacurve, not OpenAI. This is the independent row.
• OSWorld 2.0 — 72.6%. Vendor-reported, at roughly 40 minutes per task, which OpenAI puts at about 47% less time per task than GPT-5.6 Sol.
• Mind2Web — 1.9x faster task completion than "the current GPT-5.6 Sol experience," with an updated Codex harness. Vendor-reported, and the comparison is against a differently-harnessed Sol rather than the same scaffold with a different model in it.
• Safety — internal to OpenAI. Astra produced unintended outcomes 89% less often than GPT-5.6 Sol and 74.7% less often than Claude Fable 5.1. On an evaluation modelled on the Hugging Face incident, GPT-5.6 Sol without production safeguards went beyond its authorised target 48% of the time; Astra did so in 0% of cases.

Saturated means the benchmark stopped being a reason to buy
When OpenAI says FrontierMath Tier 4 and ARC-AGI-3 are saturated, it is saying something specific and worth taking literally: the tests have stopped discriminating. A benchmark earns its keep by separating models — by putting a gap between the best system and the next one, so that a buyer can read a number as a difference. Once every frontier model lands at or within noise of the ceiling, the score stops measuring the models and starts measuring the test's remaining headroom.
What that means for this page is that the 98% and the 99.9% should not be used to choose anything. They are not evidence that Astra is better than GPT-5.6 Sol or Claude Fable 5.1 on mathematics or abstract reasoning; they are evidence that all three have outgrown the tiers. The difference between 99.9% and 98% cannot be read as a 1.9-point advantage in any meaningful sense, because the run-to-run variation on evaluations of this size is larger than the gap. A vendor number at the ceiling is a statement about the ceiling.
They are not worthless. Saturation is genuine information about a tier: it tells you that a benchmark you may have used to compare models in 2025 will no longer sort them, and that you should stop weighting it in a procurement decision. It also tells you the interesting capability claim has moved somewhere else — the 96%-of-levels human action-efficiency figure is OpenAI's attempt to say something past the ceiling, and it is the sub-claim to argue with, not the 99.9%.
There is a second caution that applies to all three of the top rows. FrontierMath Tier 4 and ARC-AGI-3 are published in enough detail to be recognisable, but the harness, sampling and scoring configuration OpenAI used is not laid out in the material we read, and a 99.9% on a benchmark where the protocol is a private configuration is a number you can quote and cannot check. Where a score comes with no published methodology, say so and move on. That is what we are doing with these.
ExploitBench's 100% measures a capability most users cannot spend
A perfect score is a ceiling too, and this one has an unusual second half. Astra is the first model to reach the Critical cybersecurity capability threshold under OpenAI's Preparedness Framework, which is why its launch was gated at all. As shipped, the model refuses advanced cyber tasks such as writing proof-of-concept exploits; OpenAI says it plans to widen access with less restrictive safeguards through its Daybreak programme.
So ExploitBench is simultaneously the most impressive number on the list and the least usable one. It establishes that the capability exists and that OpenAI measured it and acted on the measurement — which is the honest reading of a Critical designation, and the reason the rollout was staggered rather than instant. It does not establish that you can do anything with that capability through the public API, because by design you mostly cannot. If offensive security is your use case, the relevant line in the documentation is not 100%, it is the refusal behaviour and the Daybreak programme's access path.
Terminal-Bench 4.0: a 24.6-point lead over Sol and a 2.1-point gap that settles nothing
Terminal-Bench 4.0 is where the vendor numbers start being useful, because the spread is wide enough to survive noise at one end and narrow enough to be uninformative at the other. Against GPT-5.6 Sol's 37.3%, Astra's 57.9% is a 24.6-point gap — large enough that harness differences are unlikely to explain it away entirely, though it is still one vendor measuring its own model and a predecessor it also built. Against Claude Fable 5.1's 55.8%, the 2.1-point lead is inside the range where we should say plainly that it settles nothing. Two models measured in two different harnesses, on a benchmark whose scoring tolerances are not published on the page we read, separated by two points, is a tie with extra steps. If your decision turns on that 2.1, you have not found a decision — you have found a marketing line.
The cost-per-task claim deserves its own sentence of suspicion. OpenAI estimates Astra at approximately 9% lower per-task API cost than Sol and 63% lower than Claude Fable 5.1 on this workload. Those are estimates, published without the task-level accounting behind them, and "cost per task" is not a property of a model — it is a property of a model, a harness, a token budget and a retry policy together. The direction is plausible: Fable 5.1 is a $10/$50 model like Astra, but a task that takes it more steps costs more. Treat the percentages as OpenAI's own accounting, not as a rate card.
DeepSWE v1.1: the independent row, and the tie inside it
The one figure not produced by OpenAI is the one most worth having — and it is also the one that most needs qualifying. On Datacurve's DeepSWE v1.1, a long-horizon software engineering benchmark of 113 tasks across 91 repositories in five languages, run through a single mini-swe-agent harness, GPT-6 Astra scored 74% ±3% Pass@1 at an average cost of $6.52 per task, producing 30k output tokens across 29 steps. Datacurve describes the result as a record.
Read the board, and the record is a tie. The same leaderboard snapshot we read on September 16, 2026 — dated September 3, 2026 — shows gemini-3.8-flash [high] also at 74%, with a tighter interval of ±1%, and claude-opus-5 [max] at 74% ±4%, with gpt-5.6-sol [max] one point back at 73% ±3%. Three models share the top score with overlapping confidence intervals, and the board itself notes that its top models cluster within a narrow band. Nothing on it marks a record holder. A record that three models hold simultaneously, within error bars, is a very different claim from "new record," and the honest version is: Astra reaches the top cluster on the strongest independent coding evaluation available, and does not separate from it.
What does separate, and what the score alone hides, is how it got there. Astra's 74% took 29 steps and 30k output tokens. The tied gemini-3.8-flash entry needed 166 steps and 143k output tokens; claude-opus-5 needed 99 steps and 118k. Same score, roughly a fifth of the work — that is a real and useful property for anyone paying per token or waiting on an agent, and Datacurve's numbers support it in a way OpenAI's vendor rows cannot. The cost line cuts the other way, though: at $6.52 per task Astra is cheapest of the three only if you ignore that gemini-3.8-flash reached the same score at $2.36. On this benchmark, Astra is efficient in steps and mid-priced in dollars. Take the tied score and the step count; do not take a "record."
OSWorld 2.0 and time per task: the number that decides whether an agent is usable
OpenAI reports 72.6% on OSWorld 2.0 at roughly 40 minutes per task, about 47% less time per task than GPT-5.6 Sol. This deserves more weight than its position in the list suggests, because for computer-use agents the score is only half the decision. A 72.6% completion rate is good; a 72.6% completion rate at 40 minutes per task is a different product from the same rate at 75 minutes, and the difference is not a percentage. It is how many tasks a team can push through in a working day, how much wall-clock a human waits while an agent works, and whether a single stuck task consumes an afternoon.
Two qualifications, both of which matter more than the headline. First, the figure is vendor-reported and we found no independent OSWorld 2.0 score for Astra to check it against — the independent index entry we read does not include OSWorld among its components, so there is no second measurement to set beside this one. Second, "roughly 40 minutes" is a mean, and means hide the part operators actually feel: the tail. Whether the slowest decile of tasks finishes in an hour or four is what determines whether an agent needs a human sitting next to it, and no vendor publishes that distribution. The Mind2Web claim — 1.9x faster task completion than the current GPT-5.6 Sol experience — has the same shape of problem, made slightly worse by the comparison being against a differently-harnessed Sol rather than the same scaffold with a different model in it. Faster is credible here; how much faster, on your workload, is unmeasured.

Safety evaluations are measurements too, and these ones are internal
The safety numbers belong in a benchmark reference rather than a footnote, because they are the same kind of claim as the performance rows: a measurement, made by an interested party, with a stated comparison. Astra produced unintended outcomes 89% less often than GPT-5.6 Sol and 74.7% less often than Claude Fable 5.1. On an evaluation modelled on the Hugging Face incident, GPT-5.6 Sol without production safeguards went beyond its authorised target in 48% of cases; Astra did so in 0%.
The direction is meaningful and the arithmetic is not symmetrical. One side of that second comparison is explicitly a model with its safeguards removed, and the other is Astra as shipped — so the 48-versus-0 gap is partly a measurement of guardrails rather than of the underlying model, and reading it as "Astra is safer than Sol" overstates what was tested. The 89% and 74.7% figures are internal to OpenAI with no external replication, on an evaluation whose construction we cannot inspect. All three point the same way and all three are the vendor's own; treat them as encouraging and unreproduced. They are also, for an enterprise buyer, the rows that most need to be true — which is exactly why they should be the rows you ask a vendor to evidence rather than the ones you accept from a launch page.
Price: $10 / $50, read September 16, 2026 — and the 2.5x that needs a denominator
A benchmark page that omits price is a page that stops halfway. Read from OpenAI's own pricing documentation on September 16, 2026, the standard short-context rate for gpt-6-astra is $10.00 per million input tokens, $1.00 per million cached input, and $50.00 per million output. Long-context requests are roughly double — around $20 input and $75 output per million. Below 1.05M tokens of context there is no long-context multiplier to plan for, but above the threshold the effective rate changes, and that is worth modelling before you assume the $10 figure applies to a large-codebase agent run.
The comparison everyone prints is Astra against GPT-5.6 Sol, and this is where an undated multiple goes wrong. Sol's rate on the same page, same day, is $4.00 input / $0.40 cached / $20.00 output — and OpenAI states that this is promotional pricing available at least through November 21, 2026. Against that promotional rate, Astra is 2.5x on both input and output. Against Sol's pre-promotion list of $5.00 / $0.50 / $30.00, Astra is 2x on input and about 1.67x on output. Both numbers are true; they are divided by different denominators. The 2.5x is what you pay today, the 2x and 1.67x are what you would pay if Sol's promotion lapses — and since OpenAI has published no post-promotion rate, no one can tell you which one your 2027 budget should use. If you see "2x" or "2.5x" without a named tier and a date, you are reading half a fact.
Two other pricing notes, because they get confused with list price. Endpoint quotes differ from OpenAI's own card: Azure endpoints have been listed at both $10/$50 and $11/$55, at least one endpoint has been listed at $20/$100, and a router listing shows $10/$50 with a batch tier at $5/$25. Those are quotes on separate endpoints, not OpenAI's list price, and they are not interchangeable. And for scale, on the same page and date: GPT-5.6 Terra is $2.00 / $0.20 / $12.00 and GPT-5.6 Luna is $0.20 / $0.02 / $1.20 — so Astra is not a model you route bulk traffic through by reflex. It is priced for the work the benchmark rows above describe: long-horizon, end-to-end tasks where 47% less wall-clock time and fewer agent steps can pay for a 2.5x token rate, and not for chat.
That is the arithmetic where a routing layer earns its place, and it is worth being concrete about why. GPT-6 Astra is in the OrcaRouter catalog as openai/gpt-6-astra, and the platform passes OpenAI's list rates through at 0% markup — so a vendor price change, including the promotional rate this section depends on, is live on our side the same day rather than on a repricing cycle. It also means the tier question above stops being hypothetical: you can put Astra on the long-horizon work and leave Sol, Terra or Luna on the volume, behind one key, without a second contract or a code change, and fail over between them when one is degraded.

Specs, the Codex change, and what OpenAI has not published
• Model id — gpt-6-astra in OpenAI's own documentation.
• Context and output — 1,050,000-token context window, 128,000 maximum output tokens.
• Knowledge cutoff — April 30, 2026. Reasoning token support: yes.
• Modalities — input listed as file, image and text. That particular listing comes from a router, not from OpenAI, and we are labelling it as router-reported rather than vendor-confirmed.
• Available in — ChatGPT Work, Codex, and the API, plus Microsoft Azure and AWS Bedrock. Enterprise workspaces are off by default; an administrator enables access under the applicable rate card, and gets controls to restrict approved websites and desktop applications, manage uploads and downloads, and control browsing history. ChatGPT Work and Codex add confirmation policies before consequential actions and automated review of unsafe or unauthorised tool calls. Four enterprise plugins shipped alongside in ChatGPT Desktop: Oracle Analytics, Power BI (a Microsoft Fabric service), Navan and Avalara. Zero Data Retention is available for eligible API customers on supported endpoints, subject to approval.
One mechanic is worth flagging because it changes how the model behaves on long jobs rather than how it scores. Codex gets a new context approach with Astra: notes persist across context windows instead of being repeatedly compacted into a single summary, and earlier context windows stay searchable, so requirements and test results from earlier messages and tool output remain findable. OpenAI calls it experimental, enabled in the Codex config.toml, and says it will become the default for Astra. On a benchmark measured in 29 steps, that is the mechanism most likely to explain the step count.
What is not published is as important as what is. OpenAI has not stated a parameter count or training compute for this model, and we are not printing one. The "more than 100,000 GPUs at Stargate" figure circulates second-hand and is not on the pages we read. Nor does OpenAI's documentation list a separately priced or separately documented "Astra Pro" or any other tier — coverage refers to one, but a router listing is not vendor documentation, and we are not presenting a tier we could not find in OpenAI's own docs.
What a reader should actually take from this
Sort the rows by how much they should move a decision. The saturated pair — 98% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3 — should move it not at all; they say the benchmarks are finished, not that Astra won them. ExploitBench's 100% is real and mostly unspendable through the public model by design, which is a deliberate safety choice rather than a limitation to route around. Terminal-Bench 4.0 is the strongest vendor performance claim, with a genuine gap over Sol and a 2.1-point gap over Claude Fable 5.1 that is too close to call. DeepSWE v1.1 is the only independently measured row, it puts Astra in a three-way tie at the top rather than alone, and its step and token counts are the part of that result worth quoting. OSWorld 2.0's 40 minutes per task is the most operationally meaningful number on the page and the least independently checked. The safety rows are internal, unreproduced, and point the right way.
Which leaves a straightforward split. If you are choosing a model this week for long-horizon agentic work — multi-hour coding, computer use, end-to-end tasks where a human would otherwise babysit — Astra is the model with the most current evidence behind it, and the cost case rests on time saved per task rather than tokens saved per call. If you are choosing for high-volume, short-interaction work, the 2.5x against Sol's promotional rate is the number that decides it, and the answer is probably no. And if you are choosing on the strength of a 98% or a 99.9%, you are choosing on the two rows that stopped carrying information.
Questions worth asking that this page has not answered
Is the 2.1-point Terminal-Bench lead over Claude Fable 5.1 real? Not on this evidence. Both figures come from OpenAI measuring its own model against a competitor, in harnesses we cannot compare, with no published scoring tolerance. It is a lead in OpenAI's own accounting and it is inside the range where a re-run could reverse it. The way to settle it is to run both on your own tasks, which is a few hours of work and worth more than the 2.1 points.
Why does Datacurve's board not crown Astra, when OpenAI's launch material quotes a record? Because the board is measuring, and the launch page is selling. Datacurve's own snapshot shows three models at 74% with overlapping confidence intervals and marks no leader. A record claim against a tied board is not a lie so much as a choice of denominator — the same failure mode as the 2.5x multiple, and the reason we have dated every number here.
If the top benchmarks are saturated, what should a buyer use instead? Time per task, cost per completed task, and step count on long-horizon work — the dimensions where the numbers are still spread out and still correlate with what a team pays. That is a harder measurement to find published, which is precisely why vendor launch pages still lead with the saturated ones.
The open question
The number that will date this page fastest is the price. Sol's promotional rate runs at least through November 21, 2026, and OpenAI has published no post-promotion rate; when that changes, the 2.5x in this article changes with it, in the direction of the 2x. The second is whether anyone reruns these evaluations independently at the same harness — DeepSWE is the model for how that should look, and OSWorld 2.0 and Terminal-Bench 4.0 are the two rows that most need the treatment. Until then, the honest summary of GPT-6 Astra's benchmarks is that the impressive numbers are saturated, the discriminating numbers are vendor-reported and close, and the one independent number is a tie that the vendor's own launch material calls a record.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
