Title card for the article, headed 'GPT-6 Astra Benchmarks' with the subtitle 'Every score, who measured it, and where the numbers stop meaning anything'. Three stat chips read 'FrontierMath Tier 4: 98% (saturated)', 'Terminal-Bench 4.0: 57.9% (vendor-reported)' and 'DeepSWE v1.1: 74% (independent, tied at top)'. A footer line reads 'Model released September 3, 2026. Figures re-read September 16, 2026.' The OrcaRouter logo sits in the bottom-right corner of the padded canvas.
Guides & Insights

GPT-6 Astra Benchmarks: Every Score, Who Measured It, and Where the Numbers Stop Meaning Anything

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GPT-6 Astra scores 98% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3, and Ope​nAI itself calls both benchmarks saturated — which makes the two most-quoted numbers on the model's page the two that tell you least about whether to use it. Against GPT-5.6 Sol and Claude Fable 5.1, the rows that actually discriminate are elsewhere: a 57.9% on Terminal-Bench 4.0, a 74% on DeepSWE v1.1 that is independent and also a tie, and a time-per-task figure that decides more about usability than any percentage here. The model itself is from September 3, 2026 — thirteen days old as we publish, not launch news. This is a reference page for what GPT-6 Astra scores, who ran each measurement, and what each number does and does not establish.

The reason a thirteen-day-old model is still worth a page this week is that the things a reader can act on arrived after the launch did. Broad availability completed: Ope​nAI said on September 8 that Astra was fully rolled out to Plus, Pro, Business and Enterprise users in Co​dex and ChatGPT Work, having opened on September 3 with the line that it was "rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the Ope​nAI API, Microsoft Azure, and AWS Bedrock." Enterprise access, off by default at launch, became something an administrator can turn on under their applicable rate card, and Ope​nAI's enterprise-work page for the model — published September 9 — shipped admin controls and four enterprise plugins alongside it. And the first independent evaluations of the model landed, which is what changes this from a scoreboard read off a vendor slide into something a buyer can weigh.

The full benchmark list, with the source on every row

Everything below is Ope​nAI-reported unless the row says otherwise. That distinction is not decoration: only one of these headline figures was produced by someone other than the company selling the model, and it is the one that turns out to be a tie.

FrontierMath Tier 4 — 98%. Vendor-reported by Ope​nAI, and described by Ope​nAI as saturating the tier.

ARC-AGI-3 — 99.9%. Vendor-reported, likewise described as saturated. Ope​nAI adds that Astra "surpassed our human action-efficiency baseline on 96% of levels, effectively reaching human parity on the benchmark" — a more specific and more interesting claim than the 99.9%, and still Ope​nAI's own measurement.

ExploitBench — 100%. Vendor-reported, and a perfect score.

Terminal-Bench 4.0 — 57.9%. Vendor-reported by Ope​nAI, against 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost per task respectively. The cost figures come with no published methodology in the material we read.

DeepSWE v1.1 — 74%. Measured by Datacurve, not Ope​nAI. This is the independent row.

OSWorld 2.0 — 72.6%. Vendor-reported, at roughly 40 minutes per task, which Ope​nAI puts at about 47% less time per task than GPT-5.6 Sol.

Mind2Web — 1.9x faster task completion than "the current GPT-5.6 Sol experience," with an updated Co​dex harness. Vendor-reported, and the comparison is against a differently-harnessed Sol rather than the same scaffold with a different model in it.

Safety — internal to Ope​nAI. Astra produced unintended outcomes 89% less often than GPT-5.6 Sol and 74.7% less often than Claude Fable 5.1. On an evaluation modelled on the Hugging Face incident, GPT-5.6 Sol without production safeguards went beyond its authorised target 48% of the time; Astra did so in 0% of cases.

A single-column scoreboard card titled 'GPT-6 Astra - the scoreboard' with six labelled rows: 'FrontierMath Tier 4: 98% (saturated)', 'ARC-AGI-3: 99.9% (saturated)', 'Terminal-Bench 4.0: 57.9% (vs GPT-5.6 Sol 37.3%)', 'DeepSWE v1.1: 74% (Datacurve, tied at top)', 'OSWorld 2.0: 72.6% at about 40 min per task', and 'List price: $10.00 in / $50.00 out per 1M'. The footer reads 'Vendor-reported by OpenAI except DeepSWE v1.1, measured by Datacurve. Read September 16, 2026.' The OrcaRouter logo is composited in the bottom-right.

Saturated means the benchmark stopped being a reason to buy

When Ope​nAI says FrontierMath Tier 4 and ARC-AGI-3 are saturated, it is saying something specific and worth taking literally: the tests have stopped discriminating. A benchmark earns its keep by separating models — by putting a gap between the best system and the next one, so that a buyer can read a number as a difference. Once every frontier model lands at or within noise of the ceiling, the score stops measuring the models and starts measuring the test's remaining headroom.

What that means for this page is that the 98% and the 99.9% should not be used to choose anything. They are not evidence that Astra is better than GPT-5.6 Sol or Claude Fable 5.1 on mathematics or abstract reasoning; they are evidence that all three have outgrown the tiers. The difference between 99.9% and 98% cannot be read as a 1.9-point advantage in any meaningful sense, because the run-to-run variation on evaluations of this size is larger than the gap. A vendor number at the ceiling is a statement about the ceiling.

They are not worthless. Saturation is genuine information about a tier: it tells you that a benchmark you may have used to compare models in 2025 will no longer sort them, and that you should stop weighting it in a procurement decision. It also tells you the interesting capability claim has moved somewhere else — the 96%-of-levels human action-efficiency figure is Ope​nAI's attempt to say something past the ceiling, and it is the sub-claim to argue with, not the 99.9%.

There is a second caution that applies to all three of the top rows. FrontierMath Tier 4 and ARC-AGI-3 are published in enough detail to be recognisable, but the harness, sampling and scoring configuration Ope​nAI used is not laid out in the material we read, and a 99.9% on a benchmark where the protocol is a private configuration is a number you can quote and cannot check. Where a score comes with no published methodology, say so and move on. That is what we are doing with these.

ExploitBench's 100% measures a capability most users cannot spend

A perfect score is a ceiling too, and this one has an unusual second half. Astra is the first model to reach the Critical cybersecurity capability threshold under Ope​nAI's Preparedness Framework, which is why its launch was gated at all. As shipped, the model refuses advanced cyber tasks such as writing proof-of-concept exploits; Ope​nAI says it plans to widen access with less restrictive safeguards through its Daybreak programme.

So ExploitBench is simultaneously the most impressive number on the list and the least usable one. It establishes that the capability exists and that Ope​nAI measured it and acted on the measurement — which is the honest reading of a Critical designation, and the reason the rollout was staggered rather than instant. It does not establish that you can do anything with that capability through the public API, because by design you mostly cannot. If offensive security is your use case, the relevant line in the documentation is not 100%, it is the refusal behaviour and the Daybreak programme's access path.

Terminal-Bench 4.0: a 24.6-point lead over Sol and a 2.1-point gap that settles nothing

Terminal-Bench 4.0 is where the vendor numbers start being useful, because the spread is wide enough to survive noise at one end and narrow enough to be uninformative at the other. Against GPT-5.6 Sol's 37.3%, Astra's 57.9% is a 24.6-point gap — large enough that harness differences are unlikely to explain it away entirely, though it is still one vendor measuring its own model and a predecessor it also built. Against Claude Fable 5.1's 55.8%, the 2.1-point lead is inside the range where we should say plainly that it settles nothing. Two models measured in two different harnesses, on a benchmark whose scoring tolerances are not published on the page we read, separated by two points, is a tie with extra steps. If your decision turns on that 2.1, you have not found a decision — you have found a marketing line.

The cost-per-task claim deserves its own sentence of suspicion. Ope​nAI estimates Astra at approximately 9% lower per-task API cost than Sol and 63% lower than Claude Fable 5.1 on this workload. Those are estimates, published without the task-level accounting behind them, and "cost per task" is not a property of a model — it is a property of a model, a harness, a token budget and a retry policy together. The direction is plausible: Fable 5.1 is a $10/$50 model like Astra, but a task that takes it more steps costs more. Treat the percentages as Ope​nAI's own accounting, not as a rate card.

DeepSWE v1.1: the independent row, and the tie inside it

The one figure not produced by Ope​nAI is the one most worth having — and it is also the one that most needs qualifying. On Datacurve's DeepSWE v1.1, a long-horizon software engineering benchmark of 113 tasks across 91 repositories in five languages, run through a single mini-swe-agent harness, GPT-6 Astra scored 74% ±3% Pass@1 at an average cost of $6.52 per task, producing 30k output tokens across 29 steps. Datacurve describes the result as a record.

Read the board, and the record is a tie. The same leaderboard snapshot we read on September 16, 2026 — dated September 3, 2026 — shows gemini-3.8-flash [high] also at 74%, with a tighter interval of ±1%, and claude-opus-5 [max] at 74% ±4%, with gpt-5.6-sol [max] one point back at 73% ±3%. Three models share the top score with overlapping confidence intervals, and the board itself notes that its top models cluster within a narrow band. Nothing on it marks a record holder. A record that three models hold simultaneously, within error bars, is a very different claim from "new record," and the honest version is: Astra reaches the top cluster on the strongest independent coding evaluation available, and does not separate from it.

What does separate, and what the score alone hides, is how it got there. Astra's 74% took 29 steps and 30k output tokens. The tied gemini-3.8-flash entry needed 166 steps and 143k output tokens; claude-opus-5 needed 99 steps and 118k. Same score, roughly a fifth of the work — that is a real and useful property for anyone paying per token or waiting on an agent, and Datacurve's numbers support it in a way Ope​nAI's vendor rows cannot. The cost line cuts the other way, though: at $6.52 per task Astra is cheapest of the three only if you ignore that gemini-3.8-flash reached the same score at $2.36. On this benchmark, Astra is efficient in steps and mid-priced in dollars. Take the tied score and the step count; do not take a "record."

OSWorld 2.0 and time per task: the number that decides whether an agent is usable

Ope​nAI reports 72.6% on OSWorld 2.0 at roughly 40 minutes per task, about 47% less time per task than GPT-5.6 Sol. This deserves more weight than its position in the list suggests, because for computer-use agents the score is only half the decision. A 72.6% completion rate is good; a 72.6% completion rate at 40 minutes per task is a different product from the same rate at 75 minutes, and the difference is not a percentage. It is how many tasks a team can push through in a working day, how much wall-clock a human waits while an agent works, and whether a single stuck task consumes an afternoon.

Two qualifications, both of which matter more than the headline. First, the figure is vendor-reported and we found no independent OSWorld 2.0 score for Astra to check it against — the independent index entry we read does not include OSWorld among its components, so there is no second measurement to set beside this one. Second, "roughly 40 minutes" is a mean, and means hide the part operators actually feel: the tail. Whether the slowest decile of tasks finishes in an hour or four is what determines whether an agent needs a human sitting next to it, and no vendor publishes that distribution. The Mind2Web claim — 1.9x faster task completion than the current GPT-5.6 Sol experience — has the same shape of problem, made slightly worse by the comparison being against a differently-harnessed Sol rather than the same scaffold with a different model in it. Faster is credible here; how much faster, on your workload, is unmeasured.

Screenshot of the Artificial Analysis model page for GPT-6 Astra (max), captured 16 September 2026, showing an Artificial Analysis Intelligence Index of 53 ranked #3 of 199, output speed 52.9 tokens per second ranked #111 of 199, a cost of $3.26 per Intelligence Index task ranked #71 of 199, $10.00 input and $50.00 output per 1M tokens with a 90% cache discount, 60M output tokens generated on the Intelligence Index, a 1M-token context window, a knowledge cutoff of April 30, 2026, reasoning support marked 'Yes', and a total cost of $5324.10 to run the full index.

Safety evaluations are measurements too, and these ones are internal

The safety numbers belong in a benchmark reference rather than a footnote, because they are the same kind of claim as the performance rows: a measurement, made by an interested party, with a stated comparison. Astra produced unintended outcomes 89% less often than GPT-5.6 Sol and 74.7% less often than Claude Fable 5.1. On an evaluation modelled on the Hugging Face incident, GPT-5.6 Sol without production safeguards went beyond its authorised target in 48% of cases; Astra did so in 0%.

The direction is meaningful and the arithmetic is not symmetrical. One side of that second comparison is explicitly a model with its safeguards removed, and the other is Astra as shipped — so the 48-versus-0 gap is partly a measurement of guardrails rather than of the underlying model, and reading it as "Astra is safer than Sol" overstates what was tested. The 89% and 74.7% figures are internal to Ope​nAI with no external replication, on an evaluation whose construction we cannot inspect. All three point the same way and all three are the vendor's own; treat them as encouraging and unreproduced. They are also, for an enterprise buyer, the rows that most need to be true — which is exactly why they should be the rows you ask a vendor to evidence rather than the ones you accept from a launch page.

Price: $10 / $50, read September 16, 2026 — and the 2.5x that needs a denominator

A benchmark page that omits price is a page that stops halfway. Read from Ope​nAI's own pricing documentation on September 16, 2026, the standard short-context rate for gpt-6-astra is $10.00 per million input tokens, $1.00 per million cached input, and $50.00 per million output. Long-context requests are roughly double — around $20 input and $75 output per million. Below 1.05M tokens of context there is no long-context multiplier to plan for, but above the threshold the effective rate changes, and that is worth modelling before you assume the $10 figure applies to a large-codebase agent run.

The comparison everyone prints is Astra against GPT-5.6 Sol, and this is where an undated multiple goes wrong. Sol's rate on the same page, same day, is $4.00 input / $0.40 cached / $20.00 output — and Ope​nAI states that this is promotional pricing available at least through November 21, 2026. Against that promotional rate, Astra is 2.5x on both input and output. Against Sol's pre-promotion list of $5.00 / $0.50 / $30.00, Astra is 2x on input and about 1.67x on output. Both numbers are true; they are divided by different denominators. The 2.5x is what you pay today, the 2x and 1.67x are what you would pay if Sol's promotion lapses — and since Ope​nAI has published no post-promotion rate, no one can tell you which one your 2027 budget should use. If you see "2x" or "2.5x" without a named tier and a date, you are reading half a fact.

Two other pricing notes, because they get confused with list price. Endpoint quotes differ from Ope​nAI's own card: Azure endpoints have been listed at both $10/$50 and $11/$55, at least one endpoint has been listed at $20/$100, and a router listing shows $10/$50 with a batch tier at $5/$25. Those are quotes on separate endpoints, not Ope​nAI's list price, and they are not interchangeable. And for scale, on the same page and date: GPT-5.6 Terra is $2.00 / $0.20 / $12.00 and GPT-5.6 Luna is $0.20 / $0.02 / $1.20 — so Astra is not a model you route bulk traffic through by reflex. It is priced for the work the benchmark rows above describe: long-horizon, end-to-end tasks where 47% less wall-clock time and fewer agent steps can pay for a 2.5x token rate, and not for chat.

That is the arithmetic where a routing layer earns its place, and it is worth being concrete about why. GPT-6 Astra is in the OrcaRouter catalog as openai/gpt-6-astra, and the platform passes Ope​nAI's list rates through at 0% markup — so a vendor price change, including the promotional rate this section depends on, is live on our side the same day rather than on a repricing cycle. It also means the tier question above stops being hypothetical: you can put Astra on the long-horizon work and leave Sol, Terra or Luna on the volume, behind one key, without a second contract or a code change, and fail over between them when one is degraded.

Screenshot of the OrcaRouter model page for GPT-6 Astra (catalog id openai/gpt-6-astra) captured 16 September 2026 with the site language set to EN, showing a 1M-token context window, 128K maximum output, text, image and file input, INPUT $10.00 and OUTPUT $50.00 per 1M tokens, p50 TTFT 6.70 s, p95 TTFT 10.00 s, 181.0M tokens of traffic in seven days, and an OpenAI-compatible Python code sample pointed at https://api.orcarouter.ai/v1 using the model id openai/gpt-6-astra.

Specs, the Co​dex change, and what Ope​nAI has not published

Model id — gpt-6-astra in Ope​nAI's own documentation.

Context and output — 1,050,000-token context window, 128,000 maximum output tokens.

Knowledge cutoff — April 30, 2026. Reasoning token support: yes.

Modalities — input listed as file, image and text. That particular listing comes from a router, not from Ope​nAI, and we are labelling it as router-reported rather than vendor-confirmed.

Available in — ChatGPT Work, Co​dex, and the API, plus Microsoft Azure and AWS Bedrock. Enterprise workspaces are off by default; an administrator enables access under the applicable rate card, and gets controls to restrict approved websites and desktop applications, manage uploads and downloads, and control browsing history. ChatGPT Work and Co​dex add confirmation policies before consequential actions and automated review of unsafe or unauthorised tool calls. Four enterprise plugins shipped alongside in ChatGPT Desktop: Oracle Analytics, Power BI (a Microsoft Fabric service), Navan and Avalara. Zero Data Retention is available for eligible API customers on supported endpoints, subject to approval.

One mechanic is worth flagging because it changes how the model behaves on long jobs rather than how it scores. Co​dex gets a new context approach with Astra: notes persist across context windows instead of being repeatedly compacted into a single summary, and earlier context windows stay searchable, so requirements and test results from earlier messages and tool output remain findable. Ope​nAI calls it experimental, enabled in the Co​dex config.toml, and says it will become the default for Astra. On a benchmark measured in 29 steps, that is the mechanism most likely to explain the step count.

What is not published is as important as what is. Ope​nAI has not stated a parameter count or training compute for this model, and we are not printing one. The "more than 100,000 GPUs at Stargate" figure circulates second-hand and is not on the pages we read. Nor does Ope​nAI's documentation list a separately priced or separately documented "Astra Pro" or any other tier — coverage refers to one, but a router listing is not vendor documentation, and we are not presenting a tier we could not find in Ope​nAI's own docs.

What a reader should actually take from this

Sort the rows by how much they should move a decision. The saturated pair — 98% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3 — should move it not at all; they say the benchmarks are finished, not that Astra won them. ExploitBench's 100% is real and mostly unspendable through the public model by design, which is a deliberate safety choice rather than a limitation to route around. Terminal-Bench 4.0 is the strongest vendor performance claim, with a genuine gap over Sol and a 2.1-point gap over Claude Fable 5.1 that is too close to call. DeepSWE v1.1 is the only independently measured row, it puts Astra in a three-way tie at the top rather than alone, and its step and token counts are the part of that result worth quoting. OSWorld 2.0's 40 minutes per task is the most operationally meaningful number on the page and the least independently checked. The safety rows are internal, unreproduced, and point the right way.

Which leaves a straightforward split. If you are choosing a model this week for long-horizon agentic work — multi-hour coding, computer use, end-to-end tasks where a human would otherwise babysit — Astra is the model with the most current evidence behind it, and the cost case rests on time saved per task rather than tokens saved per call. If you are choosing for high-volume, short-interaction work, the 2.5x against Sol's promotional rate is the number that decides it, and the answer is probably no. And if you are choosing on the strength of a 98% or a 99.9%, you are choosing on the two rows that stopped carrying information.

Questions worth asking that this page has not answered

Is the 2.1-point Terminal-Bench lead over Claude Fable 5.1 real? Not on this evidence. Both figures come from Ope​nAI measuring its own model against a competitor, in harnesses we cannot compare, with no published scoring tolerance. It is a lead in Ope​nAI's own accounting and it is inside the range where a re-run could reverse it. The way to settle it is to run both on your own tasks, which is a few hours of work and worth more than the 2.1 points.

Why does Datacurve's board not crown Astra, when Ope​nAI's launch material quotes a record? Because the board is measuring, and the launch page is selling. Datacurve's own snapshot shows three models at 74% with overlapping confidence intervals and marks no leader. A record claim against a tied board is not a lie so much as a choice of denominator — the same failure mode as the 2.5x multiple, and the reason we have dated every number here.

If the top benchmarks are saturated, what should a buyer use instead? Time per task, cost per completed task, and step count on long-horizon work — the dimensions where the numbers are still spread out and still correlate with what a team pays. That is a harder measurement to find published, which is precisely why vendor launch pages still lead with the saturated ones.

The open question

The number that will date this page fastest is the price. Sol's promotional rate runs at least through November 21, 2026, and Ope​nAI has published no post-promotion rate; when that changes, the 2.5x in this article changes with it, in the direction of the 2x. The second is whether anyone reruns these evaluations independently at the same harness — DeepSWE is the model for how that should look, and OSWorld 2.0 and Terminal-Bench 4.0 are the two rows that most need the treatment. Until then, the honest summary of GPT-6 Astra's benchmarks is that the impressive numbers are saturated, the discriminating numbers are vendor-reported and close, and the one independent number is a tie that the vendor's own launch material calls a record.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily