
Claude Opus 5.5 vs GPT-6 Astra: The $6 Gap, and the Second Price Astra Charges Above 272K
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 177 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1323 tok/s
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 108 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
Two vendors published benchmark tables for their flagships this month, each one showing its own model winning, and neither table contains a single row where the two models were run against each other. Claude Opus 5.5, out September 22, 2026 at $4 per million input tokens, and GPT-6 Astra, out September 3, 2026 at $10 per million input tokens, have between them exactly two benchmarks that overlap — and they split them, with Opus 5.5 taking AutomationBench-adjacent work in Anthropic's table and Astra taking it in OpenAI's, and Astra clearly ahead on scientific reasoning in both. That is what the vendor material can tell you. The independent picture is less ambiguous and points the other way, and the pricing picture is more lopsided than either.
What each side published
Anthropic's table for Opus 5.5 uses its own harness and its own effort settings, and where it lists Astra the figures are Anthropic's, not a re-run. Treat the whole set as vendor-reported:
• Terminal-Bench 4.0 — Claude Opus 5.5 66.4% vs GPT-6 Astra 57.9% (Anthropic notes the Opus run used xhigh effort)
• FrontierCode v1.1 — Claude Opus 5.5 54.4% vs GPT-6 Astra 53.3%
• GDPval-AA v2.1, knowledge work — Claude Opus 5.5 1,846 Elo vs GPT-6 Astra 1,542
• AutomationBench — Claude Opus 5.5 40.0% vs GPT-6 Astra 41.4% (Astra ahead)
• Terminal-Bench-Science 0.1 — Claude Opus 5.5 58.7% vs GPT-6 Astra 64.6% (Astra ahead)
• Humanity's Last Exam, with tools — Claude Opus 5.5 67.7% vs GPT-6 Astra 57.2%
OpenAI's side of the aisle is harder to line up, because OpenAI published Astra against its own predecessor rather than against Anthropic's flagship, and the benchmarks it chose — computer use, front-end engineering, general knowledge — mostly do not appear in Anthropic's table at all. That is not dishonesty, it is normal launch behaviour, and it is exactly why a comparison built from two vendor tables is a comparison of two selections rather than of two models. The one thing both tables agree on is that Astra is strong at agentic science and computer use, and that the two models are close enough on coding that the choice of harness could move the result.
The independent read, which neither vendor chose

Artificial Analysis runs every model in one published configuration, so its numbers are the only ones here that were produced by the same evaluator under the same conditions:
• Intelligence Index — Claude Opus 5.5 58, ranked #1 of 210 vs GPT-6 Astra 53, ranked #6
• Configuration label — Claude Opus 5.5 "Adaptive Reasoning, Max Effort, Default Fallback", GPT-6 Astra "max"
• Output speed — GPT-6 Astra 52.5 tokens/sec, at the lower end of its price tier vs Claude Opus 5.5 not yet published
• Time to first token — GPT-6 Astra 352.15s vs a board median of 3.83s; Claude Opus 5.5 not yet published
• Price — Claude Opus 5.5 $4.00 in / $20.00 out per million vs GPT-6 Astra $10.00 / $50.00
• Context and output — GPT-6 Astra 1M context and 128K max output vs Claude Opus 5.5 1M and 128K
A five-point index gap is not decisive on its own, and it should not be read as one — both models land in the same capability band, which is what "frontier" means this quarter. What the Astra rows do establish is that the premium is not buying a capability lead on the one board both models appear on. The latency row is the honest complication: 352 seconds to first token is the highest in this group by a wide margin, though it is measured at max effort and a reasoning model that thinks before it emits will always look slow by this metric. The 52.5 tokens/sec figure is the more portable one, and it is on the low side for a model at $10/$50.
Astra's second price, above 272,000 tokens

The rate card is where these two stop being comparable at all, because Astra's pricing has a step in it. Standard rates are $10.00 input, $1.00 cached input, $12.50 cache write and $50.00 output per million tokens. Cross 272,000 input tokens, though, and the whole request reprices — not the excess, the entire call — at 2x input and cache rates and 1.5x output. That lands long-context work at roughly $20 input and $75 output per million tokens.
Claude Opus 5.5 does not do this. Anthropic bills $4.00 and $20.00 with the full 1M-token window at standard pricing, and states explicitly that a 900K-token request is billed at the same per-token rate as a 9K-token one. So the headline ratio between these two — Astra costing 2.5x Opus 5.5 on both directions — is not the ratio that applies to the workloads both models are actually sold for. On a long-horizon agent carrying a large working context, the comparison is $20/$75 against $4/$20, which is a factor of five on input and nearly four on output. Batch pricing compounds it rather than fixing it: Opus 5.5 halves to $2.00/$10.00 while Astra's flex tier halves to $5.00/$25.00 on a base that is already five times higher in the long-context case.
This is the single most decision-relevant asymmetry in the matchup, and it is invisible in every launch table from either company, because both vendors quote the headline rate. If your prompts sit under 272K tokens, ignore it. If you are building an agent that reads a repository or a document set, price it at $20/$75 before you compare anything.
Access is part of the specification
Astra is the first OpenAI model to cross the Critical cybersecurity capability threshold under the company's own Preparedness Framework, and the rollout reflects that classification rather than capacity. Access opened to a limited set of organisations first, enterprise access requires an administrator to switch it on before users see it, and the shipped model declines some categories of work outright. Claude Opus 5.5 arrives with a version of the same problem from the other direction: it ships with the safeguard tier Anthropic previously reserved for Claude Fable 5.1, which means most cybersecurity requests are re-routed to Claude Opus 4.8 and biology work falls back to Claude Opus 5 unless the account is verified.
Both behaviours show up in the same place, which is your evaluation. Artificial Analysis labels Opus 5.5's configuration "Default Fallback" precisely because requests can land on a different model than the one you named, and on security or life-sciences prompts that happens routinely. If your workload lives in those domains, the model string in your request is not a guarantee about the model that answered, and your quality numbers will include a smaller model on an unknown fraction of calls. Neither vendor is hiding this — both document it — but neither headline mentions it either.
Choosing between them
The decision this matchup actually poses is whether a long-context workload justifies Astra's step pricing, and for most teams the answer will be no: Opus 5.5 is cheaper on every line below 272K, dramatically cheaper above it, scores higher on the one independent board, and reaches 1M tokens of context without a surcharge. Astra's case is narrower and real — scientific reasoning, computer use, and the OpenAI ecosystem's tooling — and if your evals put it ahead on those, the premium is a line item rather than an error.
Where that gets practical is in how you run the two against each other, because a fair comparison needs both models on the same key and the same bill. Both are routable through OrcaRouter at provider list price with 0% markup — Claude Opus 5.5 at Anthropic's $4.00/$20.00 and GPT-6 Astra at $10.00/$50.00 — which means the 272K decision can be tested on your own traffic rather than argued from two press releases. Pinning Astra to the task classes where it wins and letting automatic failover hold the rest is a routing rule, and a request that crosses the 272K line can be sent down a cheaper path without a code change, because the rule lives in the router rather than in your application.

What would change the answer
One thing would: a controlled head-to-head at matched effort on shared benchmarks. Neither vendor has published one and neither has an incentive to, which leaves the independent index as the only same-conditions comparison available — and that index puts Opus 5.5 five points and five ranks above Astra at less than half the token price. The counter-argument that deserves watching is that both models were scored at max effort while Opus 5.5 ships defaulting to medium; if Astra's advantage on long-horizon scientific work survives a matched-effort run, the case for the premium gets stronger rather than weaker. Until someone runs it, the defensible position is that Astra is a specialist priced as a generalist, and Opus 5.5 is the generalist that happens to score higher.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
