A hero title card reading "Claude Opus 5.5 vs GPT-6 Astra" with the subtitle "A taste test that outran the benchmarks", an overline reading "Frontier flagship comparison - September 2026", three badges reading "Taste: not on any benchmark axis", "Research workloads: Astra holds" and "Price: $4.00 / $20.00 vs $10.00 / $50.00 per 1M", a footer line reading "Vendor launch tables and independent runs disagree on the size of the lead.", and the OrcaRouter logo bottom-right.
Guides & Insights

Claude Opus 5.5 vs GPT-6 Astra: A Taste Test That Outran the Benchmarks

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Someone asked Claude Opus 5.5 to make a short video about a rival lab solving Navier-Stokes, posted the result, and then wrote the sentence that makes this comparison worth doing: the model is "very nice," it has "great taste," and it is the first Claude since Opus 4.7 they actually want to run heavily — but they are still keeping GPT-6 Astra and GPT-5.6 Sol as their main drivers, because their work is research-heavy. That is three claims in one post, and they pull in different directions. A model can be the most pleasant thing to work with and still not be the one you route production traffic to. The interesting question is not which model is better. It is which of those two verdicts survives contact with the numbers, and which turns out to be a statement about taste.

The post is one practitioner's report, not a benchmark run — treat it as a signal about how the model feels in creative and open-ended work, not as evidence about capability. Everything numeric below is labeled with who produced it.

What a taste test can and cannot tell you

The artifact matters here. Making a video about a mathematical result is not a coding task with a pass/fail gate. There is no compile step, no test suite, no Terminal-Bench harness that scores whether the thing is any good. The model had to pick an angle, decide what to show, keep it coherent, and stop. When a practitioner says a model has "great taste," that is what they are describing: judgment about scope and emphasis that shows up in work where the specification is deliberately vague.

That is a real capability, and it is the one benchmark suites are worst at measuring. It is also why the same person can praise a model and not switch to it. Taste is what you want when you are exploring. It is not what you want when you have 200,000 lines of code to audit and a bill to justify.

Anthropic's own launch framing points at the same faculty, in prose rather than video: Claude Opus 5.5 is pitched as writing less "Claudish" — less throat-clearing, less hedging, more willingness to put the point first. Independent readability scoring backs the direction of that claim even where it disagrees on the size. The vendor also reports an internal fact-checking test passing 16 of 18 attempts, against zero of 18 for both Claude Fable 5.1 and Claude Opus 5; that number is Anthropic's own, unreproduced, and should be read as a claim rather than a result.

Two scoreboards, one disagreement that matters

A two-column scoreboard titled "Claude Opus 5.5 vs GPT-6 Astra - the scoreboard". The Claude Opus 5.5 column reads: Released Sep 22 2026, Price per 1M $4.00 / $20.00, Terminal-Bench 4.0 66.4% vendor and 59.6% independent, Terminal-Bench-Science 0.1 58.7%, GDPval-AA v2.1 1,846 Elo, Tokens per problem about 119,000. The GPT-6 Astra column reads: Released Sep 3 2026, Price per 1M $10.00 / $50.00, Terminal-Bench 4.0 57.9%, Terminal-Bench-Science 0.1 64.6%, GDPval-AA v2.1 1,542 Elo, Tokens per problem about 27,000. A footer reads "Opus 5.5 figures per Anthropic unless noted; Astra and independent figures per Artificial Analysis."

On raw published benchmarks, Claude Opus 5.5 leads GPT-6 Astra more often than not. The problem is that the two most-quoted sources do not agree on how much.

Anthropic's launch table puts Claude Opus 5.5 at 66.4% on Terminal-Bench 4.0, with GPT-6 Astra at 57.9% and Claude Fable 5.1 at 55.8%. That is a vendor-reported 8.5-point lead on agentic coding — the kind of margin that ends an argument in a marketing deck. Artificial Analysis, running its own harness, measured Claude Opus 5.5 at 59.6% on the same benchmark: level with GPT-6 Astra at xhigh effort, and about 11 points above Claude Opus 5. The lead is real in both readings. The size is not.

Where the two models trade places is more useful than where they do not:

• Terminal-Bench 4.0 (agentic coding) — Claude Opus 5.5 66.4% per Anthropic's own table, 59.6% per Artificial Analysis; GPT-6 Astra 57.9%.

• Humanity's Last Exam, with tools — Claude Opus 5.5 67.7% vendor-reported, independently measured at 61.4%; GPT-6 Astra 57.2%.

• GDPval-AA v2.1 (real-world knowledge work across 44 occupations) — Claude Opus 5.5 1,846 Elo; GPT-6 Astra 1,542 Elo. Both figures come from Artificial Analysis's independent board, not the vendor.

• Terminal-Bench-Science 0.1 — GPT-6 Astra 64.6% vs Claude Opus 5.5 58.7%. This is a clean Astra win, and it is the science-flavored agentic benchmark.

• AutomationBench — GPT-6 Astra 41.4% vs Claude Opus 5.5 40.0%. Narrow, but Astra again.

• FrontierCode v1.1 — Claude Opus 5.5 54.4% vs GPT-6 Astra 53.3%. Inside a point.

Read that list the way the practitioner would. The research-heavy workloads — the ones with a science or literature component, the ones that run a long tool-using loop and need the model to keep its footing — are exactly where GPT-6 Astra holds its ground. The knowledge-work and coding boards where Claude Opus 5.5 runs away are the ones you would point at if your job were shipping software. Both people in this conversation are reading their own benchmark correctly. They are just reading different benchmarks.

The effort setting is doing more work than the model is

There is a second, less flattering reason the headline margins are soft. The vendor comparisons are not run at the same setting.

Claude Opus 5.5's published figures come from an extra-high (xhigh) reasoning-effort configuration, against GPT-6 Astra at high. Artificial Analysis's independent runs put both models at xhigh and the coding gap disappears — the 8.5-point vendor lead becomes a tie. That is not an accusation of misconduct; it is what happens when a vendor picks the configuration that shows its model best and an independent lab picks the configuration that matches the competition. Before you move a workload on the strength of a benchmark table, check which effort setting it was run at, because the answer is frequently "the flattering one."

The token bill tells the same story from another angle. Anthropic's own launch material has Claude Opus 5.5 spending on the order of 119,000 tokens per problem with roughly 84,000 of those going to thinking, where GPT-6 Astra solves comparable problems in roughly 27,000 tokens. Thinking tokens are billed as output tokens. A model that uses four times the tokens at a fifth of the output price is close to a wash — and that is before accounting for the fact that the cheaper model is often the one you would have been willing to run at a higher effort setting anyway.

One further caveat sits on the Claude Opus 5.5 side specifically: Anthropic's safeguard routing can hand some cyber- and bio-adjacent requests to a more restrictive path, which pulls the published scores on those evaluations down relative to what the underlying model would produce. If your work touches those domains, the gap between the benchmark and your experience will be real, and it will be in the direction of the benchmark being optimistic.

What a real task costs on each

Screenshot of Anthropic's Claude Platform release notes page, with "September 22, 2026" selected in the sidebar. The entry announces the launch of Claude Opus 5.5 (claude-opus-5-5) with a 1M token context window by default, 128k max output tokens and always-on adaptive thinking, at $4 / $20 USD per MTok against Claude Opus 5 at $5 / $25, and notes that thinking cannot be disabled, that depth is controlled through the effort parameter, and that Fast mode is available as a research preview.

Claude Opus 5.5 is the cheaper model on the rate card, and it is not close:

• Claude Opus 5.5 — $4.00 per million input, $20.00 per million output, $0.20 per million cached input, $5.00 per million cache writes. 1M-token context, 128k max output.

• GPT-6 Astra — $10.00 per million input, $50.00 per million output, $1.00 per million cached input, $12.50 per million cache writes. Past 272,000 input tokens in a single request, the whole request reprices at $20.00 input and $100.00 output; the batch tier is $5.00/$25.00.

So Claude Opus 5.5 is 2.5 times cheaper on input and 2.5 times cheaper on output, with cached input five times cheaper. If your tasks are token-similar, that is the whole decision and the taste question is decoration.

The token-usage caveat above is what stops it from being that simple, and GPT-6 Astra's long-context repricing is the other trap. Cross 272,000 input tokens in one request and the entire request — not just the excess — moves to the long-context rate. A research workload that stuffs a large corpus into a single call is the exact shape that triggers it. Batch at half price is the legitimate escape hatch, and it is on the slower queue.

The honest summary is that the cost picture inverts depending on what you are doing. For a task where both models use comparable tokens, Claude Opus 5.5 wins on price by a wide margin. For a long agentic research loop where the token counts diverge the way Anthropic's own launch material shows, the two converge — which is a much less exciting headline than a 40% cost cut, and closer to what the practitioner was reporting.

Why the research-heavy user did not switch

Put the pieces together and the "still prefer Astra" line stops being a contradiction of the "great taste" line. It is the same observation from a different seat.

The workloads the user named are long, open-ended, tool-using, and expensive per run. On those, GPT-6 Astra holds the science and automation boards, uses a fraction of the thinking tokens, and — per independent measurement — sits level on coding once both models run at the same effort. Claude Opus 5.5's advantages concentrate in the work where you can see the output and judge it: prose, design choices, scoping an ambiguous brief, deciding what a video about Navier-Stokes should actually show. That is taste, and taste is exactly the part that does not appear on a benchmark axis.

If you were hoping for a verdict that one model wins, the evidence does not support it. The defensible version is narrower: Claude Opus 5.5 is the better model for work where judgment is the bottleneck, GPT-6 Astra is the better model for work where sustained long-context effort is the bottleneck, and the price gap between them narrows to near nothing on the second kind of work. That is a routing decision, not a conversion.

Running both without a migration

Screenshot of the OrcaRouter models catalogue filtered by the query "astra", showing one result card: OpenAI GPT-6 Astra, model id openai/gpt-6-astra, described as OpenAI's flagship model for demanding long-horizon end-to-end work, with a 1M context window and a per-1M-token base rate of $10.00 input, $50.00 output, $1.00 cache read and $12.50 cache write.

This is the part where the two models stop being an either/or. GPT-6 Astra is on OrcaRouter today at OpenAI's own list price — $10.00 input, $50.00 output, $1.00 cached read, $12.50 cache write — with 0% markup, provider list price passed straight through. That means a vendor rate change lands on our side the same day rather than whenever a reseller syncs.

Claude Opus 5.5 is reachable through Anthropic's own API and several third-party platforms; we are not listing it as something we serve, and you should be sceptical of anyone who claims otherwise before the model has propagated. What you can do today is put both models behind one key and one endpoint shape, and route per workload instead of per preference: send the open-ended, judgment-heavy request to Claude Opus 5.5 and the long research loop to GPT-6 Astra without maintaining two contracts or two client integrations.

Automatic failover is what makes that safe to test. A new model is not yet a production path, and routing through one endpoint means a provider incident or a rate limit moves the request rather than failing it. If you want to find out whether the taste gap is real on your own work, that is the cheap way to find out — run it beside what you already have rather than replacing it, and let your own suite decide instead of the launch tables.

The question the benchmarks cannot answer

Two things are worth watching, and neither is another benchmark point.

The first is whether the effort-setting asymmetry gets resolved. Right now, a vendor table at xhigh against a competitor at high produces an 8.5-point lead that independent testing at matched settings turns into a tie. As independent labs publish matched-setting runs on the science and automation boards where GPT-6 Astra currently leads, that picture can move in either direction — and it will move on evidence rather than on anyone's framing.

The second is the token-efficiency question. If Claude Opus 5.5's thinking overhead comes down in a point release, its 2.5x rate-card advantage stops being offset and the price argument wins outright. If it does not, then the practitioner who kept GPT-6 Astra as their main driver is not behind the news cycle. They read their own workload correctly, and the launch table simply was not measuring it.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily