
GPT-6 vs Claude Opus 5: Both Sides of This Matchup Have Already Been Replaced
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 60 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 356 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The most useful thing to know about GPT-6 versus Claude Opus 5 is that both halves of the matchup are one generation behind their own vendors. Claude Opus 5 shipped on 24 July 2026 and was replaced by Claude Opus 5.5 on 22 September. GPT-6 Sol shipped on 22 September and was replaced by GPT-6.1 Sol on 29 September. If you searched for this comparison because you are choosing between them for new work, you are choosing between two models that each vendor has already moved past — and the honest version of this article is about when the older pair is still the right call, not about who wins.
What the two models actually are
Claude Opus 5 was Anthropic's flagship: a 1,000,000-token context window, 128,000 tokens of maximum output, text, image and file input, listed at $5.00 per million input tokens and $25.00 per million output tokens, with a $0.50 cached-input rate and a $10.00 cache-write rate. It is a single model with a single id, anthropic/claude-opus-5.
GPT-6 Sol is the middle tier of a three-model generation — GPT-6 Astra above it and GPT-6 Luna below it — with the same 1,000,000-plus-token context (the catalogue lists 1,050,000 for the OpenAI side against Opus 5's 1,000,000), the same 128,000-token output ceiling, and the same text/image/file input. It lists at $2.00 / $10.00, with a $0.20 cached-input rate and a $2.50 cache-write rate. Its rate card carries a step the Anthropic card does not: any request above 272,000 input tokens reprices the whole request to $4.00 / $15.00.
That is a 2.5x price gap per token in Sol's favour at the short end, and the gap holds on cache too — a 90% cached-input discount on Sol against a 98% discount on Opus 5, which narrows the difference to $0.20 versus $0.50 per million rather than eliminating it. On a long-context workload the gap halves rather than closes.
• Input — GPT-6 Sol $2.00 per million vs Claude Opus 5 $5.00
• Output — GPT-6 Sol $10.00 per million vs Claude Opus 5 $25.00
• Cached input — GPT-6 Sol $0.20 per million vs Claude Opus 5 $0.50
• Cache writes — GPT-6 Sol $2.50 per million vs Claude Opus 5 $10.00
• Above 272K input — GPT-6 Sol reprices the whole request to $4.00 / $15.00; no equivalent step on the Opus 5 card
• Context and output — 1,050,000 in and 128,000 out vs 1,000,000 in and 128,000 out
The independent board, where the gap is larger than the price
Price favours GPT-6 Sol by 2.5x. The board favours Claude Opus 5 by more than the price gap would suggest, which is the opposite of the usual cheap-model story.
• AA Intelligence Index — Claude Opus 5 50.8, ranked 11th; GPT-6 Sol 47.6, ranked 14th
• AA Coding Index — Claude Opus 5 78.0, ranked 2nd on that board; GPT-6 Sol 60.0
• GPQA Diamond — Claude Opus 5 93.2
• Humanity's Last Exam — Claude Opus 5 54.9 vs GPT-6 Sol 47.9
• Terminal-Bench 2.1 — Claude Opus 5 89.1
• Terminal-Bench 4.0 — Claude Opus 5 49.0 vs GPT-6 Sol 43.9
• τ-bench banking — Claude Opus 5 42.1
• SciCode — Claude Opus 5 56.4 vs GPT-6 Sol 57.6
• Long-context recall — Claude Opus 5 79.3 vs GPT-6 Sol 83.7
Two rows run the other way and they are worth naming, because the aggregate hides them. GPT-6 Sol beats Opus 5 on long-context recall by 4.4 points and edges it on SciCode by 1.2. So the honest shape is not "Opus 5 is better at everything" — it is that Opus 5 is decisively ahead on the coding and agentic boards (a 24-point Coding Index lead is a different class of result) while Sol holds the retrieval-over-a-long-window row and ties on one of the two science-code measures.
One methodological caveat before either number is quoted anywhere else: Artificial Analysis does not guarantee that two model pages are rendered against the same index revision on the same day, and it splits its peer groups — open-weights models are ranked against other open-weights models of the same size class, proprietary models across a proprietary price band. Both models here are proprietary, so the two figures come from the same side of that line, but treat the sub-point differences as direction rather than as a controlled measurement. Every vendor figure in this piece is the vendor benchmarking its own model against its own predecessor and is labelled as such.
What the two replacements changed
Claude Opus 5.5 arrived on 22 September at $4.00 / $20.00 — cheaper than Opus 5 on both meters, with a $0.20 cached-input rate that matches Sol's — and scores 57.6 on the same Intelligence Index. That is 6.8 points above Opus 5 for 20% less money. Any argument for keeping Opus 5 in a new build has to explain why it is worth 6.8 index points and $1.00/$5.00 per million to stay on the older checkpoint.
GPT-6.1 Sol arrived on 29 September at the same $2.00 / $10.00 as GPT-6 Sol, scoring 51.8 on the index against Sol's 47.6, with a $0.10 cached-input rate that halves the cached read. It is the same price with a better score and better cache economics. The one argument for GPT-6 Sol specifically is the same one that applies to Opus 5: a regression suite that was baselined against it, where changing the model invalidates the comparisons you already trust.
There is a second, quieter reason a team stays on the older pair, and it is not about benchmarks. Both of these are the checkpoints with the longest production history behind them. Opus 5 has been callable since 24 July and GPT-6 Sol since 22 September, and the independent evaluator's measurements on both have been running long enough to show a shape rather than a single reading. Sol's last seven days of traffic on our own routing data come to about 2.1 million tokens; Opus 5's come to about 23.0 million. Those are a rolling window, not a lifetime total, and they move between reads — but they are a reminder that Opus 5 is still doing real work in production even after its replacement shipped.
The latency and cost shape, which is where Sol earns its keep
• Median time to first token — GPT-6 Sol 6,785 ms vs Claude Opus 5 4,652 ms
• Median output speed — GPT-6 Sol 179.6 tokens/s vs Claude Opus 5 82.2 tokens/s
• Seven-day traffic observed on our side — GPT-6 Sol ~2.1M tokens, Claude Opus 5 ~23.0M tokens
Sol answers its first token roughly 2.1 seconds later but then streams at more than twice Opus 5's rate. On a long generation — the kind of run where the output is thousands of tokens rather than dozens — that second number dominates, and a cheap model that finishes sooner is worth more than its rate card suggests. On a short, latency-sensitive call, the first-token figure is the one the user feels.
Both of these are serving measurements from our own routing, taken over a rolling seven-day window, and they drift between reads. They are useful for comparing the shape of the two models and not for quoting as a service-level expectation.

Running the pair while the migration is undecided
The reason this comparison is worth answering at all is that neither replacement is a drop-in decision. If your evaluations were built against Claude Opus 5, moving to Opus 5.5 is a re-baseline, not an upgrade; the same is true of GPT-6 Sol against GPT-6.1 Sol. The useful thing to do in that window is run both generations side by side on your own traffic instead of choosing from someone else's board.
All four models sit in the same catalogue on OrcaRouter — Claude Opus 5, Claude Opus 5.5, GPT-6 Sol and GPT-6.1 Sol, called through one API in front of 200+ models, with the provider's list price passed through at 0% markup so a vendor price cut lands here the same day rather than at your next reconciliation — which makes the comparison a routing question rather than a procurement one. The routing DSL lets you split by request size or by share: send a slice of production traffic to the newer checkpoint, compare it against the baseline on your own evaluations, widen the slice when the numbers hold, and keep the older model as the fallback until they do. Automatic failover across providers covers the case where a route degrades mid-migration, which is the failure mode that makes people abandon a migration halfway.


Who should pick which, given that both are superseded
If you are starting something new: neither. Claude Opus 5.5 is cheaper than Claude Opus 5 and scores 6.8 points higher, and GPT-6.1 Sol is the same price as GPT-6 Sol with a better index and a halved cached read. The only reason to begin on the older pair is a hard dependency on a checkpoint you cannot change.
If you are already running one of them: the decision is about your regression suite, not about the boards. Claude Opus 5 remains the stronger coding and agentic model of the two by a wide margin, so a coding pipeline that is stable on it should compare against Opus 5.5 rather than migrating sideways to a GPT-6 tier. A high-volume, cost-sensitive pipeline on GPT-6 Sol has the easier upgrade — same price, better score, better cache — and the honest reason to delay is only that you have not re-run your evaluations yet.
And if the answer you wanted was simply which of the two wins: Claude Opus 5 does, on almost every board, for 2.5x the money. GPT-6 Sol's case was never that it was better. It was that it was cheap enough to be worth the difference — and that case now belongs to GPT-6.1 Sol at the same price with a better score.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
