
Claude Opus 5.5 vs GPT-5.6 Sol: Two Launch Tables That Barely Overlap
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Put Claude Opus 5.5 and GPT-5.6 Sol side by side this week and the first thing you notice is not a winner. It is that the two vendors published benchmark tables which hardly share a row. The launch material for Claude Opus 5.5, released September 22, 2026, puts it 29 points ahead of GPT-5.6 Sol on Terminal-Bench 4.0. The launch material for Sol, generally available since July 9, 2026, leads instead with Terminal-Bench 2.1, where Sol Ultra scored 91.9%, and the Artificial Analysis Coding Agent Index, where it took the top spot at 80. Both tables are real. Both were run by the vendor that wins them. The useful question is not which model is better in the abstract — it is which benchmark, run under whose harness, tells you anything about your workload. That is a harder question than either launch post wants you to ask, and it is the one this comparison is built around.
The two models, side by side

• Vendor — Anthropic vs OpenAI
• Released — Claude Opus 5.5 on September 22, 2026 vs GPT-5.6 Sol on July 9, 2026
• Price per million tokens — $4.00 input / $20.00 output vs $5.00 input / $30.00 output
• Cache read — $0.20 per million on Opus 5.5; Sol's cache rate is set per platform and is not published at a single comparable figure
• Context window — 1M tokens on both
• Max output — 128K tokens on both, with 300K on Anthropic's Batch API behind a beta header
• Knowledge cutoff — June 2026 for Opus 5.5 vs February 16, 2026 for Sol
• Reasoning control — adaptive thinking always on, default medium effort, scale runs low to max vs a max reasoning effort setting plus an Ultra mode that coordinates parallel sub-agents
• Lineup — Opus 5.5 is the first of a 5.5 family, with Sonnet 5.5 and Haiku 5.5 promised in coming weeks vs Sol as the flagship of a three-tier family alongside Terra and Luna
Two of those rows do more work than the rest. The knowledge cutoff gap is four months wide, which matters if your prompts touch anything that happened this spring. And Opus 5.5 now undercuts Sol on both input and output price — a reversal from three months ago, when Sol was the cheaper way to reach frontier-class output and Claude Opus 5 sat at $5/$25.
The only rows where both models actually appear
There is exactly one published table containing both models, and it is Anthropic's. Every figure in it is vendor-reported, produced by the company selling one of the two models:
• Terminal-Bench 4.0 — Claude Opus 5.5 66.4% vs GPT-5.6 Sol 37.3%
• Terminal-Bench-Science 0.1 — 58.7% vs 22.4%
• FrontierCode v1.1 (Main) — 54.4% vs 47.5%
• CursorBench 4.0 — 57.8% vs 41.7%
• AutomationBench — 40.0% vs 28.8%
• GDPval-AA v2.1 — 1846 Elo vs 1588
That is a clean sweep, and the margins on the two Terminal-Bench rows are large enough that they deserve scrutiny rather than acceptance. Anthropic's own footnote does some of that work for you: the Opus 5.5 Terminal-Bench run used xhigh effort rather than the default, and at default medium effort the FrontierCode advantage narrows to 54.6% against 47.5% — a seven-point gap instead of the headline numbers. Anthropic also attaches a general caution to the whole table, saying that at this capability level benchmark margins are "a less reliable guide to real-world differences" and that its internal gap to Claude Fable 5.1 is narrower than the scores suggest. A vendor discounting its own table is the most trustworthy sentence in it.
Note also what is missing from that table: any benchmark where OpenAI's model wins. A vendor table that contains only favourable rows is not evidence of a sweep; it is evidence of row selection.
What OpenAI's own numbers say
Sol's launch material tells a different story, on different benchmarks, at a different harness. Reported at its July launch:
• Terminal-Bench 2.1 — 91.9% for Sol Ultra and 88.8% for standard Sol, against Claude Mythos 5 at 88.0% and Claude Fable 5 at 84.3%
• Artificial Analysis Coding Agent Index — 80, described at the time as a state of the art 2.8 points above Claude Fable 5
• Agents' Last Exam, long-running professional workflows — 53.6, which OpenAI reported as 13.1 points ahead of Claude Fable 5
• BrowseComp — 92.2%; OSWorld 2.0 — 62.6%; SWE-Bench Pro — 64.6%
None of those rows include Claude Opus 5.5, because it did not exist in July. Sol's table was built to beat Fable 5 and Mythos 5, and it does. Whether it beats Opus 5.5 is simply not answered by either vendor's material — the two tables were assembled two months apart against different opponents.
The SWE-Bench Pro row is worth a specific note, because it is the one place OpenAI's launch material shows its model losing: 64.6% for Sol against 80% for Claude Fable 5. OpenAI responded by questioning the benchmark's validity, estimating that roughly 30% of its tasks are broken. That may well be true. It is also a useful reminder that "this benchmark is flawed" is a claim both labs make, and always about the benchmark where they lose.
The harness problem, which nobody's table solves

The reason the two tables do not reconcile is not that one vendor is lying. It is that agentic coding benchmarks measure a model plus a scaffold, and the scaffolds differ. Terminal-Bench 4.0 and Terminal-Bench 2.1 are different revisions of the same idea, run by different people, with different tool budgets and different retry rules. Anthropic's Opus 5.5 numbers were produced in Anthropic's harness; OpenAI's Sol numbers in Codex-style tooling. Move either model into the other's harness and the scores move.
Independent data, where it exists, describes a closer race than either table. One third-party agent benchmark from August 2026, run across multiple models on real application work, named GPT-5.6 Sol the best combination of accuracy, cost and speed — about $0.52 and five minutes per run — while Claude Opus 5, not 5.5, was the most accurate at 92% of runs. A separate leaderboard covering enterprise code tasks placed Claude Opus 5 first at 96.4% with GPT-5.6 Sol third at 86.5%, again comparing Sol to the previous Opus rather than the new one. Neither is a head-to-head with Opus 5.5, and neither settles anything. They do establish that when someone other than a vendor runs the comparison, the margins get small.
The practical takeaway is to stop reading the tables as a ranking and start reading them as a list of hypotheses. If your work is long-horizon terminal automation, Anthropic's Terminal-Bench gap is a hypothesis worth testing on your own tasks. If it is multi-step professional research, Sol's Agents' Last Exam result is the same kind of hypothesis. Neither number transfers to your workload without a run.
Cost per finished task, which is the only cost that matters
On the rate card, Claude Opus 5.5 is now cheaper in both directions: $4.00 against $5.00 on input, $20.00 against $30.00 on output, and a $0.20 cache-read rate that is 60% below what Claude Opus 5 charged. For an agent that re-sends a long system prompt on every turn, that cache line is the dominant cost and the reason a long-running session can get materially cheaper without any prompt change.
But per-token price has never decided an agent's bill. OpenAI's case for Sol at launch rested on token efficiency rather than sticker price — it reported a 54% improvement in coding token efficiency, with output tokens and elapsed time under half of Claude Fable 5's at roughly a third lower cost. Anthropic makes the mirror-image argument for Opus 5.5: about 40% lower cost to run on typical workloads than Claude Opus 5, on a rate card that only fell 20%, with customer statements from Box, Kiro, Factory and GitHub all citing large reductions in tokens or steps. Both claims point the same way — the model that finishes in fewer turns wins — and neither is independently audited.
Where independent cost-per-task math does exist, it currently favours Sol: one analysis published in September put GPT-5.6 Sol at $1.56 per solved issue on SWE-bench Verified at 96.2% success, against Claude Opus 4.8 at 88.6%. That is Sol measured against an older Anthropic model, not against Opus 5.5, so it is a starting point rather than a verdict — but it is a reminder that a model which solves the issue on the first attempt is cheaper than a cheaper model that needs three passes. The only measurement that settles this for you is your own task, run both ways, with the token counts in front of you.
That is a routing problem more than a procurement one, and it is worth being precise about which side of it is already solvable. GPT-5.6 Sol is on OrcaRouter today at OpenAI's own list price with 0% markup — provider list price passed straight through, so a vendor rate change is live on our side the same day rather than after a sync. Claude Opus 5.5 is not one of our routes yet; it is available through Anthropic's own API and the major clouds. Once it lands, the same key serves both, which is what turns the experiment into a routing rule instead of a second contract and a second SDK: send the workload to whichever model your own numbers favour, keep automatic failover pointed at the other, and let a routing-DSL line decide per request rather than per quarter. A price cut on either side is a config change, not a migration.

Where the two genuinely differ
Strip out the benchmarks and four differences survive, all of them checkable:
• Knowledge cutoff. June 2026 for Opus 5.5 against February 16, 2026 for Sol. Four months is a real gap for anything touching recent libraries, recent events or recently changed APIs, and no benchmark captures it.
• Safeguard routing. Opus 5.5 ships with Fable 5.1-class safeguards: most cybersecurity requests are re-routed to Claude Opus 4.8 and biology work falls back to Claude Opus 5 unless the account is verified. If your product sits in either domain, part of your traffic is being answered by a smaller model. Sol carries no equivalent documented routing, though OpenAI does note its own cyber-related safety evaluations.
• Orchestration. Sol ships an Ultra mode that coordinates multiple parallel sub-agents, four by default, and a max reasoning effort setting. Opus 5.5 has neither as a platform feature; its lever is the effort parameter, and multi-model composition is something you build or route rather than something the vendor hands you.
• Family economics. Sol is one of three tiers, with Terra and Luna below it — and Luna's price was cut 80% and Terra's 20% in a July 30 update. Anthropic's cheaper siblings, Sonnet 5.5 and Haiku 5.5, are announced but not shipped. If your workload is mixed, OpenAI currently offers the cheaper rungs today.
Choosing between them
Pick Claude Opus 5.5 if your work is long-running agentic coding or knowledge work, if the price per token matters, if your prompts touch anything from the last four months, or if a 1M-token context at $0.20 per million cached is what makes your architecture affordable. Start at medium effort, not the inherited high setting, and measure whether low holds up.
Pick GPT-5.6 Sol if you want a vendor-bundled orchestration mode, if you need a cheaper tier below the flagship today rather than in a few weeks, or if your workload is the kind that Sol's own launch evidence covers best — long-horizon professional workflows and computer use, where it reported its strongest results.
If you are already on Claude Opus 5, note that Opus 5.5 is not a drop-in: thinking can no longer be disabled, forced tool use returns an error, thinking blocks are bound to the model and conversation, and the older computer_20251124 computer-use tool is not accepted on the Claude API or Google Cloud. The first three also apply to Claude Fable 5.1. Switching to Sol is a different integration entirely.
What would actually settle this comparison is a single independent run of both models at matched effort, in a matched harness, on the same task set. As of this week nobody has published one. Until someone does, the honest position is the one the evidence supports: Anthropic's table favours Opus 5.5 and was produced by Anthropic; OpenAI's table favours Sol and was produced by OpenAI; the independent results that exist compare Sol to the previous Opus and describe a much closer race; and the two rows that will decide your bill — tokens per finished task and cache-read volume — are the two rows neither vendor publishes.
Compared in this article4
Detected from this article · Benchmarks: Artificial Analysis · updated daily
