A generated hero card for GPT-6.1 Sol vs Claude Opus 5.5 headed 'A real gap, unevenly distributed', with two info cards reading 'GPT-6.1 Sol: Intelligence Index 51.8, $0.72 per index task, long-context recall 83.0%' and 'Claude Opus 5.5: Intelligence Index 57.6, $5.98 per index task, long-context recall 84.7%', a footer reading 'Figures per Artificial Analysis, Intelligence Index v4.3.2, both models charted at their maximum effort setting.', and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

GPT-6.1 Sol vs Claude Opus 5.5: A Real Gap, Unevenly Distributed

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The five-point-eight-point gap between Claude Opus 5.5 and GPT-6.1 Sol on the Artificial Analysis Intelligence Index is the largest of any pairing in this set, and unlike most of them it will not close with a setting change. Claude Opus 5.5, released September 22, 2026 and billed at $4.00 per million input and $20.00 per million output tokens, scores 57.6. GPT-6.1 Sol, released September 29, 2026 at $2.00 and $10.00, scores 51.8. Claude Opus 5.5 is the higher-scoring model by a margin that is bigger than the one separating most frontier models from each other, and it costs 8.3× more per task to get there — $5.98 against $0.72. So far the expensive model is simply winning, which is the least interesting version of this comparison. What makes it worth reading is that the 5.8 points are not spread evenly across the suite: on the one axis that decides most long-document and long-session work, the two models are within two points of each other.

Both are flagship models in the same market slot: 1M-class context windows, 128K output ceilings, text, image and file input, extended thinking on one side and a five-level effort ladder on the other. Claude Opus 5.5 succeeds Claude Opus 5 and is positioned for demanding reasoning, coding and long-horizon agentic work; GPT-6.1 Sol sits one tier below GPT-6 Astra and is positioned at agentic coding, computer use and document-heavy professional work. They are aimed at the same jobs. The measurement says they are not equally good at all of them.

Where the 5.8 points actually live

An index composite hides its components. Pull the individual evaluations and the gap stops looking like a gap and starts looking like a profile.

• Humanity's Last Exam — Claude Opus 5.5 61.4% vs GPT-6.1 Sol 52.9%, an 8.5-point spread on abstract reasoning

• SciCode — Claude Opus 5.5 66.9% vs GPT-6.1 Sol 54.2%, a 12.7-point spread on research-grade code

• Long-context recall — Claude Opus 5.5 84.7% vs GPT-6.1 Sol 83.0%, a 1.7-point spread on retrieving facts from a long input

• Terminal-Bench 4.0 — Claude Opus 5.5 59.6% against a Sol figure not published at the same revision

• Context window — GPT-6.1 Sol 1,050,000 tokens vs Claude Opus 5.5 1,000,000 tokens, with a 128,000-token output ceiling on both

• Cost per index task — Claude Opus 5.5 $5.98 vs GPT-6.1 Sol $0.72

The distribution is the story. Where the task is reasoning in the abstract — a hard exam question, a research-grade coding problem with no existing answer to copy — Claude Opus 5.5 is decisively ahead, and a 12.7-point SciCode gap is not noise. Where the task is finding the right passage in a very long input and using it, the two models are effectively level, and GPT-6.1 Sol is holding that position with 50,000 more tokens of capacity. A large share of production LLM work is the second kind: retrieval-augmented answering, contract and filing review, log and transcript analysis, code review over a large repository. That work is measured by the third bullet, not the first two, and on that bullet the premium buys 1.7 points.

Reading the index rows honestly

Two caveats belong next to those numbers rather than at the bottom of the page.

The configurations differ. Artificial Analysis charts Claude Opus 5.5 in a "Max, Default Fallback" configuration and GPT-6.1 Sol at max effort. Both are the strongest settings each model is published under, which makes the comparison fair, but it is a comparison of two ceilings rather than two shipping defaults. Claude Opus 5.5's own defaults in a hosted API may differ from the charted run, and GPT-6.1 Sol's default effort is medium.

The Terminal-Bench row is deliberately blank on one side. Claude Opus 5.5's 59.6% comes from the Artificial Analysis revision in which it was published; the equivalent GPT-6.1 Sol figure is not available at a matching revision. Quoting a number from a different run of the same benchmark would be a false comparison, and the honest move is to leave the cell empty rather than fill it with something approximately right.

Verbosity cuts in the same direction here as it did against the cheaper siblings: Claude Opus 5.5 emits 260M output tokens on the index suite against GPT-6.1 Sol's 67M, a 3.9× difference at a rate that is already twice as high. That is where an 8.3× per-task ratio comes from when the rate card shows 2× on input and 2× on output. The premium is not 8.3× because the model costs 8.3× — it is 8.3× because it costs 2× and thinks 3.9× as long.

The case for paying it

If your work resembles the first two bullets, the calculation is straightforward and it favours Claude Opus 5.5. A 12.7-point gap on research-grade scientific code, or 8.5 points on hard abstract reasoning, is the difference between an agent that closes a ticket unattended and one that needs a human on every third step. Agentic coding over a large codebase is the workload both vendors name first, and it is the workload where the spread is widest. Paying $5.98 instead of $0.72 per task to avoid a failed run that costs an engineer twenty minutes is not a close call.

There is a second argument that is not about scores at all. Claude Opus 5.5 is fifteen days old and has generated a body of third-party evaluation, review and deployment experience that a model eight days old does not have. For a production path, maturity of the surrounding knowledge is worth something the index does not price.

The case against, and the number that decides it

The counter-argument is the long-context row. If the dominant shape of your traffic is "large input, short answer" — and for a great many teams it is — then the axis you are being charged for is not the axis you consume. On long-context recall the two models are 1.7 points apart at an 8.3× price ratio, and the cheaper model carries the larger window. That is the case where the premium is real, measured, and spent on capability the workload never exercises.

The deciding number, then, is not the index. It is the share of your tasks that look like SciCode rather than like recall. Teams that can answer that question from their own logs should answer it from their own logs; a benchmark suite is a proxy for a mix of jobs, and the mix is the variable.

Both on one key, which is what makes the question answerable

A screenshot of the OrcaRouter model page for anthropic/claude-opus-5.5, credited to Anthropic and dated 2026-09-22, showing text, image and file input with text output, a 1,000,000-token context window reading 1M tokens with a 128K maximum output, and a rate row reading $4.00 input and $20.00 output per million tokens alongside a 3.68 s median time to first token and a 36.5M seven-day traffic figure.A screenshot of OpenAI's developer model page for GPT-6.1 Sol, headed 'GPT-6.1 Sol' with the tagline 'Near-Astra performance for complex work at a lower cost', showing a 1,050,000-token context window with 128,000 max output tokens and an Apr 30, 2026 knowledge cutoff, a price row reading $2 input and $10 output per million tokens, and the note that reasoning.effort supports low, medium (default), high, xhigh and max while none and minimal are not supported.

Claude Opus 5.5 is on OrcaRouter at its own $4.00 / $20.00 with cache reads at $0.20. GPT-6.1 Sol is on OrcaRouter at $2.00 / $10.00, stepping to $4.00 / $15.00 above 272,000 prompt tokens, with cache reads at $0.10. Both at provider list price, nothing added on top, on one key and one bill.

That matters more than usual for this pairing, because the two models are not substitutes — they are a routing decision. Send the long-document and high-volume work to GPT-6.1 Sol, keep the hard reasoning and code-modification work on Claude Opus 5.5, and both live behind the same endpoint with the same credentials. If a route degrades, automatic failover moves the call rather than failing it, and because the routing is declarative it can be expressed per task class instead of per application. Testing the split is a config change, not a migration.

The last thing worth saying is about the shelf life of this comparison. Both models are days or weeks old. GPT-6.1 Sol has one independent configuration published; Claude Opus 5.5 has one. A single run per model is a snapshot, and the interesting failure mode is not that one of these numbers is wrong — it is that a medium-effort configuration, or a second evaluator, redraws the profile. Until then, the defensible reading is the one the components give: a decisive gap on hard reasoning and research code, a small one on long-context work, and an 8.3× bill that has to be justified by the part of your workload that resembles the former.

A generated scoreboard titled 'GPT-6.1 Sol vs Claude Opus 5.5 — the scoreboard' showing two columns on six shared rows: Intelligence Index 51.8 against 57.6; cost per index task $0.72 against $5.98; Humanity's Last Exam 52.9% against 61.4%; SciCode 54.2% against 66.9%; long-context recall 83.0% against 84.7%; context window 1,050,000 against 1,000,000 tokens. A footer reads 'Index, per-task cost and evaluation figures per Artificial Analysis, Intelligence Index v4.3.2.'

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily