Hero title card for Claude Opus 5.5 reading 'every score, the effort setting it came from, and who measured it', with two stat chips: 66.4% on Terminal-Bench 4.0 at xhigh effort, vendor-reported, and 59.6% on the same benchmark, independent.
Guides & Insights

Claude Opus 5.5 Benchmarks: Every Score, the Effort Setting It Came From, and Who Measured It

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ant​hropic's headline number for Claude Opus 5.5 on Terminal-Bench 4.0 is 66.4%, printed against 52.3% for Claude Opus 5 on the same table — and that 66.4% was produced at xhigh effort, not at the model's default of medium. The model shipped on September 22, 2026, two days before this page was written, so it is a GA model with a GA model's pricing and a GA model's benchmark table. The independent figure on the same benchmark, from Artificial Analysis, is 59.6%, in a configuration that carries Ant​hropic's server-side safeguard routing inside it. Between those two numbers sit a harness difference, an effort difference and a routing difference, and nobody — including us — can yet say how much of the gap each one accounts for. What this page can do is put every published figure for Claude Opus 5.5 in one place, name who produced it, name the setting it was produced at, and include the rows where GPT-6 Astra and Claude Fable 5.1 come out ahead.

Start with the effort setting, because it moves the numbers more than the models do

Claude Opus 5.5 defaults to medium effort on the Claude API. Claude Opus 5 defaulted to high. The same named level is also not the same thinking budget across the two models, so "we ran both at high" is not a controlled comparison even when the label matches. Adaptive thinking is always on and cannot be switched off on Opus 5.5; the effort parameter moves depth, latency and cost, and nothing else.

Anthropic's Terminal-Bench 4.0 figure was run at xhigh. That single fact is the most important caveat on this page, because the benchmark it headlines is the one with the widest reported margin over the previous model, and because the same table shows what happens when you leave the setting alone:

FrontierCode v1.1 Main — 54.4% at xhigh, 54.6% at default medium. Vendor-reported. The medium run is higher than the xhigh run. On this workload the effort dial bought nothing measurable; the two numbers are the same number.

CursorBench 4.0 — 57.8% at xhigh, 52.5% at medium. Vendor-reported. Same model, same harness, same week: 5.3 points, which is the largest effort delta on the table.

Those two rows sitting inside one vendor table are the whole argument. One workload is effort-insensitive and the other is strongly effort-sensitive, and the model card cannot tell you in advance which kind of workload yours is. The conclusion the data supports is narrow and worth stating narrowly: the right effort level is now a per-workload measurement, not a default you inherit. It does not support "Opus 5.5 is faster than its benchmarks suggest" or "Anthropic sandbagged the default" — for that you would need per-workload sweeps across all five settings, and nobody has published them. It does mean that if you migrate from Opus 5 and keep the effort setting you already had, you have changed the experiment, not just the model.

One more asymmetry, and it cuts the other way. The competitor column on Anthropic's Terminal-Bench row is GPT-6 Astra at 57.9% — run at high effort, per Anthropic's own chart note, while Opus 5.5's 66.4% was run at xhigh. A vendor table that runs its own model at one setting and a rival at another is not a head-to-head, whatever the two numbers look like side by side.

The vendor table, with the source and the setting on every row

Everything in this section is vendor-reported: Anthropic's launch material, Anthropic's harness, production safeguards on, adaptive thinking at max effort unless the row says otherwise. Read it as Anthropic measuring Anthropic.

Terminal-Bench 4.0 — 66.4%, at xhigh effort. Vendor-reported, standard error about ±2.6 points. On the same row: Claude Opus 5 52.3%, Claude Fable 5.1 55.8%, GPT-6 Astra 57.9% (at high effort), GPT-5.6 Sol 37.3%. Terminal-Bench 4.0 is explicitly a benchmark the vendor acknowledges is an imperfect proxy for real terminal work, and it is the model's headline.

FrontierCode v1.1 Main — 54.4% at xhigh, 54.6% at default medium. Vendor-reported. Opus 5 48.0%, Fable 5.1 50.3%, GPT-6 Astra 53.3%, GPT-5.6 Sol 47.5%. Anthropic's system card also carries the Extended variant at 63.6% and a best-at-medium Extended figure of 65.3%.

CursorBench 4.0 — 57.8% at xhigh, 52.5% at medium. Vendor-reported. Opus 5 46.6%, Fable 5.1 51.8%, GPT-5.6 Sol 41.7%. Note that Fable 5.1's and Opus 5's CursorBench rows are at max effort, so Opus 5.5 at medium is being compared to its own siblings at max.

GDPval-AA v2.1 — 1,846 Elo. This is the one row on the vendor table that the vendor did not run. GDPval-AA is an Artificial Analysis evaluation, and Artificial Analysis's own write-up on Opus 5.5 states the same 1,846, with Fable 5.1 at 1,735 and Opus 5 at 1,708. In the vendor table: Fable 5.1 1,735, Opus 5 1,708, GPT-5.6 Sol 1,588, GPT-6 Astra 1,542. It is the only row here where two organisations agree on the same number, and that is because it is the same measurement.

AutomationBench (Zapier) — 40.0%. Vendor-reported, and run with no fallback models, so every safeguard intervention was scored as a failure. GPT-6 Astra 41.4%, Fable 5.1 31.4%, GPT-5.6 Sol 28.8%, Opus 5 26.9%.

Humanity's Last Exam, with tools — 67.7%. Vendor-reported. Fable 5.1 65.6%, Opus 5 63.6%, GPT-6 Astra 57.2%. Anthropic's system card separately reports 64.4% without tools, which is the figure to reach for if you are comparing against a source that does not give its models tool access.

Terminal-Bench-Science 0.1 — 58.7%. Vendor-reported, standard error roughly ±3.5 to ±5 points, so this is a band rather than a point. GPT-6 Astra 64.6%, Fable 5.1 52.6%, Opus 5 29.0%, GPT-5.6 Sol 22.4%.

OSWorld 2.0 — 81.8% partial. Vendor-reported, and "partial" is doing real work in that sentence: Anthropic's system card gives a strict completion rate of 48.7% for the same model on the same evaluation. Both are published; the partial number is the one that gets quoted. Fable 5.1 80.7%, Opus 5 74.0%.

Chartography, with tools — 89.0%. Vendor-reported. Fable 5.1 88.4%, Opus 5 83.4%.

Scoreboard for Claude Opus 5.5 with six rows: Terminal-Bench 4.0 66.4% at xhigh effort (vendor-reported); Artificial Analysis Intelligence Index 58 at max effort (independent); Humanity's Last Exam 67.7% with tools (vendor-reported); Terminal-Bench-Science 0.1 58.7% (vendor-reported); about 119k output tokens per task at max (independent); price $4.00 input / $20.00 output per 1M. A footer separates the Anthropic launch table from Artificial Analysis.

The independent numbers, and what "Default Fallback" actually means

Artificial Analysis is the only third party with a published, current evaluation of Claude Opus 5.5 on a composite index, and its figures are the ones to hold beside Anthropic's. As read on September 24, 2026:

Intelligence Index — 58 at max effort, in a configuration labelled "Adaptive Reasoning, Max Effort, Default Fallback". Independent, measured by Artificial Analysis. Its own article calls 58 "the highest score we have measured by several points." The same index also publishes Opus 5.5 at every other effort setting, and the spread is the useful part: 56 at xhigh, 54 at high, 51 at medium, 42 at low. That is a 16-point swing across the effort dial on a single model — larger than the gap between most models on that board.

Terminal-Bench 4.0 — 59.6%. Independent. Artificial Analysis reports it as "level with the leader GPT-6 Astra (xhigh)" and about 11 points above Claude Opus 5.

Humanity's Last Exam — 61.4%. Independent, and the previous best on the same board was 59.1%, held by Claude Fable 5.1.

SciCode — 66.9%. Independent, against 63.1% for Claude Fable 5.1.

Now the clause that needs explaining rather than quoting. "Default Fallback" is not a benchmark artifact and not an Artificial Analysis setting — it is Anthropic's server-side safeguard routing, left switched on. Claude Opus 5.5 ships with the safeguard tier previously reserved for Fable 5.1: most cybersecurity requests are transparently re-routed to Claude Opus 4.8, and biology work falls back to Claude Opus 5 unless the account is verified. When Artificial Analysis runs the index with default fallback enabled, it is measuring the deployed system — the thing an unverified API key actually reaches — rather than a clean checkpoint of Opus 5.5 weights. That is the more honest measurement for a buyer, and it is the less clean measurement for a benchmark. Anthropic says as much about its own table: interventions on the Terminal-Bench 4.0 run had cyber tasks completed by Opus 4.8 and biology tasks by Opus 5, and the company states those interventions likely lowered the reported score. So the safeguard routing is a real operational fact in both directions, and no figure on this page isolates the underlying model from it.

Where the two tables disagree, and why

Put the two Terminal-Bench 4.0 numbers side by side and you get 66.4% against 59.6% — 6.8 points, on the same benchmark, for the same model. The reasons are known, and each one is large enough to matter on its own:

Different harnesses. Anthropic ran its own agent scaffold; Artificial Analysis ran its own reference agent harness. A terminal benchmark score is a property of the scaffold plus the model, not of the model alone, and neither organisation has published a scaffold-swap experiment.

Different effort settings. Anthropic's 66.4% is at xhigh. The effort setting attached to Artificial Analysis's 59.6% is not stated in the sentence that carries the number, and its reported benchmark figures are generally the max-effort ones.

Safeguards on, in both cases, counted differently. Anthropic ran with fallback and treated interventions as score-reducing but still scored. On AutomationBench it ran with no fallback and counted interventions as outright failures. Artificial Analysis runs with fallback enabled throughout.

The numbers move even inside one vendor's own table. Anthropic prints Claude Opus 5 at 52.3% on Terminal-Bench 4.0 and notes the public board's 51.8% for the same model is "within noise".

And then the strongest piece of evidence on this page about how much a harness is worth. The public Terminal-Bench 4.0 leaderboard, captured on September 24, 2026, carries no Claude Opus 5.5 row at all. Its top entry is GPT-6 Astra at max effort, Codex agent, 58.2% ±2.8; second is Claude Fable 5.1 at max effort through Claude Code at 57.9% ±3.8; third is Claude Opus 5 at xhigh through Claude Code at 53.9% ±3.2. So the 66.4% is not merely a different number from the independent 59.6% — it is not on the public board, and the board's own version of Claude Opus 5 is 53.9%, not the 52.3% the launch table prints. Until Terminal-Bench accepts and publishes an Opus 5.5 submission, the 66.4% is a vendor harness result reproduced by nobody, and the 59.6% is the number an outside party will actually stand behind. Note also that on that same board, Artificial Analysis's Terminal-Bench 4.0 leader and Anthropic's Opus 5.5 land on the same 59.6% — which is a tie in one harness, and tells you nothing about the other.

Effort-dial infographic for Claude Opus 5.5 showing the Artificial Analysis Intelligence Index moving from 42 at low effort through 51 at medium (the product default), 54 at high, 56 at xhigh and 58 at max - a 16-point swing on one model - with two comparison cards: FrontierCode v1.1 at 54.4 xhigh against 54.6 medium (effort-insensitive) and CursorBench 4.0 at 57.8 xhigh against 52.5 medium (effort-sensitive).

The rows where Opus 5.5 loses

A roundup that omits the losses is not a roundup, and Anthropic's own table contains both of them.

Terminal-Bench-Science 0.1 — 58.7% against GPT-6 Astra's 64.6%. Vendor-reported, on Anthropic's table, and a 5.9-point loss to a named competitor on the row closest to scientific reasoning. The vendor's own standard-error note (±3.5 to ±5) means the gap is near the edge of noise rather than comfortably inside it — but it is a loss on the vendor's own page, and the band does not make it a win. Claude Opus 5 sits at 29.0% on the same row, so the generational gain is real even where the competitor leads.

AutomationBench — 40.0% against GPT-6 Astra's 41.4%. Vendor-reported, a 1.4-point loss, and the run had no fallback models, so safeguard interventions were scored as failures. That means this particular comparison measures Opus 5.5 with its routing penalty included against a competitor whose routing penalty, whatever it is, is not described. It is the least interpretable row on the page in either direction.

Two more places where the lead is thinner than the headline suggests, both vendor-reported. On Humanity's Last Exam with tools the margin over Claude Fable 5.1 is 2.1 points (67.7% vs 65.6%); on Chartography with tools it is 0.6 points (89.0% vs 88.4%); on OSWorld 2.0 partial it is 1.1 points (81.8% vs 80.7%). Those are the rows where "Opus 5.5 is the best agentic model" is a claim about rank rather than about a measured gap, and a 0.6-point chart-reading difference should not decide a procurement.

Claude Fable 5.1's independent standing, on a different index revision

Placing Opus 5.5 against its own sibling needs its own section, and it needs a warning attached, because the comparison is harder than it looks.

When Artificial Analysis published its Claude Fable 5.1 write-up on September 1, 2026, it scored 66 at max effort on its Intelligence Index and called it the highest score it had measured — ahead of Claude Opus 5 at 63 and GPT-5.6 Sol at 61. On the index as it is listed on September 24, 2026, the same model appears at 53 at max effort with fallback, and Opus 5.5 appears at 58 in the equivalent configuration. The September 22 write-up on Opus 5.5 calls 58 "the highest score we have measured by several points."

Those two statements cannot both be true on one scale, and the resolution is that they are not on one scale. The index is a live object that Artificial Analysis revises; two of its published articles, five weeks apart, place the same sibling pair on opposite sides of the top spot. That is a statement about index revisions, not about either model regressing, and it means the 66 and the 58 must never be printed side by side. Read on the same day, the same listing, the same labelled configuration, the current ordering has Opus 5.5 five points ahead of Fable 5.1 — and even that is a same-index, same-vendor-of-index comparison, not an audited head-to-head. What is safe to say is narrower: on the current revision of the one independent composite that measures both, Opus 5.5 is the stronger of the two Anthropic models, and Fable 5.1's better-known 66 belongs to an earlier revision of that index.

The settings are worth one more sentence because they are easy to conflate. Fable 5.1's 66 and Opus 5.5's 58 are both max-effort rows — but Fable 5.1's five effort settings span an 11x range in output token use, from 13.1M at low to 143.7M at max, with index scores from 58 to 66. A single index point for a model with that much internal spread is a poor substitute for the row you would actually run.

Token cost is a benchmark axis of its own

Every number above is a score. None of them is a bill, and on this model the bill is where the interesting movement is.

Artificial Analysis measures Claude Opus 5.5 at max effort using on the order of 119,000 output tokens per Intelligence Index task — the highest in its index. For comparison, on the same board at the same setting: Claude Opus 5 about 73,000, Claude Fable 5.1 about 78,000, and GPT-6 Astra about 27,000. Thinking tokens bill as output tokens, and adaptive thinking cannot be turned off on Opus 5.5, so the tokens a model spends deliberating are not a footnote to its price — they are most of it.

Run the arithmetic, because the sticker price is misleading on its own. Opus 5.5 lists at $4.00 input / $20.00 output, a 20% cut from Opus 5's $5.00 / $25.00. Artificial Analysis still measures Opus 5.5 at max effort costing about $5.98 per index task against $5.86 for Opus 5 at the same setting — slightly more, not less, because roughly 1.6x the output tokens eats the entire 20% rate cut and then some. Across the whole index, Opus 5.5 produced 260M output tokens against a board median of 88M, which the evaluator labels "very verbose."

The cost story therefore lives entirely in the effort setting, and it only works if you use it. At max effort, Opus 5.5 is a more expensive model per completed task than the model it replaces. At medium — the actual product default, and where Anthropic's "about 40% lower cost to run on typical workloads" claim is scoped — the saving is real but it is a statement about an average across workloads the vendor selected, not an audited per-task result. The practical reading: the 20% price cut is a discount on the rate card and an option on the effort dial, not an automatic saving. If your migration plan is "swap the model id and keep the settings," you are buying the expensive version.

What is not established at all

This is the section the rest of the page exists to support.

No independent matched-effort, matched-harness head-to-head of Claude Opus 5.5 against any named rival has been published. Every cross-vendor comparison on this page — Opus 5.5 vs GPT-6 Astra, vs GPT-5.6 Sol, vs Claude Fable 5.1 — is a comparison of two tables produced by two different companies, on two different harnesses, at two different effort settings, one of which is usually the vendor's own model at a favourable setting. There is no experiment anywhere in the public record that holds the scaffold and the effort constant across two vendors and reports both.

The per-problem token split is not published. The aggregate output-token figure is measured independently, but how much of it is thinking versus answer text is not published by anyone we could find. Any page that gives you a precise thinking-token number for this model is giving you a number nobody measured.

Most of the system card is unbenchmarked by third parties. SWE-bench Pro 89.9%, Multilingual 93.9%, DeepSWE v1.1 74.2%, ProgramBench 91.2%, OfficeQA 78.9%, HealthBench Professional 65.6% length-adjusted, GMMLU 94.3%, ArXivMath 96.9% with tools — all vendor-reported, none independently reproduced at the time of writing.

Terminal-Bench 4.0's headline figure is not on the public board. Covered above, and it belongs in this list too: the single most-quoted number about this model is one nobody outside the vendor has run.

Anthropic reports that Claude Opus 5.5 "often suspects it is being evaluated" — the company's own words, offered as a limitation on its safety assessment. Take that seriously as a measurement problem and not as a curiosity: if the model behaves differently when it believes it is being tested, then every benchmark number on this page, vendor or independent, is a measurement of the model under observation, and the degree to which that differs from deployed behaviour is unknown and unquantified. It applies to the good rows and the bad ones equally. It is the single caveat that no amount of matched-effort methodology fixes.

Sonnet 5.5 and Haiku 5.5 are not available. Anthropic has announced them for "the coming weeks." They have not shipped, they have no published benchmarks, and nothing on this page transfers to them.

Price, envelope, and the parts of the spec that change your code

Read from Anthropic's published rate card on September 24, 2026: $4.00 per million input tokens, $20.00 per million output, $0.20 per million cache read, $5.00 per million 5-minute cache write, $8.00 per million 1-hour cache write, and 50% off both directions on the Batch API. The cache-read rate is the quiet one — it is 5% of base input, where most Claude models sit at 10% and Fable 5.1 sits at 2.5%. Fast mode is priced separately at $8.00 / $40.00 and Anthropic describes it as a research preview, not a general tier.

This is the section where a rate card stops being trivia and becomes arithmetic a router can do for you. Claude Opus 5.5 is served as anthropic/claude-opus-5.5 on OrcaRouter at the same $4.00 / $20.00, with list prices passed through at 0% markup — so the moment Anthropic moves a rate, that is the rate here, on the same day and with no repricing lag to model. The pieces that decide the real bill are the ones the catalogue carries next to the price: the $0.20 cache read at 5% of base input against Fable 5.1's 2.5%, the 1M-token context window, and the 128K output ceiling. Because the effort parameter is where this model's cost actually moves — a 16-point index swing and roughly 119,000 output tokens per task at max against about 73,000 for Opus 5 — the migration that saves money is the one that lets you dial effort per workload rather than per account, and put the Opus 5.5 route behind an automatic failover while you measure which setting your traffic wants.

Envelope — 1M-token context, 128K maximum output, 300K output on the Batch API behind the output-300k-2026-03-24 beta header, knowledge cutoff June 2026, retirement no sooner than September 22, 2027. Output is more than 30% faster than Claude Opus 5.

Breaking changes versus Opus 5 — thinking can no longer be disabled; forced tool use (a tool_choice naming a specific tool) returns an error; thinking blocks are bound to the model and the conversation that produced them; the older computer_20251124 computer-use tool is not accepted on the Claude API or Google Cloud. The first three also apply to Fable 5.1. There is a silent fifth: text between tool calls now arrives inside thinking blocks whose text is empty at the default display setting, so a UI that streams that text as a progress indicator goes quiet without anything erroring.

Safeguard routing, again, because it is a spec item and not a footnote — most cybersecurity requests go to Claude Opus 4.8 and biology work falls back to Claude Opus 5 unless the account is verified. On those evaluations the answer you receive may not come from Opus 5.5 at all, so the published scores on cyber- and bio-adjacent benchmarks are not clean measurements of this model.

Efficiency claims, all vendor-reported — about 40% lower cost to run on typical workloads against a 20% sticker cut; a worked merger-analysis example taking 63 minutes on Opus 5.5 against 93 on Opus 5 at 50% less cost; and customer statements from Box (a third of the tokens, answers about 40% less verbose), Kiro (roughly half the tokens, 40% fewer calls), Factory (20–25% fewer output tokens), and GitHub (among the fewest tokens and steps it has measured). These are averages across workloads the vendor and its launch partners selected. They are attributions, not audited results, and they sit in visible tension with the independent per-task cost figure above — which is itself a measurement at max effort rather than at default. Both can be true; neither is a general claim about your workload.

Screenshot of the OrcaRouter model page for Claude Opus 5.5, model id anthropic/claude-opus-5.5, showing the NEW, FLAGSHIP and FEATURED badges, a 1M-token context window, 128K max output, text plus image plus file input, and the $4.00 input / $20.00 output per 1M token rate with pricing passed through from Anthropic.

The state of the evidence

Claude Opus 5.5 is a real model with a real rate cut, a real envelope of 1M tokens of context and 128K of output, and a real and consequential change in how thinking and effort are configured. It also arrives with a benchmark table that mostly measures Anthropic, an independent table that mostly measures a routed deployment rather than a checkpoint, and a documented self-awareness problem that undercuts both. No independent matched-effort, matched-harness head-to-head of Claude Opus 5.5 against a named rival has been published. Until one is, read every cross-vendor comparison on this page as what it is: two vendor tables from two companies at two settings, arranged side by side by us — and re-measure your own effort setting before you re-measure the model.

Compared in this article4

Detected from this article · Benchmarks: Artificial Analysis · updated daily