Generated title card for an article on the Epoch Capabilities Index, with the headline 'Claude Opus 5.5 Tops the Epoch Capabilities Index' over a light-blue monospace line reading '167.35 vs 166.51' and a grey subline reading 'Top two on a composite of 50+ benchmarks', with the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Claude Opus 5.5 Tops the Epoch Capabilities Index by 0.84 Points

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The number is 0.84. Claude Opus 5.5 sits at 167.35 on the Epoch Capabilities Index as of the file we read on October 2, 2026, with GPT-6 Astra at 166.51 — a gap of less than one point on a scale where the published 95% intervals for the two models overlap almost entirely, 164.0 to 172.0 for Opus 5.5 against 163.14 to 171.05 for Astra. Underneath them, Claude Sonnet 5.5 posted 165.2 and Claude Fable 5.1 164.82, so the top four entries on an independent composite of more than fifty benchmarks are separated by 2.53 points and three of them carry the same vendor's name. The ordering is real in the sense that the point estimates are ordered. It is not real in the sense that a 0.84-point lead survives its own error bars, and everything worth saying about this leaderboard follows from holding both of those facts at once.

What the index is, and what 0.84 of a point is not

The Epoch Capabilities Index is not a vendor scoreboard. It is built by Epoch AI, an independent research organisation, as a composite that stitches together scores from more than fifty distinct benchmarks into one general-capability scale, inferring each benchmark's relative difficulty from the models that happen to appear on several of them at once. The methodology is set out in a paper written with researchers from Google DeepMind's AGI Safety & Alignment team and funded by DeepMind, but Epoch states plainly that the index is its own product and that it holds full rights over it — a distinction worth keeping when the funding source and the publisher of a score are not the same party. The code is public.

Two properties of the scale matter more than the headline. The first is that ECI is deliberately anchored rather than absolute: raw scores are rescaled so Claude 3.5 Sonnet lands at 130 and GPT-5 at 150, which means a number on this index is only meaningful next to another number from the same file, never on its own. The second is that the index has no ceiling. There is no 100 to approach, so "top of the index" is a statement about a date, not about a limit anyone has reached.

Which is why the honest reading of 167.35 against 166.51 is not "Claude Opus 5.5 is the most capable model in the world." It is narrower and less quotable: on Epoch's modelling of the evidence available on October 2, 2026, the two models are indistinguishable at the top, and the ordering between them is a tiebreak rather than a measurement. The published intervals say so directly — 164.0 to 172.0 and 163.14 to 171.05 share the great majority of their range. Anyone quoting the 0.84 as a verdict is quoting a rounding artifact.

The Anthropic block, and the size of a release

The more informative feature of the table is not who is first. It is the third and fourth rows. Claude Sonnet 5.5, released September 28 and scoring 165.2, lands 0.38 points above Claude Fable 5.1, released September 1 at 164.82 — close enough that the two are the same position in practice, and close enough that the pair sits only 2.15 points below the top of the board. Sonnet is the mid-tier product in Anthropic's line and Fable is the model the company has historically pointed at general reasoning work; that both of them now bracket the frontier from just underneath is the substantive change in this refresh, more than the identity of the leader.

Anthropic's own generational step is visible too. Claude Opus 5.5 is the successor to Claude Opus 5, which Epoch scores at 162.94 with an interval of 160.2 to 166.7. That is a 4.41-point improvement over roughly two months, from a July 24 release to a September 22 one — a larger move than the 0.84 points separating today's first and second place. Read that way, the interesting quantity in this table is the slope inside a single vendor's line, not the gap between two vendors at the top.

Where the open-weights line runs

Epoch separates models by accessibility, which makes the open-weights boundary legible without having to guess. The best open-weights entry is Kimi K3 at 157.61, dated July 16 and labelled non-commercial. Behind it sit DeepSeek V4 Pro 0813 at 155.38 and DeepSeek V4.1 Flash at 155.0, both unrestricted, then GLM-5.3-Flash at 151.94 and GLM-5.2 at 151.78. The nearest open model to the top is therefore 9.74 points back — a gap eighteen times wider than the one between first and second place, and wide enough that the intervals do not come close to touching.

That asymmetry is the one durable finding in the table. The closed frontier is a scrum inside three points; the distance from the closed frontier to the best downloadable weights is an order of magnitude larger. A composite index that lets you compare across benchmark generations is exactly the tool for noticing that, because it can say the same thing in July and in September instead of reporting that a saturating benchmark has gone quiet.

Some of the nearby rows are worth pinning for context, since they are the models readers actually weigh against Opus 5.5:

• Claude Sonnet 5 — 156.34, released June 30, interval 153.61 to 158.81

• Grok 4.6 — 156.61, released August 12, interval 154.66 to 158.93

• Gemini 3.8 Flash — 156.93, released September 2, interval 154.64 to 160.42

• Muse Spark 1.3 — 156.86, released September 2, interval 154.73 to 159.56

• GPT-5.6 Sol — 161.8, released July 9, interval 159.26 to 165.26

An index position is not a bill

A two-column scoreboard comparing Claude Opus 5.5 and GPT-6 Astra on the Epoch Capabilities Index. Left column Claude Opus 5.5: ECI 167.35 with interval 164.0 to 172.0, released 22 September 2026, input $4.00, output $20.00, cache read $0.20, context 1M tokens with 128K output. Right column GPT-6 Astra: ECI 166.51 with interval 163.14 to 171.05, released 3 September 2026, input $10.00, output $50.00, cache read $1.00, context 1.05M tokens with 128K output and a long-context tier above 272K tokens. A footer line reads 'Index figures per the Epoch Capabilities Index file read 2 October 2026; prices per the OrcaRouter catalogue.'

Here the leaderboard stops being the whole story, because the two models at the top do not cost remotely the same thing. From our own catalogue, which carries both at the provider's list price with no markup, Claude Opus 5.5 at $4.00/$20.00 with $0.20 cache reads and GPT-6 Astra at $10.00/$50.00 with $1.00 cache reads are 2.5 times apart on both directions of the rate card. Astra also steps up above 272,000 prompt tokens, where input goes to $20.00 and output to $75.00; Opus 5.5 holds $4.00/$20.00 flat across its full 1M-token window. Both serve 128,000 output tokens.

Run the arithmetic that a long-context workload actually incurs — a 180,000-token prefix served from cache, 5,000 tokens generated. On Opus 5.5 that is 180,000 tokens at $0.20 per million, or about 3.6 cents, plus 5,000 at $20 per million, or 10 cents: roughly 13.6 cents. On GPT-6 Astra it is 180,000 at $1.00, or 18 cents, plus 5,000 at $50, or 25 cents: roughly 43 cents. A 0.84-point edge and a 3.2-times cost gap are not the same kind of fact, and a procurement decision that weighs only the first one has weighed the smaller number.

Fable 5.1 makes the point from the other side of the same vendor's card: at $10.00/$50.00 it is priced exactly like Astra, and it scores 164.82 — 2.53 points below Opus 5.5 for the same money. The index puts Anthropic's three frontier entries within 2.53 points of each other and its rate card puts them 2.5 times apart, which is a much better description of the choice in front of a buyer than any single ranking is.

If you are going to argue about a number, argue on your own traffic

Screenshot of the Epoch Capabilities Index leaderboard page, showing the ranking list of models by capability score with the composite index chart above it.

The reason this index is worth reading at all is that it is measured by someone other than the vendors — but the composite's own structure sets the terms of the argument. Epoch's calibration, its difficulty estimates and its handling of models that skip benchmarks all live upstream of the number in the table, and none of them are things a reader can re-derive from a screenshot. The same is true of every vendor's headline benchmark, which is why the useful move is not to pick a side between them but to compare them on traffic you control.

Both of these models are reachable from one API key on OrcaRouter, which passes provider list price through at 0% markup: Claude Opus 5.5 at $4.00/$20.00 with $0.20 cache reads and GPT-6 Astra at $10.00/$50.00 with its above-272K tier applied exactly as the provider sets it. Sending the same prompt set to both is a config change rather than a second contract, and where both are routed, automatic failover sits underneath, which matters for a decision this close to the noise floor. The leaderboard can tell you the two models are indistinguishable. Only your own eval can tell you which one is indistinguishable on your work.

Screenshot of the OrcaRouter model page for Claude Opus 5.5 (anthropic/claude-opus-5.5), showing the model header, a 1M-token context window with up to 128K output, the Vision, Tools, JSON and Reasoning capability tags, and the input and output price figures per million tokens.

What would move this number next month

Epoch's file is a living one, and three things would change the read quickly. The first is any new evaluation of a frontier model that lands inside the top four, since a single additional benchmark observation moves a composite more when the cluster is this tight than when there is a clear leader. The second is the drift of confidence intervals — the most recent entries carry the widest ones, because they have been evaluated on the fewest overlapping benchmarks, so September's releases are the likeliest rows to move and July's the least. The third is a benchmark reaching saturation, which is the specific failure the index was built to survive: when a component stops discriminating, Epoch re-weights the composition rather than letting a model's advantage evaporate with the test.

For now, the defensible summary is the narrow one. Claude Opus 5.5 holds the top spot on the Epoch Capabilities Index at 167.35 on the strength of a 0.84-point margin that its own confidence interval does not support, Claude Sonnet 5.5 has climbed to within 2.15 points of the lead from the mid-tier of the same line, and the open-weights field is more than nine points behind all of them. The leader is a coin-flip between two vendors. The four-point gap in price between them is not.

Compared in this article4

Detected from this article · Benchmarks: Artificial Analysis · updated daily