A generated title card headed 'level at 165', with the subtitle 'Epoch Capabilities Index, file of 1 October 2026', three stat cards reading 165 (Claude Sonnet 5.5), 165 (Claude Fable 5.1) and 11 vs 20 (benchmark evaluations behind each figure), and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Claude Sonnet 5.5 Ties Claude Fable 5.1 on the Epoch Capabilities Index

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

A tie at the top of a leaderboard is usually the least interesting thing on it. This one is the exception, because the two models level at the top do not cost the same money. In the Epoch Capabilities Index file updated 1 October 2026, Claude Sonnet 5.5 and Claude Fable 5.1 both read 165 — rows three and four of 253 tracked models — two points behind Claude Opus 5.5 and GPT-6 Astra, who are themselves level at 167, and two points clear of Claude Opus 5 at 163. Claude Sonnet 5.5 shipped on 28 September 2026 at $2 per million input tokens and $10 per million output. Claude Fable 5.1 shipped on 1 September at $10 and $50. A single number says the two are interchangeable in general capability. Five times the output price says they are not. Both hold up, and the reason they hold up together is the part of the table nobody quotes.

What the index actually is, and who is measuring

The ECI is not a vendor scoreboard. It is built by Epoch AI, an independent research organisation, as a composite that stitches scores from more than fifty distinct benchmarks into one general-capability scale, inferring each benchmark's relative difficulty from the models that happen to appear on several of them at once. The methodology is set out in a paper, “A Rosetta Stone for AI Benchmarks”, funded by Go​ogle DeepMind and written with researchers from its AGI Safety & Alignment team — but Epoch states plainly that the index is its own product and that it holds full rights over it, and the fitting code is public. That distinction matters when the funder of the research and the publisher of the score are not the same party.

Two properties of the scale decide how a number on it may be read. The first is that ECI is anchored rather than absolute: raw values are rescaled so Claude 3.5 Sonnet lands at 130 and GPT-5 at 150, which means a figure only means something next to another figure from the same file. The second is that there is no ceiling. There is no 100 to approach, so “top of the index” is a statement about a date rather than about a limit anyone has reached.

The third property is the one that turns a tie into a story. The intervals published beside each score — 90% bootstrap intervals, constructed by resampling each model's own observed benchmark results five hundred times — are identical for the two models at 165: 162 to 169 for Claude Sonnet 5.5 and 162 to 169 for Claude Fable 5.1. The counts that produced them are not. Claude Sonnet 5.5's 165 is fitted from 11 benchmark evaluations. Claude Fable 5.1's is fitted from 20.

Why the same interval does not mean the same evidence

The ECI needs a minimum of four benchmark evaluations before a model is scored at all, so both figures clear the floor comfortably. But the interval is a statement about how much a model's own results scatter, not about how many there are — which is why two models can print the same band around the same point estimate while one of them is resting on roughly half the evidence. Treat the two 165s as equally firm and you have read the interval as a sample-size correction it is not.

The asymmetry has a practical consequence for timing. Epoch refits the whole model jointly across all data whenever new evaluations arrive, so a model's figure can move without anything about that model changing. The model carrying 11 observations has more room to move than the one carrying 20, which makes Claude Sonnet 5.5 the likelier of the pair to shift in the next revision — in either direction.

Epoch's own framing of the current reading is narrower than the headlines it produced. The organisation's summary of the update leads with Claude Opus 5.5 taking the top spot at 167 narrowly ahead of GPT-6 Astra, and gives the second sentence to the fact that Claude Sonnet 5.5 has “roughly matched” Claude Fable 5.1 at 165. Roughly matched is the operative phrase, and the published intervals are the reason it is the right one.

Where the two profiles actually diverge

A two-column comparison scoreboard titled 'Claude Sonnet 5.5 vs Claude Fable 5.1 - the scoreboard'. Left column 'Claude Sonnet 5.5': rows reading 'ECI: 165', 'Interval: 162-169', 'Benchmarks: 11', 'Mystery Game Puzzles: 65%', 'Furniture Assembly: 75%', 'Price: $2 / $10'. Right column 'Claude Fable 5.1': rows reading 'ECI: 165', 'Interval: 162-169', 'Benchmarks: 20', 'Mystery Game Puzzles: 58%', 'Furniture Assembly: 70%', 'Price: $10 / $50'. A footer line reads 'Epoch Capabilities Index, file updated 1 October 2026; prices per the vendors' own rate cards.'

Level on the composite does not mean level on the parts. Epoch publishes a per-benchmark panel on each model page, and six benchmarks appear on both the Claude Sonnet 5.5 page and the Claude Fable 5.1 page — which makes those the only six where a like-for-like comparison is available from the index itself.

• Mystery Game Puzzles — Claude Sonnet 5.5 65% vs Claude Fable 5.1 58%

• FrontierMath Tiers 1–3 (v2) — Claude Sonnet 5.5 89% vs Claude Fable 5.1 90%

• FrontierMath Tier 4 (v2) — Claude Sonnet 5.5 80% vs Claude Fable 5.1 88%

• OTIS Mock AIME 2024–2025 — Claude Sonnet 5.5 100% vs Claude Fable 5.1 100%

• Furniture Assembly Benchmark — Claude Sonnet 5.5 75% vs Claude Fable 5.1 70%

• SimpleQA Verified — Claude Sonnet 5.5 47% vs Claude Fable 5.1 71%

Four points of movement either way on a composite this crowded, and the underlying split is not close to symmetric. Claude Sonnet 5.5 is ahead where the work is games-like and multimodal — puzzle solving, physical assembly — and level on competition mathematics. Claude Fable 5.1 is ahead by 8 points on the hardest tier of FrontierMath and by 24 points on SimpleQA Verified, the world-knowledge set where a 47% score against a 71% is the largest single gap in the shared panel. A buyer choosing between these two models is not choosing between two readings of the same capability; they are choosing between a cheaper model that is stronger on some work and a dearer one that is stronger on other work, with a general-capability index that declines to separate them.

Epoch's index is also explicit that a per-benchmark lead does not have to show up in the composite. Benchmarks are not weighted by accuracy, and a model that specialises can score below a differently-shaped rival even while leading individual tests — its documentation says so directly. That is not a defect being excused; it is what a composite over fifty-odd tests is for.

The two independent instruments point opposite ways

A capture of the Epoch AI model page for Claude Sonnet 5.5, headed ECI #3, Anthropic, Sep. 28, 2026, Closed weights, and the line that it has an ECI score of 165 and ranks 3rd of 253 tracked models. Beneath it a comparison table headed 'Claude Sonnet 5.5 vs select alternatives' carries four columns - Claude Sonnet 5.5, Claude Opus 5.5, GPT-6 Astra and Kimi K3 - with an ECI score row reading 165, 167, 167 and 158, a ranking row reading #3, #1 - Best, #2 and #13, and per-benchmark rows for Mystery Game Puzzles, FrontierMath Tiers 1-3 (v2), FrontierMath Tier 4 (v2), OTIS Mock AIME 2024-2025, Furniture Assembly Benchmark, SciCode, GPQA Diamond, Text Arena (Coding) and SimpleQA Verified.

There is a second independent measurement in circulation, and it does not agree with the first about which of these two models is ahead. On the Artificial Analysis Intelligence Index, the figures our own catalogue carries for the two models read 56 for Claude Sonnet 5.5 and 53.4 for Claude Fable 5.1, drawn from the same revision of that index and therefore carrying the same sample of evaluated models. Three points apart, in favour of the cheaper model — where Epoch reads them as identical.

Neither reading is wrong, and the way to use them is to notice what they disagree about. The two indices cover different benchmark sets and weight them differently; the Epoch composite is built to survive benchmarks saturating, while the Artificial Analysis index is a straight aggregate of a different basket. When a pair of models is this close, the choice of instrument is worth more than the gap being measured. A reader who wants the stronger of two adjacent numbers should look at the panel of shared individual benchmarks above rather than at either headline, because that is the only place the two models are compared against each other on the same test.

There is one more piece of context the composite does not carry, and Epoch does: release date. Claude Fable 5.1 ranks fourth of 253 tracked models today. At the time of its own release on 1 September 2026 it ranked first of the 247 models then tracked. Its weights have not changed. The frontier simply moved past it six places inside a month — which is the shape of the whole table, and the reason a monthly index reading is a snapshot rather than a verdict.

What the difference costs, and where the cheap side is

A capture of the OrcaRouter model page for Claude Sonnet 5.5, listed under anthropic/claude-sonnet-5.5, showing a 1M-token context window with up to 128K output tokens, input pricing at $2.00 per million tokens and output at $10.00 per million tokens, and the catalogue's capability tags for the model.

Level on general capability and 2.5 times apart on input and 5 times apart on output is the entire practical content of this reading. Both models are carried on our own catalogue — anthropic/claude-sonnet-5.5 at $2.00 per million input and $10.00 per million output with $0.20 cache reads, and anthropic/claude-fable-5.1 at $10.00 and $50.00 with $0.25 cache reads — and both are passed through at the provider's own list price with no markup added, so a vendor rate change lands on our side the same day it lands on theirs.

The recurrence matters more than the headline. Take a workload that sends a 120,000-token prefix from cache and generates 8,000 tokens of output, run once a day for a month. On Claude Sonnet 5.5 that prefix costs 120,000 tokens at $0.20 per million, about 2.4 cents, plus 8,000 at $10 per million, 8 cents — roughly 10.4 cents a run, or about $3.12 over thirty days. On Claude Fable 5.1 the same prefix is 120,000 at $0.25, 3 cents, plus 8,000 at $50, 40 cents — about 43 cents a run, or $12.90 across the month. Both models serve the same 1M-token window with up to 128K output, so the difference is the card rather than the ceiling.

Against that, the six shared benchmarks say which way the extra money buys something. A team whose work looks like FrontierMath Tier 4 or open-ended factual recall is buying a real 8- to 24-point gap on those tests by paying four times more; a team whose work looks like puzzle solving or multimodal assembly is buying 5 to 7 points the other way while paying the smaller number. On a single platform these are not two procurement decisions. Both models sit behind one API key among more than 200 models, so comparing them on your own traffic is a config change rather than a second contract, and where both are routed, automatic failover sits underneath — which matters when two independent instruments disagree about which one is ahead.

What would move this reading

Three things would change it, and only one of them involves either model.

The first is more benchmark evaluations. Claude Sonnet 5.5 entered the index four days ago with 11 behind it; every additional evaluation narrows its interval and can move its point estimate, and a joint refit means the movement is not confined to the new row.

The second is the arrival of another frontier model. The top four rows span four points, with two ties inside that span, so a single well-evaluated newcomer lands inside the cluster rather than below it — and the published intervals, which overlap across all four, mean the ordering among them is a tiebreak rather than a measurement.

The third is the price of the models themselves, which no index tracks. The gap this piece turns on is a rate-card gap: $2 and $10 against $10 and $50 today. A change on either card moves the decision without moving a single capability score.

The defensible summary is the narrow one. On Epoch AI's composite, in a file updated 1 October 2026, Claude Sonnet 5.5 and Claude Fable 5.1 are the same number — 165, tied for third, with identical published intervals resting on different amounts of evidence. On Artificial Analysis's separate instrument the same two models are three points apart. On the six benchmarks where the index compares them directly, they trade leads by domain rather than sitting level. And on the rate card they are not close at all. The one thing this reading will not support is the sentence it is most likely to be quoted as: that the models are equivalent.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily