
Claude Sonnet 5.5 vs Claude Fable 5.1: The Cheap Model Sits Three Points Ahead
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 982 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 197 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Claude Sonnet 5.5 arrived on September 28, 2026 at $2.00 per million input tokens and $10.00 per million output tokens. Claude Fable 5.1, the Mythos-class model the vendor shipped on September 1, sits at $10.00 and $50.00 on the same two lines. The tempting summary is that Fable 5.1 is the expensive one and therefore the better one, and on the Artificial Analysis Intelligence Index the ordering is the other way round: Claude Sonnet 5.5 scores 56 and Claude Fable 5.1 scores 53. That is not a rounding artifact, and it is not the whole story either, because the same evaluator puts the two of them within three cents per task of each other on the same suite — $7.60 against $7.63. A five-fold difference in the rate card has almost vanished by the time a task is finished. What remains is a set of differences in shape rather than in score, and choosing between these two is mostly a question of which shape your workload has.
Two rate cards, and the one line where the gap is not five-fold
The per-token comparison is stark and needs no interpretation.
• Input — Claude Sonnet 5.5 at $2.00 per million tokens against Claude Fable 5.1 at $10.00, a five-fold difference
• Output — $10.00 against $50.00 per million, again exactly five-fold
• Cache read — $0.20 against $0.25 per million. This is the line that carries long agent runs, and the discount collapses from 5× to 20%
• Cache write — $2.50 against $12.50 per million for the standard five-minute window, the widest single gap on either card
• Batch — Anthropic's batch mode applies to neither of these cards the way it does to Fable 5.1's: Fable 5.1 launched with batch pricing halving to $5.00 input and $25.00 output. Read the batch line off Anthropic's own pricing page before you assume it carries across the family
• Context and ceiling — 1,000,000 tokens of context and 128,000 tokens of maximum output on both. On the Message Batches API, Claude Sonnet 5.5 supports up to 300,000 output tokens behind the output-300k-2026-03-24 beta header, which is a real ceiling difference for bulk generation and not a marketing one
There is a modest point in Claude Sonnet 5.5's favour that the rate card understates. Anthropic reports that the model typically needs far fewer tokens to do the same work than its predecessor, which it says translates to as much as 30% less cost per task. That is a vendor claim about a different comparison — Sonnet 5.5 against Sonnet 5 — and it is not evidence about this matchup. It is, however, the mechanism that explains what the independent measurement found.
Where the five-fold price gap goes
Artificial Analysis ran both models on 2026-09-29 under a single harness and one configuration label, "Adaptive Reasoning, Max Effort, Default Fallback", on Intelligence Index v4.3. The two headline figures are 56 for Claude Sonnet 5.5 and 53 for Claude Fable 5.1 — the cheaper model ahead by roughly three points on a scale where the same evaluator places the median model at 26. Both are labelled vendor-name models on the same page, so this is one instrument and one snapshot, not two.
Then the cost column, and the reason this pairing is interesting rather than obvious. On the same index run, Claude Sonnet 5.5 costs $7.602633 per task and Claude Fable 5.1 costs $7.629706. Twenty-seven thousandths of a dollar apart. The decomposition is where it gets strange: Claude Fable 5.1's input side alone is $3.73 per task against Sonnet 5.5's substantially smaller input component, and its reasoning tokens alone cost $2.36 per task. The model that is five times cheaper per token is not cheaper per finished task at all.
The cause is volume, and the two models fail in opposite directions.
• Output tokens across the index suite — Claude Sonnet 5.5 generated 410,577,501 against Claude Fable 5.1's 188,465,663, on a suite where the tracked median is 88 million. Both are well above median; Sonnet 5.5 is more than double Fable 5.1 and nearly five times the median
• Output tokens per task — 192,838 for Sonnet 5.5 against 78,111 for Fable 5.1, with the reasoning share 142,386 against 47,240. Sonnet 5.5 thinks roughly three times as long per task
• Input tokens across the index suite — 18.5 billion for Sonnet 5.5 against 5.6 billion for Fable 5.1. Sonnet 5.5 read three times as much, which is the quieter number in the pair and the one that eats the cheaper input rate
• Blended price at a 7:2:1 input-to-cache-read-to-output mix — $1.54 for Sonnet 5.5 against $7.17 for Fable 5.1. This is the line that matches the rate card, and it is the line that describes a workload with short, well-bounded outputs
Put those together and the honest statement is this: the blended price gap is real and the per-task gap is not. Which of the two numbers describes your bill depends entirely on whether your tasks terminate or ramble, and the index suite is built to let them ramble.
The effort ladder is the actual decision
Both models expose an effort parameter — low, medium, high, xhigh, max — and both were measured at every rung on 2026-09-29. This is the part of the comparison no rate card can express, because the cheap model at a lower rung beats the expensive model at a lower rung by a wide margin and the expensive model only separates near the top.
• Low effort — Claude Sonnet 5.5 scores 35.84 at $0.41 per task; Claude Fable 5.1 scores 46.82 at $2.37. At the bottom rung the expensive model is eleven points ahead, for nearly six times the money
• Medium effort — 40.74 at $0.59 against 48.92 at $2.98. Fable 5.1 holds an eight-point lead at five times the cost
• High effort — 46.74 at $1.08 against 51.15 at $3.91. The gap narrows to four and a half points; the cost ratio holds at about 3.6×
• Xhigh effort — 51.85 at $2.74 against 53.20 at $5.98. One and a third points apart, at 46% of the price
• Max effort — 55.98 at $7.60 against 53.35 at $7.63. The ordering inverts and the cost converges to within half a percent
Read the fourth and fifth rows together and the shape of this matchup appears. Claude Sonnet 5.5 at xhigh delivers 51.85 — within one and a half index points of Claude Fable 5.1's best score at any effort — for $2.74 a task against $7.63. If your work needs to be near the frontier rather than at it, that row is the entire answer, and it saves 64% per completed task. Push Sonnet 5.5 to max and you buy three more points for 2.8 times the money, which lands you at the same bill as Fable 5.1's max with a marginally better score and a very different reasoning profile.
That last sentence is worth pinning down, because it is counterintuitive enough to check twice. At matching max effort, the model with a five-fold cheaper rate card costs the same per task as the model with the expensive one. It gets there by emitting 2.5 times as many output tokens, of which the reasoning share is three times as large. Anthropic's own documentation explains part of the mechanism: the tokenizer changed, and the same text produces about 30% more tokens on Claude Sonnet 5.5 than on earlier models. That is one contributing factor among several, and it is a useful reminder that output-token counts are not comparable across tokenizer revisions without saying so.
Latency is where the two stop resembling each other
Score convergence does not extend to time. On the same 2026-09-29 snapshot, Artificial Analysis reports a time to first token of 254.41 seconds for Claude Fable 5.1 and 2.80 seconds for Claude Sonnet 5.5 on the long-prompt slice, with median output speeds of 67.9 and 49.4 tokens per second respectively. That is a two-order-of-magnitude difference in how long a caller waits before anything arrives, and it is the single most operationally consequential number in this comparison.
Both figures need their qualifiers stated plainly. Fable 5.1's TTFT is measured on the long prompt type; the evaluator's medium-prompt slice for the same model on the same day is 117.60 seconds, so the 254-second figure is the tail and not a constant. Sonnet 5.5's 2.80 seconds is its overall median, with the evaluator's own variance band running from 1.91 seconds at the fifth percentile to 4.23 at the ninety-fifth. Anthropic's platform documentation describes Fable 5.1's comparative latency as slower and Sonnet 5.5's as fast, which is directionally consistent without being a measurement.
Our own infrastructure gives a third reading, because we route Claude Fable 5.1 and measure what we serve. Over the seven days to September 28, on OrcaRouter's own traffic, Claude Fable 5.1 posts a median first token at 6,973 milliseconds with a 95th percentile at 10.0 seconds, an output speed of 65.8 tokens per second and an error rate of 0.349%. That p50 is more than a hundred times faster than the evaluator's long-prompt figure, which is not a contradiction so much as a definition of terms: a developer-facing first token over a short prompt and a benchmark's first token over a long one are measurements of different things. Anyone planning around the leaderboard number should know which one they are buying.

Sourcing honesty, retention, and the migration nobody mentions
Vendor benchmark tables and independent evaluations are different instruments and should never be subtracted from one another. Anthropic's own launch table for Claude Sonnet 5.5 reports Terminal-Bench 4.0 at 70.6%. Artificial Analysis, running its own harness, reports 63.6% for the same model on an evaluation of the same name. For Claude Fable 5.1 the split runs the other way: Anthropic reported 55.8% at launch and the independent harness reads 52.0%. Neither pair is a correction of the other. Judge each on its own terms, and prefer the independent one when you are deciding where to send production traffic.
The retention posture differs too, and it is the kind of difference that only shows up in a procurement review. Anthropic states that Claude Sonnet 5.5 is available with zero data retention, the same posture it took with Opus 5.5 and Sonnet 5. Claude Fable 5.1's launch material describes an incremental data-usage policy living alongside it and a separate deployment under Project Glasswing carrying a default 30-day retention window for safety monitoring, with the zero-retention arrangement described as the interim state. If your team answers to a data-processing review, that distinction is a real line item rather than a footnote, and it is worth reading Anthropic's current policy page rather than any summary of it, including this one.
The migration cost is asymmetric in a way that rarely makes it into launch coverage. Moving to Claude Sonnet 5.5 is not a model-string change. Anthropic's migration guide states that non-default temperature, top_p and top_k values now return a 400 error, that a tool_choice of any or a named tool is rejected outright, and that thinking: {"type": "disabled"} must become between_tools if you were running with up-front thinking off. The between_tools type is accepted at low, medium and high effort and returns a 400 at xhigh or max. On the Claude API and Google Cloud, computer use is supported only through the computer_toolset_20260801 toolset; the older computer_20251124 is refused. None of those apply to a move onto Claude Fable 5.1, which is a same-family upgrade from Fable 5 with three breaking changes of its own. If you are weighing the two as parallel migration targets, the cheaper model is the more expensive one to adopt.

Which one to put where
The decision that follows from the numbers is a decision about how much of your traffic is well-scoped.
• Bounded, high-volume, tool-driven work — Claude Sonnet 5.5, at xhigh effort if you need to be near the frontier. It is one and a half index points behind Fable 5.1's best for 64% less per completed task, it returns its first token in about three seconds, and its cheaper cache-write line makes a rebuilt context prefix affordable
• Long-horizon, open-ended work where a wrong answer is expensive — Claude Fable 5.1. It matches Sonnet 5.5's max score at the same per-task cost while getting there with a third of the reasoning tokens and a quarter of the input, which matters when the task is genuinely open-ended and the eval suite's shape resembles your problem
• Latency-bound interactive surfaces — Claude Sonnet 5.5, and this is not close. On the evaluator's own snapshot the first-token gap is two orders of magnitude on the long prompt type, and the medium-prompt reading narrows it only to about 118 seconds. A chat surface cannot be built on the Fable 5.1 number without batching or streaming around it
• Bulk generation with long outputs — Claude Sonnet 5.5, because of the 300,000-token output ceiling available on the Message Batches API behind the output-300k-2026-03-24 header, against a 128,000-token ceiling everywhere else. Confirm your batch pricing before committing; the batched rate is not identical across the two cards
• Work that must span accounts or hand reasoning between systems — read Anthropic's preserved-thinking documentation before either. Thinking blocks are bound to the model and the account that produced them, and Sonnet 5.5 is the first Sonnet to ship with classifiers that prevent reasoning extraction. That is a compliance feature and a constraint at the same time
• Regulated or retention-sensitive workloads — check the current policy rather than the launch coverage. Claude Sonnet 5.5 launched with zero data retention available; Fable 5.1's material describes an incremental policy with a default 30-day window on one deployment and zero retention as an interim state

On OrcaRouter, Claude Fable 5.1 and Claude Sonnet 5 both sit behind one credential at the provider's list price with 0% markup, so the cache-read line and the tokenizer change are the only things that move your bill when you switch between them. Claude Sonnet 5.5 is not yet a route on our platform — the model is one day old — so the honest statement is that today you reach it through Anthropic's own API. Anyone running Fable 5.1 through us can price the switch on the day it lands without changing anything else about their integration, and in the meantime automatic failover across providers is what makes it safe to try a model this new on a path that has to stay up.
The verdict is narrower than the price gap suggests. Claude Sonnet 5.5 is the better default for almost everything bounded — it scores three points higher on the independent index, it answers in seconds, and at xhigh it gets within a point and a half of Fable 5.1's ceiling for well under half the cost per task. Claude Fable 5.1 keeps one real argument, and it is the one Anthropic's own comparison table makes: when the work is open-ended and the cost of being wrong exceeds the cost of thinking longer, five times the input price buys a model that reaches the same score without spending three times the reasoning tokens to get there. Pick on the shape of the task, not the size of the rate card, because at the top of the effort ladder those two models bill you the same.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
