
Claude Sonnet 5.5: A Stealth-Test Leak Built From Four Numbers That Already Exist
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 180 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1277 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 110 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
Four numbers were posted on X on September 24 as the evidence that Claude Sonnet 5.5 is "already being stealth tested": a 1M-token context window, a 128K-token maximum output, $2 per million input tokens and $10 per million output tokens. All four are on Anthropic's own price list right now, attached to Claude Sonnet 5 — the model a 5.5 would replace, generally available since June 30, 2026. The same $2 / $10 headline also belongs to GPT-6 Sol, which OpenAI shipped on September 22, and that overlap is what makes the post read as an answer to OpenAI rather than a routine tier refresh: "This reads to me as a direct response to GPT-6-Sol. Same pricing and presumably similar p[erformance]."
None of that makes the claim wrong. It makes the claim unfalsifiable, which is a different and more useful problem to write down. A stealth-test sighting is only as good as the artefact behind it, and the artefact behind this one is a spec sheet that is indistinguishable from the model Anthropic has been serving for twelve weeks.
What was actually posted
The source is a single X post from the account @kimmonismus, dated September 24, 2026. It states the context window, the output ceiling and the two prices, compares the pricing to GPT-6 Sol, and concludes that Sonnet 5.5 is in testing. There is no screenshot, no model string, no API response, no console listing, no testers' description of behaviour, and no second observer attached to any of it. The post is a claim with numbers in it, and the numbers are the tier's public limits.
That is worth separating from the other kind of Anthropic pre-release story, because 2026 has produced both and they look nothing alike. The August early-access identifiers claude-marshmallow-eap and claude-melon-eap were strings that existed in requests — ugly, unglamorous, and checkable. The Opus 5.2 story running through September is a routing report: developers noticing that the backend behind the "Opus 5" label answered differently, and describing the behaviour. Both are artefacts. "Already being stealth tested" with no artefact attached is a summary of what the Sonnet tier already sells.
There is also a pattern worth noting in the source itself. The same account posted a round-up on September 20 that opened with four models it said were "already seeing testing for" — Claude Sonnet 5.2, Claude Opus 5.2, Claude Fable 5.2 and Gemini 4 Pro. Of those four names, the Opus and Fable entries have separate, corroborated reporting behind them and the Gemini entry has sightings of its own. The Sonnet entry had a position in a list. Four days later the Sonnet claim has grown a spec sheet, and the spec sheet is the model that already exists.
The four numbers are Claude Sonnet 5's numbers
This is the part that takes thirty seconds to check and almost nobody has. Anthropic's models overview page, read today, lists Claude Sonnet 5 at a 1M-token context window, 128K maximum output and $2 / $10 per million tokens. Line the leak up against it:
• Context window — 1M tokens claimed for Claude Sonnet 5.5; 1M tokens published for Claude Sonnet 5
• Maximum output — 128K tokens claimed; 128K tokens published
• Input price — $2 per million claimed; $2 per million published
• Output price — $10 per million claimed; $10 per million published
• Model ID — nothing published for Claude Sonnet 5.5; claude-sonnet-5 for Claude Sonnet 5
• Independent scores — none for the claimed model; Artificial Analysis places Claude Sonnet 5 at an Intelligence Index of 38 at rank 54 of the 210 models it measures, in the adaptive-reasoning, max-effort configuration

There is a second layer to the pricing point that the leak gets backwards. When Claude Sonnet 5 launched, $2 / $10 was announced as introductory pricing running through August 31, 2026, with a scheduled step up to $3 / $15 on September 1. That step up did not happen. Anthropic's pricing page now carries a footnote stating that $2 / $10 "is now the standard price" and that "the previously scheduled increase to $3 / $15 per million input/output tokens on September 1, 2026 will not occur." So the price the leak presents as the new model's competitive weapon is the incumbent's permanent rate — which means a Sonnet 5.5 launching at $2 / $10 would be launching at no discount at all, into a tier that is already priced against GPT-6 Sol.
The one thing Anthropic has actually said
Strip out everything from unnamed leakers and one first-party sentence survives. When Anthropic launched Claude Opus 5.5 on September 22, 2026, it said that Claude Sonnet 5.5 and Claude Haiku 5.5 would follow "in the coming weeks," bringing comparable gains in performance, efficiency and safety. That sentence is the entire confirmed record. It carries no model ID, no context window, no price, no benchmark and no date — and in particular it does not say anything about a stealth test, a gray route, or testing of any kind.
The rest of Anthropic's public surface agrees with that reading. The models overview still lists four models — Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5 and Claude Haiku 4.5 — with no fifth row, no pricing entry and no "upcoming" marker for a Sonnet successor. If a Sonnet-tier checkpoint were being served somewhere a developer could reach, the identifier would be the first thing to surface, because a served model has to have a name in a request. Nothing has.
One thing the September announcement did settle: it contradicted the earlier leak claiming the Haiku line would be retired and the 5.5 generation would restructure the lineup. Haiku 5.5 is coming, per Anthropic, which means the restructure story was wrong in the direction leakers rarely are — the cheap tier is staying.
The August leak said 2 million tokens. This one says 1 million.
Six weeks before the September 24 post, on August 9, a different X account posted a "SONNET 5.5 LEAK" spec sheet with a much bolder claim: a 2-million-token context window, double Sonnet 5's, faster inference, better browser and terminal tool use, performance approaching Claude Fable 5, and a release "next month." That leak came with a graphic styled to look like an Anthropic preview card. It also came with a fatal detail its critics fixed on immediately: the codename it gave, "Fennec," was the codename that had already shipped with Claude Sonnet 5 on June 30. A codename cannot be fresh evidence of a successor to the model that used it.
So the two leak waves disagree on the single most-cited number. August said 2M context. September says 1M. Both cannot be describing the same build. The likeliest reading is the boring one: the September post describes the Sonnet tier's published limits because that is what a Sonnet-tier model looks like from the outside, and the August post was a guess at what a refresh would plausibly double. Note which direction the number moved — down, toward what already exists.
So is the GPT-6 Sol comparison right?
Mostly, and the place where it breaks is the interesting part. GPT-6 Sol shipped on September 22, 2026, the same day as Claude Opus 5.5, and its list price starts at $2 per million input and $10 per million output — exactly the Sonnet 5 rate and exactly what the leak claims for Sonnet 5.5. It serves a 1,050,000-token context window with a 128K output ceiling, both a shade away from Claude Sonnet 5's numbers. The cheaper sibling, GPT-6 Luna, sits at $0.10 / $0.50.
The break is long context. GPT-6 Sol prices in two tiers: $2 / $10 up to 272,000 input tokens, then $4 / $15 beyond that. Anthropic does not do that on the Sonnet line — Claude Sonnet 5 bills the full 1M window at standard rates, so a 900K-token request costs the same per token as a 9K one. A real Sonnet 5.5 versus a real GPT-6 Sol is therefore a price tie for ordinary prompts and a two-times gap for the long-document and large-repo workloads both models are sold on. "Same pricing" is true at the headline and wrong at the workloads where the tier is chosen.
A second caveat sits underneath the whole comparison. Claude Sonnet 5 uses Anthropic's newer tokenizer, which produces roughly 30% more tokens for the same text than the previous one. A headline rate is not a bill; the bill is tokens per task, and that is the number no leak has ever carried.
What would make a Sonnet 5.5 matter
Set aside whether it exists and ask what would change for someone paying for tokens. Four things, and none of them is a spec sheet.
The first is a price below $2 / $10. With the September 1 increase cancelled, the incumbent's rate is now permanent — so a successor holding the same rate is a lateral move, and the only Sonnet-tier news that moves money is a cut.
The second is context past 1M. That was the August leak's actual proposition and it is the one capability gap the leak can name that the incumbent does not already cover.
The third is tokens per task. Claude Sonnet 5's adaptive thinking is on by default and independent analysis has measured it generating materially more output tokens per task than its predecessor. A version that keeps the $2 / $10 rate and stops that inflation would cut real bills further than a headline price cut.
The fourth is an identifier. An API model string, a docs row, a pricing entry, a model card — any one of those ends the argument, and the absence of all four after a leak that explicit is itself the most telling fact in this story.
What you can call today
Claude Sonnet 5 is not a placeholder while the tier sorts itself out. It is the model most Claude API traffic already runs on, with a 1M-token window, 128K output and the $2 / $10 rate that is now standard rather than promotional.

It has been on OrcaRouter since June at Anthropic's own $2 / $10, passed through at provider list price with no markup, so a vendor price change on the Sonnet line would be live here the same day rather than at the next billing cycle. The same key already reaches GPT-6 Sol at $2 / $10, which matters more than usual in a piece like this one: the comparison the leak leans on is a comparison you can run yourself on one credential instead of two contracts.

That is also the cheapest way to answer the question the leak raises. If a Sonnet 5.5 does appear, it arrives in the catalogue like anything else, and you route a slice of traffic at it behind automatic failover to Claude Sonnet 5 — no second integration, no renegotiation, and no production path riding on a model nobody has benchmarked independently. Trying an unproven tier is a routing decision, not a migration project.
What to watch instead of the next leak post
Three signals, in descending order of how much they would tell you. An identifier is decisive, because a served model has to have a name in a request and that is how every checkable Anthropic pre-release story in 2026 became checkable. A vendor acknowledgement — a model card, a docs row, a pricing line, or the model simply appearing in Anthropic's own list — ends it outright; the "coming weeks" sentence means one of those is genuinely owed. A tester report describing behaviour rather than specs is the weakest of the three, and still more than exists today.
Until one of those lands, the honest summary is that Anthropic has promised a Sonnet refresh in the coming weeks and one observer has described it using the numbers of the model it will replace. If your stack already runs claude-sonnet-5, there is nothing to wait for and nothing to plan around. If you are weighing Claude Sonnet 5 against GPT-6 Sol, the numbers you can act on are the ones already published — and they are the same numbers the leak is trying to sell you as news.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
