
Claude Sonnet 5.5 vs Grok 4.5: Two $2 Input Rates and a Six-Minute Wait
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 931 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 192 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1177 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 70 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 107 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Claude Sonnet 5.5 and Grok 4.5 charge the same price for input: $2.00 per million tokens. That is where the agreement ends, and one number makes the point better than any spec table. On Artificial Analysis's Intelligence Index revision v4.3.2, the median time to first answer token for Claude Sonnet 5.5 at its Max Effort setting is 370.8 seconds. For Grok 4.5 it is 6.9 seconds. Same input rate. Fifty-three times the wait before anything comes back at all. Claude Sonnet 5.5, released September 28, 2026, and Grok 4.5, released July 8, 2026 and since overtaken twice inside its own family, are the two cheapest ways into the front half of the frontier on paper — and the paper is hiding the decision that actually matters, which is what you are buying with the wait.
That figure needs its caveat immediately, because a latency number without a configuration is meaningless: 370.8 seconds is Anthropic's default reasoning mode at maximum effort, spending a very large thinking budget before it emits an answer token. A caller who wants a fast first token configures something else. But the direction is real and it is not an artifact — Artificial Analysis measures Claude Sonnet 5.5's output speed at 138.7 tokens per second, second only to a handful of models, and Grok 4.5's at 59.2. The model that waits longest also writes fastest. What you are choosing between is a long deliberation followed by a fast answer, and a short one followed by a slower one.
What changed on September 28
Claude Sonnet 5.5 is one day old. Anthropic shipped it on September 28, 2026 under the API identifier claude-sonnet-5-5, alongside Bedrock, Google Cloud and Microsoft Foundry, and describes it as running 30% faster than Claude Sonnet 5 and costing up to 30% less per task for most work. Those are vendor figures and have not been independently reproduced; the price it is quoting against is real, because the rate card held flat at $2.00 per million input tokens, $10.00 per million output, $0.20 for cache reads and $2.50 for cache writes. The window is 1,000,000 tokens with a 128,000-token synchronous output ceiling, and 300,000 tokens on the Message Batches API behind the output-300k-2026-03-24 header. Knowledge cutoff is June 2026. Zero data retention is available from launch, and Anthropic is not retiring the model before September 28, 2027.

Anthropic's own docs page is also unusually explicit about the cost of the upgrade: five breaking changes separate Claude Sonnet 5.5 from the model most Claude traffic currently runs on, and they are covered further down.
Grok 4.5 is the older proposition and the more complicated one. SpaceXAI — the company still shipping under the xAI brand on its model cards — released it on July 8, 2026 as a 1.5-trillion-parameter mixture of experts, trained alongside the Cursor coding editor, with a 500,000-token window, text, image and file input, and a launch pitch built almost entirely on token economy rather than top scores. It is not the flagship any more. Grok 4.6 and Grok 4.7 both exist, both carry the same $2.00 / $6.00 list price, and our catalogue lists all three. The interesting question in September 2026 is therefore not whether Grok 4.5 is good; it is whether anything still justifies reaching for the oldest of the three.
The bill is not the rate card
Put the two side by side on the independent board and the shape of the trade appears.
• Intelligence Index, revision v4.3.2 — Claude Sonnet 5.5 56 at Max Effort vs Grok 4.5 39
• Cost per Index task — Claude Sonnet 5.5 $7.60 vs Grok 4.5 $1.04, both measured, not estimated
• Output rate — Claude Sonnet 5.5 $10.00 per million vs Grok 4.5 $6.00 per million
• Output speed — Claude Sonnet 5.5 138.7 tokens/sec vs Grok 4.5 59.2 tokens/sec
• Time to first answer token, Max Effort — Claude Sonnet 5.5 370.8 s vs Grok 4.5 6.9 s
• Tokens generated across the whole Index evaluation — Claude Sonnet 5.5 410M vs Grok 4.5 77M
• Context — Claude Sonnet 5.5 1,000,000 tokens vs Grok 4.5 500,000 tokens
The verbosity line is the mechanism behind the cost-per-task gap, and it is the single most useful number on this page. A model that emits 410 million tokens across an evaluation where its rival emits 77 million does not need a 67% higher output rate to cost seven times as much per task — but it has one anyway, and both effects push the same way. Anthropic's own framing agrees with the interpretation rather than contradicting it: "up to 30% less per task" is a claim about token counts, not rates, because the rates did not change between Claude Sonnet 5 and Claude Sonnet 5.5. Artificial Analysis reaches the same conclusion from the other side and describes the model in plain language as notably fast but very verbose. Against a lean model, that same property is Claude Sonnet 5.5's largest operating risk.

Where Grok 4.5's cheap rate stops being cheap
The $2.00 / $6.00 line is a tier, not a rate. On SpaceXAI's own developer documentation, that price applies up to 200,000 input tokens and doubles to $4.00 input and $12.00 output beyond it. Grok 4.5's window is 500,000 tokens, so the second half of the window it advertises costs twice what the first half does. Loading a repository-scale prompt into Grok 4.5 is not a $6-output operation; it is a $12-output operation, and at that point the model with the "expensive" $10 output rate is the cheaper one on any request that also emits a lot of tokens.
Claude Sonnet 5.5 has no equivalent tier. The $2.00 / $10.00 rate holds across the entire 1,000,000-token window, and cache reads bill at $0.20 — a tenth of the input rate — which is what makes the long-context case workable in practice rather than only on paper.
Window size is the other half of that argument and it points the same way. Grok 4.5's window is 500,000 tokens and every one of them bills at the advertised rate. Claude Sonnet 5.5 has 1,000,000 tokens to work with, and the shape of the workload is what makes that matter: a repository to reason over, a long agent transcript, a pile of documents that has to stay in view. Twice the ceiling, no tier break at half of it, and cache reads at a tenth of the input rate are what turn carrying that much context across many calls into a line item instead of a budget.
The practical split is not subtle:
• Short prompts with dense output — Grok 4.5 wins on rate and wins again on token habit
• Prompts past 200,000 input tokens — Grok 4.5's rate doubles and Claude Sonnet 5.5's does not
• Latency-sensitive interactive work — Grok 4.5 answers first, by a wide margin, at default settings
• Long-horizon agentic runs — Claude Sonnet 5.5's composite score and 128K output ceiling are what you are paying the extra task cost for, and the 300K Batches ceiling extends it further
The migration is not a model string
Anyone moving from Claude Sonnet 5 to Claude Sonnet 5.5 should read Anthropic's migration note before writing code, because five behaviours changed and only one of them fails loudly in an obvious place.
• Non-default temperature, top_p or top_k now returns a 400 — a sampling-heavy pipeline stops working rather than degrading
• thinking: {"type": "disabled"} returns a 400 and must become between_tools, with effort restricted to low, medium or high; xhigh and max are rejected
• Forced tool use — tool_choice of "any" or a named tool — is rejected outright, so a loop that forces a schema-valid call has to move to auto with strict tool use or structured outputs
• The older computer_20251124 computer-use tool is refused on the Claude API and on Google Cloud
• Thinking blocks are bound to the producing model and account, so reasoning carries forward from Claude Sonnet 5 but not sideways into a different family — and the advisor tool no longer accepts Claude Opus 4.8, Claude Opus 4.7 or Claude Sonnet 5 as advisors behind a Sonnet 5.5 executor
There is a sixth change that does not fail anything at all, which makes it the dangerous one: text between tool calls now returns inside thinking blocks. An application that streams that text to a user goes quiet between tool calls until it either sets a display value or turns off up-front thinking entirely. Nothing errors. The output simply stops appearing.
Grok 4.5 asks for none of that. It is an OpenAI-format model with the sampling surface intact, and moving to it from Grok 4.5's own predecessors is a model-string change. The migration cost lands entirely on the Claude side, and it is a day of work rather than an afternoon.
Both on one credential, and the honest limit of that
Holding a two-vendor comparison in your head is easy; running one is where the cost hides. OrcaRouter puts 200-plus models behind a single OpenAI-compatible endpoint with provider list price passed through and no markup added, so a rate change on either side is live here the same day rather than at the next contract renewal, with automatic failover between upstream providers and a routing DSL for composing calls. The whole Grok line is in the catalogue at SpaceXAI's own rates, and Grok 4.5 is routable as grok/grok-4.5 at $2.00 input and $6.00 output per million tokens across its 500,000-token window.
Claude Sonnet 5.5 is not in our catalogue, and this page is not going to imply otherwise. The route to it today is Anthropic's own API, or Bedrock and Google Cloud if that is where your infrastructure already sits. Claude Sonnet 5 has been routable here since June 30 at Anthropic's $2.00 and $10.00 and remains the model most Claude API traffic runs on, which makes it the natural failover target while Claude Sonnet 5.5 is still a day old and outside anyone's production evidence.
One caveat belongs on all of that, and it is not a small one. The Grok 4.5 row above is the vendor's pitch as our catalogue currently reproduces it; the traffic counter on that entry reads zero over the last seven days and its throughput figure is not being measured, with time to first token sitting at the one-second floor that a barely-called model tends to show. A model with no recent traffic here is not evidence the model is bad — 4.6 and 4.7 are where the family's calls have gone — but it does mean the operational numbers on that page are thin, and the latency figure worth trusting is the reading of a live route.

Which one to reach for
Grok 4.5 is the right model when the prompt is short, the answer is dense, and the integration already speaks OpenAI — and when the request will not cross 200,000 input tokens, because past that line the arithmetic changes hands. It is also the right choice when a first token has to arrive in seconds rather than minutes. What it is not is the current Grok flagship; Grok 4.6 and Grok 4.7 sit above it at the same list price, and the reason to pick 4.5 over either needs to be a specific one.
Claude Sonnet 5.5 is the right model when the prompt is enormous, the work is agentic enough that seventeen Index points and a 128K output ceiling justify a seven-fold higher cost per task, or the composite score is simply the thing being bought. Its verbosity is a design decision, not a defect, and it is the number to measure on your own traffic before committing — because $7.60 per Index task versus $1.04 is a proxy for a figure only your workload can produce.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
