Hero card titled Claude Sonnet 5.5 vs Grok 4.6, subtitled the rate card says one thing the token bill says another, with pills reading output $10 vs $6, cost per task $7.60 vs $1.86, and Index 56 vs 44, and a footer citing Artificial Analysis v4.3.2 and vendor pricing.
Guides & Insights

Claude Sonnet 5.5 vs Grok 4.6: The Rate Card Says One Thing, the Token Bill Says Another

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The headline arithmetic looks settled before the article starts. Grok 4.6 costs $2 per million input tokens and $6 per million output tokens; Claude Sonnet 5.5 costs $2 and $10. Space​XAI's model is cheaper on output by 40%, matches on input, and has been generally available since August 12, 2026 — long enough that its quirks are documented rather than rumoured. Anthro​pic's model arrived on September 28, 2026, a day before this was written, and is the newest thing in the tier by a wide margin.

Then you reach the independent efficiency figures and the ordering flips. Artificial Analysis measures GPT-and-Claude-class models on a cost-per-task basis alongside its Intelligence Index, and on the v4.3.2 revision the two models here are $1.86 per index task for Grok 4.6 and $7.60 for Claude Sonnet 5.5. The model with the cheaper rate card is the expensive one to operate, by a factor of four. Working out why, and when that reversal does and does not hold, is the whole comparison.

Two models, two positions in their own families

Grok 4.6 is the current flagship of the Grok line — SpaceXAI's own developer documentation lists it as the latest, and it succeeded Grok 4.5 with the same 500,000-token context window, the same text, image and file input surface, and the same base pricing. It supports configurable reasoning effort, native tool calling, structured outputs and the full sampling surface, and as a first-class OpenAI Responses model it drops into existing agent frameworks without a translation layer. Artificial Analysis has it at an Intelligence Index of 44, ranked 31st of the 216 models it measures, with a GPQA Diamond figure of 94.9%.

Two-column scoreboard for Claude Sonnet 5.5 and Grok 4.6 across six rows. Left column Claude Sonnet 5.5: input $2.00 per million, output $10.00 per million, 1M-token context, AA Intelligence Index 56, cost per Index task $7.60, verbosity 410M output tokens. Right column Grok 4.6: input $2.00 per million, output $6.00 per million, 500K-token context, AA Intelligence Index 44, cost per Index task $1.86, verbosity 94M output tokens. A footer line reads Index, cost per task and verbosity per Artificial Analysis v4.3.2; prices per the vendors.

Claude Sonnet 5.5 is the second entry in Anthropic's 5.5 generation and the model the company has positioned as the best combination of speed and intelligence in its lineup. It ships as claude-sonnet-5-5/ on the Claude API and the three major clouds, carries a 1M-token context window with a 128K synchronous output ceiling, keeps Claude Sonnet 5's $2 / $10 rates for input, output and caching, and is available with zero data retention. On the same index revision it scores 56 and ranks third of 216.

So the two are not really in the same market. One is a 500,000-token frontier model that has been on sale for seven weeks at a lean rate. The other is a 1M-token general-purpose tier that costs nothing more per token than its predecessor and scores twelve points higher on the composite. The comparison is worth making anyway, because both are $2-input models and that is where procurement decisions actually start.

The four-times gap nobody puts on a slide

Rate cards price tokens. Bills are denominated in tasks, and the number of tokens a model spends reaching an answer is a design choice rather than a constant. That is where these two diverge, and the measured figures are stark:

• Output rate — Claude Sonnet 5.5 $10 per million vs Grok 4.6 $6 per million

• Cost per Intelligence Index task — Claude Sonnet 5.5 $7.60 vs Grok 4.6 $1.86

• Verbosity on the Index — Claude Sonnet 5.5 410M output tokens vs Grok 4.6 94M

• Intelligence Index — Claude Sonnet 5.5 56 vs Grok 4.6 44, both on v4.3.2

• Output speed — Grok 4.6 measured at 67.7 tokens per second against a median of 82.7; Anthropic reports Claude Sonnet 5.5 generating output 30%+ faster than Claude Sonnet 5, which Artificial Analysis timed at 78.7 tokens per second

The verbosity line is the mechanism. A model that emits four times the tokens does not need a 67% higher output rate to cost four times as much per task — but it helps, and both effects point the same way here. Anthropic's own framing supports the interpretation rather than contradicting it: the launch material claims Claude Sonnet 5.5 "costs up to 30% less per task" than Claude Sonnet 5 at identical per-token prices, which is a claim about token counts, not rates. Against an external baseline that is far leaner, the same property becomes the model's largest cost risk.

None of this makes the index gap go away. Twelve points of composite is a real capability difference, and a cheaper task that fails is not cheaper. It does mean the honest question is not "which rate is lower" but "how many tokens does the harder task take", and that number only exists once you run your own workload.

Artificial Analysis model page for Claude Sonnet 5 (Adaptive Reasoning, Max Effort), released June 2026, showing an Intelligence Index of 38, an output speed of 78.7 tokens per second against an average of 83, a cost of $2.00 per million input tokens and $10.00 per million output tokens, $5.09 per Intelligence Index task, a verbosity of 370M output tokens, and a 1M-token context window.

Context, effort and the shape of the integration

The second divergence is architectural, and it decides which of the two you can adopt without rewriting anything.

OrcaRouter model page for grok/grok-4.6, attributed to SpaceXAI and dated 2026-08-12, showing a 500K-token context window, text, image and file input, an input price of $2.00 and output price of $6.00 per million tokens, a p50 time to first token of 3.00 seconds, and OpenAI-compatible base URL https://api.orcarouter.ai/v1

Claude Sonnet 5.5 changes six behaviours relative to Claude Sonnet 5, five of them breaking. Sending thinking: {"type": "disabled"}/ now returns a 400 and must be replaced with between_tools/. Forced tool use — tool_choice/ of "any"/ or a named tool — is rejected outright, so a loop that forces a schema-valid call has to move to auto/ with strict tool use or structured outputs. The older computer_20251124/ computer-use tool is refused on the Claude API and Google Cloud. Thinking blocks are now bound to the model and the account that produced them, so a conversation can carry its reasoning forward from Claude Sonnet 5 but not sideways into a different model family. And the advisor tool no longer accepts Claude Opus 4.8, Claude Opus 4.7 or Claude Sonnet 5 as advisors behind a Sonnet 5.5 executor.

Grok 4.6 asks for none of that. It is an OpenAI-format model with the sampling surface intact, and the migration from Grok 4.5 is a model-string change. Its own cost structure has a catch, though, and it is a mirror image of Anthropic's: the $2 / $6 rate applies up to 200,000 input tokens, and everything beyond that bills at $4 / $12. Claude Sonnet 5.5 charges the standard rate across the entire 1M window.

Put the two structures next to each other and the choice stops being about which vendor is cheaper:

• Short prompts, dense output — Grok 4.6's leaner token habit and lower output rate win decisively

• Very long inputs — Claude Sonnet 5.5's flat 1M rate starts beating a doubling tier well before the context limit

• Large single outputs — Sonnet 5.5's 128K ceiling matches Grok 4.6's, and it reaches 300K on the Message Batches API behind the output-300k-2026-03-24/ header

• Existing OpenAI-shaped code — Grok 4.6 needs no changes; Sonnet 5.5 needs the tool-use and thinking work above

Getting both onto one credential

Holding a two-vendor comparison in your head is easy. Running one is where the cost hides: two accounts, two sets of limits, two billing surfaces, and a token count that has to be measured twice before it means anything.

OrcaRouter collapses that into one OpenAI-compatible endpoint covering 200-plus models, with provider list price passed through and no markup, so a vendor rate change on either the Grok or the Claude side is live here the same day rather than at the next contract renewal. The whole Grok 4.x line sits in the catalogue at the vendor's own rates, and Claude Sonnet 5 has been routable since June 30 at Anthropic's $2 / $10 — the model most Claude API traffic still runs on, measured at 150 output tokens per second with a p50 first-token time of roughly five seconds.

Claude Sonnet 5.5 specifically is not in our catalogue yet. Until it is, the route to it is the vendor's own API, and the way to test it without betting a workload on a model nobody has independently benchmarked past one composite is to send it a percentage of traffic behind automatic failover, with Grok 4.6 or Claude Sonnet 5 catching what it drops. Trying a brand-new tier is a routing decision, not a migration project.

What to do with the numbers

Grok 4.6 is the better model to reach for when the job is short, the output is dense, and the integration already speaks OpenAI. It is measurably leaner per task than anything else in this price band, its sampling controls are intact, and seven weeks of availability means its failure modes have been found by other people.

Claude Sonnet 5.5 is the better model when the input is enormous, when the composite score is the thing you are buying, or when the work is agentic enough that a twelve-point index gap and a 128K output ceiling matter more than a four-fold difference in cost per task. The migration checklist is real, and it is a day of work rather than a model-string edit.

What would settle it is a tokens-per-task measurement on your own traffic, run on both. Every number in this article that decides the outcome — $1.86 against $7.60, 94M against 410M — is a proxy for that one figure, and it is the only one of them you can produce yourself.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily