
Claude Sonnet 5.5 vs Gemini 3.1 Pro: Cheaper per Token, Eleven Times Dearer per Task
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 984 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
On the rate card, Claude Sonnet 5.5 is the cheaper model and it is not close on capability. It reached general availability on September 28, 2026 at $2.00 per million input tokens and $10.00 per million output tokens; Gemini 3.1 Pro, announced February 19, 2026 and still served from a preview endpoint, is $2.00 and $12.00. Sonnet 5.5 also scores 56 on Artificial Analysis's Intelligence Index v4.3 against Gemini 3.1 Pro Preview's 30 — a twenty-six point gap on a single harness. Gemini's 1,048,576-token context is the only headline figure it wins outright. Yet the same evaluator reports that finishing one open-ended task costs $0.67 on the Google model and $7.60 on the Anthropic one, which is a factor of eleven running the opposite way. Both numbers are correct, and the reason they coexist — a 65,536-token output ceiling on one side and a verbosity problem on the other — is the whole of this comparison.
What each endpoint actually is
Fix the identities before the arithmetic. The subject here is anthropic/claude-sonnet-5.5, released September 28, 2026, text, image and file input, 1,000,000-token context, 128,000-token maximum output, adaptive thinking with an effort parameter defaulting to high on the Claude Platform. The opponent is google/gemini-3.1-pro-preview, released February 19, 2026, text, image, video, audio and file input, 1,048,576-token context, 65,536-token maximum output. One of those two names carries a word the other does not, and it is not decoration.
The preview suffix has been attached for seven months. Google shipped several Flash releases in that window and no successor to the Pro tier; the model 3.1 Pro replaced is listed as shut down. None of that is a defect in the model — it has been live, stable and correctly priced the whole time — but it does mean the endpoint you are integrating against is not covered by a general-availability lifecycle commitment, and this family has already retired one Pro model. Treat that as a standing line item rather than a footnote, because it changes what a multi-year architecture decision costs if you get it wrong.
Two numbers that point in opposite directions
The per-token comparison first, because it is the one everybody runs and it favours the newer model.
• Input — $2.00 per million on both; an exact tie at the headline rate
• Output — $10.00 for Claude Sonnet 5.5 against $12.00 for Gemini 3.1 Pro; the Anthropic model is 17% cheaper
• Cache read — $0.20 per million on both below the tier boundary
• Cache write — $2.50 per million for the five-minute window on Claude Sonnet 5.5 against $0.375 for Gemini 3.1 Pro; Google is roughly seven times cheaper to write a cache entry with
• Context — 1,000,000 against 1,048,576; a difference with no practical consequence
• Maximum output — 128,000 tokens against 65,536; the Anthropic model has nearly twice the ceiling
• Input modality — text, image and file against text, image, video, audio and file
Then the per-task comparison, from the same evaluator, in the same week, in a maximum-effort configuration on both sides: $7.60 for Claude Sonnet 5.5 against $0.67 for Gemini 3.1 Pro. Eleven times. The mechanism is documented on the same page rather than inferred — Sonnet 5.5 generated 410 million tokens across the index suite against a median of 88 million for the field, which the evaluator describes as very verbose, while Gemini 3.1 Pro is described as fairly concise. When one model writes several times as much as the other, a 17% advantage on the output rate does not survive contact with the invoice.
The lesson generalises past these two models: a rate card prices a token, and you are billed for finished work. Any comparison that stops at the price sheet is measuring the wrong quantity.
The tier boundary at 200,000 tokens
Gemini 3.1 Pro's pricing is not flat, and the step is large enough to invert parts of the comparison above. Below a 200,000-token prompt the rates are $2.00 input and $12.00 output with a $0.20 cache read. Above that line the whole call reprices at $4.00 input, $18.00 output and $0.40 cache read — the input rate doubles to exactly twice Claude Sonnet 5.5's, and output goes from 17% cheaper on the Anthropic side to 80% more expensive on the Google side.
The boundary is not exotic. A 200,000-token prompt is a small codebase, a long contract set, a few hours of transcript, or an agent that has been working for a while. Long-context work is precisely what a 1M-token window is sold for, so the tier that most rewards the window is also the tier where the price advantage disappears. If your prompts routinely cross 200,000 tokens, evaluate Gemini 3.1 Pro at $4.00 and $18.00, not at the rates on the landing page.
Cache writes deserve their own sentence, because they invert in the other direction and survive the tier change. At $0.375 per million against $2.50, Gemini is far cheaper to populate a cache with, and that rate does not move when you cross the boundary. For an agent that rebuilds a large prefix every turn rather than reusing one, that single line can be a larger structural advantage than the input rate — and it is the one line in this comparison that no tier boundary touches.
The 65,536-token ceiling, and why it decides architecture
The figure that most changes what you can build is the maximum output: 65,536 tokens on Gemini 3.1 Pro against 128,000 on Claude Sonnet 5.5. A ceiling is not a performance metric; it is a constraint on whether a task fits in one call.
Anything that emits a large artifact — a complete source file, a long report, a full document, a structured dataset — above 65,536 tokens in Gemini's case must be decomposed into a chain of calls with state passed between them. Decomposition costs more input tokens on every step, more orchestration code, and a new failure mode at the seams where one chunk ends and the next begins. A second constraint compounds it: Anthropic's own material puts Sonnet 5.5's practical threshold for reliable long work materially higher, and Gemini's ceiling is the tighter of the two by a factor of two. If your longest artifact is 40,000 tokens, this section is irrelevant and you should compare on per-task cost. If it is 100,000, the ceiling has already made the decision for you.
Modality, the one axis Gemini owns outright
Gemini 3.1 Pro accepts video and audio; Claude Sonnet 5.5 accepts text, images and files and cannot take either. For a workload built around media — call recordings, screen captures, product video, meeting audio — there is no substitution available in this pairing, and the price comparison is moot because only one of the two endpoints can accept the input at all. For text-and-image workloads the capability sets overlap and the price arithmetic above governs.
Gemini's other operational number is worth stating because it is easy to miss on a page that leads with intelligence: our catalogue measures a median first-token latency of 10,000 ms at p50 and an output rate near 1,283 tokens per second, with a seven-day error rate of 56.8%. That is an extremely fast model with a very high observed failure rate — a combination that suits batch work with generous retry budgets far better than it suits anything a user is waiting on. Claude Sonnet 5.5 is the slower, steadier profile by comparison. Error rate does not appear on either vendor's landing page; it is the kind of figure that only shows up when the candidates sit behind a single endpoint that measures them, which is one reason to evaluate both from one place rather than two dashboards.
What to route where
Send it to Claude Sonnet 5.5 when the work is text or image, when the quality bar is high enough that a twenty-six-point index gap matters, when output artifacts are long enough to need the 128,000-token ceiling, or when the workload is interactive and a 56.8% error rate is unacceptable. On structured, bounded tasks the verbosity penalty that produces the $7.60 figure largely evaporates, and the per-token advantage becomes real money.
Send it to Gemini 3.1 Pro when the input contains video or audio, when prompts stay under 200,000 tokens and the volume is high, when you are writing large caches frequently and the $0.375 cache write compounds, or when the task is batch-shaped and per-task cost is the only column that matters — at $0.67 a task against $7.60 you can afford a great deal of retry.
Both are routable on OrcaRouter today. google/gemini-3.1-pro-preview and anthropic/claude-sonnet-5 are both live catalogue entries on one credential at the provider's list price with 0% markup, which means the split above can be expressed as a routing rule rather than as two integrations — the media-shaped requests to one endpoint, the text-and-image requests to the other, with failover if either starts erroring. Claude Sonnet 5.5 is not yet a route on our platform; when it is listed, the same rule covers it without a new key.
The honest verdict for this matchup: Claude Sonnet 5.5 is the better model by a wide margin and the cheaper one per token, and Gemini 3.1 Pro is roughly eleven times cheaper per finished task on open-ended work, six times cheaper to write a cache with, and the only one of the two that can see a video. Anyone quoting the rate card at you is describing a comparison that does not exist; the one that does is about which requests you can bound.



Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
