
Claude Haiku 5.5 vs Gemini 3.5 Flash-Lite: Cheaper per Token, Dearer per Answer
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 65 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Gemini 3.5 Flash-Lite costs $0.30 and $2.50 on the same two meters. The Anthropic model is a third of the price going in and a fifth of the price coming out, and it scores 43.40 against 22.2 on the Artificial Analysis Intelligence Index v4.3.2 — a gap of more than twenty points in a field where ten is a generational jump. On the rate card this is not a close comparison, and on capability it is not a close comparison either.
On the invoice it is a close comparison, and it goes the other way. Put both models on the same independent board and charge each of them for a finished task rather than a token, and Gemini 3.5 Flash-Lite comes out cheaper per answer. Artificial Analysis puts Claude Haiku 5.5 at $0.21 per Intelligence Index task and Gemini 3.5 Flash-Lite at $0.1235 — a 1.7-to-one advantage for the model that pays five times more for every token it emits.
That inversion is the whole of this comparison. Claude Haiku 5.5 replaces the cheapest tier Anthropic has ever shipped and beats everything near it on the shared board. Gemini 3.5 Flash-Lite was released on July 21, 2026 and has been the value pick in this band since the summer. The question a buyer actually has is not which one is better, because the answer to that is settled. It is which one to put in front of a request that has to come back cheaply.
Everything below is per million tokens in US dollars. The list prices are the vendors' own published rates. The independent scores, the cost-per-task figures and the speed measurements attributed to Artificial Analysis are that evaluator's. The latency figures for Gemini 3.5 Flash-Lite come from our own metered traffic. Where a number is carried on the model record we hold rather than read from the evaluator's page on the day, we have treated that evaluator as the source and not ourselves.
The sticker and the invoice disagree
The cost-per-task gap is not explained by the rate card, and the reason it exists is worth knowing before either model goes near production.
Claude Haiku 5.5 is verbose, and the evaluator says so in as many words. Running the Intelligence Index, Artificial Analysis recorded 440 million output tokens from the model against a 100 million median across the board, and flagged it very verbose. Thinking bills as output, thinking is on by default with an effort dial that starts at medium, and a cheap input rate does not help a workload whose bill is mostly output.
Gemini 3.5 Flash-Lite is not terse either, but it is measurably less expensive per job. The same board records 17,507 output tokens per completed index task for it — 10,405 of them reasoning and 7,102 answer. Set that against the Anthropic model's token volume and the arithmetic resolves: a five-to-one advantage on the output rate is partly consumed by a model that emits the larger response, and what survives at the task level is a small disadvantage rather than a large advantage.
There is a second, quieter effect running the same way. Anthropic's documentation says Claude Haiku 5.5 uses the tokenizer shared with Claude 4.7 and later, so the same text counts as roughly 30% more tokens than it did on the model this one replaces. The rate fell by a factor of three against Google's tier. The token count did not fall, and on the same text it rose.
Both models, on the numbers that matter
Cached rates and list prices from Anthropic's and Google's own documentation; independent figures attributed to Artificial Analysis; live latency from our own seven-day traffic window. Cached 2026-10-08.

• Input — Claude Haiku 5.5 $0.10 up to 100,000 tokens and $0.50 above; Gemini 3.5 Flash-Lite $0.30 flat across a 1,048,576-token window.
• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens and $2.50 above; Gemini 3.5 Flash-Lite $2.50 flat.
• Cached input — Claude Haiku 5.5 $0.01 up to 100,000 tokens and $0.05 above; Gemini 3.5 Flash-Lite $0.03. Cache writes are $0.125 and $0.625 per million tokens for Claude Haiku 5.5.
• Context and output — a million tokens for each. Maximum output is 128,000 tokens for Claude Haiku 5.5 against 65,536 for Gemini 3.5 Flash-Lite, so Anthropic's ceiling is roughly double.
• Input modality — text and images for Claude Haiku 5.5; text, images, video, audio and files for Gemini 3.5 Flash-Lite. Both return text.
• Independent score — 43.40 for Claude Haiku 5.5 (Max), second of 182 models on Intelligence Index v4.3.2, against 22.2 for Gemini 3.5 Flash-Lite, ninety-sixth of 147 and in the thirty-fourth percentile of that board.
• Independent cost per task — $0.21 against $0.1235 on the same board.
• Throughput — 241.9 output tokens per second measured for Claude Haiku 5.5 by Artificial Analysis, against 280.0 for Gemini 3.5 Flash-Lite measured on our own playground over the last seven days. Those are two harnesses, not one, so read them as an order of magnitude and not as a race.
Twenty-one points, and what they are made of
A twenty-one-point gap on an index that runs into the sixties is not a margin. It is the difference between a model that can carry an agentic loop and a model that should be handed a single well-specified call.
The two vendors do not measure each other's harnesses, so their own suites cannot be lined up. Anthropic publishes GDPval-AA v2.1, OSWorld 2.1, Terminal-Bench 4.0 and FrontierCode results for Claude Haiku 5.5; Google publishes its own set for the Flash-Lite tier; neither ran the other's. The only apples-to-apples comparison available is the independent board, and there the Anthropic model is ahead by a margin that is not close.
The evaluator's sub-scores for Gemini 3.5 Flash-Lite put the shape of that tier on the record. It posts 83.8% on GPQA Diamond and 54.2% on SWE-Bench Pro, 54% on Terminal-bench 2.1, 41.3% on SciCode and 39.2% on MLE-Bench, with 18.8% on Humanity's Last Exam. Those are the numbers of a competent, cheap, general-purpose tier — strong on knowledge questions, visibly behind on the hardest reasoning and research evaluations. Claude Haiku 5.5 at 43.40 on the composite sits in a different band entirely, and the practical reading is that work which needs multi-step planning over a long horizon is not a job for the Google tier at all.
A million tokens of window, and 65,536 of answer
The two context windows are the same size and they are not the same product.
Long context on Gemini 3.5 Flash-Lite degrades with distance in a way that is worth pricing before it is trusted. On GDM-MRCR v2, the eight-needle retrieval evaluation, it averages 72.2% when the needles sit inside 128,000 tokens and 21.3% on the one-million-token pointwise variant. Its Long-Context Recall figure is 76%, and CharXiv Reasoning comes in at 74.5% without tools and 76.5% with them. A million-token window is a real capability; retrieving reliably from the far end of it is a smaller one, and the two numbers above are the distance between them. The Claude model's window has not been measured the same way on the same board, so no equivalent figure is available for it, and this piece will not invent one.
The output ceilings are the more immediately useful difference. Claude Haiku 5.5 will emit 128,000 tokens in one response; Gemini 3.5 Flash-Lite stops at 65,536. If the artefact you want back is a long file, a full migration, a generated test suite or a transcript-scale rewrite, one of these two models can produce it in a single call and the other cannot — and a job split across two calls costs more than twice one call, because the shared prefix is sent twice.

Audio and video is not a price question
Gemini 3.5 Flash-Lite accepts video, audio and files at the same rate as text. Claude Haiku 5.5 accepts text and images.
For a pipeline that ingests recorded calls, screen capture or surveillance footage, the capability difference that matters is not the twenty-one index points. It is that one of these models takes the input directly, and the other needs a transcription or frame-extraction stage in front of it — a second model, a second bill, and a new failure mode sitting between the source and the answer. Teams in that position are not choosing between two cheap tiers. They are choosing between one model and a pipeline.
The reverse holds for images and documents. Both models read images natively, and Claude Haiku 5.5 does it at 43.40 on the composite, so a page-image extraction or diagram-understanding workload with no audio in it belongs on the Anthropic side on the numbers rather than on a modality argument.
Running one of the two from here
Gemini 3.5 Flash-Lite is on the OrcaRouter catalogue at Google's list price with 0% markup, which means the provider's rate is passed through rather than blended and a change to Google's card reaches your invoice on the day it reaches theirs. That is one key and one SDK for Google's tier alongside Claude Sonnet 5.5 and Claude Opus 5.5, which are on the catalogue as well — so an A/B between a Google Flash-Lite leg and an Anthropic leg is already a configuration rather than an integration project. Automatic failover sits underneath, so neither leg is a single point of failure, and the routing DSL will express the escalation in the route if that is where the logic belongs.
The latency profile of that leg is on our own measurement rather than the vendor's. Across the last seven days of live traffic, Gemini 3.5 Flash-Lite has held a p50 of 2,426 milliseconds and a p95 of 8,600 milliseconds at 280 output tokens per second, with an error rate of 0.313%. Read it as a rolling window and not a promise: it moves as traffic and provider conditions move, and the number to trust for a given pipeline is the one you meter yourself after a week.

Claude Haiku 5.5 is not on the catalogue. It shipped on October 7, 2026 — a day before this was written — and it is not routable from us today, so the Anthropic side of this comparison is currently reachable only at Anthropic. Saying that plainly is better than leaving it implied. A gap between a model shipping and a model being callable through us is normal, it is not a promise about when it closes, and a reader planning a migration should plan around the calendar they can see rather than the one they would prefer.
Which one, and for whom
If the work is text in and text out, if the prompts fit under 100,000 tokens, and if the task needs multi-step planning or long-horizon agents, Claude Haiku 5.5 is the stronger model by twenty-one points and the cheaper one per token by a factor of three to five. Budget on cost per task rather than on the rate card, expect roughly $0.21 for the kind of job the independent board measures, and recount your prompts before you forecast anything, because the tokenizer change moves the count by about thirty percent on its own.
If the work is short, high-volume and answer-shaped — a tag, a category, an extraction, a guardrail decision — then the arithmetic favours Gemini 3.5 Flash-Lite and does so for an unglamorous reason. The bill for a call that emits very little is dominated by input, the two input rates are $0.30 against $0.10, and the Google model's much shorter task footprint means the cheaper rate wins the finished job even though it loses the sticker. Add video, audio or file input and the choice stops being about price at all.
What the pairing actually settles is narrower than "the new Haiku is cheap." It settles that the headline rate on a cheap tier is a poor guide to what that tier costs you, and that the model which looks three times dearer on the pricing page can send you the smaller invoice. Whichever way you land, measure the tokens you emit and not the ones you send, because on both of these models that is the meter the bill is really written on.
