A generated hero card titled 'Claude Haiku 5.5 vs Gemini 3.6 Flash' with the kicker '7.5x THE COST PER ANSWER', showing Claude Haiku 5.5 at Input $0.10, Output $0.50 and Index 43.40 against Gemini 3.6 Flash at Input $0.75, Output $3.75 and Index 33.98, over a bar reading 'Cost per finished task $0.21 vs $1.60'.
Engineering & Research

Claude Haiku 5.5 vs Gemini 3.6 Flash: 7.5x the Cost per Answer, Nine Points Less Model

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Haiku 5.5 and Gemini 3.6 Flash sit in the same tier of their respective lineups — the small, fast, cheap one — and that is where the resemblance stops. On Artificial Analysis Intelligence Index v4.3.2, Claude Haiku 5.5 scores 43.40 in its Max configuration against 33.98 for Gemini 3.6 Flash in its High configuration. On the same run, one finished Index task costs $0.21 on Claude Haiku 5.5 and $1.60 on Gemini 3.6 Flash.

Read that pair together, because either alone misleads. The cheaper model is not marginally cheaper; it is 7.5 times cheaper per finished task. And the model it is beating is not marginally weaker; it is 9.42 index points weaker, which on a 182-model board is the distance between the upper quartile and the middle.

Both facts come from the same independent run, which matters because the vendor cards here do not tell you either one. Anthropic launched Claude Haiku 5.5 on October 7, 2026; Google shipped Gemini 3.6 Flash on July 21, 2026. The three months between them are the interesting part.

The two cards, side by side

Per million tokens in US dollars. Anthropic's and Google's rates come from their own pricing documentation; the shared rows are Artificial Analysis Intelligence Index v4.3.2. Cached 2026-10-08.

A generated two-column scoreboard titled 'Claude Haiku 5.5 vs Gemini 3.6 Flash — the scoreboard' with rows Index v4.3.2 43.40 against 33.98, cost per task $0.21 against $1.60, input price $0.10 against $0.75, output tokens per task 162,164 against 41,606, Terminal Bench 4.0 0.3283 against 0.0707 and input modality 'text, image' against 'text, image, video, audio', footed 'All figures per Artificial Analysis Intelligence Index v4.3.2; Gemini 3.6 Flash is flagged deprecated on the board.'

• Input — Claude Haiku 5.5 $0.10 up to 100,000 tokens and $0.50 above; Gemini 3.6 Flash $0.75 flat today, rising to $1.50 on January 1, 2027.

• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens and $2.50 above; Gemini 3.6 Flash $3.75 flat today, rising to $7.50 on the same date.

• Cached input — Claude Haiku 5.5 $0.01 and $0.05; Gemini 3.6 Flash $0.075, plus a storage charge of $0.50 per million tokens per hour.

• Context and output ceiling — 1M tokens for both. Claude Haiku 5.5 caps a response at 128,000 tokens; Gemini 3.6 Flash at 65,536.

• Modality — Claude Haiku 5.5 takes text and images. Gemini 3.6 Flash takes text, images, video, audio and files. This is the one row where Google's side is unambiguously ahead, and for a media-understanding pipeline it can decide the whole question.

• Independent score — 43.40 for Claude Haiku 5.5 (Max) against 33.98 for Gemini 3.6 Flash (High).

• Cost per finished task — $0.21 against $1.60, on the same evaluation run.

A capture of Anthropic's Claude Platform models documentation showing the Claude model comparison table — Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5 and Claude Haiku 5.5 'This model', the latter reading 1M context, 128K max output, 'From $0.10/$0.50', latency 'Fastest', default effort medium — above the Claude Haiku 5.5 specifications block with model ids claude-haiku-5-5, the pricing lines $0.10/$0.50 up to 100,000 tokens and $0.50/$2.50 over it, and the release line 'Released October 7, 2026'.

Why per-task, not per-token, is the number to use

The rate-card multiple looks like 7.5 on input and 7.5 on output already, so the per-task figure is not a surprise here — but it is worth showing why the two models land in the same place on one measure and so far apart on the other, because the same arithmetic cuts the other way in this batch's other matchups.

Gemini 3.6 Flash is the less verbose of the two. It emits 41,606 output tokens per Intelligence Index task — 20,061 reasoning and 21,545 answer — against Claude Haiku 5.5's 162,164, split 129,047 reasoning and 33,118 answer. Anthropic's model spends nearly four times as many tokens doing the same job, and it is the one whose output rate is a seventh of Google's, so the token gap and the rate gap multiply rather than cancel.

That combination — cheap meter times heavy use — is the honest description of Claude Haiku 5.5's economics, and Anthropic's own launch post acknowledges the other side of it: the model is "especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model." Past 100,000 tokens the input meter becomes $0.50 and the output meter $2.50, at which point the multiple over Gemini 3.6 Flash narrows to roughly 1.5x.

What Google's own tier did while this comparison sat still

Here is where the piece stops being a two-model comparison and becomes a warning about which Gemini you are actually comparing to.

Google has kept the Flash price fixed across three generations. Gemini 3.6 Flash, Gemini 3.7 Flash and Gemini 3.8 Flash all list at $0.75 input and $3.75 output on Google's own pricing documentation, with the same $0.075 cache rate and the same January 1, 2027 step-up to $1.50 and $7.50. Google's card for the 3.6 line describes it as "our previous generation Flash model," which is the vendor's own words for it.

What moved was the score, not the price. On the independent board, Gemini 3.6 Flash scores 33.98, Gemini 3.7 Flash 39.06, and Gemini 3.8 Flash 40.93 — all in their highest-intensity published configuration. Nine index points arrived inside Google's flash tier in six weeks, for no change in rate.

The consequence for anyone mid-evaluation is concrete: if you benchmarked this matchup three weeks ago against Gemini 3.7 Flash, your result is stale, and if you benchmark it against Gemini 3.8 Flash the gap to Claude Haiku 5.5 narrows to 2.47 points. Gemini 3.6 Flash is the version of this comparison that flatters Anthropic the most, and it is also the version Google is furthest from selling.

Artificial Analysis flags the 3.6 entry as deprecated. Google's pricing documentation still publishes the rate, and its model list still routes to a 3.6 section, so the honest reading is not "gone" but "superseded twice inside its own family in ten weeks" — and that is a fact about how fast this tier moves, not a reason to skip the comparison.

Who this actually decides for

There is a real case for Gemini 3.6 Flash, and it is not the score.

If you need audio or video in the prompt, Claude Haiku 5.5 cannot do it at all — it reads text and images and nothing else. Gemini 3.6 Flash reads text, images, video, audio and files, and our own catalogue records it at a measured 582 output tokens per second with a 4.1-second median time to first token over the last week, which is a fast model even by Flash standards. For a media-understanding pipeline the modality list is the entire decision and the index gap is a footnote.

If your work is text and images, the arithmetic is not close enough to argue about. At a 7:2:1 input-heavy blend — the shape most production traffic has — Claude Haiku 5.5 comes to $0.077 per million tokens against $0.63 for Gemini 3.6 Flash. Over a million blended tokens a day that is 77 cents against $6.30, and the second number gets worse on January 1, 2027 while the first one does not move.

The 100,000-token cliff is the one condition that flips it. A retrieval or long-document pipeline that routinely assembles 150,000-token prompts pays Anthropic's $0.50/$2.50 tier on the whole request, and at that point the two models are inside 1.5x of each other on price while Gemini 3.6 Flash still wins on modality and on the 91.8% figure Google publishes for multi-needle retrieval at 128k average. That is vendor-reported and unreproduced, and it belongs in exactly that framing.

A capture of the OrcaRouter model page for Gemini 3.6 Flash under google/gemini-3.6-flash, showing a 65K maximum output, text, image, video, file and audio input, public benchmarks published by Google dated 2026-07-21, a p50 time to first token of 4.1 seconds, and the page's own description of the model as Google's fast, cost-efficient multimodal model in the Gemini 3 family with a 1M-token context window.

Calling both, or calling one

Gemini 3.6 Flash is on the OrcaRouter catalogue today at google/gemini-3.6-flash, at Google's list price with 0% markup — $0.75 and $3.75 with a $0.075 cache read, carried as the provider charges it. That matters for this particular model because Google's January 1 step-up is a rate change, and a pass-through catalogue takes it the day it happens rather than at the next contract renewal. One API key and one SDK call either this model or any of the 200-plus others on the catalogue, so an A/B between a Google leg and an Anthropic leg is a model-string change with automatic failover underneath rather than a second integration.

Claude Haiku 5.5 is not on the catalogue. Anthropic's model is reachable through Anthropic's own API and through AWS, Google Cloud and Microsoft Azure, and that is the whole of the honest answer. What we do carry from that lineup is Claude Sonnet 5.5 and Claude Opus 5.5 — which is what makes the escalation path from a cheap classification call up to a mid-tier agentic call a routing decision rather than a project.

Which leaves the verdict simple enough to state plainly. For text and image work at short prompt lengths, Claude Haiku 5.5 wins on every axis except the response ceiling. For anything with audio or video in it, Gemini 3.6 Flash is the only one of the two that can answer at all — and if you are going to pay Google's rate for that capability, ask which Gemini you are paying it for, because on this tier you can get nine more index points for the same money from a newer sibling.

One API for 200-plus models, one key, automatic failover underneath, so an A/B between a Google leg and an Anthropic leg is a model-string change rather than a second integration.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily