A generated hero card for 'Six times the cost for two points of index', badged 'COST PER TASK' with the kicker 'TWO FLASH-TIER MODELS, ONE BOARD' and the subtitle 'Claude Haiku 5.5 at 43.40 against Gemini 3.8 Flash at 40.93 on Intelligence Index v4.3.2.' Four chips read '43.40 vs 40.93', '$0.21 vs $1.24', '$0.10 vs $0.75' and 'rises Jan 1 2027'. Two cards below are labelled 'ANTHROPIC'S EDGE' ('higher on the shared board, and six times cheaper per finished task') and 'GOOGLE'S EDGE' ('speech and video input, and a flat rate across the window'). A footer reads 'Independent measurements from Artificial Analysis; list prices read October 8, 2026.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Claude Haiku 5.5 vs Gemini 3.8 Flash: Two Points of Index, Six Times the Cost per Task

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Claude Haiku 5.5 and Gemini 3.8 Flash are closer in measured capability than any other pair in this price band — 43.40 against 40.93 on the Artificial Analysis Intelligence Index v4.3.2, both placed inside the top fifty of a 225-model board. What separates them is what one of those index points costs. On the same independent measurement, a finished Intelligence Index task costs $0.21 on Claude Haiku 5.5 and $1.24 on Gemini 3.8 Flash.

That is a sixfold gap for a two-and-a-half-point difference in score, and it is the number that should decide this comparison. Everything else — the sticker rates, the modality list, the deadlines attached to Google's pricing — is an argument about the edges.

Both models, on the numbers that matter

Per million tokens in US dollars, from Ant​hropic's and Google's own pricing documentation, with the independent measurements attributed to Artificial Analysis. Cached 2026-10-08.

A generated scoreboard titled 'Claude Haiku 5.5 vs Gemini 3.8 Flash - on the numbers that matter'. Claude Haiku 5.5 reads input $0.10 under 100K and $0.50 above, output $0.50 and $2.50, cached input $0.01 and $0.05, maximum output 128,000 tokens, input modality text and images, Intelligence Index v4.3.2 43.40, cost per task $0.21. Gemini 3.8 Flash reads input $0.75 now and $1.50 after January 1 2027, output $3.75 and $7.50, cached input $0.075 plus storage, maximum output 65,536 tokens, input modality text images speech and video, Intelligence Index v4.3.2 40.93, cost per task $1.24. A footer reads 'Vendors' published list rates; Google's lower rate runs through December 31, 2026; independent scores from Artificial Analysis.'

• Input — Claude Haiku 5.5 $0.10 up to 100,000 tokens and $0.50 above; Gemini 3.8 Flash $0.75 flat, rising to $1.50 on January 1, 2027.

• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens and $2.50 above; Gemini 3.8 Flash $3.75 flat, rising to $7.50 on the same date.

• Cached input — Claude Haiku 5.5 $0.01 and $0.05; Gemini 3.8 Flash $0.075, plus a storage charge of $0.50 per million tokens per hour.

• Context — 1M tokens for both. Maximum output — 128,000 tokens for Claude Haiku 5.5 against 65,536 for Gemini 3.8 Flash.

• Input modality — text and images for Claude Haiku 5.5; text, images, speech and video for Gemini 3.8 Flash. This is the one line where Google's list is strictly longer.

• Independent score — 43.40 for Claude Haiku 5.5 (Max) against 40.93 for Gemini 3.8 Flash (High), on Intelligence Index v4.3.2.

• Independent cost per task — $0.21 against $1.24 on the same board.

• Speed — 241.9 output tokens per second against 125.2, both measured by Artificial Analysis.

Where the six-times figure comes from

The cost-per-task gap is larger than the sticker gap, and that is the interesting part, because it means the rate card alone does not explain it.

The sticker arithmetic is simple enough: $0.10 against $0.75 on input is seven and a half times cheaper, and $0.50 against $3.75 on output is the same ratio. If both models generated the same tokens for the same work, that ratio would carry straight through to the invoice.

It does not carry through, and both effects run in Google's favour. Claude Haiku 5.5 is verbose in a way its own vendor acknowledges: Artificial Analysis recorded 440 million output tokens while running the Index, against a 100 million median, and calls the model very verbose. Gemini 3.8 Flash generated 170 million against an 81 million median — also flagged verbose, but a third of the token count. Ant​hropic's own documentation adds the second effect: Claude Haiku 5.5 uses the newer tokenizer shared with Claude 4.7 and later, so the same text counts as roughly 30% more tokens than on the model it replaces.

Put together, the honest version of the pricing claim is this: a ten-to-one input rate advantage and a seven-and-a-half-to-one output rate advantage, partly consumed by a model that writes more tokens per answer, arriving at a six-to-one advantage in the measurement that charges for the whole job. Still a rout, and a smaller one than the rate card suggests.

Google's price has a date on it

The single most useful thing on Google's pricing page is a sentence most comparison tables drop.

Gemini 3.8 Flash is listed at $0.75 input and $3.75 output "through December 31, 2026," and at $1.50 and $7.50 "starting January 1, 2027." Cached input moves from $0.075 to $0.15 on the same date. That is an introductory rate with an expiry printed next to it, roughly eleven weeks out from this writing.

Nothing in Ant​hropic's card reads that way. Claude Haiku 5.5's rates are stated as the rates, with a two-tier structure keyed to prompt length rather than to a calendar, and a retirement commitment of not sooner than October 7, 2027. If you are modelling a pipeline that runs into next year, Google's January step is a doubling and it belongs in the model; Ant​hropic's does not have an equivalent event to plan for.

That cuts both ways, and the fair reading is not "one vendor is honest and the other is not." A promotional rate is a real rate for as long as it lasts, and Google is entitled to price a launch that way. It is simply a fact about the comparison that the cheaper-looking side of the January numbers is the one with the deadline.

A screenshot of Anthropic's model documentation showing the Claude Haiku 5.5 specifications row and the published pricing table beneath it, with the per-million-token input, output and cache rates and the prompt-length threshold at which the second tier begins.

What two points of index is worth

A 2.47-point gap on an index that runs to the sixties is not a rounding error and it is not a chasm. Both models sit in the same working band, and the practical question is whether the gap shows up on the tasks you actually run.

On the vendor's own figures the two are not measured on the same harness, so they cannot be lined up directly — Ant​hropic publishes GDPval-AA v2.1, OSWorld 2.1, Terminal-Bench 4.0 and FrontierCode results for Claude Haiku 5.5, and Google publishes its own set for Gemini 3.8 Flash, and neither vendor ran the other's suite. The only apples-to-apples comparison available is the independent board, and there the answer is that Ant​hropic is ahead by a margin measurable but small.

The capability difference that will matter more to some teams is on the modality line. Gemini 3.8 Flash accepts speech and video input; Claude Haiku 5.5 accepts text and images. A pipeline that ingests call recordings or screen capture natively is not choosing between these two models on price at all — one of them cannot do the job without a transcription or frame-extraction step in front of it, and that step has its own cost and its own failure modes.

Running the cheaper side of the pair

Gemini 3.8 Flash is on the OrcaRouter catalogue at Google's list price with 0% markup, which matters more here than usual: the provider's rate is passed through rather than blended, so if you want to see what the January step actually does to a bill, the number moves when Google's does. Claude Sonnet 5.5 and Claude Opus 5.5 are on the catalogue as well, which makes a three-way routing test — Google's Flash leg, an Ant​hropic mid tier, and an Ant​hropic small tier later — a configuration rather than a project. One key, one SDK, automatic failover underneath, and the routing DSL if you would rather express the escalation in the route than in application code.

Claude Haiku 5.5 itself is not on the catalogue yet, and it is better to say that than to leave it ambiguous. The Ant​hropic leg of this comparison is currently reachable only at Ant​hropic, while the Google leg can be called from us today at the price shown above.

A screenshot of the OrcaRouter catalogue page for Gemini 3.8 Flash, showing the model identifier in provider-slash-model form, the provider Google, the model description and a stat strip carrying the per-million-token input and output rates alongside the context window and maximum output.

Which one, and for whom

If your work is text in and text out, your prompts sit under 100,000 tokens, and you are paying per token for real volume, Claude Haiku 5.5 is the cheaper model by every measure in this piece and the higher-scoring one on the only shared board — with the caveat that the advantage is sixfold on cost per task rather than the tenfold the sticker implies.

If you need audio or video in, or you need a 65K-plus response, or your prompts regularly run long, the calculus changes: Gemini 3.8 Flash takes inputs Ant​hropic's small tier cannot, and past 100,000 tokens Ant​hropic's own second tier is priced well above Google's flat rate.

The one thing to hold on to is the dates. Google's rate is good through the end of the year and doubles after it. Ant​hropic's is stated flat with a length step. Whichever side of that you land on, build the calendar into the spreadsheet before you build the pipeline.