A hero card headed 'Claude Haiku 5.5 vs Gemini 3.7 Flash' showing Claude Haiku 5.5 released October 7, 2026 on the left against Gemini 3.7 Flash marked deprecated on the right, an Intelligence Index v4.3.2 row reading 43.40 against 39.06, and a Terminal Bench 4.0 row reading 0.3283 against 0.1364.
Guides & Insights

Claude Haiku 5.5 vs Gemini 3.7 Flash: One of These Two Is Already Retired

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

There is an awkward fact sitting under any Claude Haiku 5.5 versus Gemini 3.7 Flash comparison written today: Gemini 3.7 Flash is deprecated. The vendor shipped it on August 13, 2026, replaced it with Gemini 3.8 Flash on September 2, and Artificial Analysis now carries the 3.7 page with a deprecation flag on it. Claude Haiku 5.5, by contrast, launched on October 7, 2026 and carries a retirement commitment of not sooner than October 2027.

That does not make the comparison useless. It changes the question. This is no longer "which of these two should I call" — for most teams the answer to that is neither, because the 3.7 line has been superseded in its own family. What it is good for is something subtler and arguably more useful: a clean read on how fast Google's flash tier moved in three weeks, and on which parts of a flash-tier model's spec sheet actually determine whether it gets used.

What you can still measure

Both models have been run on the same independent board, which is the only place the two can be lined up at all. On Artificial Analysis Intelligence Index v4.3.2, Claude Haiku 5.5 scores 43.40 in its Max configuration against 39.06 for Gemini 3.7 Flash in its High configuration — a 4.33-point gap.

Both vendors also publish their own evaluation sets, and they do not overlap in a way that permits direct comparison. The shared independent rows are the four Artificial Analysis ran on both:

• Terminal Bench 4.0 — 0.3283 for Claude Haiku 5.5 against 0.1364 for Gemini 3.7 Flash. A 19-point absolute gap, and the widest asymmetry in the set.

• SciCode — 0.5498 against 0.5718. Gemini 3.7 Flash edges it by 2.2 points.

• Humanity's Last Exam — 0.4439 against 0.4787. Gemini 3.7 Flash by 3.5 points.

• AutomationBench, partial score — 0.3541 against 0.6203. Gemini 3.7 Flash by 27 points.

Read that set carefully, because it contains the interesting result. Gemini 3.7 Flash wins three of the four shared benchmarks and still loses the composite index by 4.3 points. An index is a weighted blend, and the weighting puts the terminal-and-code rows — where the gap runs the other way by a much larger margin — ahead of the aggregate of the others. That is not a scoring artefact so much as a statement about what the index is for: it is trying to predict whether a model can finish a job, and on that question the 3.7 line was weak in a way its respectable science-benchmark scores did not hide.

A two-column scoreboard titled 'Claude Haiku 5.5 vs Gemini 3.7 Flash - the scoreboard' with rows Index v4.3.2 43.40 against 39.06, Terminal Bench 4.0 0.3283 against 0.1364, SciCode 0.5498 against 0.5718, Humanity's Last Exam 0.4439 against 0.4787, cost per task $0.21 against $0.93 and input price $0.10 against $0.75, footed 'Gemini 3.7 Flash is listed as deprecated; all figures per Artificial Analysis v4.3.2.'

What the 3.7 line was actually bad at

Two rows tell the story better than the index does.

Terminal Bench 4.0 at 0.1364 is a low number, and it sits next to a 0.8577 on the previous Terminal Bench generation from the same model. That is not a model that got worse; it is a benchmark that changed what it measures, and Gemini 3.7 Flash did not make the transition. Claude Haiku 5.5, on the same revised benchmark, scores 0.3283 — not spectacular in absolute terms, but nearly two and a half times the 3.7 Flash figure, on the row that most directly predicts agentic coding performance.

The second row is Tau-Bench Banking, where Gemini 3.7 Flash scores 0.3278. Tool-use reliability in a structured domain is exactly what a cheap model gets deployed for, and a sub-0.33 score is the kind of number that quietly kills a pilot. Claude Haiku 5.5 does not have a published Tau-Bench figure in the same table, so no comparison is available there — the honest thing is to say so rather than infer one.

Where Gemini 3.7 Flash held up is where it always held up: GPQA at 0.9455 and MMMU-Pro at 0.8549 are both genuine strengths, and its multimodal input was never in question. The failure was in the part of the spec sheet that decides whether an agent loop works, not the part that wins leaderboard rows.

What changed in the three weeks after it

Gemini 3.8 Flash arrived on September 2, 2026, and the movement inside Google's own flash tier is the more useful comparison for anyone planning a 2027 pipeline.

On the same index, Gemini 3.8 Flash measures 40.93 against the 3.7 line's 39.06 — a 1.87-point gain in three weeks. Terminal Bench 4.0 moved from 0.1364 to 0.1970. Tau-Bench Banking moved from 0.3278 to 0.4495. Those are the exact rows the 3.7 model was worst at, which suggests the successor was aimed at precisely this weakness rather than at a general capability bump.

The price did not move at all. Gemini 3.7 Flash was listed at $0.75 input and $3.75 output per million tokens, and Gemini 3.8 Flash launched at the same $0.75 and $3.75 — the identical rate card, attached to a model two index points better on the rows that matter for agents.

A capture of Google's Gemini API pricing documentation page, headed 'Gemini Developer API pricing', with the banner 'Gemini 3.8 Flash is now available. Try it out.' and the documentation's left-hand model list showing Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.6 Flash and Gemini 3.5 Flash.

The rate card is where this comparison really turns

Against Claude Haiku 5.5, Google's price is the whole story and it is not close.

• Input — Claude Haiku 5.5 $0.10 per million tokens up to 100,000, $0.50 above; Gemini 3.7 Flash $0.75 flat, or seven and a half times the Anthropic rate inside the first tier.

• Output — $0.50 for Claude Haiku 5.5 inside the first tier, $2.50 above it; $3.75 for Gemini 3.7 Flash.

• Cached input — $0.01 and $0.125 for Claude Haiku 5.5; $0.075 read and $0.75 write for Gemini 3.7 Flash.

• Context — one million tokens for both. Maximum output — 128,000 tokens for Claude Haiku 5.5 against 65,536 for Gemini 3.7 Flash.

• Input modality — text and images for Claude Haiku 5.5; text, images, speech and video for Gemini 3.7 Flash. This remains the one line where Google's spec is strictly longer, and it survived the succession to 3.8 unchanged.

• Independent cost per finished Index task — $0.21 for Claude Haiku 5.5 against $0.93 for Gemini 3.7 Flash. A 4.4× gap that is wider than the seven-and-a-half-to-one input ratio would suggest, because Anthropic's model is the verbose one here: 162,164 output tokens per Index task against 58,387 for Gemini 3.7 Flash. Anthropic also documents a tokenizer shared with Claude 4.7 and later that counts roughly 30% more tokens for the same text, which pushes the same way.

• Speed — 212.3 seconds per Index task for Gemini 3.7 Flash against 424.6 for Claude Haiku 5.5. Google's model is the faster of the two by a factor of two, which is the one line in this section where the cheaper model is not also the better one.

What this means if you are choosing today

The practical reading is that neither of these two is the right answer, for different reasons.

Gemini 3.7 Flash is deprecated, priced at the same rate as its successor, and worse than that successor on exactly the rows a cheap production model needs. Recommending it would be recommending a rate card attached to a superseded model — the successor costs the same and does more. Nobody choosing between these two in October 2026 should land on the 3.7 line, and the fact that its rate card is unchanged is what settles it.

Claude Haiku 5.5 is the live model, and on text-in, text-out work it is cheaper by every measure, higher on the composite, and vastly better on the two rows that predict agentic success. It is also the slower one and it cannot accept speech or video. Whether that matters is a question about your pipeline, not about the models.

The third option, which the index point gap between 3.8 and 3.7 makes visible, is that Google's flash tier is moving fast enough that a three-week-old comparison is already historical. Treat any flash-tier head-to-head as having a shelf life measured in weeks, not quarters.

Running whichever side you land on

Gemini 3.8 Flash — the successor, not the deprecated 3.7 — is on the OrcaRouter catalogue at Google's list price with 0% markup, so the rate Google publishes is the rate that reaches your bill and a change on their side is live on ours the same day. Claude Sonnet 5.5 and Claude Opus 5.5 are on the catalogue as well, which makes a two-vendor routing test a configuration rather than a project: one API for 200-plus models, one key, automatic failover underneath, and the routing DSL when you would rather hold the escalation in the route than in code.

Gemini 3.7 Flash itself is not on our catalogue, and Claude Haiku 5.5 is not either. Both remain reachable through their vendors' own APIs and, in Google's case, several third-party platforms. That is worth stating plainly rather than leaving a reader to discover it.

A capture of the OrcaRouter model page for Gemini 3.8 Flash under google/gemini-3.8-flash, showing a 65K maximum output, text, image, video, file and audio input, benchmarks published by Google on 2026-09-02, a p50 time to first token of 3.838 seconds, and the rates $0.75 input and $3.75 output per million tokens.

The number to take away

It is not the 4.33-point index gap. It is that Gemini 3.7 Flash won three of the four shared benchmarks and still lost the composite by four points, because the benchmark the index weights most heavily was the one it scored 0.1364 on. A flash-tier model's science scores can look fine while it fails at the thing cheap models are bought for.

The corollary is that the successor fixed precisely those rows within three weeks, at an unchanged price. On this evidence, the competitive pressure in this tier is not price. It is whether the model can finish a multi-step task, and that is now moving faster than rate cards do.