A generated hero card for 'The American small model is now the cheaper one', badged 'PRICE FLOOR' with the kicker 'OCTOBER 7 2026' and the subtitle 'Claude Haiku 5.5 at $0.10 and $0.50 against DeepSeek V4 Flash - until 100,000 tokens.' Four chips read '$0.10 vs $0.15 in', '$0.50 vs $0.60 out', '100K step' and 'peak windows'. Two cards below are labelled 'UNDER 100,000 TOKENS' ('Anthropic is cheaper on both meters') and 'ABOVE IT' ('DeepSeek is cheaper by three to four times'). A footer reads 'Vendors' published list rates read October 8, 2026.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Claude Haiku 5.5 vs DeepSeek V4 Flash: Anthropic Just Undercut the Price Floor

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

For two years the cheap tier of the market had a shape everyone knew: the American labs charged more and the Chinese labs charged less, and you picked your side accordingly. Claude Haiku 5.5 broke that on October 7, 2026, by launching at $0.10 per million input tokens and $0.50 per million output tokens — which is below the published rate for DeepSeek V4 Flash, the model that has been the reference point for cheap since the summer.

It is a narrower win than it sounds, and the reasons are worth more than the headline. The Ant​hropic rate holds only for prompts up to 100,000 tokens. Above that it becomes $0.50 and $2.50, which is several times Dee​pSeek's price. And Dee​pSeek's own card has a shape nobody copies: it doubles during two weekday windows. Two vendors, two completely different ideas of what "cheap" means.

What each vendor actually publishes

Per million tokens, in US dollars. The Ant​hropic figures are from Ant​hropic's pricing documentation; the Dee​pSeek figures are from Dee​pSeek's own Models & Pricing page, which is the authoritative card for the model.

A generated scoreboard titled 'What each vendor actually publishes'. Claude Haiku 5.5 reads input $0.10 under 100K and $0.50 above, output $0.50 and $2.50, cached input $0.01 and $0.05, maximum output 128,000 tokens, modality text and images, peak windows none. DeepSeek V4 Flash reads input $0.15 off-peak and $0.30 peak, output $0.60 and $1.20, cached input $0.003 and $0.006, maximum output 384,000 tokens, modality text only, peak windows two weekday UTC ranges. A footer reads 'Vendors' published list rates; DeepSeek's peak windows are weekday UTC ranges.'

• Input — Claude Haiku 5.5 $0.10 up to 100,000 tokens, $0.50 above it. DeepSeek V4 Flash $0.15 off-peak, $0.30 in the peak windows.

• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens, $2.50 above it. DeepSeek V4 Flash $0.60 off-peak, $1.20 in peak.

• Cached input — Claude Haiku 5.5 $0.01 up to 100,000 tokens and $0.05 above. DeepSeek V4 Flash $0.003 off-peak and $0.006 in peak, which is cheaper by more than a factor of three.

• Peak windows — Ant​hropic has none; the rate changes with prompt length and nothing else. Dee​pSeek doubles the rate between 01:00 and 04:00 and again between 06:00 and 10:00 UTC, Monday to Friday, excluding Chinese public holidays.

• Context and output — 1M tokens and 128,000 maximum output for Claude Haiku 5.5; 1M tokens and 384,000 maximum output for DeepSeek V4 Flash.

• Modality — Claude Haiku 5.5 takes text and images; DeepSeek V4 Flash is text-only, and vision sits on the sibling model rather than this one.

• Independent score — Artificial Analysis Intelligence Index v4.3.2 puts Claude Haiku 5.5 (Max) at 43.40, second of 182 models, and DeepSeek V4 Flash 0731 (Max) at 34.33, tenth of 117.

• Speed — the same source measures 241.9 output tokens per second for Claude Haiku 5.5 and 222.0 for DeepSeek V4 Flash, second-fastest on its own board.

The crossover is the whole story

Line the two cards up on a single request and the comparison inverts twice.

Under 100,000 tokens, Ant​hropic is cheaper on both meters and on the sticker that matters: $0.10 against $0.15 in, $0.50 against $0.60 out. That is the first time an Ant​hropic small model has come in under a Dee​pSeek Flash rate on like-for-like terms, and it is the fact the launch post leads with.

Over 100,000 tokens, it reverses hard. Ant​hropic's $0.50 input is more than three times Dee​pSeek's off-peak rate, and the $2.50 output is more than four times it. A retrieval or long-document pipeline that routinely assembles 150,000-token prompts is not a customer for Claude Haiku 5.5 at all; it is a customer for the model it was already using.

And in Dee​pSeek's two weekday windows the second comparison narrows from the other direction, because Dee​pSeek's own rate doubles while Ant​hropic's does not move. Neither vendor's structure is better in the abstract. One prices by how much you send, the other prices by when you send it, and only one of those is under your control at request time.

What a finished task costs, which is the only fair number

Sticker rates are per token, and the two models do not emit the same number of tokens for the same work. That is where the launch arithmetic gets uncomfortable for Ant​hropic.

Artificial Analysis puts Claude Haiku 5.5 (Max) at $0.21 per Intelligence Index task and DeepSeek V4 Flash 0731 (Max) at $0.22 per task. A tenth of the input rate, a wider window and a 26-point Intelligence Index gap — 43.40 against 34.33 — and the measured cost of getting one task done lands within a cent of the model it was supposed to undercut.

The reason is on the same page. Evaluating the Index, Claude Haiku 5.5 generated 440 million output tokens against a 100 million median, and the evaluator calls it very verbose. DeepSeek V4 Flash generated 240 million against a 140 million median and gets called very verbose too. Both models think at length, both bill thinking as output, and output is the meter that decides the bill. The ninety-percent input cut is real and it is charged on the smaller half of the invoice.

There is a second, quieter effect. Ant​hropic's own documentation says Claude Haiku 5.5 uses the newer tokenizer shared with Claude 4.7 and later, so the same text counts as roughly 30% more tokens than it did on Claude Haiku 4.5. The rate fell by a factor of ten; the token count did not fall at all, and on the same text it rose.

A screenshot of Anthropic's model documentation showing the Claude Haiku 5.5 specifications and pricing rows, listing the context window, maximum output and thinking behaviour beside the published per-million-token input, output and cache rates with their prompt-length thresholds.

The model behind the Dee​pSeek name changed too

Anyone comparing these two on the strength of a months-old benchmark table should know this before they do.

Dee​pSeek's pricing page now lists deepseek-flash as the model name, backed by Dee​pSeek-V4.1-Flash, and states that the legacy deepseek-v4-flash id is still accepted but that "the corresponding models have been retired" — requests under the old name are served by the newer model and billed at the Flash price. In other words, the id kept its name, its price and its slot in the lineup, and the weights underneath it moved.

That is not a criticism; it is a warning about what a benchmark table is dated to. The Artificial Analysis entry this piece quotes is for the 0731 release of DeepSeek V4 Flash in its Max configuration, which is the model the board measured, and it is not automatically a statement about whatever answers to that id today. Ant​hropic's Claude Haiku 5.5, by contrast, is a fixed id with no date suffix and no alias, and its Index placement is dated to its own launch week.

Two models, one endpoint

The reason this comparison is easy to settle for yourself rather than on a spec sheet is that both sides are callable from the same place — with one exception worth stating.

DeepSeek V4 Flash is on the OrcaRouter catalogue today, and automatic failover sits underneath it, so the Dee​pSeek leg of this comparison is callable from us this morning rather than read about. One caution worth stating rather than leaving for you to find: our own listed rate for this model is $0.22 in and $0.66 out per million tokens, which sits above Dee​pSeek's off-peak card and below its peak one. Treat Dee​pSeek's page, not ours, as the authority on the peak and off-peak structure, and ours as a single blended number to start a metered run against. Claude Sonnet 5.5 and Claude Opus 5.5 are on the catalogue as well. So an A/B between a Dee​pSeek Flash leg and a Claude leg is one key and one SDK, and the routing DSL will compose them into a single call if the escalation logic belongs in the route rather than the application.

Claude Haiku 5.5 is not on the catalogue yet. That is not a hedge — it means the Ant​hropic side of this comparison currently has to be reached at Ant​hropic directly, and the Dee​pSeek side does not. If you want the cheap tier of the Dee​pSeek line on the same key as the rest of your stack today, that part is already live.

A screenshot of the OrcaRouter catalogue page for DeepSeek V4 Flash, showing the model identifier in provider-slash-model form, the provider DeepSeek, the model description and a stat strip carrying the per-million-token input and output rates alongside the context window and maximum output.

Which one to pick

If your calls are short — classification, extraction, routing, summarisation, subagent turns — and they fit under 100,000 tokens, Claude Haiku 5.5 is now the cheaper model on both meters and the stronger one on the independent board, and the case is straightforward. It is also the only one of the two that takes an image.

If your calls are long, if you re-send a large shared prefix, or if you can schedule, the arithmetic points the other way and does so by a wide margin: Dee​pSeek's cached input at $0.003 off-peak is an order of magnitude below Ant​hropic's $0.01, its output ceiling is three times higher, and its off-peak discount is a scheduling problem rather than a product limitation.

What the launch actually settled is narrower than "Ant​hropic is cheap now." It settled that the American labs will price below the Chinese ones on the short-prompt case, and that the short-prompt case is where most production traffic lives. Everything past 100,000 tokens is still the other model's market.