
Claude Haiku 5.5 vs DeepSeek V4.1 Flash: The Cheaper Sticker Is Not the Cheaper Model
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 127 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 58 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 350 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
On the rate card this is not close. Claude Haiku 5.5, which the vendor shipped on October 7, 2026, lists at $0.10 per million input tokens and $0.50 per million output tokens. DeepSeek V4.1 Flash, released September 10, 2026, lists at $0.30 and $1.20 on the independent board, or $0.15 and $0.60 on the vendor's own off-peak card. The vendor is between two and three times cheaper on both meters, and that is the first time a Haiku-class model has come in under a DeepSeek Flash rate.
Then you measure what a finished task costs, and the gap collapses. On the same evaluation run, one Intelligence Index task costs $0.21 on Claude Haiku 5.5 and $0.27 on DeepSeek V4.1 Flash — a 20% edge, not a 200% one. The reason is on the same page: Claude Haiku 5.5 generated 162,164 output tokens per task against DeepSeek's 88,574. You pay less per token and buy nearly twice as many of them.
That inversion is the whole comparison, and everything below is an argument about which side of it your workload lands on.
Both rate cards, read as written
Per million tokens in US dollars. Anthropic's figures are from Anthropic's own pricing documentation; DeepSeek's are from DeepSeek's Model & Pricing page, which is the authoritative card for its peak and off-peak structure. Independent measurements are Artificial Analysis Intelligence Index v4.3.2, read 2026-10-08.

• Input — Claude Haiku 5.5 $0.10 up to 100,000 tokens and $0.50 above that. DeepSeek V4.1 Flash $0.30 flat on the board's card, or $0.15 in DeepSeek's off-peak windows.
• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens and $2.50 above. DeepSeek V4.1 Flash $1.20 flat, or $0.60 off-peak.
• Cached input — Claude Haiku 5.5 $0.01 below 100,000 tokens and $0.05 above, a 90% discount. DeepSeek V4.1 Flash $0.006, a 98% discount. This is the one meter where the DeepSeek card is cheaper in absolute terms, and it is cheaper by more than people expect.
• Time-of-day pricing — Anthropic has none. DeepSeek doubles both meters between 01:00 and 04:00 and again between 06:00 and 10:00 UTC on weekdays, excluding Chinese public holidays. Every other hour, including all weekend, is off-peak.
• Context and output ceiling — 1M tokens each. Claude Haiku 5.5 caps a single response at 128,000 tokens; DeepSeek V4.1 Flash caps it at 384,000.
• Weights — Claude Haiku 5.5 is proprietary, API-only, model id claude-haiku-5-5. DeepSeek V4.1 Flash is open under MIT, 552B total parameters with 16B active, on Hugging Face.
• Modality — both take text and images and return text. Neither is video-native.
• Independent score — 43.40 for Claude Haiku 5.5 (Max) against 39.46 for DeepSeek V4.1 Flash (Max). A 3.94-point gap on a 182-model board.

Why the per-task number beats the per-token number
Sticker rates are quoted per token because that is what gets metered. They are a poor guide to what a workload costs, because a reasoning model decides how many tokens a task needs, and the two models here disagree by a factor of 1.8.
Claude Haiku 5.5 emits 162,164 output tokens per Intelligence Index task, split 129,047 reasoning and 33,118 answer. DeepSeek V4.1 Flash emits 88,574, split 62,670 reasoning and 25,904 answer. Both bill thinking as output. So the model with the 60%-lower output rate spends almost twice as much of the expensive meter, and the arithmetic lands them within six cents of each other.
There is a second effect pulling the same way. The index run is dominated by input as well as output because the harness re-sends context: Claude Haiku 5.5's $0.2128 per task breaks down as $0.1317 input and $0.0811 output, and within the input line, $0.0954 of that is cache reads and $0.0337 is cache writes against only $0.0026 of fresh input. DeepSeek's $0.2652 splits $0.1589 input and $0.1063 output, with $0.0966 of the input in cache writes. The two models read roughly the same amount of cached context; DeepSeek pays more for the same reads and writes.
Anthropic's own launch post leads with a similar caution from the other direction. It says Haiku 5.5 is "especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model." That qualifier is doing real work — see the cliff below.
The 100,000-token cliff is the real fork
Anthropic does not price Claude Haiku 5.5 flat. The $0.10 and $0.50 meters hold for prompts up to 100,000 tokens and then become $0.50 and $2.50 — a fivefold jump on input and a fivefold jump on output, applied to the whole request rather than the excess.
DeepSeek's structure is completely different. Its rate depends on when you send, not on how much. A 150,000-token prompt costs the same per token as a 500-token prompt, doubled inside two weekday windows.
Put a long-prompt pipeline on both and the comparison inverts hard. At 150,000 input tokens with 20,000 output, Claude Haiku 5.5 bills $0.125 of input and $0.05 of output — $0.175 for the call — while DeepSeek off-peak bills $0.0225 and $0.012, which is under four cents. That is DeepSeek winning by a factor of roughly 4.5, on the same two models where the short-prompt case had Anthropic winning by three.
Neither structure is better in the abstract. One prices by how much you send, the other by when you send it, and only one of those two variables is under your control at request time.
DeepSeek's other lever nobody prices in
The cache line deserves its own look, because it is the cheapest meter anywhere in this comparison and it is the one that grows with production volume rather than shrinking.
DeepSeek's cache-hit rate is $0.006 per million tokens with a stated 98% discount — roughly a sixth of Anthropic's $0.01 and a fraction of what fresh input costs on either side. Anything that re-sends a stable prefix — a long system prompt, a retrieved document set, a tool schema — pays that rate on every turn after the first. The index run above already shows how large that share is: $0.0954 of Claude Haiku 5.5's per-task input cost is cache reads, from just $0.0026 of uncached input.
So the practical split is narrower than "Anthropic is cheaper." If your requests are short, single-turn, and cache-warm, Claude Haiku 5.5 is ahead on price, ahead on capability by 3.94 index points, and ahead on the composite agentic rows. If your requests are long, or prefix-heavy, or schedulable into DeepSeek's off-peak windows, DeepSeek V4.1 Flash dominates on cost and it is not particularly close.
What you can actually call today
Both models are reachable, but not symmetrically, and that asymmetry is worth stating rather than leaving for you to find.
DeepSeek V4.1 Flash is on the OrcaRouter catalogue now, passed through at the provider's list price with 0% markup — which matters more here than usual, because DeepSeek's rate has a peak structure. We list $0.15 in and $0.60 out with a $0.003 cache read and the two weekday doubling windows carried in the pricing record, so a request routed to the DeepSeek leg is billed the way DeepSeek bills it, not at a blended average. Automatic failover sits underneath the route, so a peak-window 429 or a provider brownout moves the call rather than failing it.
Claude Haiku 5.5 is not on the catalogue. Anthropic's model is reachable through Anthropic directly and through its three cloud partners — AWS, Google Cloud and Microsoft Azure, which the launch post names — and that is the only honest way to say it. What the catalogue does carry from the Anthropic line is Claude Sonnet 5.5 and Claude Opus 5.5, which is what makes an escalation path from a cheap classification call to a mid-tier agentic call a routing configuration rather than a second integration.

Which one to pick
If your traffic is classification, extraction, routing decisions, summarisation and subagent turns — short prompts, high volume, images in the mix — Claude Haiku 5.5 is the better model and, under 100,000 tokens, the cheaper one. It scores 3.94 index points higher, it takes images, and Anthropic ships an adjustable effort dial so you can trade capability for cost per call without changing models.
If your traffic is long-document, retrieval-heavy, or you can schedule it, DeepSeek V4.1 Flash wins on arithmetic: three-to-four-times cheaper per long call, a cache rate a sixth of Anthropic's, a 384,000-token response ceiling, and open weights you can hold if the per-token model ever stops fitting.
The thing this comparison actually settles is narrower than "Anthropic is the cheap option now." It settles that a small Anthropic model can undercut a DeepSeek Flash rate on the sticker while landing within 20% of it on the invoice — and that the distance between those two facts is a verbosity setting.
One API for 200-plus models, one key, automatic failover underneath, so a peak-window 429 or a provider brownout moves the call rather than failing it.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
