A hero title card headed 'The Chinese Flash Tier, Audited' with an 'October 2026' claim-check tag, three price chips reading Claude Haiku 5.5 $0.10 in / $0.50 out per 1M tokens, GLM 5.3 Flash $0.15 / $0.50 and DeepSeek V4.1 Flash $0.30 / $1.20, a line reading 'Terminal Bench 4.0 0.32828 - tie', and an Index v4.3.2 line reading 43.40 - 41.81 - 39.46.
Guides & Insights

Claude Haiku 5.5 Was Aimed at China's Flash Tier. Only Half of That Claim Survives Scrutiny.

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

A reader writing in Chinese made a pointed argument this week: Claude Haiku 5.5 looks built to take the mid-tier Chinese market, because it undercuts DeepSeek V4.1 Flash and GLM 5.3 Flash on price and scores above GLM 5.3 Flash on Terminal Bench 4.0 and the Artificial Analysis Intelligence Index. It is a sharper read than most launch commentary. It is also two claims wearing one coat, and they do not carry the same amount of truth.

Anthropic shipped Claude Haiku 5.5 on October 7, 2026 — a full public launch with a dated announcement, a system card and a published rate card. That gives the claim something secondhand market commentary rarely gets: numbers you can check.

The audit below finds one half stronger than the reader said and the other half wrong on the specific benchmark it cites. Both findings are worth knowing, because they point in opposite directions.

The claim, stated precisely

Stripped to its parts, the argument has three pieces.

• Price — Claude Haiku 5.5 is cheaper than DeepSeek V4.1 Flash and cheaper than GLM 5.3 Flash.

• Independent score — it sits about a point above GLM 5.3 Flash on the Artificial Analysis Intelligence Index.

• Coding benchmark — it also scores a point above GLM 5.3 Flash on Terminal Bench 4.0.

Two of those are checkable against published tables and hold up, with adjustments. The third is checkable against the same table and comes out differently than described.

A three-column scoreboard titled 'The claim, audited - Claude Haiku 5.5 against the flash tier' comparing Claude Haiku 5.5, GLM 5.3 Flash and DeepSeek V4.1 Flash across six rows: input price, output price, Intelligence Index v4.3.2 (43.40, 41.81 and 39.46), Terminal Bench 4.0 (0.32828, 0.32828 and 0.2677), cost per finished task ($0.21, $0.25 and $0.27) and weights (proprietary, MIT open, MIT open), footed 'Vendor list rates; Index and benchmark figures per Artificial Analysis v4.3.2.'

The price half: true, and understated in one direction, overstated in another

Anthropic's published rate for Claude Haiku 5.5 is $0.10 per million input tokens and $0.50 per million output tokens up to a 100,000-token prompt, rising to $0.50 and $2.50 above that threshold. Cached input reads at $0.01 and writes at $0.125. Context window is one million tokens; maximum output is 128,000.

GLM 5.3 Flash, from Z.ai, is listed at $0.15 and $0.50, with cached input at $0.026. DeepSeek V4.1 Flash is listed at $0.30 and $1.20, with cached input at $0.006.

On the headline rate the reader is right. A $0.10 input rate is a third below GLM 5.3 Flash and two thirds below DeepSeek V4.1 Flash, and a $0.50 output rate matches GLM 5.3 Flash exactly while coming in at well under half of DeepSeek's. That is a deliberate-looking landing: match the cheaper rival's output, beat it on input, and let the ratio do the talking.

Where it gets better for the claim is the measured cost per finished task. On Artificial Analysis's Intelligence Index v4.3.2, the whole-job cost comes out at $0.21 for Claude Haiku 5.5, $0.25 for GLM 5.3 Flash and $0.27 for DeepSeek V4.1 Flash. The rate card says Claude Haiku 5.5 is 1.5 times cheaper than GLM 5.3 Flash on input; the measurement says the finished job is about a sixth cheaper. Still the cheapest of the three — just not by the margin the sticker implies.

The reason is the one most price comparisons drop. Claude Haiku 5.5 is verbose. Artificial Analysis recorded 162,164 output tokens for it while running the Index, against 68,673 for GLM 5.3 Flash and 88,574 for DeepSeek V4.1 Flash. Anthropic's own documentation adds a second effect: the model uses the newer tokenizer shared with Claude 4.7 and later, so the same text counts as roughly 30% more tokens than on the model it replaces. A cheaper rate applied to more tokens is a cheaper rate applied to more tokens.

One wrinkle that only shows up on a routed bill: our own catalogue lists DeepSeek V4.1 Flash at $0.15 input and $0.60 output with a 2× multiplier applied during two daily windows, 01:00–04:00 and 06:00–10:00 UTC. Eighteen hours a day that is the base rate; six hours a day the output rate is $1.20. Batch traffic that can be scheduled away from those windows gets a materially better DeepSeek price than the vendor's headline list rate suggests, and interactive traffic does not. If you are modelling this comparison for real, the calendar matters as much as the rate card.

A capture of the Artificial Analysis model page for Claude Haiku 5.5 (Max), showing its standing chips of Intelligence #21/182, Speed #91/182, Cost #46/182 and Verbosity #62/182, the rates $0.10 in and $0.50 out with a 90% cache discount, a cost of $0.21 per task, 440M generated tokens, and the summary line that Claude Haiku 5.5 scores 43 on the Intelligence Index against a median of 13.

The score half: "about a point" is 1.6

On Artificial Analysis Intelligence Index v4.3.2, Claude Haiku 5.5 is measured at 43.40 in its Max configuration. GLM 5.3 Flash is measured at 41.81. DeepSeek V4.1 Flash sits at 39.46.

So the index half of the claim is directionally right and rounds downward: the gap is 1.59 points, not one. That is a real lead on an index that runs into the sixties, and it is the number that gets quoted in every summary of this matchup.

What it is not is a claim about the benchmarks underneath. Both vendors publish their own evaluation sets, and they are not the same sets — Anthropic reports GDPval-AA v2.1, OSWorld 2.1, Terminal-Bench 4.0 and FrontierCode for Claude Haiku 5.5; Z.ai reports its own suite for GLM 5.3 Flash. The independent board is the only apples-to-apples row available, which is exactly why the index gap gets leaned on so hard and why the next section matters.

The coding half: this one is a dead tie, not a win

Terminal Bench 4.0, on the shared independent measurement, reads 0.32828 for Claude Haiku 5.5 and 0.32828 for GLM 5.3 Flash. Identical to five decimal places.

That is the single most quotable number in this matchup and the claim is wrong about it. The likely source of the confusion is the neighbouring row: the full GLM-5.3, the non-Flash sibling, scores 0.4192 on the same benchmark. If you have seen "GLM" and "Terminal Bench" in one sentence next to a number a point above Claude Haiku 5.5, that is the sibling, not the Flash.

The other three rows Artificial Analysis publishes for both models are more interesting than the tie, and they do not point one way:

• SciCode — 0.5498 for Claude Haiku 5.5 against 0.5162 for GLM 5.3 Flash. A 3.4-point edge to Anthropic.

• Humanity's Last Exam — 0.4439 against 0.3985. A 4.5-point edge to Anthropic.

• AutomationBench, partial score — 0.3541 against 0.6037. A 25-point edge to GLM 5.3 Flash, and by far the largest gap in the set.

That last row is the one procurement teams will care about most, because agentic automation is where the mid-tier models are actually being deployed. On a measure of whether the model can carry a multi-step task to completion, the cheaper-to-run open-weights model wins by a margin that dwarfs the 1.6-point index difference in the other direction.

Speed splits the same way the price does, only more so. Artificial Analysis measured 424.6 seconds per Index task for Claude Haiku 5.5 against 1035.4 seconds for GLM 5.3 Flash. Anthropic's model gets there in less than half the wall-clock time, which on an agent loop is a bigger operational number than the token rate.

What the claim leaves out entirely

The comparison is framed as a price-and-score story, which quietly skips the dimension where the two models are least alike. GLM 5.3 Flash is open weights, published under MIT, at 320 billion total parameters and 18 billion active. DeepSeek V4.1 Flash is likewise open weights at 552 billion total and 16 billion active. Claude Haiku 5.5 is proprietary, API-only, with no published parameter count and no repository.

That is not a footnote for a lot of buyers; it is the first line of the requirements document. A team that needs to run inference in its own tenancy, fine-tune, or hold a model version fixed for years is not choosing between these three on index points at all — two of the three are disqualified before the price column is read. The reader's framing is a market-positioning observation and it is a fair one; it just happens that the structural difference runs against the model the framing is defending.

Running the comparison yourself

GLM 5.3 Flash and DeepSeek V4.1 Flash are both on the OrcaRouter catalogue at the provider's list price with 0% markup, which makes the Chinese side of this comparison something you can call today rather than model in a spreadsheet. Because the rate is passed through rather than blended, a vendor price move lands on our side the same day it lands on theirs — and the calendar-based DeepSeek window described above is visible in the same pricing block you would be routing against. One key, one SDK, automatic failover underneath, and the routing DSL if you would rather express a cost ceiling in the route than in application code.

Claude Haiku 5.5 is not on our catalogue. It is reachable through Anthropic's own API, and it is better to say that plainly than to leave it implied.

A capture of the OrcaRouter models index, headed 'Models' with the subtitle '207 models, 16 providers - one API key, one bill' and an OpenAI-compatible cURL example showing a POST to the chat completions endpoint with the model field.

What the reader got right, and what to do with it

The instinct is sound. If you wanted the wave of cheap, capable Chinese mid-tier models to look expensive, this is roughly the card you would play: match the cheaper rival's output rate, halve its input rate, land 1.6 index points ahead, and let the price-per-finished-task number carry the argument. Claude Haiku 5.5 does that, and it does it while being more than twice as fast per task as the model it is aimed at.

What the reader got wrong is narrower and more interesting. Terminal Bench 4.0 is not a win; it is an exact tie. The price edge, once you charge for the whole job rather than the token, is about a sixth rather than the 1.5× the rate card suggests. And the one benchmark where the two models are furthest apart is the agentic one, where GLM 5.3 Flash leads by 25 points.

So the honest version of the claim is this: Claude Haiku 5.5 is aimed at the Chinese flash tier, it wins that comparison on the composite index, on cost per finished task and on latency, it does not win it on agentic automation or on openness, and the specific coding-benchmark edge that made the argument feel decisive does not exist.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily