A hero card headed 'Dead Even at 0.32828' with Claude Haiku 5.5 and GLM 5.3 Flash either side of the number, a row reading Terminal Bench 4.0 0.32828, an Index row reading 43.40 against 41.81, and the subtitle 'Same number, different machines'.
Guides & Insights

Claude Haiku 5.5 vs GLM 5.3 Flash: Dead Even on the Benchmark Both Vendors Quote

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Run Claude Haiku 5.5 and GLM 5.3 Flash through Terminal Bench 4.0 and you get the same number: 0.32828 for both. Not close — identical, to five decimal places, on the benchmark that the vendor leads with in its own launch materials and that Z.ai's launch materials lead with too. For two models built on entirely different stacks, by vendors in different countries, released six weeks apart, that is a coincidence worth stopping on.

It is also the wrong number to decide on. The two models separate sharply everywhere else, and the pattern of that separation is unusual enough that it should change how you read either vendor's headline claims. One of them wins the composite index and the cost-per-task figure. The other wins almost everything that looks like an actual deployment.

They are not separated by capability. They are separated by shape.

Start with the one number that says "these are the same class of model."

• SciCode — 0.5498 for Claude Haiku 5.5 against 0.5162 for GLM 5.3 Flash. Two points to Anthropic, on a benchmark where two points is noise.

• Terminal Bench 4.0 — 0.32828 against 0.32828. A tie, exactly.

• Context window — one million tokens for both.

• Output price, first tier — $0.50 per million tokens for both.

Four rows, four near-matches. If you stopped here you would conclude these models are interchangeable and pick on price. That conclusion would be wrong, and the next set of rows is why.

Where they diverge, the direction is consistent

The rows that do separate them are not scattered. They cluster into two groups, and each model owns one group outright.

The agentic group belongs to GLM 5.3 Flash. AutomationBench partial score reads 0.6037 for GLM 5.3 Flash against 0.3541 for Claude Haiku 5.5 — a 25-point gap, the largest single difference between these two models on any shared measurement. Tau-Bench Banking reads 0.4722 for GLM 5.3 Flash; Claude Haiku 5.5 has no published figure on that row. GPQA reads 0.9121 for GLM 5.3 Flash; again no Anthropic figure to compare against. On the measures of whether a model can hold a multi-step task together and use tools reliably inside a structured domain, Z.ai's model is ahead by a margin that dwarfs the composite-index difference running the other way.

The composite-and-cost group belongs to Claude Haiku 5.5. Artificial Analysis Intelligence Index v4.3.2 reads 43.40 against 41.81 — 1.59 points, and the number every summary of this matchup quotes. Cost per finished Index task reads $0.21 against $0.25. Wall-clock time per Index task reads 424.6 seconds against 1,035.4 seconds: Anthropic's model finishes in less than half the time.

That is the shape of the choice. Claude Haiku 5.5 is faster, cheaper per finished job, and higher on the aggregate. GLM 5.3 Flash is more reliable at the specific class of work that cheap models get bought for.

A two-column scoreboard titled 'Claude Haiku 5.5 vs GLM 5.3 Flash - the scoreboard' with rows Terminal Bench 4.0 0.32828 against 0.32828, SciCode 0.5498 against 0.5162, Humanity's Last Exam 0.4439 against 0.3985, AutomationBench 0.3541 against 0.6037, Index v4.3.2 43.40 against 41.81 and cost per task $0.21 against $0.25, footed 'Both figures per Artificial Analysis v4.3.2; Terminal Bench 4.0 is an exact tie.'

Which is the one to believe?

Neither, on its own — and the disagreement itself is the information.

An Intelligence Index is a weighted blend across reasoning, science, coding and agentic categories. It is designed to answer "roughly how capable is this model" in a single number, and it does that job. What it cannot do is tell you whether a 1.59-point lead comes from the categories your workload lives in or from the ones it never touches. Here, the lead comes substantially from rows like SciCode and Humanity's Last Exam — 0.5498 against 0.5162, and 0.4439 against 0.3985 — while a 25-point deficit sits on AutomationBench with the opposite sign.

A team building a coding assistant on a terminal loop will notice the SciCode edge. A team building an agent that has to complete a banking workflow end to end will notice the AutomationBench gap, and it will notice it every single day. The two teams are looking at the same index and should reach opposite conclusions.

The practical move is to read the composite as a filter rather than a ranking. If two models are within about two index points, the composite has told you they are in the same band and has told you nothing about which one to pick. At that point the only rows worth reading are the ones that match your workload — and if a vendor has not published a row you need, the absence is itself a data point, not a gap to be filled in by assumption.

The tokenizer line nobody puts in the comparison table

Anthropic's own documentation carries one sentence that changes every price comparison involving Claude Haiku 5.5, and it is easy to read past. The model uses the newer tokenizer shared with Claude 4.7 and later, which counts the same text as roughly 30% more tokens than the model it replaces.

Set that against the rate card and the arithmetic gets less flattering than the sticker suggests. Claude Haiku 5.5's input rate is $0.10 per million tokens against GLM 5.3 Flash's $0.15 — a 1.5× advantage. Take 30% more tokens for the same prompt and the effective advantage on identical text narrows considerably, though it does not vanish. Output rates are identical at $0.50 in the first tier, so the tokenizer effect applies with nothing offsetting it there.

The measured result is the honest version of the claim: on Artificial Analysis's own accounting, Claude Haiku 5.5 generated 162,164 output tokens per Index task against 68,673 for GLM 5.3 Flash — 2.4 times as many — and still came out at $0.21 per finished task against $0.25. The rate advantage survives the verbosity, and what is left of it is smaller than the rate card says.

Only one of these two has its rates stated as flat. Claude Haiku 5.5's price steps up past 100,000 tokens of prompt to $0.50 input and $2.50 output; GLM 5.3 Flash does not have an equivalent threshold published. For most chat and code-assist traffic that step never binds. For long-document or large-context agent work it binds constantly, and at that point the comparison inverts well past anything the first-tier numbers suggest.

A capture of Anthropic's models-overview documentation table headed 'Lifecycle and reference', listing Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5 and Claude Haiku 5.5 with context windows, maximum outputs, prices per million tokens, latency and default effort, and showing Claude Haiku 5.5 as 'This model' at 1M context, 128K maximum output and 'From $0.10/$0.50'.

The line that decides it before any of the numbers do

Everything above assumes you can choose between these two models freely. A large fraction of teams cannot, and the reason is a single spec-sheet row that the benchmark discussion tends to bury.

GLM 5.3 Flash is open weights, released under the MIT licence, at 320 billion total parameters with 18 billion active. It can be downloaded, self-hosted, fine-tuned, quantised, pinned to a version and run inside a tenancy you control. Claude Haiku 5.5 is proprietary and API-only, with no published parameter count, no licence and no repository.

If your requirements document says the model must run on your own hardware, or that a specific checkpoint must remain callable for the next three years, this comparison is over before the price column is read. That is not a technicality. It is the most common reason teams end up on an open-weights mid-tier model at all, and it is the reason a 25-point AutomationBench advantage is not the only thing pulling in Z.ai's direction.

It cuts the other way too, and the same honesty applies. Anthropic publishes a retirement commitment for Claude Haiku 5.5 of not sooner than October 2027, and carries the model on a first-party API with a system card and a documented evaluation set. An open-weights release is not automatically more durable — you inherit the operational burden along with the weights, and a licence file is not a support contract. Which of those two you want is a question about your team, not about the models.

Both sides on one key, if you want to settle it empirically

GLM 5.3 Flash is on the OrcaRouter catalogue at Z.ai's list price with 0% markup, passed through rather than blended, so the rate above is the rate on the bill and any change Z.ai makes lands on our side the same day. Claude Haiku 5.5 is not on our catalogue — it is reachable through Anthropic's own API — and it is better to state that than to leave it implied.

What the catalogue does make easy is the shape of test this matchup actually calls for. The disagreement between the composite index and the agentic rows is a hypothesis about your workload, and the only way to resolve it is to run both models over the same trace. With GLM 5.3 Flash on the route and an Anthropic leg on the other side of the same gateway, that becomes a config change rather than two integrations — one API for 200-plus models, one key, automatic failover underneath, and the routing DSL when you want to split traffic by task class and read the cost off a single invoice.

A capture of the OrcaRouter model page for GLM 5.3 Flash under z-ai/glm-5.3-flash, showing a 1M-token context window, 128K maximum output, text, image and video input, benchmarks published by Z.ai on 2026-08-26, the description of its 320B total and 18B active parameters, and a price strip of $0.15 input, $0.50 output and $0.030 cache read per million tokens.

The verdict, stated the way the numbers support it

If your work is text in, text out, your prompts stay inside the first pricing tier, and you are paying per token for real volume, Claude Haiku 5.5 is the better model here: higher composite, cheaper per finished task, and more than twice as fast per task. The terminal-benchmark headline is a wash and should not be used to settle anything.

If your work is an agent that has to complete a multi-step workflow without supervision, GLM 5.3 Flash is ahead on the one shared benchmark that measures exactly that, by 25 points, and it is the cheaper of the two to modify because you can hold the weights.

The tie at 0.32828 is the right image for this pair, but not because the models are equal. It is the one row where two models with almost nothing else in common happen to land on the same digit — and treating that digit as a summary of the comparison is how a procurement decision gets made on the least informative number in the table.