A generated title card for 'Claude Haiku 5.5 vs GLM-5.3-FlashX — paying 2.5x for the same weights', showing the two model names either side of a divider and the caption 'Two mid-price models, four rate cards, one 100K threshold', with the footer 'Anthropic and Z.ai list prices, October 2026.' The real OrcaRouter logo is composited bottom-right.
Guides & Insights

Claude Haiku 5.5 vs GLM-5.3-FlashX: Paying 2.5x for the Same Weights, and Whether It Is Worth It

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The most useful thing to know about GLM-5.3-FlashX before comparing it to Claude Haiku 5.5 is that it runs the same weights as GLM-5.3-Flash and costs 2.5 times as much to call. Z​.ai launched FlashX on its own API on September 18, 2026, at $0.37 per million input tokens and $1.25 per million output tokens, against $0.15 and $0.50 for the base model it is derived from. Nothing about the model changed; what changed is where it runs and how fast it is promised to run there. Understanding that pairing is the whole game here, because it means this is not a comparison between two capability levels — it is a comparison between a price tier and a serving tier, and Haiku 5.5 lands in the middle of that argument from a completely different direction.

Two models, four prices, one threshold

Line the four relevant rate cards up and the shape of the decision gets clearer than any single pair does:

• Claude Haiku 5.5, under 100,000 tokens — $0.10 input, $0.50 output per million (A​nthropic list)

• Claude Haiku 5.5, over 100,000 tokens — $0.50 input, $2.50 output per million (A​nthropic list)

• GLM-5.3-Flash — $0.15 input, $0.50 output per million (Z​.ai list)

• GLM-5.3-FlashX — $0.37 input, $1.25 output per million (Z​.ai list)

• Context — 1 million tokens on all four (vendor-published)

• Open weights — GLM-5.3-Flash and FlashX inherit the base model's MIT-licensed weights; Claude Haiku 5.5 is closed (vendor-published)

• Independent score — GLM-5.3-Flash measured at an Artificial Analysis Intelligence Index of 41.807; Claude Haiku 5.5 measured at 43.395 (Artificial Analysis)

The first thing to notice is that the FlashX price sits almost exactly on top of Claude Haiku 5.5's long-context tier — $0.37 against $0.50 on input, $1.25 against $2.50 on output. That is not a coincidence of the market; it is what a premium serving tier costs when the underlying model is mid-weight and the vendor is charging for throughput rather than for capability. The second thing is that GLM-5.3-FlashX is the expensive version of a model A​nthropic's new Haiku is cheaper than at short context and comparable to at long. The third is that all four numbers sit inside a factor of five of each other, which is a much tighter band than the "cheap model versus expensive model" framing of most head-to-heads implies.

A generated two-column scoreboard titled 'Claude Haiku 5.5 vs GLM-5.3-FlashX — the scoreboard'. Left column Claude Haiku 5.5: input price (short) $0.10 per million, output price (short) $0.50 per million, input price (over 100K) $0.50 per million. Right column GLM-5.3-FlashX: input price $0.37 per million, output price $1.25 per million, and a serving-premium row. A footer labels the sources. The real OrcaRouter logo is composited bottom-right.

What FlashX actually is

GLM-5.3-FlashX is not a new model. It is GLM-5.3-Flash served on Z​.ai's own infrastructure, advertised at up to 200 output tokens per second at peak against the base model's measured rate, and priced at 2.5 times the base model's list. Z​.ai's pitch is straightforward: same answers, more of them per second, and you pay for the serving rather than the weights. For a throughput-bound workload — batch classification, high-volume extraction, anything where latency is the constraint rather than the cost — that is a coherent thing to sell and a coherent thing to buy.

What it is not is a capability upgrade, and the pricing reflects that honestly: if FlashX were a better model, Z​.ai would not be able to charge for the base model's weights separately. The 2.5x is a serving premium. Whether it is worth paying is a question about your workload's bottleneck, and it has a different answer for a latency-bound pipeline than for a cost-bound one.

Artificial Analysis has no separate page for FlashX at the time of writing, which is worth stating plainly rather than papering over: the independent figure the board publishes for GLM-5.3-Flash — 41.807 on the Intelligence Index, 52.14 output tokens per second measured, 0.328 on Terminal-Bench Hard — is a figure for the base model, not for the served-at-2.5x version. The vendor's own throughput claim for FlashX is a peak figure and vendor-reported. Nobody independent has published throughput for FlashX itself, so any comparison of Haiku 5.5's measured speed against FlashX's is a comparison against a number Z​.ai supplied about its own serving.

Z​.ai's 2.5x multiplier is not the only thing that makes FlashX expensive relative to its base. Because the base weights are MIT-licensed and hosted by third parties, a workload that genuinely needs more throughput than one vendor's endpoint provides can often get it by spreading across several providers of the same weights at or near the base price, rather than by paying one provider's premium tier. That is a real alternative that the FlashX rate card is competing against, and Z​.ai knows it — which is presumably why the pitch leans on peak throughput rather than on uniqueness.

A screenshot of Z.ai's documentation pricing page, showing the API pricing table for the GLM model line with per-million input and output rates, captured in English.

Where Haiku 5.5 and FlashX part ways

Set the two against each other on the dimensions that decide a real workload and the split is clean enough to choose on:

• Price under 100K context — Haiku 5.5 $0.10 / $0.50 vs FlashX $0.37 / $1.25; Haiku wins on both sides

• Price over 100K context — Haiku 5.5 $0.50 / $2.50 vs FlashX $0.37 / $1.25; FlashX wins on both sides

• Throughput — vendor peak of 200 tok/s for FlashX vs Anthropic-published serving for Haiku 5.5 that does not advertise a peak of that kind

• Independent index — Haiku 5.5 43.395 vs GLM-5.3-Flash 41.807 (Artificial Analysis; no independent figure for FlashX itself)

• Weights — FlashX inherits MIT-licensed open weights; Haiku 5.5 is closed and API-only

• Self-hosting — possible for FlashX via the open weights; not possible for Haiku 5.5

• Tool/agent feature set — adjustable effort dial on Haiku 5.5, absent on the G​LM Flash line

The interesting cell is the second one. Claude Haiku 5.5 is a five-times-cheaper model than GLM-5.3-FlashX on short requests and slightly more expensive than it on long ones, because the Haiku price list steps up by 5x at 100,000 tokens while FlashX charges one flat rate. If your traffic is mostly short, the Haiku is the obvious pick on price alone. If it is mostly long, FlashX is not competing with the Haiku's short-tier price at all — it is competing with the Haiku's long tier, and it wins that comparison by roughly a third.

The capability comparison is closer than either vendor would like. 41.807 against 43.395 on the same independent index is a gap of under four percent, well inside the noise of most application-level evaluations and far too small to justify a pick on its own. Neither model dominates the other on general reasoning; the decision is being made on the rate card and on throughput, not on which one is smarter.

A screenshot of the Artificial Analysis model page for Claude Haiku 5.5, headlined 'Proprietary model — Released October 2026', showing an Intelligence Index of 43.395 and speed rank #10 of 182, with the cost comparison block reading $0.10 input.

The licence difference is a real difference

FlashX being open-weight in origin matters more than the price gap on one axis: it can be self-hosted, and it can be served by anyone. Claude Haiku 5.5 cannot be either. For a team with a compliance requirement that the inference happen on its own hardware, or with enough steady volume that amortising a GPU is cheaper than paying per token, the G​LM line is the only one of the two that answers the question at all. For everyone else — the majority, who want a model behind an API and would rather not run the serving layer — the licence is a footnote and the price is the argument.

It also means the two models have different failure modes when a provider has a bad day. A closed model served only by its vendor and that vendor's cloud partners goes down when those go down. An open model with several independent hosts can be failed over between them, provided something is doing the failing over. That is a genuine operational difference and it is the sort of thing that only shows up on the day it matters.

Routing both without a second contract

If the decision above does not collapse to a single winner for your traffic — which it usually does not, because the two models beat each other on different halves of the price list — the practical question is how to run both without maintaining two integrations. OrcaRouter serves z-ai/glm-5.3-flash at $0.15 and $0.50 per million at the vendor's list price, alongside qwen/qwen3.8-flash and the rest of a 200-plus model catalogue, on one key and one bill. An A​nthropic price move on Claude Haiku 5.5, or a Z​.ai move on the Flash line, is passed through at zero markup the same day rather than lagging behind a reseller's own rate card, so the numbers above stay the numbers you actually pay.

We do not serve Claude Haiku 5.5, and we do not serve GLM-5.3-FlashX — neither is in our catalogue; for those, call A​nthropic and Z​.ai directly. What we do serve is GLM-5.3-Flash, the base model underneath FlashX, which is the right call for the many workloads where the 2.5x serving premium buys throughput nobody asked for. Where a production path genuinely depends on either of the two models we do not host, our failover is the mechanism for trying it without betting the whole path on it: point the request at it, keep a working model as the fallback, and let the routing layer decide which one answers.

Which one to pick

Below 100,000 tokens per request, Claude Haiku 5.5 wins on price and matches FlashX closely enough on capability that the choice is not close. Above 100,000 tokens, GLM-5.3-FlashX is cheaper than Haiku 5.5's long-context tier and is the model to reach for if the request size is a property of the task rather than a choice — though the base GLM-5.3-Flash is cheaper still, and the only reason to pay the 2.5x is if you specifically need FlashX's advertised throughput.

The framing to discard is that one of these is the budget option and the other is the performance option. They are both mid-price models with flat, well-published rate cards, and the interesting variable is not which is better but which half of your traffic each one is cheaper for. Pull your request-size distribution first; the two-column price comparison above will tell you the rest.

a 200-plus model catalogue

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily