
GLM-5.3-FlashX vs GLM-5.3-Flash: Same Weights, 2.5× the Price, One Different Number
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
You cannot buy GLM-5.3-FlashX and get a better model. That is not a criticism — it is the entire premise. GLM-5.3-FlashX, launched on Z.ai's API on September 18, 2026, runs the identical weights as GLM-5.3-Flash: the same 320B-total, 18B-active mixture-of-experts backbone, the same 1M-token context, the same native text, image, video and file input, the same 131,072-token output cap. Z.ai has published no FlashX benchmark because there is nothing new to benchmark. What differs is how fast the tokens come out and what they cost, and those two numbers are the only things a buyer has to reason about.
That is a cleaner decision than most model comparisons, and a harder one, because there is no capability story to hide behind. Either the latency is worth 2.5× to you or it is not.
One model, two price tags
• What FlashX is — a serving tier, not a model release; same weights, same scores, same limitsbr> • Input price — $0.37 per 1M tokens on FlashX vs $0.15 per 1M on Flash at listbr> • Output price — $1.25 per 1M vs $0.50 per 1M, the multiplier that actually moves a billbr> • Cached input — $0.075 per 1M vs $0.03 per 1Mbr> • Speed — up to 200 output tokens per second (vendor-reported peak) vs 98 tok/s measured independently by Artificial Analysisbr> • Subscription access — GLM-5.3-Flash is included in Z.ai's GLM Coding Plan with 3× quota; FlashX is not includedbr> • Self-hosting — GLM-5.3-Flash is MIT-licensed open weights on Hugging Face; FlashX has no weights to download

In Z.ai's home market the same ratio is stated plainly: ¥2 in and ¥7 out for FlashX against ¥0.8 and ¥2.8 for Flash. Multiple Chinese outlets describe the launch the same way — 5× the speed at 2.5× the price — and every one of them notes that the intelligence is unchanged.
Latency is a line item, and it is not priced per token
The reason the 2.5× is not automatically a bad deal is that token price and task cost are different quantities. A batch job that generates a million output tokens overnight pays $0.50 on Flash and $1.25 on FlashX for output alone, and gets nothing back for the extra $0.75 because nobody is waiting. An agent loop that makes ten sequential calls pays the same per-token premium, but each call returns faster, and the savings land in wall-clock rather than in the invoice. If FlashX really does return in a fraction of the time, ten turns can hand back tens of seconds of latency per task — and if that task is blocking a human or an upstream service, the seconds are worth more than the tokens.
Z.ai's own positioning follows this line exactly. It points FlashX at high-concurrency coding, multi-step agent calls, real-time dialogue, code completion and long-context multimodal work, and points price-sensitive, latency-insensitive workloads — offline batch, bulk generation, anything scheduled — back at the base Flash. That is the correct split, and it is worth noting that the vendor is the one drawing it.
Where the reasoning breaks down is at the throughput end. FlashX's 200 tokens per second is a peak the vendor reports; the base model's 98 is a median an independent lab measured. Peaks are reached on short outputs with warm caches and an idle cluster. A long-context multimodal request — the exact workload FlashX is pitched at — is the least likely to approach it, and nobody outside Z.ai has published numbers to settle the question. If your workload is throughput-bound rather than latency-bound, a peak figure tells you very little about what you will actually get.
The escape hatch that FlashX does not have
There is a third option in this comparison, and it exists only because the base model is open. GLM-5.3-Flash shipped on August 26, 2026 under the MIT licence, with weights on Hugging Face and ModelScope, and it is served today by inference stacks including SGLang, vLLM, TokenSpeed and KTransformers. If your volume is high enough to justify the hardware, the honest comparison is not Flash versus FlashX — it is FlashX's $1.25 per million output tokens against your own amortised cost per token on your own accelerators.

That option has a real entry price. The FP8 release needs on the order of 328 GB of GPU memory to serve, which is a multi-node deployment for most teams and a non-starter for the rest. But it is worth stating because it is the asymmetry that makes FlashX's pricing a deliberate test rather than an inevitability: Z.ai is charging a premium for a serving configuration that, for the base model, customers can in principle build themselves. The premium is for convenience, not for capability, and convenience has a market rate that goes down over time.
What the price change is really about
The economics behind the launch are unusually legible. GLM-5.3-Flash was released at one-tenth the price of GLM-5.3 — half that during a launch promotion — and it was heavily used, including a stretch of anonymous pre-release testing that Z.ai has since confirmed. FlashX is the first time this family has moved in the other direction: the same model, priced up. Analysts quoted in the Chinese financial press read it as a deliberate commercial test of whether developers will pay for speed on its own, after a year of buying market share with near-frontier capability at commodity prices.
Two details support that reading. FlashX is excluded from the GLM Coding Plan while the base Flash is included with triple quota, so subscription users cannot absorb the new tier into an existing commitment — they pay per token. And the price is flat, with no off-peak discount, unlike the time-of-day structure some competitors use to fill idle capacity. A flat 2.5× on a speed tier is a price for people who cannot wait, and Z.ai is finding out how many of those there are.
Choosing between them
Take FlashX if a human or a service is blocked on the next token and the model itself is already the right choice — inline completion, interactive agents, real-time chat, anything where the pause is the product. Take GLM-5.3-Flash for batch, offline and bulk work, for anything you can schedule, and for any workload where you are paying for output volume rather than output latency. If your volume is large and your team can run accelerators, price the self-hosted base weights before you price FlashX; the comparison is more favourable than the token rates suggest.
Whichever way you go, treat the two as interchangeable at the call site. They take the same prompts, return the same format and produce the same quality — the only thing that changes is a model string, which is the cheapest kind of A/B test to run. On OrcaRouter the vendor's list price is passed through at 0% markup, so a Z.ai pricing change or a promotion applies the day it goes live upstream, and both tiers sit behind one endpoint in front of 200+ models. For a decision that hinges entirely on measured latency, being able to switch between them mid-experiment without touching your integration is most of the work.
The verdict is narrower than a usual model comparison and that is the point. FlashX is not a better model and does not pretend to be. It is the same model with a shorter queue, sold at a price that only makes sense if the queue was your problem. Measure your own time to first token on both before you commit — the vendor's 200 is a ceiling, your workload is the only benchmark that matters, and the difference between the two tiers is a config change away.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
