GLM-5.3-Flash Revealed — The Anonymous "Ox Alpha" Was Zhipu's Open Multimodal Frontier Model All Along Regenerated illustration.
Guides & Insights

GLM-5.3-Flash, Five Weeks On: an Independent 42, a Mythos-Level Exploit Run, and the GLM Agent That Rebuilt Its Own Serving Stack

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GLM-5.3-Flash has been serving production traffic for five weeks, and on September 17, 2026 Zhipu AI said something about where: the company disclosed that every production request for GLM-5.3-Flash now runs on a fleet of more than 100,000 domestically produced Chinese AI accelerators, that it went from first boot on that hardware to carrying its entire live workload in under two weeks, and that end-to-end throughput rose 3.2x across the same stretch. The part that travelled furthest is who did the work. Much of it was done not by the infrastructure team but by an agent driven by GLM-5.3 itself, which read bottlenecks, proposed fixes and wrote inference-system code, with engineers setting objectives, building the feedback environment and reviewing anything that touched numerical semantics, concurrency or production risk. GLM-5.3-Flash is the 320-billion-total, 18-billion-active (320B-A18B) multimodal model Zhipu launched and open-sourced on August 26, 2026 — the anonymous "Ox Alpha" the company revealed that day — and its bigger sibling GLM-5.3 is the 743B flagship it was priced against. None of the three numbers in today's disclosure came from us or from anyone else: they are Zhipu's own. The flagship is no longer the only comparison that matters either: on ExploitBench, an independent evaluation of how far a model climbs a real Chrome exploit ladder, GLM-5.3-Flash has now been measured level with Claude Mythos Preview, the closed frontier model its developer released only to vetted defenders, at a fraction of the cost per attempt.

That is the good news, and it is real but it is vendor-stated. Three other things have moved since this post first ran on launch day, and two of them cut the other way. Artificial Analysis has now published its own measurement of the model, and its number is not Zhipu's. The launch discount that made this the cheapest capable model on the market expired on schedule on September 9 and was not extended. And a third-party evaluation outfit, Generality Labs, has published its own ExploitBench run in which GLM-5.3-Flash scored above Claude Mythos Preview's three-run average on the same ladder — for roughly six cents on the dollar. For anyone who bookmarked this model in late August as “the cheap one,” the arithmetic is different today than it was, and the honest headline is a mixed one: the serving economics genuinely improved, the price went up because a promotion ended, the widely-quoted intelligence figure did not survive first contact with an independent harness, and the capability comparison that did go the model's way is a single unrefereed run whose own authors say should not be over-read.

Zhipu is the only source for the cluster size, the two-week timeline and the 3.2x — there is no independent audit of any of them, and the company itself says it has not achieved recursive self-improvement, describing the result as a minimum-scale loop rather than the real thing. What can be checked, we checked: the one concrete artifact in the account that lives outside Zhipu's walls — a precision fix in a Flash Linear Attention kernel — is genuinely merged upstream, and we confirmed it. Where a figure below comes from Zhipu, it says so. Where it comes from Artificial Analysis, it says that instead. Where it comes from Generality Labs — the safety-evaluation outfit behind the exploit numbers in this post — it says that, along with the caveats its authors attached themselves: one run, 41 bugs, a self-published method, no third-party reproduction, and a contamination question they raise on their own initiative. Where a number is Anthropic's — the only figures here that come from a lab with a competitive stake in the answer — the body says which model it actually measured, because not all of them are this one.

What actually shipped

GLM-5.3-Flash is a 320B-total / 18B-active mixture-of-experts model with 45 layers, which puts its active footprint a little over half of GLM-5.2's roughly 32B. The headline capability is modality: it natively takes text, image and video input — the first GLM-5 model to do so — rather than routing vision through a separate adapter. Context runs to 1 million tokens, and Zhipu claims long-context serving cost lands at roughly a third of GLM-5.3's.

On architecture, Zhipu describes it as the first open-source frontier model to pair hybrid sparse-plus-linear attention with Manifold-Constrained Hyper-Connections (mHC) for scaling and an IndexPool mechanism that compresses four indexer key vectors into one weighted pool to cut long-context latency. Pretraining ran on a 30-trillion-token multimodal corpus. None of that is independently audited — it is Zhipu describing its own model — but the serving-efficiency claims are specific enough to test now that the weights sit in front of the open-source inference stack, and the September 17 account adds the first detailed picture of how the serving side was actually built.

The launch-day spec sheet, at a glance:

• Parameters — 320B total / 18B active, 45 layers, MoE.

• Modality — native text, image and video input; text output.

• Context window — up to 1,000,000 tokens.

• License — MIT, open weights on Hugging Face since launch day.

• Price — $0.15 / $0.50 per million tokens (input / output), $0.03 per million cached input.

A generated single-model scoreboard titled 'GLM-5.3-Flash — the scoreboard' with rows for AA Index 57 (vendor-reported), Active params 18B, Context 1M tokens, Modality text/image/video, Open weights MIT live now, and Price ~$0.14/$0.44 per 1M (discount $0.07/$0.22); footer notes benchmarks are vendor-reported and pricing is derived from Zhipu's stated ratios.

That card is a launch-day snapshot, and two of its rows have moved since August 26. The AA Index figure is still Zhipu's own claim and is now contradicted by independent measurement — more on that below. The discounted price it shows is gone; the promotion ended on September 9. The parameter counts, context window, licence and modality are unchanged current fact, and they are what the picture is mostly for.

The Ox Alpha week, and what the free giveaway was really testing

The reveal closed a week of speculation. On August 20, an anonymous model appeared on a third-party AI API platform with a 1-million-token context window, text/image/video input and a price of zero, and it quickly became the most-called model on that platform's weekly charts. The community fingerprinted it back to Zhipu's GLM family within 48 hours — the same anonymous-then-claimed pattern that had previously identified models from other Chinese labs — and analysts widely guessed a Flash-variant GLM-5.3. The August 26 announcement confirmed the guess: the free week was an intentional test, and the company says the anonymous run moved more than 62 trillion tokens in six days across the coding platforms where it was offered, all of it served from domestic Chinese chips.

That last clause looked like a footnote at the time. It turned out to be the point. The free week was not primarily a marketing stunt or even a load test of the model; it was a load test of the accelerator fleet underneath it, run at a scale where failure would have been public and free. Three weeks later Zhipu is describing the same fleet as carrying 100% of the model's production traffic.

A screenshot of the Z.ai blog page for the GLM-5.3-Flash launch, titled 'GLM-5.3-Flash: Frontier Intelligence, Flash Cost', describing the 320B-total / 18B-active model as the first natively multimodal model in the GLM-5 series, at one-tenth the price, approaching Claude Opus 4.8 on coding benchmarks.

The pricing logic of that week still holds, with one correction. A model given away at $0 for a week and then launched at a tenth of the flagship's rate with a further 50% off was a deliberate market-making move in the budget tier, and the "limited-time" qualifier was always the part to watch. We flagged it as a thing that could be walked back. It was walked back on schedule — which is worth saying plainly, because the discount's expiry is the single change in this story that shows up on a bill.

The serving stack an agent rebuilt

Zhipu's September 17 technical account is the first detailed description of how GLM-5.3-Flash is served on Chinese silicon, and its central claim is that a model helped optimise the system that runs it. The company calls the mechanism dense feedback: correctness tests, execution logs, traces, microbenchmarks and end-to-end metrics were wired up so that an agent could validate a hypothesis locally instead of waiting for a full deployment and load test after every change. For that to work, Zhipu says the feedback had to be local, cheap and objectively checkable — bound to a specific kernel or code path, fast enough to iterate on, and true or false without a human in the loop.

With that loop in place, the agent found and fixed three things the company has described in detail:

• A context-parallel path in the KDA kernel was losing precision as sequences grew, because tl.dot defaulted to TF32 despite FP32 inputs, compounding rounding error with context length. The fix pins input precision to tf32x3, with a fallback to ieee on platforms without TF32 support. This one is the most interesting of the three, because it is the only claim in the account that leaves Zhipu's own infrastructure: the change is merged upstream in Flash Linear Attention as PR #1180, which we checked and confirmed merged on August 27, 2026.

• KV transfer was not overlapping with DeepEP's scheduling. Neither intranode dispatch nor intranode combine was releasing the Python GIL in the DeepEP version in use, which stalled the transfer thread behind them; English-language and Chinese-language accounts of the same writeup put the resulting overhead above 20% and above 30% respectively. After the fix, Zhipu says the gap against its prefill-only baseline dropped below 1%.

• A decode kernel was recomputing the same FP32 normalization and gating four times over, once per V-dimension tile. Splitting the division work cut kernel time 9.6%, and merging the tiles into a single thread block — so the normalization runs once — produced a 1.71x speedup.

Around those fixes sit the structural changes: ReplaySSM, which trades compute for on-chip memory because the accelerators have less of the latter than the GPUs the stack was originally written for; intra-node tensor parallelism to relieve the same pressure; mixed INT8, FP8 and BF16 caching to stretch capacity; and an Encode-Prefill-Decode disaggregated architecture that schedules the three inference stages separately rather than as one pipeline. Zhipu also names the constraints honestly — limited on-chip memory and interconnect bandwidth, immature software, incomplete kernel coverage and thin documentation — and says it has not achieved recursive self-improvement, describing the result as a minimum-scale loop.

Read the account sceptically, because the incentives are obvious. Zhipu will not say how much of the work the agent did versus how much the engineers did, and there is no outside audit of the throughput figure, the two-week timeline or the cluster size. The dense-feedback framing is also self-serving in a specific way: it makes the most impressive-sounding claim ("our model improved itself") rest on the least checkable evidence (an internal loop nobody can inspect). What keeps the story from being pure marketing is that its most concrete claim is falsifiable and turns out to be true — the tf32x3 kernel fix really is in the upstream repository, dated August 27, three weeks before the press cycle around it.

What it costs, in actual dollars

The rate card is Zhipu's, and it is simple: $0.15 per million input tokens, $0.50 per million output tokens and $0.03 per million cached input tokens. That is the list price as of today. The 50%-off launch promotion ran through September 9, 2026 and was not extended, so the $0.075 / $0.25 / $0.015 numbers that circulated widely in late August are no longer available on the vendor's own API. Against GLM-5.3's published $1.40 / $4.40, the Flash still lands at roughly a tenth at list — the discount was a bonus on top of a price that was already the point, not the price itself. Artificial Analysis computes a blended rate of about $0.10 per million tokens on a 7:2:1 cache-hit/input/output mix, and lists output speed at roughly 117 tokens per second, which it describes as notably fast, though AA notes that speed figures describe the model's first-party API rather than an AA-run test.

Zhipu's broader cost story is the reason the price can sit there. In its half-year results, published August 31, the company said unit-token inference cost on its 100,000-accelerator domestic fleet had fallen 80% since the start of 2026, and that API gross margin went from -0.4% to 24.6% as a result of scale, pricing and cheaper serving. Those are company figures from a company reporting to shareholders, and they are unaudited by us; treat them as directionally clear rather than precise. The serving account above is the mechanism behind the claim — fewer redundant kernel passes, less memory pressure, better overlap — which is a more credible story than a bare percentage, even if the percentage is the one that will get quoted.

The efficiency claims against the flagship are unchanged: attention compute down 3.0x and KV cache down 4.4x versus GLM-5.3, vendor-reported, with Zhipu conceding its KV cache remains slightly larger than DeepSeek V4 Flash's and Kimi K3's. The launch-time claim of a 3x end-to-end serving improvement on domestic chips is now stated as 3.2x, measured from first boot to full production traffic. Both numbers are Zhipu's, and the two accounts that carry them are three weeks apart — nothing here has been reproduced outside the company.

Why the reveal matters — and where the headline number went

Here is the correction that matters most to anyone who read the launch coverage and kept the number. Zhipu's materials put GLM-5.3-Flash at 57 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8. Artificial Analysis has since published its own measurement of the model on Intelligence Index v4.3, and it reads 42 — against 47 for GPT-5.6 Sol (max) on the same version. The 57 was never an independent figure; it was Zhipu's, and independent measurement puts the model roughly five points below the frontier-adjacent band and well below the vendor's headline. On the same current index, GLM-5.3 itself measures 45, so the figure of 60 for the flagship that appeared in our launch coverage no longer matches what Artificial Analysis publishes today. The gap of 15 points between claim and measurement is not a rounding difference, and it is the single most useful thing a reader can take from this update — with one qualification the section below adds, because the second independent measurement in this story points the other way.

What survives that correction is narrower but still real. GLM-5.3-Flash is an open-weights multimodal model with 1M context and native image and video input at roughly a tenth of its flagship sibling's list price — a combination with no direct equivalent from the US frontier labs, whose comparable multimodal models cost materially more and ship closed. Zhipu's own framing, "frontier intelligence, flash cost," is the part that took the hit: at a measured 42 it is a strong budget model, not a frontier one at a discount. Independent measurement also puts it comfortably above the median for open-weight models of similar size, which is a real achievement for an 18B-active model with a tenth of the inference cost — it is just a different achievement than the launch coverage implied.

The other independent number — and it went the model's way

If the Artificial Analysis result took GLM-5.3-Flash down a peg, this one goes the other way, and it is the strangest fact in the story. ExploitBench, built at Carnegie Mellon, does not ask whether a model can crash a target. It scores the climb: reaching the vulnerable code, triggering the bug, building reusable primitives, and finally full control-flow hijack with arbitrary code execution, graded mechanically against production V8 with the security sandbox switched on. Generality Labs ran GLM-5.3-Flash across 41 bugs with a 1-billion-token budget per bug rather than the benchmark's usual 300-turn cap, on the argument that turns are a poor proxy for what a buyer actually spends and dollars are the honest constraint. GLM-5.3-Flash reached full arbitrary code execution on 13 of the 41, and a mean capability score of 64.8% of the ladder's maximum. Claude Mythos Preview's three published runs scored 9, 10 and 11 on the same ladder — a mean of 10, or 62.2%. Generality's own framing, printed under the chart, is that GLM-5.3-Flash “may be Mythos level at cyber” on this measurement. Anthropic's own ExploitBench run, published September 29, reports the same ordering at much lower absolute rates — though on the flagship GLM-5.3, not on the Flash: it counts end-to-end exploits per attempt rather than ladder progress, and puts GLM-5.3 at 50 of 410 against Mythos Preview's 56 of 410. Different counting rule, similar ordering — which is exactly why an ExploitBench figure should never be quoted without the rule attached.

The reason that matters is the denominator underneath it. The run priced GLM-5.3-Flash's output tokens at $0.25 per million — its 50%-off launch rate, which has since expired — against Claude Mythos Preview's historic $125 per million, a 500-fold gap. At cost parity the run spent $16.81 per vulnerability against Mythos' $297.17, which is where the roughly 6% figure comes from. Read the claim carefully, because the version that circulates is the wrong one. It is not that GLM-5.3-Flash is a better exploit developer than Claude Mythos Preview. It is that a dollar of GLM-5.3-Flash tokens buys about what a dollar of Mythos tokens buys at this specific task, and that you can afford many more dollars of the former.

Generality's own caveats are worth more than the headline. They say plainly that a single run does not establish parity across cyber as a whole, that contamination is an open question — ExploitBench may have been trained against in some way — and that four of the 41 samples errored and were scored by carrying the last observed result forward, which if anything understates the model. They also published a follow-up on September 28 that complicates the whole comparison. Running GLM-5.3-Flash through seven different agent harnesses at a 250-million-token budget, on ten samples selected to span the difficulty range, the same weights scored 10.6 under the Codex harness — above Claude Mythos Preview's 10.02 reference — and 6.2 under the Claude Code harness, below even the benchmark's original turn-limited configuration at 6.4. The harness moved the score more than the model comparison did, and the authors say the ten-sample subset is probably not representative. That the plot circulated widely on October 2 with the argument that test-time compute scaling is erasing the gaps between model tiers is a claim about the industry, not a finding of the evaluation: what the evaluation found is that a cheap open model, given budget instead of a turn limit, can reach a frontier lab's exploit specialist on one ladder.

It is no longer one lab's story. Anthropic's own Frontier Red Team assessment, published September 29, ran GLM-5.3-Flash — the smaller sibling — in a human-in-the-loop session: handed public details of a patched Chrome flaw plus one other, with no significant direction, the model chained them into a reliable ARM64 exploit that bypassed pointer-authentication hardening, in 20 minutes of human attention and eight hours of model work — $20.40 at Zhipu's API rates. The same assessment is the source of the 50-of-410 ExploitBench figure above, and that part of it measured GLM-5.3, the 743B flagship, not the Flash — a distinction the coverage of the report has largely flattened. Two independent parties, different harnesses, different budgets, same direction. Both are also single lab runs and neither is a reproduced result, and only the Generality run reports a dollar figure for the Flash rather than for the flagship. The practical reading is the one Anthropic states outright: the weights are downloadable, abliterating GLM-5.3-Flash took roughly 600 GPU hours, and a model that costs under a dollar an hour to run is a different kind of availability problem than one behind a vetting program.

Try it, and how the routing layer fits

Access is straightforward. Zhipu's own API serves it at the list rate above, the MIT-licensed weights are on Hugging Face at zai-org/GLM-5.3-Flash in FP8 and BF16, and Zhipu's coding products — ZCode and the GLM Coding Plan — bundle it. For self-hosting, vLLM and SGLang both support it. Two practical notes if you go that route: SGLang's cookbook carries an open change warning that vision input can be silently dropped on transformers builds before 5.13, which will not raise an error and will simply give you a text-only reading of an image request; and the AMD path is still not a recipe. The original launch-day gfx950 enablement PR (sgl-project/sglang #37653) was closed unmerged on September 6 and superseded by a broader set of gfx950 day-zero pull requests — mHC through AITER, k-pool DSA indexer, FP8 and MXFP4 serving — several of which were still open on September 17. Enablement on MI350X-class hardware is in progress, not shipped, and there is still no independent AMD throughput figure.

On the routing side the situation has changed since launch: GLM-5.3-Flash is now on OrcaRouter's roster, alongside the rest of the GLM line. We pass provider list price through with 0% markup and no router surcharge, which is exactly the mechanism that matters for a model whose price just moved — the discount expiring on Zhipu's side is the price you pay here, the same day, with no second contract to renegotiate and no code change. That pass-through cuts both ways, and it is worth being explicit about it: an expired promotion is a real price increase on your bill, and it happens here at the same moment it happens upstream. If you are weighing the Flash against GLM-5.3 or another budget model, both sit behind one key with automatic failover between providers, which is the cheap way to try a model whose independent scores are still settling without betting a production path on it. It is also the right shape for the result above, and for a blunter reason than convenience: the same weights scored 10.6 or 6.2 on the same benchmark depending on which agent harness drove them, so what a buyer actually tests is a scaffold-plus-model combination, and swapping the model behind a fixed scaffold under one key is cheaper than committing to either.

A screenshot of the Hugging Face model card for zai-org/GLM-5.3-Flash showing the MIT license badge, the Text Generation tag, and the 320B total / 18B active parameter counts.

What to watch next

Four things settle the rest of this story. First, whether the 42 moves: the model has been in wide public use since August, so if Zhipu's internal evaluation and AA's harness disagree for a structural reason — reasoning-token accounting, verbosity, an evaluation the vendor weighted heavily — that should show up as the index is revised rather than as a one-off gap. Second, whether the exploit result reproduces. Generality Labs ran 41 bugs once, on its own harness, and says so; the harness follow-up shows the same weights moving four points on harness choice alone, so the number that would settle it is a second team re-running the 41 with the budget fixed in dollars and the harness held constant. Third, whether the domestic-silicon account gets a second source. Zhipu is the only party able to state the cluster size, the two-week timeline and the 3.2x, and the single independently verifiable claim in it — the merged kernel fix — is a small one. Fourth, the price: the discount is gone, and the interesting question is whether it returns as a promotion when Zhipu needs volume, or whether the list rate is now the floor.

The through-line is that GLM-5.3-Flash is being asked to prove two different things at once — that a cheap model can be good enough, and that Chinese accelerators can serve it at scale — and the September 17 disclosure is Zhipu's answer to the second question, delivered by a model that helped write it. The first question now has two independent numbers attached and they point in opposite directions: a 42 on the intelligence index, well under the vendor's headline, and an exploit run that lands level with a frontier lab's specialist at a twentieth of the cost. Both can be true at once, and the gap between them — not either number — is where the decisions actually get made.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily