
Ox Alpha: The Complete Spec Sheet for the Stealth Model Now Shipping as GLM-5.3-Flash
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 517 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 196 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Ox Alpha is the anonymous name that GLM-5.3-Flash was previewed under, it was revealed on August 26, 2026, and it is not a model you can call today. The preview listing that carried the Ox Alpha name no longer serves it; the model underneath has a vendor page, MIT-licensed weights, a published rate card and a documented API surface, all of them Z.ai's. This page is the reference for that model — the envelope, the modalities, the rate card, the endpoint surface, every published benchmark figure with the revision or harness attached to it, and, because the alias still returns a thin and confusing search result, a plain statement of what is not documented anywhere. Everything below was read on 2026-09-28 from Z.ai's own model documentation and pricing page, from the Z.ai model card on Hugging Face, from Artificial Analysis, and from our own catalogue; figures the vendor publishes and figures an independent party measures are labelled separately, and nothing here is inferred from a resemblance.
What the name Ox Alpha refers to today
The alias is retired, and the vendor is the one who retired it. Z.ai's own documentation for GLM-5.3-Flash says the model was tested anonymously under the ox-alpha name before release, to gather user feedback, and that the anonymous run became the most popular model of that week. The datable anchors are short and specific: the stealth listing showed Ox Alpha as released on August 20, 2026, and the reveal landed on August 26 — the same date Z.ai's documentation page carries and the same release date our own catalogue records for z-ai/glm-5.3-flash.
Two things follow from that, and both matter to a reader who searched the name.
• The alias is not a route. Our catalogue returns "model not found" for a stealth/ox-alpha lookup, read 2026-09-28. There is nothing to call under that name, because the vendor's product carries the vendor's name.
• There is no vendor-owned weights repository under the alias either. Searching Hugging Face for ox-alpha returns community uploads created during the stealth window in late August 2026, plus a card-only repository from September 27, 2026 that carries no weights and points at Z.ai's official release. A repository name is a label its uploader chose; it is not documentation of anything. Z.ai's own organisation publishes the model as zai-org/GLM-5.3-Flash, with the repository created on 2026-08-25 in UTC.
The provider question that dominated the anonymous week is closed, and it closed in the ordinary way: a lab stepped forward and put its name on the model. What is left is a specification, which is what this page covers. The discovery story — the fingerprinting that preceded the reveal, the anonymous-model lineage it belongs to, and the serving-hardware disclosure that came with it — is in our earlier piece on the Ox Alpha investigation, and this page does not re-litigate any of it.
The envelope, as the vendor documents it

These are Z.ai-reported figures, read from the vendor's own documentation page on 2026-09-28, and three of the four envelope dimensions are independently corroborated — Artificial Analysis's model record carries the same 320B total, 18B active, 1,000,000-token context and MIT licence, and our own catalogue card carries the context and output limits.
• Context window — 1,000,000 tokens (Z.ai documentation; Artificial Analysis; our own catalogue card, which lists 1,000,000)
• Maximum output — 128K tokens (Z.ai documentation); our card records 128,000
• Input — text, image, video and file (Z.ai documentation, read 2026-09-28); our card lists text, image and video, and does not claim file input
• Output — text only, on every source
• Parameters — 320B total with 18B activated per token (Z.ai documentation and model card; Artificial Analysis independently records 320B and 18B)
• Licence — MIT, with weights published on Hugging Face (vendor model card; Artificial Analysis classifies the model as open weights under a permissive licence)
• Architecture — a hybrid of sparse and linear attention, which Z.ai describes as the first open-source frontier model to combine the two, reducing attention computation by 3.01× and KV cache by 4.44× against GLM-5.3, plus an IndexPool step that compresses four indexer key vectors into one at long context, and Manifold-Constrained Hyper-Connections for scaling efficiency. All vendor-stated; there is no independent architectural audit to cite.
• Pre-training — a 30-trillion-token multimodal corpus (Z.ai-stated). The vendor publishes no total training-compute figure and no parameter-by-layer breakdown beyond the 320B/18B headline.
One envelope claim has narrowed since launch, and the current documentation is the one to quote. The serving-hardware disclosure that circulated in August put the anonymous week's capacity at roughly 100,000 domestic Chinese chips. Z.ai's documentation as it reads today says the model was served "across tens of thousands of domestically developed accelerators," with a 3× end-to-end serving improvement over the vendor's own baseline on the same hardware and a per-token cost the vendor describes as comparable to mainstream NVIDIA GPUs. Tens of thousands is what the page says now; the larger figure is the launch-era one. Both are company claims, neither has been independently audited, and the direction of the revision is downwards.
The rate card, and why the served number is the one that matters
Z.ai's pricing page lists GLM-5.3-Flash at $0.15 per million input tokens, $0.03 per million cached input tokens and $0.50 per million output tokens, and the sibling FlashX tier at $0.37 / $0.075 / $1.25. That is the vendor's list price, read from the vendor's own page on 2026-09-28.
The rate actually served on our route today is lower, and this is the practical point of the section. Read on 2026-09-28, the OrcaRouter card for z-ai/glm-5.3-flash prices input at $0.075 per million tokens, output at $0.25 per million and cache reads at $0.0173 per million. That is half the vendor's list rate, on the same model, on the same day — the discounted figure rather than the headline, with no promotional label on the card. Our earlier coverage of this model noted the same gap after the stated promotional window closed; the observation still holds today, and we still cannot source the reason, so we report the served number and leave the cause open.
Cached input is where the three sources visibly disagree, which is a useful reminder that a "$0.03" in a table is not necessarily a price anyone pays. The vendor's page says $0.03 per million cached input tokens; Artificial Analysis's model record carries a cache-hit price of $0.026; our route serves cache reads at $0.0173. All three read on or shortly before 2026-09-28, all three vendor- or platform-reported. We quote the vendor's list figure as the list figure and the served figure as what a caller is actually billed.
Run the arithmetic on a representative agent job — 100,000 input tokens and 10,000 output tokens, no cache hits:
• At the vendor list rate — 100,000 tokens × $0.15/1M = $0.015, plus 10,000 × $0.50/1M = $0.005, so $0.020 per call
• At the rate our route serves today — $0.0075 plus $0.0025, so $0.010 per call, exactly half
That is the whole reason to check the served price rather than the list price before you size a workload: on a long-horizon agent that runs thousands of calls a day, the difference between those two numbers is the difference between two budgets. OrcaRouter passes the provider's list price through with 0% markup, which is why the discount the vendor is running shows up on our card rather than being absorbed somewhere in the middle.
Modalities and endpoint surface
The vendor documents one interface shape and two model codes: glm-5.3-flash and glm-5.3-flashx, both served through the Chat Completion API. Images go in as a content block with type image_url on the message content array, taking either a URL — which the vendor recommends — or a base64 data URL, and multiple image blocks can be sent in one message. Streaming is supported, and Z.ai recommends enabling stream and tool_stream together for streaming requests. Tool calling, context caching and structured JSON output are all documented capabilities; thinking is supported but, as the next section explains, not optional.
FlashX is a separate serving tier of the same model rather than a separate capability: Z.ai documents it at 200 tokens per second, and states that it is not yet available on the GLM Coding Plan. It is not the tier our catalogue carries, and it is not what the figures on this page describe.
On our route, the endpoint surface is narrower and worth stating exactly. The card for z-ai/glm-5.3-flash exposes POST /v1/chat/completions — chat completions only, with no second protocol endpoint — and accepts reasoning, reasoning_effort, include_reasoning, tools, tool_choice, response_format, stream, temperature, top_p, max_tokens and stop. The capability chips on the card are vision, tools, JSON and reasoning. Context is 1,000,000 tokens and maximum output 128,000, matching the vendor's documentation.
Effort and thinking: the settings that pin every benchmark number
This is the part of the specification that most changes how you read the scores, and the vendor documents it precisely. GLM-5.3-Flash takes a reasoning_effort parameter with three levels — low, high and max — and it defaults to max if you pass nothing, or if you pass anything that is not low or high. Thinking cannot be switched off: the vendor states that thinking.type supports enabled only. The chat template's clear_thinking flag defaults to false, and Z.ai recommends passing it explicitly as true for chat scenarios. The vendor's recommended sampling settings are temperature 1 and top_p 0.95, and it says outright that benchmark and leaderboard reproduction should keep the default max effort.
Read those two facts together and the consequence for anyone comparing numbers is unavoidable: a score quoted without an effort level is not a complete measurement, because the low, high and max settings are different operating points of the same model, and the default is the most expensive of the three. Our card exposes reasoning and reasoning_effort among its supported parameters, so the setting is reachable through the route, not only through the vendor.
The benchmark sheet, with the revision attached
Z.ai publishes a benchmark table for GLM-5.3-Flash, and it is entirely vendor-reported — Artificial Analysis's independent record does not reproduce these rows. The vendor's headline coding and agentic figures are DeepSWE v1.1 at 63.4 against GLM-5.2's 46.2, and AutomationBench v1.0.6 at 48.8 against GLM-5.2's 26.2. The rest of the sheet, as our own catalogue carries it with Z.ai as the named source: Humanity's Last Exam with tools 55.3, Agents' Last Exam 26.3, NL2Repo 56.3, Chartography with tools 78, CharXiv reasoning with tools 89.4, and on the vision side MMVU 80.5, MVBench 77.8 and BabyVision 53.4.
Two vendor rows deserve their footnotes carried into the sentence, because they are the ones a reader is most likely to quote without them. Z.ai's in-house coding evaluation, Z.ai Code Bench v1.0, puts GLM-5.3-Flash at 29.0 at max effort against Claude Opus 4.8 at 29.5 — a near-tie on the vendor's own harness, run on Claude Code 2.1.207, reported by the vendor. That is not an independent comparison and it is not a matched-effort comparison run by a third party; it is Z.ai benchmarking its model against a competitor on its own evaluation. And the vendor quotes an Artificial Analysis Intelligence Index of 57 at $0.045 per task, naming the revision as v4.1.1.
That last figure needs separating from what Artificial Analysis publishes today, because the two cannot be placed side by side as a movement. Artificial Analysis's own page for GLM-5.3-Flash, read on 2026-09-28, reports an Intelligence Index of 41.8 on Index v4.3.2, recorded at max effort. v4.3.2 incorporates ten evaluations — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1 — and v4.1.1 did not. Two index revisions, two task sets, two numbers. Neither is wrong; they are not the same measurement, and quoting the 57 next to the 41.8 as though the model had moved would be a mistake in either direction.
What the independent record does corroborate is the envelope rather than the scores: Artificial Analysis independently records 320B total parameters, 18B activated, a 1,000,000-token context window, MIT licensing, open weights and a 2026-08-26 release date, and lists vendor pricing of $0.15 in and $0.50 out per million tokens with a cache-hit price of $0.026. Where the vendor publishes a score and the independent party does not, we say so; where the independent party confirms a spec, that is the strongest class of fact on this page.
There is no matched head-to-head here, and saying so is more useful than producing one. Nobody has published a run of GLM-5.3-Flash against Claude Opus 4.8, GPT-6 Sol or Grok 4.7 at a stated effort on a shared harness — Z.ai's own comparison is on its own bench, at max effort, against a competitor's default. Until that changes, the honest comparison between this model and anything else is on price, envelope and endpoint surface, which are the three things on this page that everyone can read off the same numbers.
What is not documented, stated plainly
A reference page for a model with a thin information surface is only useful if it is honest about the gaps, and this one has several.
• No vendor knowledge cutoff. Z.ai's documentation and the Hugging Face model card do not state one, and Artificial Analysis's record leaves the field null. Estimates from the stealth window are community work, not a vendor figure, and this page does not repeat one as though it were.
• No harness footnotes for most of the benchmark sheet. The model card does carry footnotes for the HLE with-tools run — temperature 1.0, top_p 0.95, a 163,840-token maximum generation length, evaluation at up to a 300,000-token context with a context-management strategy, and GPT-5.6 Luna at medium effort as the judge model — and for NL2Repo and DeepSWE, which the vendor says was run with the mini-swe-agent harness. The remaining rows are bare percentages with no harness or effort attached.
• No explanation of the cached-input price spread. Three figures for the same quantity on the same model, and no source states why they differ.
• No independent benchmark of the anonymous serving period. The stealth listing is gone; whatever it measured cannot be measured again, and no third party published a run against the alias while it was live.
• No matched-effort third-party comparison, for the reason given above.
• No vendor confirmation of the launch-era chip count. The documentation says tens of thousands of domestic accelerators; the 100,000 figure is not on the page.
Calling it: what our catalogue serves, verified today

The model behind Ox Alpha is live in our catalogue under its own name, checked today rather than assumed. A read of the OrcaRouter model API for z-ai/glm-5.3-flash on 2026-09-28 returned the card in full: 1,000,000-token context, 128,000 maximum output, text, image and video input with text output, the capability chips vision, tools, JSON and reasoning, a release date of 2026-08-26, not deprecated and not delisted, priced at $0.075 per million input tokens, $0.25 per million output tokens and $0.0173 per million cached input tokens, over the single endpoint POST /v1/chat/completions.
Our own seven-day playground measurement on that card, for the same period, reads 61.8 output tokens per second, a p50 time to first token of 7.39 seconds and an error rate of 1.47%. Those are first-party numbers about our serving, not vendor figures and not independent benchmarks, and they are the only speed figures on this page for that reason.
The alias itself is not on the route. If you arrive here having searched for Ox Alpha, the model to call is z-ai/glm-5.3-flash, and the request shape is the one described above. It sits on the same key as everything else in the catalogue — one API across 200+ models, the provider's price passed through with no markup on top, and automatic failover if a provider stalls, so a model you are auditioning does not have to be a model your production path depends on.

One clarification about what this page is, since a reader arriving from the model's own name deserves it. This is a reference page for a model that is already in the catalogue, not a launch report: GLM-5.3-Flash is from August 26, 2026, and its place here rests on the fact that the Ox Alpha name keeps getting searched with no dedicated page to land on, not on the freshness of the release. The date above is used to date the model, not as news. What would change this page is one of four things — an independent benchmark run at a stated effort, a vendor-stated knowledge cutoff, a move in the served rate, or a decision by Z.ai to state what the launch-era chip count actually was. Until one of those lands, the specification above is the whole of what the vendor documents, and the gaps listed in the previous section are as much a part of the answer as the numbers are.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
