Hero title card for the Union Alpha vs GLM-5.3-Flash comparison, reading "Free until someone claims it" over a stat line contrasting Union Alpha's 262K context, $0 price and missing benchmarks against GLM-5.3-Flash's 1M context, $0.075 / $0.25 pricing and MIT weights.
Guides & Insights

Union Alpha vs GLM-5.3-Flash: Free Until Someone Claims It

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Only one of these two models can be taken away from you. Union Alpha is an unclaimed stealth listing that appeared on 16 September 2026 and became routable on 17 September 2026. It costs nothing and belongs to nobody who will admit it. GLM-5.3-Flash costs $0.075 per million input tokens and $0.25 per million output tokens, and Z.ai published its weights under MIT on 26 August 2026 — which means if the API ever gets worse, more expensive or disappears, you can download it and serve it yourself. Everything else about this matchup is close. That one difference is not, and it is the reason the answer to "which should I call today" splits cleanly in two.

The pairing is not arbitrary. GLM-5.3-Flash spent the first six days of its life under exactly the conditions Union Alpha is in now: an anonymous listing called Ox Alpha, a provider field reading "Stealth," no lab, no card, no weights. Z.ai revealed itself on 26 August, and the anonymous listing gave way to a published model with published pricing. Union Alpha is running the same playbook on the same platform, a month later, with a smaller advertised surface — and this time nobody has published a capacity figure or a "near-unlimited" promise to go with it.

So the free endpoint is not a cheaper version of the paid one. It is a different kind of thing, with a different expiry, and the specs that differ are the ones that tell you what it was built for.

What the two spec sheets actually say

Both figures below are read off the live model pages on our routing layer on 17 September 2026, plus the vendor's own release material for the parts a routing layer does not publish.

• Context window — Union Alpha 262,144 tokens (262K) vs GLM-5.3-Flash 1,048,576 tokens (1M). A quarter of the window.

• Maximum output — Union Alpha 131K tokens vs GLM-5.3-Flash 128K. Effectively level, and the one row where the stealth model is nominally ahead.

• Input modalities — Union Alpha takes text and images vs GLM-5.3-Flash takes text, images and video. Union Alpha outputs text only, as does GLM-5.3-Flash.

• Price — Union Alpha is $0 on a free tier that is explicitly rate-limited, with over-limit requests returning HTTP 429. GLM-5.3-Flash is $0.075 input / $0.25 output per 1M tokens, with cache reads at $0.017.

• Reasoning controls — Union Alpha exposes none. Its accepted parameters are tools, tool_choice, response_format, temperature, top_p and max_tokens, and that is the whole list. GLM-5.3-Flash exposes a reasoning control with low, high and max effort, defaulting to max.

• Weights — Union Alpha has no repository, no licence, no parameter count and no architecture note published. GLM-5.3-Flash is a 320B-parameter sparse mixture-of-experts with roughly 18B active per token, MIT-licensed, downloadable.

• Benchmarks — Union Alpha's benchmark field reads "pending." GLM-5.3-Flash has a published set.

• Trailing-week traffic on our layer — Union Alpha 1.2M tokens vs GLM-5.3-Flash 18,128.8M tokens. Read that one carefully: GLM-5.3-Flash has been generally available for three weeks and Union Alpha only since 05:32 UTC on 17 September, so the gap measures how much production traffic has had time to land, not how good either model is.

Read down that list and the shape of Union Alpha is coherent. A 262K window, image input, no video, no reasoning knob and no weights is the profile of a fast-tier or sibling variant rather than a frontier flagship — the same structural position GLM-5.3-Flash occupies in its own family. That is a hypothesis drawn from an advertised surface, not a finding, and it is worth naming the competing explanation: an operator can also trim a preview's advertised surface specifically to frustrate fingerprinting. Neither reading has evidence behind it yet.

Two-column scoreboard comparing Union Alpha and GLM-5.3-Flash: Union Alpha shows $0 rate-limited, 262K tokens, text + image, 10.00 s p50 first token, no reasoning control and benchmarks pending; GLM-5.3-Flash shows $0.075 in / $0.25 out, 1M tokens, text + image + video, 6.49 s p50 first token, low / high / max reasoning control and published benchmarks, eight of them measured independently of the vendor.

The latency story inverts, and one number is doing something odd

This is where the matchup gets more interesting than the spec sheet suggests, and where you should be suspicious of the headline figure.

Union Alpha's page reports p50 time to first token of 10.00 seconds — and p95 of 10.00 seconds. Identical percentiles are what a thin sample looks like: when almost every request in the window lands in the same bucket, the median and the tail collapse onto each other. Treat that 10-second figure as "slow to start, exact value unsettled" rather than as a precise median. GLM-5.3-Flash, measured over the same kind of window, reports 6.49 seconds at p50 against 10.00 seconds at p95 — a real distribution with a real spread.

Then there is output speed, and here the sources genuinely disagree. Our routing layer measured 225 tokens per second for Union Alpha over the trailing week. Independent testers calling the endpoint directly on launch day described something much slower — queueing, and waits measured in minutes for trivial prompts. We are not going to pretend those reconcile. The most likely explanation is a rolling window that is mostly off-peak against launch-hour congestion, and the honest position is that Union Alpha's throughput is unsettled until it has been up long enough to measure. GLM-5.3-Flash reports 77.8 tokens per second with a 0.27% error rate, against Union Alpha's 1.1%.

What is not in dispute is the direction of the trade. GLM-5.3-Flash starts producing sooner and produces steadily. Union Alpha's free tier may burst faster when it is not congested, and it will not tell you which regime you are in until you have already sent the request.

Free is not a price — it is a set of terms you have not read

"$0, rate-limited, HTTP 429 over the limit" is the entire pricing section on Union Alpha's model page. There is no published rate limit, no requests-per-minute figure, no capacity number and no end date. A stealth preview has no data processing agreement, no stated retention period, no jurisdiction, no SLA, no support path and no notice period, because there is no named counterparty to hold to any of it. That is not a reason to avoid it. It is a reason to know exactly what you are trading.

The specific risk is not hypothetical, because we watched it happen. When Ox Alpha's operator revealed itself on 26 August, the free tier did not survive the announcement — the anonymous listing went away and GLM-5.3-Flash's paid pricing took its place. A production path pointed at Union Alpha today has an expiry date that nobody has published. Build for the 404.

There is a second, quieter cost: no reasoning parameter means no way to dial thinking down. On a model that always reasons at maximum effort, you pay for the deliberation in output tokens on every call — which is a real line item on GLM-5.3-Flash and structurally impossible to incur on Union Alpha, because the knob does not exist. Whether that is a discount or a missing control depends on whether the model needs the deliberation, and with no benchmark data published for Union Alpha, nobody can currently tell you.

OrcaRouter model page for stealth/union-alpha-free captured 17 September 2026: context 262K tokens, max output 131K, input text + image, output text, p50 TTFT 10.00 s, price Free and rate-limited, P50 and P95 TTFT both 10.00 s over 7 days, traffic 1.2M tokens / 7d, supported parameters limited to max_tokens, response_format, temperature, tool_choice, tools and top_p, and an over-limit note reading HTTP 429 when a limit is hit.

What GLM-5.3-Flash's $0.075 actually buys

At $0.075 in and $0.25 out per million tokens, with cache reads at $0.017, GLM-5.3-Flash is priced like a budget model and specified like something else: a 1M-token window, video input, and MIT weights. Our own model page's cost calculator, on its default assumptions, puts a representative workload at roughly $1.07 a month with prompt caching against $1.28 without — the kind of gap that only matters at volume, but that costs nothing to collect since caching is a flag rather than a re-architecture.

The part that is not on the pricing page: because the weights are MIT and downloadable, $0.075/$0.25 is a ceiling, not a floor. If Z.ai raises the API price, or the endpoint degrades, or you simply outgrow metered inference, the same model runs on your own hardware with no one's permission. That option does not exist for Union Alpha at any price, and it is why "free" and "cheap" are not the same axis in this comparison.

Both models sit behind one key on OrcaRouter, which routes 200+ models at 0% markup — provider list price passed through unchanged. That matters more for the paid half of this matchup than the free half: if Z.ai cuts or raises GLM-5.3-Flash pricing, the number on our page is the new one the same day, and there is no second contract, no separate SDK and no migration to chase it. And because both are reachable through the same endpoint, you can evaluate Union Alpha and fall back to GLM-5.3-Flash without changing anything but a model string.

OrcaRouter model page for z-ai/glm-5.3-flash captured 17 September 2026: released 2026-08-26 by Z.ai, 320B total / 18B active parameters, 1M-token context, 128K max output, text + image + video input, text output, p50 TTFT 6.49 s and p95 10.00 s over 7 days, traffic 18,128.8M tokens / 7d, input $0.07 per 1M, output $0.25 per 1M and cache read $0.017 per 1M.

The benchmark question is the honest differentiator

Union Alpha has no benchmark number. Its model page's benchmark field says "pending," and nothing outside the operator's own listing has produced a reproducible score. Any figure you see circulating for this model right now is uncorroborated — including the ones that sound most impressive. That is not a knock on the model. It is simply the state of the evidence one day in, and it is the single biggest reason not to build on it yet.

GLM-5.3-Flash's published set is not uniformly vendor-reported, and the distinction changes what you can do with it. Our model page tags every entry with the source it came from, and eight of those entries come from Artificial Analysis rather than from Z.ai: GPQA Diamond 91.2, AA Coding 71.5, AA Intelligence 41.9, SciCode 51.6, Long-Context Recall 80, Humanity's Last Exam 39.9, tau_banking 47.2 and terminalbench_v2_1 84.27. Those are measured by a third party rather than reported by the lab that built the model, which puts them in a different category from everything else on the list.

The rest of the set is Z.ai's own reporting: DeepSWE v1.1 63.4, Terminal-Bench 2.1 84.3, AutomationBench v1.0.6 48.8, Agents' Last Exam 26.3, BabyVision 53.4, Chartography 78, CharXiv Reasoning 89.4, HLE with tools 55.3, MMVU 80.5, MVBench 77.8, NL2Repo 56.3, OfficeQA Pro 62.4 and Toolathlon Verified 78.4. We have not re-run any of those, and they are a claim with a methodology behind it — stronger than no claim at all, weaker than a harness you control. Set both lists against Union Alpha and the asymmetry is the whole story: GLM-5.3-Flash has scores a third party measured, and Union Alpha has none at all — not vendor-reported, not independently measured, nothing.

The practical consequence is that you cannot currently compare these two models on quality. You can compare them on price, context, modality, latency and terms, and those comparisons are what the decision actually rests on. If a quality comparison is what you need, the answer today is that it does not exist.

Where Union Alpha genuinely wins

Three places, and they are not small.

The first is cost of experimentation, which is exactly zero. If you want to know how a 262K-window model handles your repository, your prompt shape or your tool schema, Union Alpha will tell you for nothing, today, with no card on file. On GLM-5.3-Flash the same exploration is metered — cheap, but not free, and it adds up across a few hundred failed agent runs.

The second is output ceiling. Union Alpha's 131K-token maximum output is nominally larger than GLM-5.3-Flash's 128K, and for long single-shot generation — a full file, a long structured document, an agent transcript that has to finish in one turn — that is the row that decides whether you get a truncation.

The third is the absence of a reasoning tax. If your workload is high-volume, low-difficulty and latency-tolerant, a model with no deliberation knob cannot bill you for deliberation. It also cannot be tuned, which is the trade.

Where GLM-5.3-Flash wins, which is most of the rest

Four times the context window, which is the difference between fitting a mid-size repository and fitting a large one. Video input, which Union Alpha does not have at all. A reasoning control you can turn down, which is a cost lever as much as a quality lever. Weights you can hold, which converts a vendor relationship into an asset. Published benchmarks, eight of them measured independently of the vendor. A 6.49-second median start against a flat 10-second one. And an identifiable counterparty with a licence, a jurisdiction and a support path — the unglamorous thing that turns a prototype into a system.

How to take both without betting on either

The mistake available here is treating the choice as exclusive. It is not, and the tooling makes it cheap to stop treating it that way.

Point your evaluation traffic at Union Alpha, because it is free and it might turn out to be excellent. Put GLM-5.3-Flash behind it as the fallback, because it is cheap, specified, and cannot be withdrawn from you. Then make the switch a configuration rather than a rewrite. That is what the routing DSL is for — composing several models behind one endpoint so a request can try the unproven one first and land on the proven one when it throttles, returns a 429, or stops existing. Automatic failover across providers does the same job without you writing the retry logic, and it is the specific mechanism that makes an anonymous endpoint safe to evaluate: you are not committing to a model, you are committing to a socket.

Is Union Alpha secretly GLM-5.3-Flash?

Almost certainly not the same model, and the spec sheet is why. Union Alpha has a quarter of GLM-5.3-Flash's context, no video input and no reasoning parameter. A model that was GLM-5.3-Flash would not arrive with those three things removed. A model from the same family, positioned below it, is a much better fit for what is advertised — but nobody has published a fingerprint, the name "Union" breaks the animal-codename convention that has otherwise held for a year, and there is no evidence either way. Treat any confident attribution you read this week, including a confident denial, as a guess.

What happens to Union Alpha when the preview ends?

The precedent is the most useful thing in this article. Ox Alpha appeared on 20 August, was revealed on 26 August, and its free tier did not outlive the reveal. Six days, start to finish. If Union Alpha follows the same arc, the free window is measured in days, and the most likely outcomes are a named vendor, a price, published weights, or the listing simply disappearing. Any of those four changes what you can do with it. None of them is announced.

Is it safe to send Union Alpha work you care about?

Only if you would be comfortable sending that work to an anonymous operator with no stated retention policy, no jurisdiction and no agreement. The model page says plainly that the provider is not disclosed while the model is being evaluated and that capabilities and availability may change without notice — which is a fair warning rather than a red flag. The sensible boundary is code and data you can afford to lose: public repositories, synthetic fixtures, throwaway branches. Customer data, credentials and anything under a compliance obligation stay on the paid, attributable side of this comparison.

Which to call today

If you want to know whether a 262K-window image-capable model is useful to you, call Union Alpha right now and pay nothing. Do it behind a fallback, expect a 429 eventually, assume the free tier dies without notice, and do not let anything you cannot lose touch it. That is a genuinely good deal and it expires on a schedule nobody has published.

If you need a model you can commit to — 1M of context, video, a reasoning lever, published scores, an MIT licence and weights that make the price a ceiling rather than a floor — GLM-5.3-Flash is the one, and at $0.075/$0.25 with $0.017 cache reads it is not an expensive commitment. Half the benchmark set is a claim rather than a measurement, but the other half was measured by someone other than Z.ai, and the rest you can check against your own workload for a few cents.

The asymmetry worth carrying out of this article is not free versus paid. It is that one of these two models is a test and the other is a purchase. Union Alpha is free until somebody claims it. GLM-5.3-Flash you can keep.