
Union Alpha vs GLM-5.3-Flash:誰かが手に入れるまで無料
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 100万トークンあたり
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040知能
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453知能77コーディング
- googleGoogle: Gemini 3.8 Flash2026-09-0241知能76コーディング
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245知能76コーディング
- anthropicAnthropic: Claude Fable 5.12026-09-0153知能82コーディング
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 100万トークンあたり
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642知能72コーディング
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 100万トークンあたり
- z-aiZ.ai: GLM 5.32026-08-1845知能75コーディング
- obsidianQwen3.8 27B2026-08-1534知能68コーディング
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236知能69コーディング
- grokSpaceXAI: Grok 4.62026-08-1244知能77コーディング
- metaMeta: Muse Spark 1.22026-08-0540知能72コーディング
- qwenQwen: Qwen3.8 Max2026-08-0345知能76コーディング
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135知能69コーディング
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 100万トークンあたり
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451知能78コーディング
Only one of these two models can be taken away from you. Union Alpha is an unclaimed stealth listing that appeared on 16 September 2026 and became routable on 17 September 2026. It costs nothing and belongs to nobody who will admit it. GLM-5.3-Flash costs $0.075 per million input tokens and $0.25 per million output tokens, and Z.ai published its weights under MIT on 26 August 2026 — which means if the API ever gets worse, more expensive or disappears, you can download it and serve it yourself. Everything else about this matchup is close. That one difference is not, and it is the reason the answer to "which should I call today" splits cleanly in two.
The pairing is not arbitrary. GLM-5.3-Flash spent the first six days of its life under exactly the conditions Union Alpha is in now: an anonymous listing called Ox Alpha, a provider field reading "Stealth," no lab, no card, no weights. Z.ai revealed itself on 26 August, and the anonymous listing gave way to a published model with published pricing. Union Alpha is running the same playbook on the same platform, a month later, with a smaller advertised surface — and this time nobody has published a capacity figure or a "near-unlimited" promise to go with it.
So the free endpoint is not a cheaper version of the paid one. It is a different kind of thing, with a different expiry, and the specs that differ are the ones that tell you what it was built for.
2つの仕様書に実際に書かれていること
Both figures below are read off the live model pages on our routing layer on 17 September 2026, plus the vendor's own release material for the parts a routing layer does not publish.
• Context window — Union Alpha 262,144 tokens (262K) vs GLM-5.3-Flash 1,048,576 tokens (1M). A quarter of the window.
• Maximum output — Union Alpha 131K tokens vs GLM-5.3-Flash 128K. Effectively level, and the one row where the stealth model is nominally ahead.
• Input modalities — Union Alpha takes text and images vs GLM-5.3-Flash takes text, images and video. Union Alpha outputs text only, as does GLM-5.3-Flash.
• 価格 — Union Alpha は、明示的にレート制限された無料枠で $0 です。上限を超えたリクエストは HTTP 429 を返します。GLM-5.3-Flash は 100万トークンあたり入力 $0.075 / 出力 $0.25 で、キャッシュ読み取りは $0.017 です。
• Reasoning controls — Union Alpha exposes none. Its accepted parameters are tools, tool_choice, response_format, temperature, top_p and max_tokens, and that is the whole list. GLM-5.3-Flash exposes a reasoning control with low, high and max effort, defaulting to max.
• Weights — Union Alpha has no repository, no licence, no parameter count and no architecture note published. GLM-5.3-Flash is a 320B-parameter sparse mixture-of-experts with roughly 18B active per token, MIT-licensed, downloadable.
• Benchmarks — Union Alpha's benchmark field reads "pending." GLM-5.3-Flash has a published set.
• 当社レイヤーにおける直近1週間のトラフィック — Union Alpha 1.2Mトークン 対 GLM-5.3-Flash 18,128.8Mトークン。この点は注意深く読んでください。GLM-5.3-Flash は3週間前から一般提供されていますが、Union Alpha は9月17日 05:32 UTC 以降に提供が始まったばかりです。したがって、この差は、どちらのモデルが優れているかではなく、本番トラフィックがどれだけ着地する時間を得たかを測っているものです。
Read down that list and the shape of Union Alpha is coherent. A 262K window, image input, no video, no reasoning knob and no weights is the profile of a fast-tier or sibling variant rather than a frontier flagship — the same structural position GLM-5.3-Flash occupies in its own family. That is a hypothesis drawn from an advertised surface, not a finding, and it is worth naming the competing explanation: an operator can also trim a preview's advertised surface specifically to frustrate fingerprinting. Neither reading has evidence behind it yet.

The latency story inverts, and one number is doing something odd
This is where the matchup gets more interesting than the spec sheet suggests, and where you should be suspicious of the headline figure.
Union Alpha's page reports p50 time to first token of 10.00 seconds — and p95 of 10.00 seconds. Identical percentiles are what a thin sample looks like: when almost every request in the window lands in the same bucket, the median and the tail collapse onto each other. Treat that 10-second figure as "slow to start, exact value unsettled" rather than as a precise median. GLM-5.3-Flash, measured over the same kind of window, reports 6.49 seconds at p50 against 10.00 seconds at p95 — a real distribution with a real spread.
Then there is output speed, and here the sources genuinely disagree. Our routing layer measured 225 tokens per second for Union Alpha over the trailing week. Independent testers calling the endpoint directly on launch day described something much slower — queueing, and waits measured in minutes for trivial prompts. We are not going to pretend those reconcile. The most likely explanation is a rolling window that is mostly off-peak against launch-hour congestion, and the honest position is that Union Alpha's throughput is unsettled until it has been up long enough to measure. GLM-5.3-Flash reports 77.8 tokens per second with a 0.27% error rate, against Union Alpha's 1.1%.
議論の余地がないのは、このトレードオフの方向性だ。GLM-5.3-Flash はより早く生成を開始し、安定して生成し続ける。Union Alpha の無料枠は、混雑していないときにはより速くバーストするかもしれないが、リクエストを送信してしまった後になるまで、自分がどちらの状態にあるのかを教えてくれない。
無料は価格ではない — それはあなたが読んでいない条件の集合だ
"$0, rate-limited, HTTP 429 over the limit" is the entire pricing section on Union Alpha's model page. There is no published rate limit, no requests-per-minute figure, no capacity number and no end date. A stealth preview has no data processing agreement, no stated retention period, no jurisdiction, no SLA, no support path and no notice period, because there is no named counterparty to hold to any of it. That is not a reason to avoid it. It is a reason to know exactly what you are trading.
The specific risk is not hypothetical, because we watched it happen. When Ox Alpha's operator revealed itself on 26 August, the free tier did not survive the announcement — the anonymous listing went away and GLM-5.3-Flash's paid pricing took its place. A production path pointed at Union Alpha today has an expiry date that nobody has published. Build for the 404.
There is a second, quieter cost: no reasoning parameter means no way to dial thinking down. On a model that always reasons at maximum effort, you pay for the deliberation in output tokens on every call — which is a real line item on GLM-5.3-Flash and structurally impossible to incur on Union Alpha, because the knob does not exist. Whether that is a discount or a missing control depends on whether the model needs the deliberation, and with no benchmark data published for Union Alpha, nobody can currently tell you.

What GLM-5.3-Flash's $0.075 actually buys
At $0.075 in and $0.25 out per million tokens, with cache reads at $0.017, GLM-5.3-Flash is priced like a budget model and specified like something else: a 1M-token window, video input, and MIT weights. Our own model page's cost calculator, on its default assumptions, puts a representative workload at roughly $1.07 a month with prompt caching against $1.28 without — the kind of gap that only matters at volume, but that costs nothing to collect since caching is a flag rather than a re-architecture.
The part that is not on the pricing page: because the weights are MIT and downloadable, $0.075/$0.25 is a ceiling, not a floor. If Z.ai raises the API price, or the endpoint degrades, or you simply outgrow metered inference, the same model runs on your own hardware with no one's permission. That option does not exist for Union Alpha at any price, and it is why "free" and "cheap" are not the same axis in this comparison.
両モデルはOrcaRouter上の1つのキーの背後にあり、200以上のモデルを0%のマークアップでルーティングします — プロバイダーの定価がそのまま適用されます。このことは、この対決の無料の半分よりも有料の半分にとってより重要です。Z.aiがGLM-5.3-Flashの価格を引き下げたり引き上げたりすれば、当社のページに載る数字は同じ日に新しいものになり、それを追いかけるための2つ目の契約も、別個のSDKも、移行も必要ありません。そして、どちらも同じエンドポイントを通じて利用できるため、モデル文字列以外は何も変更することなく、Union Alphaを評価し、GLM-5.3-Flashにフォールバックできます。

The benchmark question is the honest differentiator
Union Alphaにはベンチマークの数値がない。モデルページのベンチマーク欄には「保留中」と表示されており、運営者自身の掲載情報以外に、再現可能なスコアを示したものは何もない。このモデルについて今出回っている数値は、どれも裏付けが取れていない——最も印象的に聞こえるものも含めて。これはモデルを貶すものではない。公開から一日時点での証拠の状況にすぎず、まだこれを土台に進めるべきでない最大の理由なのだ。
GLM-5.3-Flash's published set is not uniformly vendor-reported, and the distinction changes what you can do with it. Our model page tags every entry with the source it came from, and eight of those entries come from Artificial Analysis rather than from Z.ai: GPQA Diamond 91.2, AA Coding 71.5, AA Intelligence 41.9, SciCode 51.6, Long-Context Recall 80, Humanity's Last Exam 39.9, tau_banking 47.2 and terminalbench_v2_1 84.27. Those are measured by a third party rather than reported by the lab that built the model, which puts them in a different category from everything else on the list.
The rest of the set is Z.ai's own reporting: DeepSWE v1.1 63.4, Terminal-Bench 2.1 84.3, AutomationBench v1.0.6 48.8, Agents' Last Exam 26.3, BabyVision 53.4, Chartography 78, CharXiv Reasoning 89.4, HLE with tools 55.3, MMVU 80.5, MVBench 77.8, NL2Repo 56.3, OfficeQA Pro 62.4 and Toolathlon Verified 78.4. We have not re-run any of those, and they are a claim with a methodology behind it — stronger than no claim at all, weaker than a harness you control. Set both lists against Union Alpha and the asymmetry is the whole story: GLM-5.3-Flash has scores a third party measured, and Union Alpha has none at all — not vendor-reported, not independently measured, nothing.
実際的な帰結として、現時点ではこれら2つのモデルを品質で比較することはできません。価格、コンテキスト、モダリティ、レイテンシ、契約条件で比較することはできます。そして、意思決定が実際に依拠するのはそうした比較です。品質の比較が必要なのであれば、今日の答えは、それは存在しない、ということになります。
Where Union Alpha genuinely wins
Three places, and they are not small.
The first is cost of experimentation, which is exactly zero. If you want to know how a 262K-window model handles your repository, your prompt shape or your tool schema, Union Alpha will tell you for nothing, today, with no card on file. On GLM-5.3-Flash the same exploration is metered — cheap, but not free, and it adds up across a few hundred failed agent runs.
二つ目は出力上限です。Union Alphaの最大出力131Kトークンは、名目上はGLM-5.3-Flashの128Kより大きく、長い単発生成——ファイル全体、長い構造化ドキュメント、1ターンで終わらせる必要のあるエージェントのトランスクリプト——においては、打ち切りが起きるかどうかを決めるのはこの行です。
第三は、推論税が存在しないことだ。ワークロードが大量で低難易度、かつレイテンシを許容するなら、熟考の調整ノブを持たないモデルは、熟考分をあなたに課金できない。また、チューニングもできない。それがトレードオフだ。
Where GLM-5.3-Flash wins, which is most of the rest
Four times the context window, which is the difference between fitting a mid-size repository and fitting a large one. Video input, which Union Alpha does not have at all. A reasoning control you can turn down, which is a cost lever as much as a quality lever. Weights you can hold, which converts a vendor relationship into an asset. Published benchmarks, eight of them measured independently of the vendor. A 6.49-second median start against a flat 10-second one. And an identifiable counterparty with a licence, a jurisdiction and a support path — the unglamorous thing that turns a prototype into a system.
How to take both without betting on either
The mistake available here is treating the choice as exclusive. It is not, and the tooling makes it cheap to stop treating it that way.
Point your evaluation traffic at Union Alpha, because it is free and it might turn out to be excellent. Put GLM-5.3-Flash behind it as the fallback, because it is cheap, specified, and cannot be withdrawn from you. Then make the switch a configuration rather than a rewrite. That is what the routing DSL is for — composing several models behind one endpoint so a request can try the unproven one first and land on the proven one when it throttles, returns a 429, or stops existing. Automatic failover across providers does the same job without you writing the retry logic, and it is the specific mechanism that makes an anonymous endpoint safe to evaluate: you are not committing to a model, you are committing to a socket.
Is Union Alpha secretly GLM-5.3-Flash?
Almost certainly not the same model, and the spec sheet is why. Union Alpha has a quarter of GLM-5.3-Flash's context, no video input and no reasoning parameter. A model that was GLM-5.3-Flash would not arrive with those three things removed. A model from the same family, positioned below it, is a much better fit for what is advertised — but nobody has published a fingerprint, the name "Union" breaks the animal-codename convention that has otherwise held for a year, and there is no evidence either way. Treat any confident attribution you read this week, including a confident denial, as a guess.
What happens to Union Alpha when the preview ends?
この記事で最も役立つのは前例だ。Ox Alphaは8月20日に登場し、8月26日に公開され、その無料枠は公開後も存続しなかった。最初から最後まで6日間。もしUnion Alphaが同じ経過をたどるなら、無料で使える期間は日単位で測られ、最も可能性の高い結末は、ベンダー名の公表、価格、重みの公開、あるいは掲載が単に消えることだ。これら四つのどれが起きても、それでできることは変わる。どれも告知はされない。
Is it safe to send Union Alpha work you care about?
Only if you would be comfortable sending that work to an anonymous operator with no stated retention policy, no jurisdiction and no agreement. The model page says plainly that the provider is not disclosed while the model is being evaluated and that capabilities and availability may change without notice — which is a fair warning rather than a red flag. The sensible boundary is code and data you can afford to lose: public repositories, synthetic fixtures, throwaway branches. Customer data, credentials and anything under a compliance obligation stay on the paid, attributable side of this comparison.
Which to call today
262Kウィンドウの画像対応モデルが自分にとって役立つかどうかを知りたいなら、今すぐUnion Alphaを呼び出して、何も支払わないでください。フォールバックを用意したうえで行い、最終的には429が返ってくることを想定し、無料枠は予告なく終了すると仮定し、失っても構わないもの以外はそれに触れさせないでください。それは本当に良い取引で、誰も公表していないスケジュールで失効します。
コミットできるモデルが必要なら——1Mのコンテキスト、動画、推論レバー、公開スコア、MITライセンス、そして価格を下限ではなく上限にする重み——GLM-5.3-Flashがそれであり、$0.075/$0.25でキャッシュ読み取りが$0.017なら、高価なコミットではない。ベンチマークセットの半分は測定ではなく主張だが、残りの半分はZ.ai以外によって測定されており、その他は数セントで自分のワークロードと照らし合わせて確認できる。
この記事から持ち帰る価値のある非対称性は、無料対有料ではない。それは、この2つのモデルの一方がテストで、もう一方が購入だということだ。Union Alpha は誰かがそれを手に入れるまで無料だ。GLM-5.3-Flash はあなたが持ち続けられる。
