
Union Alpha vs Qwen3.8 Flash: One Point Is the Whole Story
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Union Alpha and Qwen3.8 Flash are one point apart on the only test that has measured both. On the SMF Clearinghouse Official A suite — 157 tests, reasoning off — Union Alpha scored 136 and a local run of Qwen3.8-Flash-Next scored 137. That is the closest head-to-head result either model has, and it is also the most misleading number in this comparison, because the two entries did not run under the same conditions and only one of them is a model you can actually rent right now.
What the one-point gap does and does not tell you
Start with what it establishes. Union Alpha is not a bluff in the trivial sense — a broad 157-test suite with zero errors, perfect reasoning (30/30) and perfect tool use (2/2) is not something a hollow endpoint produces. Its coding result of 26/30 and math of 24/30 place it in credible territory, and its writing score of 2/5 places it nowhere near general-purpose prose work.
Now the caveats, which are load-bearing.
The Qwen3.8 Flash entry was a local run of Qwen3.8-Flash-Next — the open-weight checkpoint — not the hosted Qwen3.8 Flash API. Local serving means your own quantisation, your own inference stack, your own tool-call parser. None of that is guaranteed to match what the hosted endpoint returns, and the people running the test flagged the comparison as non-definitive for exactly that reason.
One suite is not a profile. Official A is wide but shallow per category: tools is 2 tests. A two-test tool category scoring 2/2 is a good sign, not evidence of agentic reliability, and Union Alpha is explicitly positioned as an agentic and coding model. The instrument is measuring something adjacent to the claim.
And the one point itself is inside the noise of a 157-test suite. Treat 136 and 137 as "these are in the same band", not as a ranking.
Side by side
• Context window — Union Alpha 262,144 tokens vs Qwen3.8 Flash 1,000,000 tokens served (262,144 native, extended via YaRN).
• Max output — Union Alpha 131,072 tokens vs Qwen3.8 Flash 131,000 tokens. Effectively level.
• Inputs — text and image on Union Alpha vs text, image, and video on Qwen3.8 Flash.
• Pricing — $0 during Union Alpha's preview vs Qwen3.8 Flash at $0.150/M input, $0.470/M output, $0.018/M cached read, $0.230/M cache write.
• Reasoning controls — none on Union Alpha vs a thinking mode on Qwen3.8 Flash that defaults on and can be disabled.
• Time to first token — Union Alpha 10.00 s p50 vs Qwen3.8 Flash 6.62 s p50, both measured on our endpoint over the same seven-day window. Raw decode speed is closer than the reputation suggests: 170 tokens per second against 103.
• Provenance — anonymous operator, no weights vs Alibaba/Qwen, with Qwen3.8-Flash-Next weights open-sourced on Hugging Face and ModelScope.

Qwen3.8 Flash: a 6B-active bet on efficiency
Qwen3.8 Flash launched on 26 August 2026 as the production, managed version of the open-weight Qwen3.8-Flash-Next checkpoint, which Alibaba released simultaneously. It is a 125B-parameter mixture-of-experts model activating roughly 6B parameters per token, plus 51B N-gram embedding parameters, built on hybrid attention that combines Gated DeltaNet with a sparse-attention indexer.
The vendor framing is about training economics: Alibaba reports the run cost roughly one-ninth of Qwen3.7-Plus, with claimed prefill up to 7.6× faster and decode up to 4.9× faster on 1M-token workloads with high cache hits. Those are vendor figures and have not been independently reproduced.
Its benchmark table is also vendor-reported and unreproduced: SWE-bench Pro 62.5, DeepSWE v1.1 58.7, SWE-bench Multilingual 81.0, NL2Repo 48.1, CoWorkBench 73.9, JobBench 55.7, RealWorldQA 88.5, LVBench 76.6, AndroidWorld 84.5. The number worth quoting back at the marketing is Humanity's Last Exam at 35.9, which sits below Claude Opus 4.6's 40.0 in the same comparison.
On our own model page the public benchmark section still reads "pending" — no third party has scored it yet. What we do have is our own traffic: a 6.62-second median time-to-first-token, a 103 tokens-per-second output rate, and a 4.5% error rate. That error rate is high enough to be worth watching, and it is the kind of figure that only shows up when you measure a live endpoint rather than a launch post.

Union Alpha: the model with no owner
Union Alpha appeared on third-party catalogues on 16 September 2026 and is served on OrcaRouter's free tier. It has no named lab, no weights, no licence, no parameter count, and no architecture note. Its provider field reads "Stealth".
Its endpoint, at least, is legible: 262,144-token context, 131,072-token maximum output, text and image input with text-only output, tool calling, JSON output, temperature, top_p and max_tokens, and no reasoning controls at all. Tool choice effectively only honours "auto". Its free window has been described as about a week, rate-limited, with no post-preview pricing announced.
The stated claim — "frontier-level performance across a broad range of general-purpose tasks" — is operator-reported and unaudited. The 136/157 Official A result is the only independent measurement that exists.

Speed is where they stop looking alike
The headline throughput numbers do not separate them the way you would expect, and the reason is worth understanding.
On our own routing telemetry over the last seven days, Qwen3.8 Flash streams at 103 output tokens per second with a p50 time-to-first-token of 6.62 seconds and a 4.5% error rate. Union Alpha, measured the same way, streams at 170 tokens per second. On raw decode speed the anonymous model is the faster of the two.
What it does not do is start. Union Alpha's p50 and p95 time-to-first-token are both pinned at 10.00 seconds, meaning at least half of all requests wait ten seconds or more before the first token appears, and its error rate over the same window is 7.7% — about 1.7 times Qwen's. The 157-test Official A run that the earlier sections lean on took 4.6 hours of wall clock, with a median latency of 91.9 seconds and a mean of 104.8 seconds per test, and four tests hit the harness's 300-second timeout. Independent testers who ran it on launch day reported far lower throughput than our endpoint shows, which is what you would expect from a service still absorbing an enormous first-day load.
So the practical difference is not decode speed. It is time-to-first-token and reliability. Union Alpha is not usable for anything interactive — no autocomplete, no inline assistance, no tight agent loop — because ten seconds of silence is indistinguishable from a hang. It fits asynchronous work: overnight batch jobs, long document passes, offline evaluation runs where nothing is waiting on the first token and a 7.7% failure rate gets absorbed by a retry.
Qwen3.8 Flash at 103 tokens per second with a 6.62-second median time-to-first-token is not fast either. The honest framing is that both are slow, and only one of them fails roughly one request in thirteen.
The efficiency argument, and who gets to make it
Alibaba makes a training-cost argument: one-ninth the cost of Qwen3.7-Plus to train, six billion active parameters per token, so cheap to serve. That is a claim about a model whose weights exist and whose lineage is documented.
Union Alpha's efficiency story is implied rather than stated — a 256K anonymous model that is free to call. Free is not the same as cheap. Someone is paying for those GPUs, the preview has a stated end, and the operator's own admission that they were tuning for speed suggests the serving economics are not yet settled. A model that is free for a week and unpriceable afterwards has an efficiency argument that expires with the window.
Trying the unknown without betting on it
The reason this specific matchup is worth holding in your head is that it is the cheapest possible test of an unproven model. Qwen3.8 Flash is the known quantity: named vendor, open weights, published (if vendor-flavoured) benchmarks, a real per-token price of $0.150/$0.470. Union Alpha is the unknown, priced at zero, whose entire competitive claim rests on one 157-test result.
Rather than picking, put the unknown behind a failover rule. OrcaRouter passes provider list price through at 0% markup and supports automatic failover across providers, so a rate-limited or vanished Union Alpha endpoint falls through to a model that is still answering instead of surfacing as a failed job at 3am. Both models sit on the same key, so the experiment costs a config change rather than a second integration — and neither has to be a decision you defend, because you can flip the routing when the free window closes.
Verdict
Qwen3.8 Flash is the model to build on. Six billion active parameters, open-sourced sibling weights, a 1M-token window, video input, a working 103 tokens per second, and a published price you can forecast. Its benchmarks are vendor-reported and its error rate is worth monitoring, but it belongs to somebody and that somebody is shipping.
Union Alpha is the model to measure. The one-point gap on Official A is genuinely interesting — it suggests that whatever this thing is, it is not a toy — and at $0 with a 131K output ceiling and vision input, it is worth a week of your attention. But a one-point gap on one suite, produced by an anonymous operator with a 7.7% error rate, is a reason to investigate rather than a reason to switch.
The thing to watch is the same thing that has always settled this: a name. Ox Alpha, the previous entry in this series, was revealed six days after appearing as Zhipu's GLM-5.3-Flash. If Union Alpha gets one, the comparison stops being about a point and starts being about a price.
The known quantity is priced and documented: Qwen3.8 Flash model page shows the $0.150/$0.470 rates, the 1M-token served context and the 131K output ceiling, with public benchmarks still pending.
