Generated title card titled "Union Alpha vs Gemini 3.8 Flash" with the subtitle "One has a score, the other has a claim", above three rounded cards reading "Gemini 3.8 Flash: dated eval table, published price", "Union Alpha: no benchmarks from its operator", and "Max output: 131,072 vs 65,536 tokens". Footer reads Union Alpha figures unaudited with Gemini figures per Google and Artificial Analysis.
Guides & Insights

Union Alpha vs Gemini 3.8 Flash: One Has a Score, the Other Has a Claim

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Union Alpha and Gemini 3.8 Flash are the same pitch delivered at two completely different levels of proof. Union Alpha's operator describes it as having "frontier-level performance across a broad range of general-purpose tasks" and has published no benchmark, no parameter count, no licence, and no name. Gemini 3.8 Flash shipped with a dated evaluation table, an Artificial Analysis Intelligence Index, per-task cost figures from independent testers, and a documented price schedule that runs into 2027. If you are choosing between them on anything other than price, you are choosing between a number you can check and a sentence you cannot.

Side by side, before the argument

• Context window — Union Alpha 262,144 tokens vs Gemini 3.8 Flash 1,048,576 tokens.

• Max output — 131,072 tokens vs 65,536 tokens. Union Alpha wins this one, and it is the only structural axis where it does.

• Inputs — text and image on both, but Gemini 3.8 Flash additionally accepts video, audio, and files.

• Pricing — Union Alpha $0 during its preview vs Gemini 3.8 Flash $0.75/M input, $3.75/M output, $0.075/M cached input, doubling to $1.50/$7.50 on 1 January 2027.

• Reasoning controls — none exposed on Union Alpha vs thinking levels on Gemini 3.8 Flash that directly move both quality and cost.

• Published evaluation — none from Union Alpha's operator vs a full dated table on Gemini 3.8 Flash.

• Provenance — anonymous operator, no weights vs Gemini 3.8 Flash's dated release and documented lineage.

Generated two-column comparison scoreboard for Union Alpha and Gemini 3.8 Flash. Union Alpha column: context window 262,144 tokens, max output 131,072 tokens, text and image input, $0 preview price, no published evals from its operator, 10.00 s p50 time to first token. Gemini 3.8 Flash column: context window 1,048,576 tokens, max output 65,536 tokens, text, image, video and audio input, $0.75 input and $3.75 output per 1M tokens, a full dated evaluator table, 3.46 s p50 time to first token. Footer reads latency per OrcaRouter over the last 7 days with Union Alpha benchmarks unaudited.

What Gemini 3.8 Flash's numbers actually say

Gemini 3.8 Flash launched on 2 September 2026 — Google's third Flash release in roughly six weeks, which tells you how much of this model's story is cadence rather than breakthrough. The published evaluator table credits artificialanalysis.ai and deepmind.google between them, so read it as a mix of independent and vendor figures rather than a clean third-party audit.

• AA Coding — 76.3, which is a genuinely strong coding score.

• AA Intelligence — 41.2. Note that Artificial Analysis separately reports 59 on the Intelligence Index at high reasoning effort, so the number moves a lot with the effort setting; quote the setting or the number is meaningless.

• GPQA Diamond — 95.3, the highest single figure in this comparison by a wide margin.

• HLE-Verified — 54.9, with plain Humanity's Last Exam at 47.8.

• Terminal-Bench 2.1 — 89.4 on Google's own table, fractionally ahead of Claude Opus 5's 89.1.

• Long-context recall — 81.3; SciCode 56.6; tau_banking 44.9.

• Observed throughput — 323 output tokens per second over a recent seven-day window, with a 0.63% error rate, a 3.46-second median time-to-first-token, and 498.5M tokens of measured traffic.

The weaknesses are as documented as the strengths. On Terminal-Bench 4.0 it scores 19.1 against Claude Opus 5's 51.8. On OSWorld 2.0 it manages 59.0% against 75.4%. On GDPVal-AA v2 it trails at 1545 Elo against 1824. Gemini 3.8 Flash is excellent at a specific band of work and falls off a cliff on the hardest general-agent evaluations — which is exactly the kind of shape a published table lets you see.

The pricing twist that catches people out

Google left the per-token price unchanged from Gemini 3.7 Flash. The bill still went up.

The model works harder: more reasoning steps, more iterative tool calls. Independent testing by Artificial Analysis measured roughly 30% more output tokens per task and about 40% higher cost per task than 3.7 Flash. At high reasoning, that lands around $0.58 per task; medium sits near $0.41 and low near $0.24.

This is the single most useful thing in the comparison for anyone budgeting, and it generalises well beyond these two models: unit price and cost-per-task are different numbers, and only one of them appears on a pricing page. Google has kept 3.7 Flash available as the cheaper option, which is a tacit acknowledgement of the shift.

What Union Alpha offers instead

Union Alpha appeared on third-party catalogues on 16 September 2026 and runs on OrcaRouter's free tier. It is anonymous — no lab, no weights, no licence, no knowledge cutoff — and free for a stated window of about a week, rate-limited, with no post-preview price announced.

Its endpoint is verifiable even where its provenance is not: 262,144-token context, 131,072-token maximum output, text and image input with text-only output, tool calling and JSON output, temperature, top_p, max_tokens, and no reasoning controls of any kind. Tool choice effectively only supports "auto".

There is exactly one independent measurement. SMF Clearinghouse ran 157 Official A tests with reasoning off and scored it 136/157 — 86.6%, zero errors, tenth of twenty-six. Reasoning was perfect at 30/30, tools 2/2, coding 26/30, math 24/30. Writing came in at 2/5.

That is not a bad result. It is simply not a comparable one. Official A is a broad 157-test general suite; AA Coding and GPQA Diamond are different instruments measuring different things at different difficulties. Putting 136/157 next to 95.3 on GPQA Diamond and declaring a winner would be exactly the kind of laundering that makes vendor numbers useless.

The OrcaRouter model page for Gemini 3.8 Flash, showing a 1M-token context window, a 65K max output, text plus image, video, file and audio input, $0.75 per million input tokens and $3.75 per million output tokens, a p50 time to first token of 3.46 s, and 498.5M tokens of traffic over seven days.The OrcaRouter model page for Union Alpha (Free), listed under the stealth vendor, showing a FREE badge, a 262K-token context window, 131K max output, text and image input with text output, and pricing shown as Free and rate-limited with model usage at $0.

The cost comparison nobody can actually finish

Here is the worked example, and it is deliberately incomplete.

Take a task you run a thousand times a month. On Gemini 3.8 Flash at high reasoning, the published per-task figure of about $0.58 puts that at roughly $580 a month. That is a real, defensible number built from measured token consumption.

On Union Alpha, that same thousand tasks cost $0 today. What they cost in three weeks is unknown, because no post-preview price exists. What they cost in reliability is also unknown, because the model is rate-limited, runs on a single endpoint with no fallback provider, and answers with HTTP 429 once you cross a cap nobody has published.

So the honest answer is that Union Alpha wins the only cost comparison available and loses the ability to forecast it. For a one-off bake-off, that is fine. For a budget line, it is not a number — it is a hope with a URL.

Where each one belongs

Gemini 3.8 Flash belongs wherever you need a measured floor: long-context document work with video or audio input, coding tasks where the AA Coding score of 76.3 is the relevant instrument, and any workload with a budget attached, because you can compute the cost before you commit and you know it doubles on 1 January 2027.

Union Alpha belongs in the exploratory slot — trying a vision-capable 256K model whose output ceiling is double Gemini's, on work where a vanished endpoint costs you nothing.

The reason to run them behind one endpoint rather than two integrations is that this comparison is going to change shape repeatedly. A routing DSL lets you compose the two into a single call — Gemini 3.8 Flash as the measured path, Union Alpha as the experimental branch — instead of hard-coding a decision you will want to revisit when the free window closes or when Google's introductory pricing expires. Neither of those dates is far away, and both will move the arithmetic.

What would settle it

One dated entry for Union Alpha on a leaderboard that also carries Gemini 3.8 Flash. That is the entire missing piece. Until then, the defensible position is not "Union Alpha is as good as Gemini 3.8 Flash" — it is "Union Alpha is free for a week and nobody has measured it against anything Google publishes." Those are very different sentences, and only one of them is currently true.

The measured half of the pair is documented here: Gemini 3.8 Flash model page carries the $0.75/$3.75 rates, the 90% cached-input discount, and the 65,536-token output ceiling that decides long-report work.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily