
Ling-3.1-flash vs Gemini 3.8 Flash: A 256K Trial Model Against a 1M-Token Live One
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 219 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 117 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 41 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Put Ling-3.1-flash and Gemini 3.8 Flash side by side and the mismatch is not about intelligence — it is about state. Ling-3.1-flash, Ant Group's ~560B-parameter mixture-of-experts announced on September 29 to 30, 2026, is a free two-week trial with a 256K served context, a 32,768-token output cap, text-only input, and no published weights, licence, or price. Gemini 3.8 Flash, Google's Flash-tier reasoning model released September 2, 2026, is a general-availability API with a full 1,048,576-token window, 65,536-token output, multimodal input, and a public rate card. On paper that is a comparison of two models; in practice it is a comparison of a model you can build on today against one you can only test this month.
That is not a reason to dismiss Ling-3.1-flash. A free 560B MoE with an agentic bent is worth two weeks of anyone's evaluation time. But it changes what a head-to-head is for: not "which should I standardise on," but "is the trial model good enough that I should hedge on the answer fifteen days from now." Those are different questions, and they have different answers on every axis below.
The three things that decide it before any benchmark does
Most comparisons start with scores. This one should start with availability, because the gap there is not a matter of degree.
• Pricing — Ling-3.1-flash is free for the length of its trial and has no published post-trial rate card at all. Gemini 3.8 Flash lists at $0.75 per million input tokens and $3.75 per million output during its promotional period, reverting to $1.50 / $7.50 on January 1, 2027. Vendor-reported list prices; both are subject to change.
• Context — Ling-3.1-flash serves 262,144 tokens in the trial, with 1,000,000 cited as the eventual target. Gemini 3.8 Flash serves the full 1,048,576 tokens today. If your workload depends on the million-token window, only one of these two has it now.
• Weights — Ling-3.1-flash has none published: no repository, no licence, no checkpoint, just a stated intention to open-source after the trial. Gemini 3.8 Flash is closed and always will be, but it is a managed API you can call without ambiguity today.
Read together, those three make one point: the Ling model's headline advantages are either promotional (free) or prospective (1M context, open weights), while the Gemini model's advantages are contractual. That asymmetry runs through everything that follows.
Where the Ling model genuinely wins
Two places, and both are real.
The first is cost during evaluation. A two-week free trial on a model this size is not a marketing gimmick to wave away — it is the cheapest way to find out whether a 560B mixture-of-experts handles your prompts. Gemini 3.8 Flash, whatever its price, is not free, and a serious evaluation across a few thousand calls costs real money.
The second is the openness trajectory. Ling-3.1-flash is the successor to a model that did ship MIT-licensed weights, and Ant has said publicly it intends to open-source this one too. If that lands with a permissive licence, you get something Gemini 3.8 Flash can never offer: the option to self-host, fine-tune, or distil a 560B base model. The catch is the one that matters most — you are betting on a promise, and the previous generation's promise was kept on a two-week delay, not a guarantee.
Where Gemini 3.8 Flash is not close to losing
Strip the trial and the openness question away and the operational comparison is lopsided in Google's direction.
• Input — Gemini 3.8 Flash takes text, image, video, file, and audio. Ling-3.1-flash takes text and returns text. For document, screenshot, or video workflows this is not a preference, it is a capability boundary.
• Output ceiling — 65,536 tokens versus 32,768. Relevant the moment you generate reports, code files, or long structured output in one pass.
• Throughput — independent testing puts Gemini 3.8 Flash's median output speed near 249 tokens per second, with OrcaRouter's own seven-day telemetry measuring 552 tokens per second at a p50 of 4.95 seconds to first token. Ling-3.1-flash's trial hosting has been reported around 63 tokens per second. The Gemini model is in a different speed class.
• Reference depth — Gemini 3.8 Flash has a body of independent measurement behind it; Ling-3.1-flash's numbers are all first-party, on a brand-new endpoint, with no third party yet re-running them.
• Multimodal agents — anything that has to read a screenshot, a PDF, or a recorded call is a Gemini task by construction, not by quality.

Benchmarks, and why the two sets do not really compare
The vendor-reported case for Ling-3.1-flash is an 81.0 aggregate across five agent benchmarks, with 40.4% on Terminal-Bench 4.0, 55.9% on SWE-Atlas, and 65.35% on HealthBench Professional; Ant's own account adds 1,673 Elo on GDPVal-AA v2.1 and 75.16 on FrontierSWE. Every one of those is Ant's own measurement. There is no harness published and, more to the point, no weights for anyone to re-run the suite on.
Gemini 3.8 Flash's picture is different in kind. On Artificial Analysis's current Intelligence Index build (v4.3.2), it scores 40.9 at high reasoning effort, 39.8 at medium, and 33.5 at low — and it is worth flagging that the index has been revised recently, so figures published about this model earlier in the autumn were computed on a different build and are not directly comparable to these. Google's own claims for the model include 54.9 on HLE-Verified and strong Terminal-Bench 2.1 results in the high 80s. Independent and vendor numbers, side by side, both legitimate, neither interchangeable.
There is one clean cross-comparison worth making, because both sit on the same benchmark family. On Terminal-Bench 4.0, Ling-3.1-flash's vendor figure is 40.4%. Claude Sonnet 5.5 — a mid-tier model, not a flagship — scores 70.6% on that same version. That single row is the strongest argument against reading Ant's card as "frontier-adjacent": the model is competitive on the agentic measures Ant chose to highlight and plainly behind on the terminal-coding measure where an outside number exists.

The practical shape of the choice
If you need to build something in the next two weeks, the answer is not close and it is not a knock on Ling-3.1-flash. Gemini 3.8 Flash is the only one of the two that gives you a stable endpoint, a published price, a full million-token window, multimodal input, and a licence you can read. It is a production Flash-tier model.
Ling-3.1-flash is an evaluation target. Its value right now is that the evaluation is free and the model is interesting: a 560B MoE with only ~25B active per token is a bet on sparsity, and Ant positioned it for exactly the agent, search, and office-automation workloads that Flash-tier models compete for. Running your own prompts through it during the trial window is worthwhile precisely because nobody outside Ant has done it yet — you would be among the first to have a non-vendor opinion.
The decision tree that makes sense this month: route production to Gemini 3.8 Flash, spend the free trial running Ling-3.1-flash against your hardest prompts, and hold the open-weight question as the real fork. If the weights arrive permissive and a price lands near the Flash tier, Ling-3.1-flash becomes a genuine second option — one you can self-host, which Gemini never will be. If they arrive with strings, the trial ends and so does the comparison.
Running both from one key, and being straight about what is missing
Gemini 3.8 Flash is on the OrcaRouter catalogue today under the model ID google/gemini-3.8-flash, at Google's own list price with zero markup — so when the promotional rate reverts in January, the price you pay here tracks it automatically rather than needing a new contract. DeepSeek-V4.1-Flash, the other cheap Flash-tier model worth measuring in the same experiment, sits on the same catalogue at its provider's list price, so a three-way test of trial model against established Flash-tier options runs behind a single API key with automatic failover between them.
Ling-3.1-flash is not on our catalogue, and a lookup against our public model API returns not-found for it and for the previous generation — so the honest framing is that the trial runs where it runs, through the vendor's own surface and third-party inference platforms, and we make no claim to host it. That is the part of this comparison that could change fastest: a model in a free trial with an unannounced price is precisely the kind of thing that lands on a routing layer the moment Ant and its hosts publish a rate card, and zero-markup pass-through means the day it does, the price here is the price there.
So the verdict is narrower than the parameter counts suggest. Ling-3.1-flash is the more ambitious model and the more interesting research result; Gemini 3.8 Flash is the one you can ship on. Until a licence and a price exist, that is the whole of the difference — and it is not one a benchmark row will settle.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
