
Ling-3.1-flash vs Qwen3.8-Flash: Two Ways to Build a Fast Tier, and Only One of Them Ships
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 262 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 114 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 969 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 49 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Two Chinese labs looked at the same problem this autumn — how do you build a cheap, fast model that is good enough to sit inside an agent loop — and arrived at opposite answers. Ling-3.1-flash, announced by Ant Group on September 30, 2026, is a roughly 560-billion-parameter mixture-of-experts activating about 25 billion parameters per token, with a context window quoted at up to 1 million tokens and a stated intention to open-source "soon." Qwen3.8-Flash, released by Alibas Qwen team on August 26, 2026, is a 125-billion-parameter multimodal mixture-of-experts activating about 6 billion parameters per token, with weights already on Hugging Face under the name Qwen3.8-Flash-Next, a production model live on the QwenCloud API at $0.15 per million input tokens and $0.47 per million output, and day-0 support in SGLang. One lab scaled up and announced. The other shrank down and shipped. That difference is more instructive than any benchmark either lab published.
It is also the reason this comparison has to be read carefully. Ling-3.1-flash has no weights, no licence, no model card, no API price, and no independent evaluation; the entire published record is an announcement post, a vendor benchmark chart, and a slot in Ant's own Ling Studio product. Qwen3.8-Flash has all of those things and a four-week head start on being run by people who do not work for Alibaba. Comparing an architecture choice is fair. Comparing capabilities is not, yet.
The two design decisions, and why they diverge
Qwen's bet is that the fast tier should be structurally unlike the flagship, not merely smaller than it. Qwen3.8-Flash activates 6 billion of its 125 billion parameters — roughly 95% sparsity — and gets its capability from where the compute is routed rather than from raw breadth. Alibaba has been explicit that the architecture underneath, which pairs Gated DeltaNet with Qwen Sparse Attention and an N-gram embedding table, is an early preview of the Qwen4 generation, published early on purpose so that inference engines, quantisation tooling and deployment patterns can mature around it before the expensive models are built on top. The result is a model that is cheap per token by construction: about 6 billion parameters of work on every token, at a price Alibaba set at $0.15 in and $0.47 out per million.
Ant's bet, as far as the announcement reveals it, is closer to conventional scaling with a very long context. Ling-3.1-flash keeps a large active fraction — roughly 25 billion parameters per token out of 560 billion, so about one in twenty-two — and quotes a window of up to 1 million tokens, four times what the previous Ling generation serves. The Ling Studio app describes it as built for "general-purpose agents, search, routine office work, and software/code development," which is fast-tier positioning, but the parameter economics are those of a much heavier model. On identical input, a token passing through 25B active parameters costs roughly four times the compute of one passing through 6B, before any pricing decision enters the picture.
Neither approach is wrong, and both are defensible answers to "what limits a cheap model." Qwen is betting that sparsity and a new attention stack are the ceiling. Ant is betting that the ceiling is context and world knowledge, and that it can afford the compute. What separates them for a reader is not the design philosophy but the delivery: a design you can download and measure versus a design you can only read about.

What each one costs you, and what that means in practice
The price gap is not subtle and it is not close. Qwen3.8-Flash is published at $0.15 per million input tokens, $0.47 per million output, and $0.018 per million on cache reads — the same numbers on Alibaba's announcement, on the vendor's production model card, and on the routed model page we serve, which is what a 0%-markup pass-through looks like when you check it against three sources. Ant has published no rate for Ling-3.1-flash at all — no API price, no cache discount, nothing comparable. The only published figure for the Ling line is the previous generation's, which lists at $0.075 in and $0.22 out per million, roughly half of Qwen's input rate and less than half its output rate. If the new Ling tier prices near where the old one sat, it would undercut Qwen3.8-Flash on paper; if it prices anywhere near the compute its 25B active parameters imply, it will not. That is a guess either way, because the number does not exist.
Context is where Ling's claim is genuinely larger. Ant quotes up to 1 million tokens for Ling-3.1-flash; Qwen3.8-Flash ships with a 1M-token default context on the production API and a 262,144-token native window on the open weights, extendable. Two caveats sit on the Ling figure and neither is closed. First, "up to 1M" is a specification, not a measurement — nobody has published long-context retrieval behaviour for a model no one outside Ant can call. Second, the previous Ling generation had exactly the same gap between the advertised window and the architectural story behind it, and the answer only came when the weights, format, and context limits were finally published. The headline 1M is a promise. The 1M on Qwen3.8-Flash is a served default you can hold Ant's claim up against.
Why "preview" is the honest word on one side of this
Alibaba's own framing of Qwen3.8-Flash is that it is an architecture preview shipped under a production name — the open weights carry the "Next" suffix precisely so that nobody confuses the prototype with the tuned serving model. That framing has been validated by events rather than by marketing: SGLang shipped day-0 support for Qwen3.8-Flash-Next on the release day, with the model also configured for Transformers, vLLM and TokenSpeed, and the weights have since been picked up across quantisation communities. Versions quickly became ordinary, which is what a shipped model looks like.
Ant's announcement has the same shape one step earlier: a specification drop ahead of the release, with the explicit statement that the model will be open-sourced soon. This is a legitimate way to launch — Ling-3.0-flash followed exactly that path in July, and the weights landed under MIT on August 7, about two weeks later. But the state of the two projects today is not comparable, and no amount of chart-reading closes the gap. It is worth being specific about what an outsider can bring to each.
What is independently checkable for Ling-3.1-flash today: that the announcement exists and is dated September 30; that the parameter, active-parameter and context figures are Ant's own; that the model appears in Ant's Ling Studio product; and that no weights, licence or model card have been published. What is independently checkable for Qwen3.8-Flash: that the weights are downloadable in BF16 and FP8; that the licence terms are readable; that the API price is published; and that Alibaba's benchmark claims have been open to reproduction since late August. The second list is longer, and that is the entire comparison in miniature.

How the two would actually be evaluated
Ant published a benchmark card with Ling-3.1-flash that places it against GPT-5.6 Sol, Claude Opus 5, Kimi K3, two GLM variants and DeepSeek-V4.1-Flash, with headline figures of 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE and 65.35 on HealthBench Professional. Two problems sit on that card before anyone argues about the numbers. The first is the comparison set: both frontier models Ant chose are a generation behind, since Claude Opus 5.5 and GPT-6 Sol went live on September 22, 2026 and are routable now, which means the card benchmarks an October model against July's frontier rather than the current one. The second is a footnote on the card itself, stating that the HealthBench Professional result came from the "AQ environment" and that the model's healthcare capabilities can currently be experienced only in AQ — an environment no public documentation defines. A headline score whose evaluation environment is unnamed is not reproducible even in principle.
Qwen3.8-Flash's comparable claims, by contrast, were made when the weights were already out, and Alibaba staked its positioning on a claim anyone can test: that the sparsity ratio, not the parameter count, is where the capability comes from. Four weeks of use is not a verdict, but it is long enough that the claims have stopped being uncontested — which is the useful state for a model.
What to actually do with this, today
For builders, the practical split is simple. Ling-3.1-flash is not a component you can adopt this week: there is no endpoint to call, no weights to host, no licence to check, and no price to budget against. Whatever the final model turns out to be, the useful move right now is to set a trigger — weights published, licence permissive, an independent score on FrontierSWE that is anywhere near 75.16 — and revisit. The prior generation's own history shows the weights can arrive two weeks after the announcement, so that trigger may fire quickly.
Qwen3.8-Flash is already a component. It is live on our catalogue as a routed model at provider list price with zero markup — the routed card carries $0.15 in, $0.47 out and $0.018 per million cache reads, the same figures Alibaba publishes, which is what a 0%-markup policy means in practice rather than in a slide. Behind one key it sits beside Claude Opus 5, GPT-5.6 Sol and DeepSeek-V4.1-Flash, so the comparison Ant is inviting — a candidate against the current frontier — can be run by pointing a config at two routes instead of maintaining two integrations, with automatic failover handling the weekend a provider wobbles. That is the difference between an architecture discussion and an experiment.
The verdict, as far as one can be given: Qwen made the bolder architectural bet and had the confidence to publish it early and let people pull it apart. Ant made the heavier bet — more parameters, more active compute per token, more context — and has so far published a chart instead of a model. The chart looks good. The chart is also the only thing anyone has.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
