Generated editorial hero card titled "Ling-3.1-flash vs Qwen3.8-Flash" with the subtitle "Two answers to the same question: how do you build a fast tier?" Left panel headed "Scale up and announce" carries the caption "~560B total · ~25B active · 1M context" and a tag reading "weights: coming soon". Right panel headed "Shrink down and ship" carries "125B total · 6B active · previews Qwen4" and a tag reading "weights: on Hugging Face".
Engineering & Research

Ling-3.1-flash vs Qwen3.8-Flash: Two Ways to Build a Fast Tier, and Only One of Them Ships

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Two Chinese labs looked at the same problem this autumn — how do you build a cheap, fast model that is good enough to sit inside an agent loop — and arrived at opposite answers. Ling-3.1-flash, announced by Ant Group on September 30, 2026, is a roughly 560-billion-parameter mixture-of-experts activating about 25 billion parameters per token, with a context window quoted at up to 1 million tokens and a stated intention to open-source "soon." Qwen3.8-Flash, released by Aliba​s Qwe​n team on August 26, 2026, is a 125-billion-parameter multimodal mixture-of-experts activating about 6 billion parameters per token, with weights already on Hugging Face under the name Qwen3.8-Flash-Next, a production model live on the QwenCloud API at $0.15 per million input tokens and $0.47 per million output, and day-0 support in SGLang. One lab scaled up and announced. The other shrank down and shipped. That difference is more instructive than any benchmark either lab published.

It is also the reason this comparison has to be read carefully. Ling-3.1-flash has no weights, no licence, no model card, no API price, and no independent evaluation; the entire published record is an announcement post, a vendor benchmark chart, and a slot in Ant's own Ling Studio product. Qwen3.8-Flash has all of those things and a four-week head start on being run by people who do not work for Alibaba. Comparing an architecture choice is fair. Comparing capabilities is not, yet.

The two design decisions, and why they diverge

Qwen's bet is that the fast tier should be structurally unlike the flagship, not merely smaller than it. Qwen3.8-Flash activates 6 billion of its 125 billion parameters — roughly 95% sparsity — and gets its capability from where the compute is routed rather than from raw breadth. Alibaba has been explicit that the architecture underneath, which pairs Gated DeltaNet with Qwen Sparse Attention and an N-gram embedding table, is an early preview of the Qwen4 generation, published early on purpose so that inference engines, quantisation tooling and deployment patterns can mature around it before the expensive models are built on top. The result is a model that is cheap per token by construction: about 6 billion parameters of work on every token, at a price Alibaba set at $0.15 in and $0.47 out per million.

Ant's bet, as far as the announcement reveals it, is closer to conventional scaling with a very long context. Ling-3.1-flash keeps a large active fraction — roughly 25 billion parameters per token out of 560 billion, so about one in twenty-two — and quotes a window of up to 1 million tokens, four times what the previous Ling generation serves. The Ling Studio app describes it as built for "general-purpose agents, search, routine office work, and software/code development," which is fast-tier positioning, but the parameter economics are those of a much heavier model. On identical input, a token passing through 25B active parameters costs roughly four times the compute of one passing through 6B, before any pricing decision enters the picture.

Neither approach is wrong, and both are defensible answers to "what limits a cheap model." Qwen is betting that sparsity and a new attention stack are the ceiling. Ant is betting that the ceiling is context and world knowledge, and that it can afford the compute. What separates them for a reader is not the design philosophy but the delivery: a design you can download and measure versus a design you can only read about.

Ant Group's own Ling-3.1-flash benchmark card, published 30 September 2026. Panels compare Ling-3.1-flash with GPT-5.6 Sol, Claude Opus 5, Kimi K3, GLM 5.3, GLM 5.3 Flash and DeepSeek-V4.1-Flash across GDPval-AA v2.1 (1,673 Elo for Ling-3.1-flash), tau-Banking, SkillsBench, Automationbench, Terminal-Bench 4.0, CyberGym, FrontierSWE (75.16), SWE Atlas Codebase QnA, Finance Agent v2, HealthBench Professional (65.35), DRACO and MultiChallenge. A footnote states HealthBench Professional was evaluated in the AQ environment.

What each one costs you, and what that means in practice

The price gap is not subtle and it is not close. Qwen3.8-Flash is published at $0.15 per million input tokens, $0.47 per million output, and $0.018 per million on cache reads — the same numbers on Alibaba's announcement, on the vendor's production model card, and on the routed model page we serve, which is what a 0%-markup pass-through looks like when you check it against three sources. Ant has published no rate for Ling-3.1-flash at all — no API price, no cache discount, nothing comparable. The only published figure for the Ling line is the previous generation's, which lists at $0.075 in and $0.22 out per million, roughly half of Qwen's input rate and less than half its output rate. If the new Ling tier prices near where the old one sat, it would undercut Qwen3.8-Flash on paper; if it prices anywhere near the compute its 25B active parameters imply, it will not. That is a guess either way, because the number does not exist.

Context is where Ling's claim is genuinely larger. Ant quotes up to 1 million tokens for Ling-3.1-flash; Qwen3.8-Flash ships with a 1M-token default context on the production API and a 262,144-token native window on the open weights, extendable. Two caveats sit on the Ling figure and neither is closed. First, "up to 1M" is a specification, not a measurement — nobody has published long-context retrieval behaviour for a model no one outside Ant can call. Second, the previous Ling generation had exactly the same gap between the advertised window and the architectural story behind it, and the answer only came when the weights, format, and context limits were finally published. The headline 1M is a promise. The 1M on Qwen3.8-Flash is a served default you can hold Ant's claim up against.

Why "preview" is the honest word on one side of this

Alibaba's own framing of Qwen3.8-Flash is that it is an architecture preview shipped under a production name — the open weights carry the "Next" suffix precisely so that nobody confuses the prototype with the tuned serving model. That framing has been validated by events rather than by marketing: SGLang shipped day-0 support for Qwen3.8-Flash-Next on the release day, with the model also configured for Transformers, vLLM and TokenSpeed, and the weights have since been picked up across quantisation communities. Versions quickly became ordinary, which is what a shipped model looks like.

Ant's announcement has the same shape one step earlier: a specification drop ahead of the release, with the explicit statement that the model will be open-sourced soon. This is a legitimate way to launch — Ling-3.0-flash followed exactly that path in July, and the weights landed under MIT on August 7, about two weeks later. But the state of the two projects today is not comparable, and no amount of chart-reading closes the gap. It is worth being specific about what an outsider can bring to each.

What is independently checkable for Ling-3.1-flash today: that the announcement exists and is dated September 30; that the parameter, active-parameter and context figures are Ant's own; that the model appears in Ant's Ling Studio product; and that no weights, licence or model card have been published. What is independently checkable for Qwen3.8-Flash: that the weights are downloadable in BF16 and FP8; that the licence terms are readable; that the API price is published; and that Alibaba's benchmark claims have been open to reproduction since late August. The second list is longer, and that is the entire comparison in miniature.

Headless capture of the OrcaRouter model page for Qwen3.8 Flash, showing the routed model ID qwen/qwen3.8-flash, a 1M-token context window, 131K maximum output, text, image and video input with text output, capability badges for Vision, Tools, JSON and Reasoning, a production date of 2026-08-26, a pricing block reading $0.15 per 1M input tokens, $0.47 per 1M output tokens, $0.018 cache read and $0.230 cache write, and a performance panel with a 4.47 s p50 time-to-first-token, 10.00 s p95, 105 tok/s output and a 2.9% error rate over the last seven days.

How the two would actually be evaluated

Ant published a benchmark card with Ling-3.1-flash that places it against GPT-5.6 Sol, Claude Opus 5, Kimi K3, two GLM variants and DeepSeek-V4.1-Flash, with headline figures of 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE and 65.35 on HealthBench Professional. Two problems sit on that card before anyone argues about the numbers. The first is the comparison set: both frontier models Ant chose are a generation behind, since Claude Opus 5.5 and GPT-6 Sol went live on September 22, 2026 and are routable now, which means the card benchmarks an October model against July's frontier rather than the current one. The second is a footnote on the card itself, stating that the HealthBench Professional result came from the "AQ environment" and that the model's healthcare capabilities can currently be experienced only in AQ — an environment no public documentation defines. A headline score whose evaluation environment is unnamed is not reproducible even in principle.

Qwen3.8-Flash's comparable claims, by contrast, were made when the weights were already out, and Alibaba staked its positioning on a claim anyone can test: that the sparsity ratio, not the parameter count, is where the capability comes from. Four weeks of use is not a verdict, but it is long enough that the claims have stopped being uncontested — which is the useful state for a model.

What to actually do with this, today

For builders, the practical split is simple. Ling-3.1-flash is not a component you can adopt this week: there is no endpoint to call, no weights to host, no licence to check, and no price to budget against. Whatever the final model turns out to be, the useful move right now is to set a trigger — weights published, licence permissive, an independent score on FrontierSWE that is anywhere near 75.16 — and revisit. The prior generation's own history shows the weights can arrive two weeks after the announcement, so that trigger may fire quickly.

Qwen3.8-Flash is already a component. It is live on our catalogue as a routed model at provider list price with zero markup — the routed card carries $0.15 in, $0.47 out and $0.018 per million cache reads, the same figures Alibaba publishes, which is what a 0%-markup policy means in practice rather than in a slide. Behind one key it sits beside Claude Opus 5, GPT-5.6 Sol and DeepSeek-V4.1-Flash, so the comparison Ant is inviting — a candidate against the current frontier — can be run by pointing a config at two routes instead of maintaining two integrations, with automatic failover handling the weekend a provider wobbles. That is the difference between an architecture discussion and an experiment.

The verdict, as far as one can be given: Qwen made the bolder architectural bet and had the confidence to publish it early and let people pull it apart. Ant made the heavier bet — more parameters, more active compute per token, more context — and has so far published a chart instead of a model. The chart looks good. The chart is also the only thing anyone has.

Generated three-lane information card. Lane one, "Shipped, priced, licensed", holds "Qwen3.8-Flash — weights on Hugging Face as Qwen3.8-Flash-Next" and "$0.15 in / $0.47 out per 1M tokens, cache $0.018". Lane two, "Announced, unpriced, unlicensed", holds "Ling-3.1-flash — no repository, no licence", "no published API price" and "no independent evaluation". Lane three, "What a reader should do", holds "test the shipped model this week" and "set a trigger for the announced one".

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily