Generated editorial hero card titled "Ling-3.1-flash: announced, not open" with the subtitle "~560B total params · ~25B active per token · up to 1M context". Three labelled cards read "Weights: not shipped yet", "Benchmarks: vendor-reported only" and "No licence · no model card · no download".
Guides & Insights

Ling-3.1-flash Is Announced, Not Open: ~560B Parameters, a 1M-Token Context, and a Chart Only Ant Has Run

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ant Group's Ling team put numbers on its next fast-tier model on September 30, 2026, and in the same breath said you cannot have it yet. Ling-3.1-flash is a roughly 560-billion-parameter mixture-of-experts that activates about 25 billion parameters per token and carries a context window of up to 1 million tokens, according to the model announcement Ant published from its own @AntLingAGI account. The sentence that defines this release is the one that follows the specs: "We plan to open-source the model soon." As of today there is no Ling-3.1-flash repository on Hugging Face, no licence, no downloadable weight file, and no evaluation from anyone outside Ant Group. What exists is a benchmark chart, a live product surface, and a promise.

That distinction is the whole story, and it is worth being blunt about it, because the framing that reaches most people first — a 500B-class open-weight release from China that lands near the frontier — describes something that has not happened. Ling-3.1-flash is not an open release. It is an announced one. The two are not close, and the gap between them is exactly where a reader can get burned: if you plan capacity, budget a pilot, or write a procurement note around open weights that do not exist yet, you are planning against a press release.

What the announcement actually contains

The post is short and the numbers in it are specific. Ant states a total parameter count of about 560 billion against roughly 25 billion active per token, a ratio near 1-in-22, marginally denser than the Ling-3.0-flash generation it succeeds — that model runs 124B total against 5.1B active, about 1-in-24, on a body less than a quarter the size. Context is quoted as "up to 1M tokens," four times the 262,144-token window the Ling-3.0-flash family serves today. Three headline scores are attached: 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional.

Every one of those figures is vendor-reported, first-party, and unpublished in any form a third party can re-run. There is no evaluation harness, no prompt set, no run configuration, and no seed — and, more fundamentally, no weights to run them on. If Ant's numbers are right, the model sits in the top tier of Chinese open-model releases; there is presently no way to find out whether they are right.

The chart, and the caveat baked into it

Ant Group's own Ling-3.1-flash benchmark card, published from the @AntLingAGI account on 30 September 2026. Multi-panel bar charts compare Ling-3.1-flash against GPT-5.6 Sol, Claude Opus 5, Kimi K3, GLM 5.3 and GLM 5.3 Flash, and DeepSeek-V4.1-Flash across GDPval-AA v2.1 (1,673 Elo), tau-Banking, SkillsBench, Automationbench, Terminal-Bench 4.0, CyberGym, FrontierSWE (75.16), SWE Atlas Codebase QnA, Finance Agent v2, HealthBench Professional (65.35), DRACO and MultiChallenge. A footnote states that HealthBench Professional was evaluated in the AQ environment.

Ant published Ling-3.1-flash alongside a multi-panel benchmark card that places it in a field of frontier and near-frontier models: GPT-5.6 Sol, Claude Opus 5, Kimi K3, two entries from the GLM family, and DeepSeek-V4.1-Flash. Across the panels the Ling bar sits at or near the top on the agentic and coding measures and mid-pack on the multi-turn and domain-specific ones. Read it as what it is — the vendor's own selection of tasks, drawn from a suite that includes General-purpose agent work, coding, banking-domain tasks, and healthcare — and it is a credible-looking card.

Look at who is on it, though, and one thing stands out. The frontier peers Ant chose are GPT-5.6 Sol and Claude Opus 5. Both of those models are now a generation behind. Opus 5.5 and GPT-6 Sol both went live on September 22, 2026, on the same day, and both are routable today. The newest models Ant did not put on the card are precisely the ones a reader would want it measured against, which is the same complaint that surfaced in the first wave of commentary on the release — the wish that a model arriving in October had been benchmarked against October's frontier rather than July's. That is not a knock on the model, which may well hold up. It is a reason to treat the card as a directional claim rather than a verdict, and to note that the comparison Ant chose flatters the result less than it dates it.

Where you can touch it, and one unexplained footnote

There is exactly one place Ling-3.1-flash runs for a user today: Ant's own product. The Ling Studio interface at ling.tbox.cn lists Ling-3.1-flash in its model selector, and the app's own description of the model is broader than the benchmark card — it pitches it at "general-purpose agents, search, routine office work, and software/code development," which is the familiar positioning for this tier: not the deepest reasoner in the family, but the one meant to be called constantly and cheaply.

That surface also carries the release's most awkward detail. The benchmark card's own footnote says HealthBench Professional — the 65.35 figure, one of the three numbers in the announcement text — was "evaluated in the AQ environment," and adds that Ling-3.1-flash's healthcare capabilities "can currently be experienced only in AQ." Nothing on the card explains what AQ is, and no public documentation we could find defines it. A headline score with an environment qualifier attached, where the environment is unnamed, is not a reproducible result; it is a product demo with a number beside it.

What is knowable, and what is not

Sorting this release into those two buckets is the most useful thing a reader can do with it today.

Knowable now: the parameters, the active count, and the context window as stated by the vendor; the three vendor benchmarks; the existence of the model in Ant's own product; the comparison set Ant selected; the fact that no weights, licence, or model card has been published; and the fact that Artificial Analysis, which independently scored the previous generation, has no Ling-3.1-flash page. As of this writing, the independent scoring pipeline that made Ling-3.0-flash's 20-point Intelligence Index figure legible has not published anything on 3.1 at all.

Not knowable now: whether the weights will actually open, and under what licence; whether the architecture is a scaled-up version of the Ling-3.0 hybrid stack or something new; what the model costs through Ant's API, if it has an API price at all; how it behaves under long context in practice, as opposed to how the window is specified; and whether the benchmark gaps hold up when someone who is not Ant runs the same tasks. Each of those is a question a released model answers on day one and an announced model answers on some later day that has not been scheduled.

Generated two-column information card. Left column headed "Knowable today" lists five cards: "~560B total / ~25B active (vendor)", "up to 1M-token context (vendor)", "1,673 Elo GDPval-AA v2.1 (vendor)", "available in Ant's Ling Studio only" and "no independent score exists". Right column headed "Not knowable" lists five cards: "whether the weights will open", "the licence terms", "API price per million tokens", "long-context behaviour in practice" and "whether the benchmark gaps hold".

How to read a day like this

The pattern is now common enough to name. A Chinese lab with a strong track record drops a model card and a benchmark chart, the numbers land near the frontier, the community reacts to the numbers, and the weights arrive — or do not — weeks later. Ling-3.0-flash ran exactly that play this summer, and it is worth remembering how it resolved: Ant announced it on July 24, 2026 as an API-only model with no licence statement and no independent benchmark, and the weights under MIT did not appear on Hugging Face until August 7, two weeks after the announcement. Ling-3.1-flash is at the equivalent point in that arc on day one, with the additional difference that this time the promise to open-source is explicit in the announcement rather than inferred.

None of the models Ant chose as comparison points are on our catalogue as subject-of-the-piece — Ling-3.1-flash, Ling-3.0-flash, and Ling-3.0-flash-VL all return not-found on the OrcaRouter catalogue today, and we will not pretend otherwise. But three of the peers on that chart are routable right now: Claude Opus 5, GPT-5.6 Sol, and DeepSeek-V4.1-Flash all sit behind a single key, at provider list price with zero markup, with automatic failover between them. That matters more than it sounds in a week like this one, because the honest way to evaluate an announced model is to run the comparison Ant is inviting — against the models you can actually call — while the subject of the comparison is still a chart.

When the weights land, the useful question will not be whether Ling-3.1-flash beat Opus 5 on one Elo rating. It will be whether a ~560B MoE at ~25B active can hold a 1M-token context at a price that makes the sparsity worth it, and whether the licence is permissive enough to self-host. Those are the two questions the announcement does not answer, and they are the only ones that will decide whether this becomes a model people build on or a slide in someone's next deck. For now: the name is real, the numbers are Ant's alone, and the weights are still a sentence.

Generated horizontal timeline with four milestone nodes. "Jul 24, 2026 — Ling-3.0-flash announced, API-only, no weights, no licence"; "Aug 7, 2026 — MIT weights land on Hugging Face, 14 days later"; "Sep 30, 2026 — Ling-3.1-flash announced: ~560B, ~25B active, 1M context"; and a dashed open node reading "Weights: not yet — 'we plan to open-source soon'", captioned "the same point in the same arc, on day one".