Hero title card for 'Ling-3.1-flash hits 41 on the AA Intelligence Index', subtitled '560B total / 25B active MoE from Ant Group'. A large rounded card carries the number '41' with the label 'Artificial Analysis Intelligence Index v4.3.2'; below it a second small card reads 'vs 20' with the label 'Ling-3.0-flash'. A footer line reads 'Index figure independently measured by Artificial Analysis; model released 2026-10-01.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Ling-3.1-flash Hits 41 on the Artificial Analysis Intelligence Index: The 560B MoE Doubles Its Predecessor

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ling-3.1-flash, Ant Group's 560-billion-parameter mixture-of-experts, now carries a number its own lab did not write: 41 on the Artificial Analysis Intelligence Index. The index is run by the independent evaluator, on hardware the evaluator controls, against a fixed suite — and it places the model 21 points above its own direct predecessor, Ling-3.0-flash, which still sits at 20 on the same index. Ant announced Ling-3.1-flash at the end of September 2026, Artificial Analysis dates the release to October 1, and it is three months younger than the model it doubles.

That gap is the story. A 21-point move on a composite that ten different evaluations feed is not the sort of thing a vendor chart can fake, and it is not the sort of thing this blog could have written about a fortnight ago, when the only existent measurement of Ling-3.1-flash was Ant's own benchmark card: 1,673 Elo on GDPval-AA v2.1, 75.16 on FrontierSWE, 65.35 on HealthBench Professional. Those numbers were real claims by a real lab, but unreproduced. The 41 is different in kind. It is the first figure in this release that a third party can be held to.

What the 41 actually measures

The Artificial Analysis Intelligence Index v4.3.2 is a weighted composite of ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. Reading a composite in isolation is a mistake, so the per-evaluation rows matter more than the headline:

• GDPval-AA — 1,621.75 for Ling-3.1-flash against 948.35 for Ling-3.0-flash. This is the largest single jump in the set and the one that most of the index move rests on.

• GDP.pdf — 0.100 all-pass, up from 0.054. Still low in absolute terms, but a near-doubling.

• AA-LCR v1.1 long-context reasoning — 0.83 against 0.73 for the predecessor, and level with DeepSeek V4.1 Flash (Max) at 0.84.

• SciCode — 0.541 against 0.420. A coding-science gain, not a chat-quality gain.

• CritPt physics reasoning — 0.180, against 0.017 before. Ten times the predecessor on a small base.

• AA-Omniscience — +2.22, against −17.88 for Ling-3.0-flash. This one is discussed below, because it cuts against the headline.

Ling-3.0-flash did not merely lose to Ling-3.1-flash on these rows; it collapsed on two of them. It scored 0.032 on AutomationBench-AA and exactly 0.000 on Terminal-Bench 4.0. That is the context in which the 41 should be read — the predecessor was an efficient, cheap, open-weight model that had essentially no agentic-terminal capability at all.

A single-column generated scoreboard titled 'Ling-3.1-flash — the scoreboard' with six rows: 'AA Index: 41', 'Predecessor index: 20', 'Terminal-Bench 4.0: 0.333', 'SciCode: 0.541', 'Price: $0.30 / $0.90', 'Weights: not published', and a footer reading 'All figures per Artificial Analysis, independently measured.' The OrcaRouter logo is composited in the bottom-right corner.

Where the gain came from: agentic work, not conversation

Compare Ling-3.1-flash's index contributions against the two Flash-tier models it is most often lined up beside, and a pattern shows up that the aggregate hides. DeepSeek V4.1 Flash (Max) posts 39 on the same index and GLM 5.3 Flash posts 42, so all three land inside a three-point band. But they get there from different places.

• Agentic terminal work — Ling-3.1-flash 0.333 on Terminal-Bench 4.0, DeepSeek V4.1 Flash (Max) 0.268, GLM 5.3 Flash 0.328.

• Agentic real-world work — Ling-3.1-flash 1,621.75 Elo on GDPval-AA, DeepSeek V4.1 Flash (Max) 1,600, GLM 5.3 Flash 1,646.86.

• AutomationBench-AA — Ling-3.1-flash 0.617, DeepSeek V4.1 Flash (Max) 0.689, GLM 5.3 Flash 0.604.

• Coding science — Ling-3.1-flash 0.541 on SciCode, DeepSeek V4.1 Flash (Max) 0.519, GLM 5.3 Flash 0.516.

The narrow reading is that Ling-3.1-flash is competitive on agentic benchmarks and slightly ahead on SciCode. The useful reading is that it is the strongest of the three on the two agentic-terminal measures most people mean when they say "can it drive a tool loop" — and it is behind both of them on the knowledge-reliability measure. That is a specific shape, not a general win, and it is worth naming before anyone buys a workload on the index number alone.

The bill that arrives with a bigger model

Ling-3.0-flash was priced at $0.07 per million input tokens and $0.22 per million output tokens on Ant's own API, and it was open weights under MIT. Ling-3.1-flash is neither of those things yet. Artificial Analysis records its running cost at $0.30 per million input and $0.90 per million output — with a cache-hit rate of $0.06, an 80% discount off input — measured against InclusionAI's own API as the serving provider. That is roughly a four-times increase in input price over the model it replaces, for a 21-point index gain.

Whether that trade is good depends entirely on the workload, and the index itself can price it: running all ten Intelligence Index evaluations against Ling-3.1-flash costs about $960 at those rates, against roughly $477 for DeepSeek V4.1 Flash (Max) and $280 for GLM 5.3 Flash. The reason is not only the rate card. Ling-3.1-flash is a very talkative reasoner — it burned 220 million output tokens working through the index, well above the median for its price tier, and its evaluation runs average about 357 seconds per task against 285 seconds for DeepSeek V4.1 Flash (Max) and 979 seconds for GLM 5.3 Flash. On a per-task basis Ling-3.1-flash costs about $0.99 against DeepSeek's $0.27 and GLM 5.3 Flash's $0.25.

None of this is visible from the 41, and all of it lands on the invoice. A model can win an index and lose a budget.

The two numbers that argue with the headline

The first is knowledge reliability. Ling-3.1-flash scores +2.22 on AA-Omniscience, which measures whether a model knows what it knows: correct answers earn points, hallucinations cost points, and refusals cost nothing. Its accuracy rate is 29.1% and its hallucination rate is 37.9%. Put beside its stablemates, that is the weakest profile in the group — GLM 5.3 Flash scores +7.47 with a 27.6% hallucination rate, and DeepSeek V4.1 Flash (Max) scores −5.30 with a staggering 96.5% hallucination rate among answered questions. The composite rewards Ling-3.1-flash for the capability it demonstrates, not for the reliability with which it demonstrates it, and a model that invents a third of what it asserts needs retrieval around it in production.

The second is context. Artificial Analysis lists Ling-3.1-flash with a 1,000,000-token context window. Ant's own release material also cites 1M as the target. But Ant's developer documentation, which enumerates the family's windows model by model, gives Ling-3.0-flash a native 256K extendable to 1M and does not yet carry an entry for Ling-3.1-flash at all. The distinctive long-context row that did get measured — AA-LCR at 0.83 — is a reasoning-over-long-context score, not a served-window measurement. Treat 1M as the declared ceiling and validate the delivered window with your own long-context request before you design around it.

How it stacks up on the specification sheet

The dimension-by-dimension picture, subject first, on each line:

• Total parameters — Ling-3.1-flash 560B vs DeepSeek V4.1 Flash (Max) 552B vs GLM 5.3 Flash 320B.

• Parameters active per token — Ling-3.1-flash 25B vs DeepSeek V4.1 Flash (Max) 16B vs GLM 5.3 Flash 18B. Ling-3.1-flash activates the most, which is one reason its serving cost sits where it does.

• Weights and licence — Ling-3.1-flash not published (proprietary, no licence text) vs DeepSeek V4.1 Flash (Max) open under MIT vs GLM 5.3 Flash open under MIT.

• Declared context — Ling-3.1-flash 1M vs DeepSeek V4.1 Flash (Max) 1M vs GLM 5.3 Flash 1M. All three claim the same ceiling; only two of them have a downloadable checkpoint you could verify it against.

• Modality — Ling-3.1-flash text in, text out vs DeepSeek V4.1 Flash (Max) text and image in vs GLM 5.3 Flash text, image and video in. This is the one axis where Ling-3.1-flash is simply behind, and no benchmark row compensates for it.

• Price on the vendor's API — Ling-3.1-flash $0.30 / $0.90 per million input / output vs DeepSeek V4.1 Flash (Max) $0.30 / $1.20 vs GLM 5.3 Flash $0.15 / $0.50.

• Index score — 41 vs 39 vs 42. Three points of spread across a three-way comparison is not a ranking; it is a tie broken by whichever evaluation the reader's workload happens to resemble.

Where you can call it

Ling-3.1-flash is not on the OrcaRouter catalogue today — a lookup against our public model API for the model under InclusionAI returns not-found, and we will not imply otherwise. It currently runs on Ant's own API and on third-party platforms carrying the release, which is the honest extent of its availability. What is worth noting for anyone comparing the three models above is that the other two are routable right now: DeepSeek V4.1 Flash and GLM 5.3 Flash both sit on OrcaRouter's catalogue at their providers' list prices with zero markup, behind a single API key, with automatic failover between them. If the comparison above says your workload cares more about knowledge reliability than about agentic terminal work, GLM 5.3 Flash is the one to try — and you can try it against DeepSeek V4.1 Flash on the same key without a second contract or a line of new client code.

What the 41 changes, and what it does not

What it changes is the release's standing. Ling-3.1-flash is no longer a specification with a vendor chart attached; it is a model with an independently measured position, and that position is a genuine generational leap in agentic capability over Ling-3.0-flash — the kind of jump that comes from a design change rather than more compute on the same recipe. What it does not change is anything about procurement. The weights are still closed, the licence is still unwritten, the modality is still text-only, and the price is roughly four times its predecessor's for a gain that is concentrated in agent and terminal tasks.

If your work is agent loops, tool driving and long tool-chains, the 41 is the most meaningful number released about this model so far and it earns a serious look while availability is still expanding. If your work is multimodal, or knowledge-critical question answering, or budget-sensitive high-volume inference, the 41 does not reach your problem — and GLM 5.3 Flash, already routable and already open under MIT at half the output price, is the better-positioned model of the three. The index tells you where Ling-3.1-flash stands. It does not tell you whether you should care, and that question is answered by what you are building.

Screenshot of the Artificial Analysis model page for Ling 3.1 Flash, in English, captured 2026-10-07, headed 'Ling 3.1 Flash Intelligence, Performance & Price Analysis' and marked 'Proprietary model' and 'Released October 2026'. The page shows an Intelligence Index of 41, 211 output tokens per second, a cost of $0.99 per Intelligence Index task, 220M output tokens generated across the index, input $0.30 and output $0.90 per 1M tokens with an 80% cache discount, a 1M context window, text input and text output, and the summary line that the model 'scores 41 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 13)'.Screenshot of the Artificial Analysis model page for Ling 3.0 Flash, in English, captured 2026-10-07, headed 'Ling 3.0 Flash Intelligence, Performance & Price Analysis' and marked 'Open weights model' and 'Released August 2026'. It shows an Intelligence Index of 20, 332 output tokens per second, input $0.075 with an 80% cache discount and output $0.22 per 1M tokens, a 262k context window, and 124B total parameters with 5.1B active parameters.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily