
Ling-3.1-flash vs Ling-3.0-flash: What a 1M-Token Window and 4.5x the Parameters Actually Change
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 262 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 114 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 969 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 49 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Ant Group's Ling family has a new fast tier, and this time it is only a sentence old. Ling-3.1-flash was announced on September 30, 2026 — roughly 560 billion total parameters, about 25 billion active per token, a context window quoted at up to 1 million tokens, and a stock phrase attached: "We plan to open-source the model soon." The model it replaces at the top of the fast line, Ling-3.0-flash, is a different animal in every sense: 124 billion total parameters, 5.1 billion active per token, a 262,144-token served context, MIT-licensed weights on Hugging Face since August 7, and an independent Artificial Analysis Intelligence Index of 20. One of these models you can download today. The other one you can read about. Comparing them is genuinely useful, but only if the comparison starts there.
The headline arithmetic is easy and slightly misleading. Ling-3.1-flash is about 4.5× the total parameter count of Ling-3.0-flash and about 4.9× the active parameter count — 25B against 5.1B per token. A casual reading says "much bigger model, therefore much smarter model." The more careful reading is that Ant has roughly preserved the sparsity ratio while scaling the body, and what that buys you is capacity, not necessarily per-token reasoning depth: a model that routes 25B of its 560B through each token is doing more work per token than one routing 5.1B, but it is also paying roughly five times the compute per token to do it. Where that trade lands depends entirely on whether Ant's architecture handles the extra parameters efficiently — and that is a question the announcement does not answer.
The two models, side by side
Set the published facts next to each other and the generations separate cleanly. Each line below states what Ant or an independent evaluator has actually published; where a number is missing, it is missing because nobody has published it.
• Total parameters — Ling-3.1-flash: ~560B (vendor). Ling-3.0-flash: 124B (model card).
• Active parameters per token — Ling-3.1-flash: ~25B (vendor). Ling-3.0-flash: 5.1B (model card).
• Context window — Ling-3.1-flash: up to 1M tokens (vendor). Ling-3.0-flash: 256K native, served at 262,144 (model card and inference recipes).
• Weights and licence — Ling-3.1-flash: not published; no repository, no licence (checked September 30 and again on October 1, 2026). Ling-3.0-flash: MIT, on Hugging Face as inclusionAI/Ling-3.0-flash and on ModelScope since August 7, 2026.
• Weight formats — Ling-3.1-flash: none available. Ling-3.0-flash: BF16 (~255GB) and FP8 (~128GB) from the lab, plus community quants.
• Benchmark provenance — Ling-3.1-flash: entirely vendor-reported, no public harness, no weights to re-run. Ling-3.0-flash: independently scored by Artificial Analysis at an Intelligence Index of 20, with measured output speed and token-generation volume published alongside.
• Availability — Ling-3.1-flash: Ant's own Ling Studio product only. Ling-3.0-flash: self-hostable, plus callable through Ant's developer API.

What the extra capacity is probably for
Ant's own one-line description of Ling-3.1-flash is the most informative thing in the release. The Ling Studio app pitches it at "general-purpose agents, search, routine office work, and software/code development" — an office-work and agentic framing, not a reasoning-benchmark framing. That is consistent with how the family is organised: Ling is the general-purpose text-and-tool line, Ring is the reasoning line, and Ming handles multimodal. Ling-3.0-flash occupied the same slot one generation earlier, and it was explicitly the fast, cheap tier rather than the flagship.
Read the scale-up in that light and it makes sense. Agentic and tool-calling workloads are long, repetitive, and context-hungry: a single agent loop can burn tens of thousands of tokens of accumulated tool output before it takes an action, so a 1M-token window is a functional requirement rather than a spec-sheet flourish. Scaling the total parameter count while keeping a large active fraction is how you add world knowledge and instruction-following depth to a model that still has to be fast enough to sit inside a loop. Ant is not building a slower, smarter flagship here; it is building a bigger fast tier.
The counter-argument is cost, and it is the same arithmetic arriving from the other direction. The extra capacity is not free at inference time — it is paid for on every token, and Ant has published no price for Ling-3.1-flash at all: no API rate card, no cache discount, nothing comparable to the $0.075/$0.22 per million tokens the previous generation lists. That is the missing number that would decide this comparison, and it is missing.
Benchmarks, and who ran them
Ling-3.1-flash arrives with three scores that look strong in isolation: 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional. All three are Ant's own, taken from the announcement, and none has been reproduced by an outside lab. The practical consequence is that they can be compared to other vendor numbers but not to independent ones — and Artificial Analysis's 20-point Intelligence Index for Ling-3.0-flash is an independent number on a different scale, so putting the two side by side would be comparing a measurement to a claim. There is no Ling-3.1-flash page on Artificial Analysis at the time of writing.
One footnote on Ant's own chart is worth flagging on its own terms. The card states that the HealthBench Professional result was evaluated in the "AQ environment" and that Ling-3.1-flash's healthcare capabilities can currently be experienced only in AQ. No public documentation explains what AQ is. A headline score carrying an unnamed environment qualifier is not something an outside team can reproduce, even once the weights exist.
Which one to actually use this week
The generational question resolves differently depending on what you need today, and for almost everyone the answer right now is not the newer model.
If you need to run something this week, Ling-3.0-flash is the only one of the two that exists in a usable form: MIT weights you can self-host, FP8 if you want the smaller footprint, first-party API pricing published, and an independent score that tells you what you are getting. It is a genuinely good model in its tier — the independent evaluation put it on the open-weights Pareto frontier for intelligence against total parameters, and rated its speed highly, with the caveat that it is unusually verbose and generated about 260M tokens in the course of being evaluated, against a median near 100M.
If you are planning for the next six months, Ling-3.1-flash is the direction of travel and not yet a component. The specific things to watch before it becomes one: whether the weights actually open and under what licence, whether the served window is genuinely 1M or an extension of a shorter native context, what it costs per million tokens, and whether an independent lab reproduces the 75.16 on FrontierSWE. Any of those landing would be a reason to re-run the comparison. None has landed.

Where this fits in a routing stack
Neither Ling model is on our catalogue today — Ling-3.0-flash, Ling-3.0-flash-VL and Ling-3.1-flash all return not-found against the OrcaRouter model API, so the honest answer to "can I call it through you" is no, for now. What a routing layer is actually good for in a week like this one is the other half of the problem: the models you would benchmark a candidate against. Claude Opus 5, GPT-5.6 Sol and DeepSeek-V4.1-Flash are all live routes at provider list price with zero markup, and a router that holds them behind one key with automatic failover is how you keep an evaluation harness pointed at a moving target without maintaining four integrations. When Ling-3.1-flash's weights land and the licence is readable, that is also the shape of the decision: add one endpoint, not rebuild the stack.
The generational summary is short. Ling-3.0-flash is a shipped, independently scored, permissively licensed model that self-hosts today. Ling-3.1-flash is a larger, longer-context, more expensive-to-run promise that Ant has not yet delivered in any form a developer can use. Treating the second as an upgrade to the first is the mistake this release invites, and the numbers do not yet support it.

