Generated editorial hero card titled "Ling-3.1-flash Is Live and Free for Two Weeks" with the subtitle "What the 560B MoE actually serves today". Four rounded cards read "Parameters: 560B total / ~25B active", "Context: 262,144 tokens", "Output cap: 32,768 tokens" and "Weights: not published", above a footer line reading "Vendor-stated specs; no published post-trial price."
Guides & Insights

Ling-3.1-flash Is Live and Free for Two Weeks: What the 560B MoE Actually Serves

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ling-3.1-flash, the roughly 560-billion-parameter mixture-of-experts that Ant Group's inclusionAI lab announced on September 29 to 30, 2026, stopped being a press release this week: it now answers requests on a live commercial API. The model is running as a two-week free trial on a third-party inference gateway, and it also appears in Ant's own Ling Studio product — which moves the question from "does it exist" to "what does it actually serve." The answer is narrower than the headline numbers suggest. Ling-3.1-flash is text-in, text-out; it serves a 262,144-token (256K) context even though the announcement cites 1M as the eventual target; its output is capped at 32,768 tokens; and nobody has published the price you will pay on day fifteen, when the free window closes and the weights are supposed to follow.

That gap between the announcement card and the live endpoint is the useful thing to hold onto, because most coverage of this release reads the specifications as if the model were shipping them today. It is not. Ling-3.1-flash is the second model from this lab in three months to arrive as what the independent tracker LLM Releases bluntly files under "weights not released," and the difference between what Ant says the model is and what a developer can call right now is where a budget or a pilot can quietly go wrong.

What changed between the announcement and today

The announcement itself was a specification post: ~560B total parameters, ~25B active per token, up to 1M tokens of context, three headline evaluation numbers, and a promise — "we plan to open-source the model soon." None of that has changed. What has changed is availability. Within a day of the announcement the model was reachable on a commercial edge, and the free trial exposes the same real endpoint a paid one would: same API surface, same text-only modality, same thinking-mode switch for long-horizon agent work.

For anyone deciding whether to spend engineering time on it, the trial is the whole story this week, and it cuts both ways. Free inference for two weeks is a genuinely cheap way to see how the model behaves on your own prompts — a rarity for a model this size. But free is a promotional state, not a price. There is no published post-trial rate card for Ling-3.1-flash anywhere yet, so any cost model you build on the free tier is a placeholder until Ant and its hosts put numbers on it.

The spec sheet, and the three figures that are still promises

• Parameters — 560B total, ~25B active per token. Vendor-stated, and roughly quadruple the previous generation's 124B total / 5.1B active. The sparsity ratio holds near 1-in-22, so the activation is not dramatically leaner — the model is just much larger than the one it succeeds.

• Context — served at 262,144 tokens today. "Up to 1M" is the target once the window is enabled; it is not what the trial endpoint gives you, and it is the single most common misreading of this release.

• Output — 32,768 tokens maximum. Reasonable for agent loops, tight if you intend to generate long documents in one pass.

• Modality — text in, text out. This is a text model. Vision, audio, and file input are not part of the launch, and no date has been given for them.

• Reasoning — a hybrid stack with an explicit thinking-mode switch, the same pattern the previous generation used, aimed at multi-step agent and tool-calling work rather than single-turn chat.

• Weights and licence — not published. This is the figure that matters most and the one with no value at all yet: no Hugging Face repository, no licence text, and no downloadable checkpoint exist as of today.

Ant Group's own Ling-3.1-flash benchmark chart, published 30 September 2026. Twelve bar-chart panels cover GDPval-AA v2.1 (1,673 for Ling-3.1-flash), tau-squared Banking (47.60), SkillsBench (69.70), Automationbench public (52.50), Terminal-Bench 4.0 (40.40), CyberGym (87.90), FrontierSWE (75.16), SWE Atlas Codebase QnA (55.92), Finance Agent v2 (57.87), HealthBench Professional (65.35), DRACO (85.49) and MultiChallenge (69.78), with GPT-5.6 Sol, Claude Opus 5, Kimi K3, GLM 5.3, GLM 5.3 flash and DeepSeek-V4.1-Flash as the comparison set. A footnote states HealthBench Professional was evaluated in the AQ environment and that Ling-3.1-flash's healthcare capabilities can currently be experienced only in AQ.

The benchmark card, read with the discount it needs

Ant's own reporting puts Ling-3.1-flash at an 81.0 aggregate across five agent benchmarks, with 40.4% on Terminal-Bench 4.0, 55.9% on SWE-Atlas, and 65.35% on HealthBench Professional; the company's own account adds 1,673 Elo on GDPVal-AA v2.1 and 75.16 on FrontierSWE. Every one of those numbers is first-party. There is no harness, no prompt set, and no weights for anyone else to re-run the suite, which is the standard caveat for a model at this stage — and here it is worth stating plainly: an unreproduced vendor score is a claim about a model, not a measurement of one.

The card is also informative in a way its authors may not have intended. The 40.4% Terminal-Bench 4.0 result sits well behind the frontier tier; Claude Sonnet 5.5, a mid-weight model rather than a flagship, scores 70.6% on the same version of that benchmark. The honest reading is not that Ling-3.1-flash is weak — 40.4% is a respectable agentic-terminal number for a model positioned on cost — but that the "close to the frontier on key benchmarks" framing that circulated on social media on announcement day does not survive contact with the specific rows. The benchmark where the model looks strongest, HealthBench Professional at 65.35, is also the one carrying the awkward footnote: Ant states it was evaluated in an environment it calls AQ and that the model's healthcare capability "can currently be experienced only in AQ," without defining what AQ is anywhere public. A headline score with an unnamed environment attached is a demo, not a reproducible result.

Why a large model arriving as a trial is the actual news

The more interesting move here is the sequence, not the parameters. Ling-3.0-flash went straight to open weights. This one is landing trial-first, with weights, licence, and official pricing all held back and a public commitment to open-source only after the free period ends. That is a commercial-release playbook wearing an open-model announcement, and it is a real change worth tracking: if the weights arrive under a permissive licence, the community gets a 560B base model to fine-tune and distil; if they arrive with commercial strings attached, the two-week trial was the product and the "open source" line was marketing. Either way, the choice lands inside the trial window, and it is the thing to watch rather than any single benchmark row.

Where it runs, and the honest routing picture

For now, Ling-3.1-flash is reachable through the vendor's own product surface and through third-party inference platforms running the trial — and that is the extent of it. It is not on OrcaRouter's catalogue today; a lookup against our public model API for Ling-3.1-flash, Ling-3.0-flash, and the Ling-3.0 vision variant all return not-found, and we will not pretend otherwise. When an Ant endpoint does reach general availability with a published list price, the pass-through model OrcaRouter runs on — provider list price forwarded with zero markup, so a vendor price cut is live on our side the same day — is exactly the arrangement that suits a model whose pricing is still unwritten.

What you can do today, without waiting on the licence decision, is run the comparison that actually matters for an unproven model: measure it against the Flash-tier models you can call right now. Gemini 3.8 Flash and DeepSeek-V4.1-Flash are both on our catalogue at their providers' list prices, behind one API key, with automatic failover between them. For a model in a free trial whose weights and rate card are both unknown, that is the sane way to hold it — one more candidate you can route to while the trial lasts, and a production path that does not depend on what happens on day fifteen.

Headless capture of Ant's Ling Studio interface with Ling-3.1-flash selected in the model picker. The centre panel shows the model-experience card headed "Ling-3.1-flash Model Experience", describing the model as strong at general-purpose agents, search, routine office work and software/code development, above three starter tasks (Hot Spot Scout, Interaction Design Master, Product Assistant). The Model Settings panel on the right shows Max Length 262144, temperature 0.6, top_k 20, top_p 0.95 and Thinking enabled.

What to do this week

If the model is relevant to your work, the trial is the reason to act now rather than wait for the open-weight question to resolve — two weeks of free inference on a 560B MoE is the cheapest evaluation you will get, and the answer to "does it hold up on my prompts" is available today even though the answer to "can I self-host it" is not. Three things to check while the window is open, in order:

• Run your own workload, not the benchmark card. The published numbers are Ant's selection of tasks; your prompts are the ones that decide your stack.

• Test the context you will actually use. The endpoint serves 262,144 tokens, not 1M. If your design depends on the larger window, you are designing against a promise.

• Do not build a cost model on "free." Free is a two-week state. Until a rate card is published, plan the switch back to a priced Flash-tier model as the default and treat Ling-3.1-flash as additive capacity.

The announcement is real and the model is running, which is more than could be said a week ago. But released-and-open and running-on-a-trial are different states, and this one is firmly the second — a 560B specification, a 256K endpoint, a vendor-only benchmark card, and a two-week clock that is the only firm deadline in the entire release.

Generated editorial card titled "Ling-3.1-flash on trial - what to verify this week" listing six rounded rows: "Your workload: run it, not the card"; "Context: 262,144 served, not 1M"; "Cost model: free is a two-week state"; "Weights: none published, promise only"; "Fallback: keep a priced Flash-tier route"; "Deadline: the trial window is the only firm date", above a footer reading "Vendor-stated specs; no published post-trial rate card."

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily