A title card for Ling-3.0-flash-VL summarizing the model as a 124B-parameter, 5.5B-active native multimodal mixture-of-experts from Ant Group's inclusionAI lab, released September 2026 under the MIT license, scoring 25 on the Artificial Analysis Intelligence Index v4.3 and serving 142.9 output tokens per second.
Guides & Insights

Ling-3.0-flash-VL: Ant's 5.5B-Active Vision Model Lands at 25 on the Intelligence Index

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ling-3.0-flash-VL now has a benchmark number that nobody at Ant Group typed, and that is the news. The first natively multimodal model out of Ant's inclusionAI lab is scored at 25 on the Artificial Analysis Intelligence Index — an independent run, on the current v4.3 scale, published alongside a model page listing its speed, latency and price. It arrives three weeks after the vendor's own model card claimed 42 on the same index, and the most useful thing a reader can do with those two numbers is understand why they are not a contradiction: Ant measured against the Intelligence Index protocol v4.1.1, Artificial Analysis measured against v4.3, and those are different rulers built from different evaluation sets. The vendor number did not survive contact with a third party so much as get re-taken on a scale that did not exist when it was written.

The model underneath the score is a sparse mixture-of-experts with 124B total parameters and roughly 5.5B active per token, released under the MIT license with BF16 and FP8 weights, accepting text, images and video and returning text. That active-parameter figure is the whole argument for the thing. At 5.5B active, Ling-3.0-flash-VL sits on what has become the most interesting curve in open models — the frontier of intelligence plotted against activated parameters rather than total size — and it is there because it does a decent amount of work per unit of inference compute, not because it is enormous.

This piece covers what actually shipped and when, what the independent run says, why the Pareto claim holds up better than most, what the model costs and where you can actually call it. The short version, for anyone who wants only that: Ling-3.0-flash-VL is real, open-weight, unusually cheap to serve, measurably above average for its size class, and noticeably more verbose and slower-to-ship than the raw score suggests. Its vendor-claimed 42 was never independently reproduced and is not comparable to the 25 now on the board.

What actually shipped, and the timeline that keeps getting flattened

Coverage of this release tends to collapse several separate events into one launch date, and the dates matter because they explain why early write-ups described a model that nobody outside Ant could score. The Hugging Face repository inclusionAI/Ling-3.0-flash-VL was created on September 4, 2026 — 64 shards of BF16 weights, an MIT license, a custom-code stack for the bailing_moe_v3_vl architecture, and, for a few days, almost nobody downloading it. The FP8 quantization followed on September 8, roughly 126 GB across 64 shards, with the vision tower kept at higher precision than the expert weights. Artificial Analysis lists the release as September 10, 2026, which is the date its own page anchors to, and the model page carrying the score went live around then.

So "released September 2026" is accurate but useless. The practical sequence was: weights on the 4th, a second precision format on the 8th, and independent measurement on the 10th. The reason this matters is that a model with weights on disk but no hosted endpoint and no third-party score is, for most teams, not yet a thing you can evaluate — it is a download you have to provision against. Ling-3.0-flash-VL crossed that line when the API and the independent score both existed.

Serving support arrived alongside the weights rather than after them. Ant published an SGLang deployment cookbook and maintains a Ling-tuned fork of vLLM for the architecture, which is the part of an open release that usually lags by weeks. It did not lag here.

The number on the board, and the one on the card

Here is the independent picture from Artificial Analysis, which is the only third-party measurement of this model that currently exists:

• Intelligence Index — 25 on v4.3, ranked #2 of 64 in its size class, against a class median of 8br /> • Total parameters — 124Bbr /> • Active parameters — 5.5B per tokenbr /> • Output speed — 142.9 tokens/second, #10 of 64, against a peer median of 105.1br /> • Time to first token — 2.21 seconds, against a peer median of 2.29br /> • Context window — 262K tokens as measured by Artificial Analysis, against the 256K the model card itself statesbr /> • License — MIT, open weights

And here is the vendor figure it is so often printed next to: 42 on the Artificial Analysis Intelligence Index protocol v4.1.1, four points above what the text-only Ling-3.0-flash scored on the same protocol. That claim is in the model card, has no published evaluation trace, and is not the number Artificial Analysis reports. Two things are true at once: the 42 is unreproduced, and it is not evidence of anything dishonest, because v4.1.1 and v4.3 do not measure the same thing. Artificial Analysis restated its index in September 2026 and re-scored its whole board; v4.3 incorporates ten evaluations including AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0 and Humanity's Last Exam. Any article still quoting a 42 next to a 25 without naming the version is comparing two different tests and calling it a decline.

The Artificial Analysis model page for Ling-3.0-flash-VL showing an Intelligence Index of 25 on v4.3 with a class ranking of #2 of 64, 124B total and 5.5B active parameters, a 262K-token context window, 142.9 output tokens per second, a 2.21 second time to first token, 160M output tokens from the Intelligence Index ranked #8 of 64, and a listed price of $0.00 per 1M tokens in and out.

The detail worth flagging beyond the headline score is verbosity. Artificial Analysis records Ling-3.0-flash-VL spending 160M output tokens to complete the Intelligence Index, ranked #8 of 64 against a median of 84M, and labels it "very verbose." For a model whose pitch is cheap inference through a small active-parameter count, that cuts directly against the pitch. You pay per token, not per parameter. A model that activates 5.5B but emits nearly twice the median token count to answer the same questions gives some of that structural advantage back at the till — which is exactly the kind of interaction a per-token price table hides and a cost-per-task measurement exposes.

The Pareto claim, put under a little pressure

The claim that Ling-3.0-flash-VL sits on the intelligence-versus-active-parameter Pareto frontier is, unusually, one the independent data supports rather than merely permits. The class median on this index is 8; Ling-3.0-flash-VL scores 25; and it does so while activating 5.5B parameters per token. For teams whose binding constraint is servable throughput — a single accelerator, a fixed request budget, an edge deployment — that ratio is the entire decision, and it is a genuinely strong showing.

Two caveats keep it from being a clean win. The first is the parameter-count discrepancy: the model card says approximately 5.5B active, while the SGLang cookbook describes a 5.1B-active model. Both are Ant documents, the difference is small, and it changes the shape of the curve slightly rather than the direction. The second is that "frontier" is a relative claim about a field that moves weekly. Being the best intelligence-per-active-parameter at a given size is a position you hold until the next model in that band ships, and this band is crowded.

The vendor's own headline demonstration — that Ling-3.0-flash-VL, competing in the Image-to-WebDev Arena under the handle "linthium," scored 1,441 against GPT-5.4's 1,440 — should be read as an attention-getter rather than a result. A one-point margin is inside the noise of an arena Elo, it is vendor-reported, and it is not the kind of gap that decides anything. The capability claim behind it is more interesting than the number: the model is built for a vision feedback loop in which it generates front-end code from a layout, renders its own output, looks at the render, and corrects — observe, act, verify, correct, rather than caption once and stop.

A single-column scoreboard card for Ling-3.0-flash-VL listing six rows: Intelligence Index 25 on v4.3, active parameters 5.5B, total parameters 124B, context window 262K tokens, output speed 142.9 tokens per second, and license MIT open weights, with a footer noting the index figure is per Artificial Analysis and the vendor card claims 42 on the retired v4.1.1 protocol.

What it costs, which is less clear than it looks

The Artificial Analysis page lists Ling-3.0-flash-VL at $0.00 per 1M input and $0.00 per 1M output tokens, against peer medians of $0.14 and $0.40. That is not a price. It is the appearance of one. Artificial Analysis prices the endpoint it can see, and the endpoint it can see is free — the model is running on Ant's Ling Studio as a free trial, and free promotional access is what the listing reflects. Ant's own multimodal documentation does not publish per-token pricing for the vision model at all, so there is no vendor rate card to check the zero against. For a sense of scale, the text-only sibling Ling-3.0-flash has been listed around $0.075 per 1M input and $0.22 per 1M output, and Ant's domain-specific Ling variants launched as free APIs too.

Treat the zero as a testing window rather than a tariff. Free tiers on freshly released models are a customer-acquisition instrument, and the honest planning assumption is that Ling-3.0-flash-VL will eventually carry a price somewhere in the neighbourhood of its text sibling. There is also a live ambiguity the arrival of a paid endpoint will not resolve: Ant's own API examples use the model identifier Ling-3.0-flash-VL-rc1, a release-candidate marker, and it remains unclear whether the served checkpoint is byte-identical to the September 4 open weights. If you are benchmarking your own prompts against the hosted API and expecting the results to transfer to a self-hosted BF16 build, that is an assumption you should test before you build on it.

The cost picture inverts depending on which side you are on. For a team with accelerators already, the FP8 build at roughly 126 GB fits within an eight-way 80GB-class node, and the 5.5B active parameter count means you are paying that node's amortized cost against a very small per-token compute footprint — the cheapest possible version of this model is the one you serve yourself. For a team without accelerators, the question is whether a paid endpoint arrives before the free tier closes.

Where you can actually call it, including the part where we say no

Ling-3.0-flash-VL is reachable through Ant's own platform — an OpenAI-compatible endpoint at api.ant-ling.com/v1/chat/completions and an Anthropic-compatible one alongside it — plus free trial access on Ling Studio, plus open weights you can serve yourself. Independent gateways have begun listing it as well, and that is genuinely the fastest way to try it without provisioning anything.

The thing worth stating plainly is that OrcaRouter does not route Ling-3.0-flash-VL today. It is not on our model catalogue, so anything you read here about calling it through us would be fiction, and we would rather you knew that up front. What we do route is the vision model most directly comparable to it: DeepSeek V4 Flash Vision Exp, Deep​Seek's experimental multimodal build, which is live on our catalogue with a 1M-token context window and a published rate. If your actual requirement is a capable vision model behind one OpenAI-compatible endpoint — with a fallback chain configured so a provider hiccup retries before your response starts, and provider list pricing passed through with nothing added on top, so a vendor price change reaches you the same day it lands — that is the model to try first while Ling-3.0-flash-VL remains off-catalogue.

A release-timeline card for Ling-3.0-flash-VL covering four dated events: the Hugging Face repository created September 4 2026 with MIT-licensed BF16 weights, FP8 weights published September 8 at roughly 126 GB, the release anchored to September 10 by Artificial Analysis, and an independent Intelligence Index score of 25 published the same week, with a footer noting the model is not carried on the OrcaRouter catalogue.

What to watch from here

Three things will decide whether this release looks significant in six months. The first is a second independent run: one score from one evaluator is a data point, and the interesting question is whether Ling-3.0-flash-VL's placement on the intelligence-versus-active-parameter curve survives a different harness. The second is real pricing. Until Ant publishes a rate card for the vision model, every cost comparison involving it is an estimate wearing a number's clothing, and the free tier is doing the work a price would otherwise do. The third is the FP8 quality delta — the vision tower was deliberately held at higher precision than the expert weights in that build, and whether the quantized version matches BF16 on fine-grained visual tasks is exactly the kind of question that only shows up once people are running it in production rather than benchmarking it.

The verdict that is available today is narrower than the launch coverage implied and more useful than the score alone. Ling-3.0-flash-VL is a real, MIT-licensed, self-hostable vision model that activates a small fraction of its 124B parameters, scores 25 on a current independent index against a class median of 8, and gives some of that advantage back through verbosity. That is a good model with an honest shape. The 42 on the card is a number from a ruler that no longer exists.