Hero title card for 'Ling-3.1-flash vs GLM-5.3-Flash', subtitled 'Four times the throughput against one point of index'. Two rounded cards sit side by side with a clear gutter: the left, labelled 'Ling-3.1-flash', reads 'Speed: 211 tok/s', 'AA Index: 41', 'Licence: none published'; the right, labelled 'GLM-5.3-Flash', reads 'Speed: 53 tok/s', 'AA Index: 42', 'Licence: MIT'. A footer line reads 'Throughput on each vendor's own API; index per Artificial Analysis.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Ling-3.1-flash vs GLM-5.3-Flash: Four Times the Throughput Against One Point of Index

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ling-3.1-flash and GLM-5.3-Flash land one point apart on the Artificial Analysis Intelligence Index — 41 against 42 — which makes them look like the same decision twice. They are not. GLM-5.3-Flash generates output at 53.0 tokens per second on Z.ai's own API; Ling-3.1-flash generates it at 211.0 tokens per second on InclusionAI's. That is a fourfold difference in throughput, and it is far larger than anything the index separates them by. One model is the more capable of the two on nearly every quality axis and open under MIT. The other is the one that finishes first, is closed, and has no published price, licence or checkpoint. Picking between them is a question about what your workload is waiting on.

The axis Ling-3.1-flash wins outright

Speed is not a footnote when you are choosing between models at the same capability level, and here it is the entire case for the Ant model.

• Output throughput — Ling-3.1-flash 211.0 tokens per second on InclusionAI's API vs GLM-5.3-Flash 53.0 tokens per second on Z.ai's. Four times the generation rate.

• Time to first token — Ling-3.1-flash 1.80 seconds vs GLM-5.3-Flash 3.18 seconds. The Ant model starts answering while the Z.ai model is still thinking, which matters more than throughput in interactive and streaming use.

• Time per Intelligence Index task — Ling-3.1-flash about 357 seconds vs GLM-5.3-Flash about 979 seconds. A single evaluation task takes GLM-5.3-Flash nearly three times as long end to end, because it is both slower per token and slower to start.

For an agent loop that makes a dozen sequential calls, or a product where a person is watching tokens arrive, that is not a tiebreaker — it is the whole reason to prefer one model. GLM-5.3-Flash is the better model on paper and the slower one in the request path.

Where GLM-5.3-Flash is simply better

Everywhere else, and the list is long enough that the throughput advantage has to earn its place.

• Knowledge reliability — GLM-5.3-Flash scores +7.47 on AA-Omniscience, the best of the three Flash-tier models, with a 27.6% hallucination rate on answered questions. Ling-3.1-flash scores +2.22 with a 37.9% hallucination rate. On a measure that penalises invention more than refusal, the gap is real and points the same way as the index.

• Scientific knowledge — GLM-5.3-Flash posts 0.912 on GPQA Diamond. Ling-3.1-flash has no GPQA Diamond row published. So does DeepSeek V4.1 Flash (Max). A model with a 0.91 on graduate-level science questions is a different instrument from one that has not been measured on them.

• Modality — GLM-5.3-Flash takes text, images and video. Ling-3.1-flash takes text only. This is not a performance gap to be closed by a benchmark; it is a capability that is absent.

• Weights and licence — GLM-5.3-Flash is open weights under MIT, downloadable and self-hostable. Ling-3.1-flash has published no checkpoint, no licence text, and is filed as proprietary.

• Agentic knowledge work — GLM-5.3-Flash 1,453.73 Elo on AA-Briefcase v1.1 against Ling-3.1-flash's 1,400.35, on a benchmark built specifically around knowledge-work tasks rather than terminal driving.

• Price — GLM-5.3-Flash is published at $0.15 per million input and $0.50 per million output on Z.ai's API, against Ling-3.1-flash's $0.30 and $0.90. Half the input rate and roughly half the output rate, from a model with a higher index score and open weights.

A two-column generated scoreboard titled 'Ling-3.1-flash vs GLM-5.3-Flash — the scoreboard'. The left column, 'Ling-3.1-flash', reads 'AA Index: 41', 'Speed: 211 tok/s', 'First token: 1.80 s', 'Omniscience: +2.22', 'Weights: closed', 'Price: $0.30 / $0.90'; the right column, 'GLM-5.3-Flash', reads 'AA Index: 42', 'Speed: 53 tok/s', 'First token: 3.18 s', 'Omniscience: +7.47', 'Weights: MIT', 'Price: $0.15 / $0.50'. A footer line reads 'Index and quality rows per Artificial Analysis; throughput per each vendor's own API.' The OrcaRouter logo is composited in the bottom-right corner.

The rows where the Ant model answers back

Ling-3.1-flash is not beaten on all of it. On agentic terminal work it is marginally ahead and on coding-science it is level:

• Terminal-Bench 4.0 — Ling-3.1-flash 0.333 vs GLM-5.3-Flash 0.328. Effectively a tie, and both sit below the frontier tier.

• SciCode — Ling-3.1-flash 0.541 vs GLM-5.3-Flash 0.516. A narrow win.

• AutomationBench-AA — 0.617 vs 0.604. Another narrow win.

• GDPval-AA — Ling-3.1-flash 1,621.75 Elo vs GLM-5.3-Flash 1,646.86. GLM takes this one back.

Read the four lines together and the shape is clear. Where the two models can be separated at all, the separation is inside the noise of a single evaluation run. The one-point index difference between them is not a ranking; it is a coin toss weighted slightly toward GLM by the reliability rows. What is not a coin toss is the 4× throughput gap and the 2× price gap, both of which point the other way.

There is also a serving detail worth flagging, because it changes the size of that throughput gap depending on where you call a model. Z.ai's own API measures at 53.0 tokens per second. Our own catalogue's seven-day measurement for GLM-5.3-Flash on our routes shows 150.6 tokens per second of output at a 4.2-second median first-token latency. The same weights served from a different deployment run close to three times faster. Vendor throughput figures are a property of a provider, not of a model.

Specification, line by line

Subject first on every line:

• Size — Ling-3.1-flash 560B total / 25B active vs GLM-5.3-Flash 320B total / 18B active.

• Declared context — 1M tokens on both. GLM-5.3-Flash's long-context reasoning row measures at 0.80 on AA-LCR; Ling-3.1-flash's at 0.83. Both are close to DeepSeek V4.1 Flash (Max)'s 0.84, which is to say long-context reasoning is no longer a differentiator at this tier.

• Output ceiling — GLM-5.3-Flash 128,000 tokens on our catalogue vs Ling-3.1-flash with no published maximum. One of these can be sized for long generation and one cannot.

• Index — 41 vs 42.

• Throughput — 211.0 tokens per second vs 53.0 on the vendor APIs, 150.6 measured on our routes.

• Weights — proprietary with no licence vs open under MIT.

• Modality — text only vs text, image and video in.

• Price — $0.30 / $0.90 per million input / output vs $0.15 / $0.50 on Z.ai's API.

How to actually run this comparison

GLM-5.3-Flash is on the OrcaRouter catalogue now, listed at $0.075 per million input, $0.25 per million output and $0.0173 per million cache reads — read the live card rather than this paragraph, because serving rates move and the pass-through arrangement means a provider price change is reflected on our side the same day. The value of that arrangement for this specific matchup is that GLM-5.3-Flash's openness and price advantage are the two things that make it the sane default, and a closed 0%-markup route is how you keep testing Ling-3.1-flash against it without adding a vendor relationship.

Ling-3.1-flash is not on our catalogue — a lookup against our public model API returns not-found for the model under InclusionAI, and this piece will not imply we serve it. It runs through Ant's own API and the third-party platforms carrying the release, which is the honest extent of its availability.

The productive shape of this comparison is not either-or. GLM-5.3-Flash is open, cheapest, multimodal, more reliable on knowledge questions and available through 23 providers — that is a default path. Ling-3.1-flash's case is throughput and TTFT on agentic loops, and its cost is that it is closed, text-only, more expensive and reachable through fewer doors. Route the interactive and latency-sensitive work toward the fast model, keep the batch, knowledge-heavy and multimodal work on the open one, and put a fallback between them so a slow upstream does not become a slow product.

Which one to pick

If your evaluation criteria are capability, cost, openness and modality, GLM-5.3-Flash wins this matchup and it is not especially close — a higher index score, the best knowledge-reliability profile in the tier, a 0.912 GPQA Diamond, MIT weights, video input, half the price, and 23 providers behind it. If your criteria include how long a user waits or how many sequential calls a task can make, Ling-3.1-flash is the model that finishes, and its fourfold generation rate is not a rounding error in an agent loop.

What would settle it is a serving detail neither vendor publishes in a comparable form: delivered throughput on your workload, through your provider. Both models answer the same OpenAI-compatible chat-completions format, which means switching between them is a configuration change rather than a migration — and for two models separated by one point of index and a factor of four in speed, measuring that one number on your own traffic is a better use of an afternoon than reading another benchmark row.

Screenshot of the OrcaRouter model page for z-ai/glm-5.3-flash, in English, captured 2026-10-07. The header reads 'GLM 5.3 Flash', 'ctx 1M tokens', 'Max output 128K', inputs 'text + image + video' and output 'text', with vision, tools, JSON and reasoning capability icons and a release date of 2026-08-26. Four rounded stat tiles read '$0.07', '$0.25', '4.20 s' and '10.00 s' — the input and output rates rendered to two decimals, then the median first-token latency and another latency figure. The description states '320B total / 18B active parameters, 1M-token context, text + image + video in, text out.' The PRICING panel lists 'Input / 1M tokens $0.075', 'Output / 1M tokens $0.250' and 'Cache read / 1M $0.017'. A code sample shows base_url https://api.orcarouter.ai/v1 with model 'z-ai/glm-5.3-flash'.Screenshot of the Artificial Analysis model page for GLM 5.3 Flash, in English, captured 2026-10-07, marked 'Open weights model' and 'Released August 2026'. It shows an Intelligence Index of 42, 53.0 output tokens per second, $0.25 cost per Intelligence Index task, 180M output tokens generated across the index, input $0.15 with an 83% cache discount and output $0.50 per 1M tokens, a 1M context window, text and image input, 320B total parameters with 18.8B active parameters, and a licence of MIT with weights on Hugging Face. The summary line calls the model 'notably slow (75)'.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily