A generated title card titled 'Qwen 4.5 and Qwen 5: 5 to 10 trillion parameters' with the subhead 'A roadmap band, not a specification' and three cards reading 'Roadmap target / 5-10T parameters', 'Shipping flagship / Qwen3.8-Max 2.4T' and 'Release date / not announced'. A footer line reads 'Per Alibaba's September 22, 2026 press release; no model card published.' Flat vector editorial styling on a white background with a blue-to-cyan gradient wash and the OrcaRouter logo composited bottom right.
Guides & Insights

Qwen 4.5 and Qwen 5 target 5 to 10 trillion parameters: what Alibaba did and did not say

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

At its Apsara Conference in Hangzhou on September 22, 2026, the company put a size on the generation after next. The Qwen 4.5 and Qwen 5 series are "projected to scale up to 5 to 10 trillion parameters," in the wording of the company's own English press release — a roadmap line, not a specification. Neither model has a release date, a model card, an API identifier, a price or a published benchmark score, which means nothing about either one is callable today. The flagship you can actually buy is Qwen3.8-Max, generally available since August 3, 2026, at 2.4 trillion parameters. In the days since the conference, that roadmap sentence has been flattened in a lot of coverage into a claim that a 10-trillion-parameter model is coming — a stronger, simpler claim than the one the company made. The distance between those two sentences is roughly a year of planning.

The claim, and the version of it that spread

Two things were said on stage, and they are worth separating because they carry very different amounts of information. The first is a status update: Qwen 4 is in training, on a new architecture. The second is a roadmap: the Qwen 4.5 and Qwen 5 series are projected to reach a parameter band of 5 to 10 trillion.

Alibaba's English press release attributes both to the company rather than to a named executive. The conference coverage that goes further attributes the roadmap to Liu Da Yiheng, who runs the Qwen large-language-model project at Alibaba Token Hub, and reports his framing that scaling is a key path toward artificial superintelligence. That framing is the context the parameter number was delivered in, and it explains why the number is a band rather than a target.

What then travelled was the top of the band. The English headlines that followed include at least one describing a teased 10-trillion-parameter model, and several describing a plan for a model of 5 to 10 trillion parameters. Both are defensible readings of a slide. Neither is a specification, and the difference matters to anyone who has to plan against it.

5 trillion is one generation's target and 10 trillion is another's

The single most useful thing to get right about this announcement is that 5 to 10 trillion is a band spanning two generations, not one model with a wide tolerance. Alibaba's own sentence attaches the band to "the upcoming Qwen 4.5 and Qwen 5 model series" — the series, plural, taken together.

Chinese-language reporting from the same conference is precise in a different direction again, with headlines describing models after Qwen 4.5 expanding into the 5-to-10-trillion range. That reading puts the 10-trillion end beyond Qwen 4.5 as well. The one statement every source agrees on is that the band covers the Qwen 4.5-to-Qwen 5 span. The assignment of 10 trillion to Qwen 5 specifically is an inference that became a headline, not a quote.

For scale, and these are the only numbers in this story that are published rather than projected:

• 5 to 10 trillion — the parameter band Alibaba's press release attaches to the Qwen 4.5 and Qwen 5 series

• 2.4 trillion — total parameters of Qwen3.8-Max, the shipping flagship, per its official model card

• 95 billion — Qwen3.8-Max's active parameters per token, the figure that drives what a token costs to serve

• 10 trillion — the top of the band, and the only part of it that reached most English headlines

• 0 — the number of published specifications, release dates, prices, context windows or benchmark scores for Qwen 4.5 or Qwen 5

A generated scoreboard titled 'Shipped vs projected - the scoreboard'. Left column 'Qwen3.8-Max (shipping)' reads GA August 3, 2026; 2.4 trillion total parameters; 95 billion active parameters; 1,000,000-token context window; $2.00 in / $6.00 out; benchmarks vendor table plus AA Index 45. Right column 'Qwen 4.5 / Qwen 5 (roadmap)' reads release none given; 5 to 10 trillion series; active parameters not published; context window not published; price not published; benchmarks none published. Footer reads 'Qwen3.8-Max specs per its model card; AA Index per Artificial Analysis; Qwen 4.5 and Qwen 5 per Alibaba's September 22, 2026 roadmap.'

The parameter number that matters is the one nobody published

Total parameters are the number a lab chooses to say. Active parameters are the number a buyer needs, and the roadmap does not include one.

Qwen3.8-Max is a mixture-of-experts model with 2.4 trillion parameters in total and 95 billion active per token. That is an activation ratio of roughly 4 percent: the model holds a great deal of knowledge in weights it does not read on any given token, and pays compute for a small slice of it. Total parameters tell you how much memory the thing occupies and how big a box you need. Active parameters tell you what you pay every time it answers.

Run that ratio forward and the shape of the problem becomes visible. A 10-trillion-parameter model built at the same 4 percent activation would activate on the order of 400 billion parameters per token — more than four times what Qwen3.8-Max activates now. That is arithmetic on a stated assumption, not a claim about Qwen 5: push sparsity up and the active count falls, push it down and it climbs. But it is the arithmetic that decides whether a 10-trillion-parameter model is a served product or a research artefact, and it is the calculation the 5-to-10-trillion headline invites and then leaves undone.

This is also where routing economics stop being abstract. On OrcaRouter the provider's list price is passed through with 0% markup, so a vendor price cut on a Qwen tier is live on your key the same day rather than at a renewal. A generation whose serving cost is genuinely uncertain is exactly the case where that pass-through is worth more than a discount negotiated in advance against a number nobody has published.

The chip announcement is the actual timeline

Alibaba announced the model roadmap and a new accelerator in the same press release, and the two are more connected than the release says.

T-Head's Zhenwu V900 is a training and inference processor with 216 GB of memory and 1,200 GB/s of inter-chip bandwidth, native support for FP8 and FP4, and a claimed three times the performance of its predecessor, the Zhenwu M890, released in May. Mass production and commercial release are scheduled for the first quarter of 2027. A companion supernode server supports clusters of up to 500,000 cards, and T-Head says the Zhenwu line has served more than 650 customers.

Read the two announcements together and the sequence is legible. Training and serving a mixture-of-experts model at 10-trillion scale is an interconnect problem before it is a compute problem, because expert parallelism means moving activations between chips constantly, and 1,200 GB/s of inter-chip bandwidth is the number that decides whether that is tractable. Silicon on that schedule puts the earliest plausible window for a frontier-scale run of this kind well into 2027 — and even then, a chip shipping says nothing about a training run finishing. Alibaba connected the two subjects by putting them in one release; the connection itself is our reading, and it is labelled as one.

The company's own time horizon is consistent with that. Chief executive Eddie Wu set a target of more than 20 gigawatts of global Alibaba Cloud data-centre capacity by 2032 — a capacity target, not deployed capacity, and a five-year-plus horizon rather than a next-quarter one.

The self-improvement claims holding the roadmap up

The reason a 5-to-10-trillion plan is arguable at all rather than merely expensive is recursive self-improvement: if the training pipeline improves itself, the compute required per generation falls, and the frontier moves closer to affordable. That is the load-bearing claim under the roadmap, and it is also the least verifiable thing Alibaba said.

What the company reported: Qwen3.8-Max completed 33 iterative cycles across roughly a month of fully automated work spanning pipeline design, data validation and error diagnosis, and lifted its Artificial Analysis score from 40 to 45. In a chip-design experiment the model spent more than 60 hours of self-improvement across a design lifecycle, made more than 10,000 electronic-design-automation tool calls, and reduced chip area by 42 percent without a performance loss. One report from the conference also carries a 96 percent inference-throughput improvement from autonomously tuning a new T-Head GPU.

These are vendor-reported and have not been independently replicated. That is not a technicality — an autonomous loop that grades its own experiments is measuring itself against its own scorer, and the outside world has no way to check the rubric. Coverage of the conference has generally said so, and readers should assume it.

There is a second trap in the same claim, and it is about labels rather than honesty. Alibaba did not say which revision of the Artificial Analysis Intelligence Index its 40-to-45 movement is measured against, and that index has been rebuilt since this model launched. The current Artificial Analysis page for Qwen3.8 Max (0902) reads 45 — an independent figure, on the revision in force today. The reading recorded for the same model in mid-August, under the revision then in force, was 56. A 45 and a 56 are not the same measurement in the same units, and subtracting one from the other produces a number that means nothing. What the page does not carry is the 40 that Alibaba says it improved from: the starting point of that delta is the company's own, and only the endpoint happens to sit where the independent reading now does. Read the movement as a direction, not a measurement.

A screenshot of the Artificial Analysis model page for Qwen3.8 Max (0902), captured September 23 2026, showing an Artificial Analysis Intelligence Index of 45 (ranked 24 of 212, badge 'Updated'), speed of 39.2 output tokens per second (159 of 212), $2.00 per 1M input and $6.00 per 1M output tokens with an 88% cache discount and $5.41 cost per Intelligence Index task (90 of 212), verbosity of 190M output tokens (88 of 212), and a technical specification listing a 984k context window with reasoning enabled.

What Alibaba did ship at the same conference

Strip out the roadmap and the conference still contained real releases, all of them multimodal:

• Qwen3.8-LiveTranslate — simultaneous interpretation, with latency (LAAL) cut nearly 20 percent, from 2.8 seconds to 2.3 seconds

• Qwen-Audio-3.1-ASR, Qwen-Audio-3.1-TTS and Qwen-Audio-3.1-Realtime — a refreshed speech suite

• Qwen-Audio-3.1-TTS-Next — cinematic soundscapes that blend dialogue and ambience from a text script, aimed at audiobooks, film and television, podcasts and games

• Qwen-Image 3.1 — image generation tuned for creative design and e-commerce marketing, due later this year, with native transparent-background generation

• Qwen Intelligence — a full-stack agentic platform aimed at smartphone makers

Every item on that list is a product with a delivery window. The frontier-scale claim is the one item on the agenda with no date attached, and the contrast is not an accident: shipping the multimodal tier is what funds a roadmap year, and the roadmap is what makes the next funding round or the next cloud contract a story rather than a spreadsheet. Both things are real. Only one of them is something you can call.

What is callable today, and what has to change for that to change

The Qwen 3.8 generation remains the whole of the buyable line. Qwen3.8-Max has been generally available since August 3, 2026 at $2.00 per million input tokens and $6.00 per million output, flat across the full 1,000,000-token context with no long-prompt surcharge. Its weights were published on August 12, 2026 as Qwen3.8-2.4T-A95B in BF16 and FP8 under a custom Qwen licence. Its vendor benchmark table is strong and unaudited — Terminal-Bench 2.1 at 86.6, GPQA Diamond at 92.6, OSWorld-Verified at 86.1, all Alibaba's own figures — The independent picture is less flattering than the vendor one: Artificial Analysis puts Qwen3.8 Max (0902) at an Intelligence Index of 45, 24th of 212 models, at $5.41 of measured cost per Intelligence Index task and 39.2 output tokens per second — 159th of 212, near the slow end of the field. The same page lists the context window as 984k against the 1,000,000 on our own model page, a gap worth knowing about before you size a long-context job. Frontier-class score, mid-pack speed: that is the honest summary of the shipping flagship, and it is the baseline any Qwen 4.5 has to beat on both axes at once.

A screenshot of the OrcaRouter model page for Qwen3.8 Max (qwen/qwen3.8-max), captured September 23 2026, showing the model id, 1M-token context, text image and video input with text output, vision tools JSON and reasoning badges, a dated snapshot label of 2026-08-03, input at $2.00 and output at $6.00 per 1M tokens, p50 TTFT 3.09s and p95 TTFT 10.00s, trailing-7-day traffic of 29.4M tokens, and a Python code sample posting to the OpenAI-compatible base_url https://api.orcarouter.ai/v1.

OrcaRouter carries the Qwen 3.8 family — qwen/qwen3.8-max, qwen3.8-flash and qwen3.8-27b — on one OpenAI-compatible endpoint alongside more than 200 other models, which means a Qwen 4.5 tier would be a model-id change on an existing key rather than a new vendor contract, a new bill and a new SDK. When one ships, that catalogue is where it would appear. Today, neither Qwen 4.5 nor Qwen 5 is on it, because neither has been released — and no amount of roadmap press changes that.

The practical use of a roadmap like this one is the opposite of a migration plan. If a Qwen 4.5 tier eventually looks plausible for your workload, the cheap way to find out is to route a slice of real traffic to it behind an automatic fallback, so a model that turns out to be slow, expensive or worse than the incumbent costs you a percentage of one workload rather than a quarter. That is what failover is for, and an unproven, days-old frontier model is the clearest case for using it.

What to watch, and what not to believe yet

A roadmap line becomes a release when a small set of concrete things appears. A model card carrying a total parameter count, an active parameter count and a context window. An API identifier you can pass in a request. Weights on a model hub. An entry on an independent leaderboard with a date attached. A price. Any one of those is a signal; most of them together is a launch.

Things that are not a release, however technical they look: a conference mention; a serving framework's pull request adding an architecture branch for an experimental Qwen build, which is engine work on somebody's inference stack rather than a model you can call; a dated snapshot id for a generation that already shipped; and a social post paraphrasing a stage remark. The signal that produced this article was the last of those, and the reason it was worth writing up is that the underlying claim is checkable against Alibaba's own press release — which is where the 5-to-10-trillion band, the Qwen 4 status and the RSI figures all come from. The signal pointed at the story; it is not the story.

The near-term question is narrower than the headlines suggest. Alibaba has said Qwen 4 is in training, which puts Qwen 4 ahead of both 4.5 and 5 in the sequence — so the next thing worth watching is not a 10-trillion-parameter model but whether Qwen 4 arrives with a model card, a price and an endpoint, and when. Until something in that chain lands, the honest position on the 5-to-10-trillion band is that it is a direction of travel, disclosed on the record, with no specification attached to it and no date. Plan for Qwen3.8-Max, watch the 4.5 and 5 band as a scaling thesis rather than a product, and treat every parameter count you see for a model with no model card as a forecast.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily