A hero title card for "Qwen 4 Max announced at Apsara 2026" with the subtitle "What Alibaba confirmed, and what it did not", three pill badges reading "4 tiers previewed", "0 specs published" and "September 22, 2026", and the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Qwen 4 Max announced at Apsara 2026: what Alibaba confirmed, and what it did not

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The Yunqi conference in Hangzhou — the Apsara Conference — was used on September 22, 2026 to show four models that have not shipped. The lab's LLM lead Liu Da Yiheng previewed Qwen 4 Max, Qwen 4 Flash, Qwen 4 Plus and Qwen 4 27B on stage, described the family as training on a new-generation architecture, and said only that it would arrive soon. As of this writing there is no model card for any of the four, no open weights, no API identifier, no context window, no price and no benchmark score. The newest flagship you can actually call today is Qwen3.8-Max, generally available since August 3, 2026, and that distance between a stage announcement and a callable endpoint is the part of this story worth reading carefully.

The four tiers, and what each one is meant to be

Alibaba's Qwen line has been organised into capability tiers since the Qwen 3 generation, and the Qwen 4 announcement kept that shape rather than replacing it. What was shown on stage was a roster, not a specification sheet:

• Qwen 4 Max — the flagship tier, and the one Alibaba positions against the top model from every other frontier lab

• Qwen 4 Flash — the high-throughput tier, built for latency-sensitive and high-volume work rather than peak reasoning

• Qwen 4 Plus — the balanced middle tier, historically the one with the widest spread of multimodal support

• Qwen 4 27B — the open-weight local tier, the one that gets downloaded, fine-tuned and quantised by people who never touch a hosted API

• Architecture — described as a new generation, and described as being trained, which is not the same as being finished

• Primary workload — agentic and tool-using tasks, per the conference framing, rather than chat

• Availability — none. No tier has a release date, a price, a context window, a modality list or a published score

The 27B tier is the one to watch for a reason that has nothing to do with benchmarks. It is the only tier in the family that has historically shipped under open weights, and the Qwen 3.8 generation showed how much independent work that unlocks: within days of the 27B weights landing, the community had produced MLX conversions, NVFP4 quantisations and uncensored fine-tunes. An announced 27B is a promise of a whole downstream ecosystem, and it is the part of this roadmap with the most predictable delivery.

The 5-to-10-trillion-parameter number does not belong to Qwen 4

Coverage of the conference has been loose about one figure. The "5 to 10 trillion parameters" target that circulated alongside the Qwen 4 news is not a Qwen 4 specification — it is a forward-looking statement about generations after it, the Qwen 4.5 and Qwen 5 timeframe. Alibaba has not attached any parameter count to Qwen 4 Max, and on the evidence available it would be a mistake to print one.

This matters because parameter count is the number readers anchor on and the number that propagates. A 5T-parameter claim attached to a model with no published architecture is unfalsifiable in both directions: nobody can confirm it, and nobody can correct it once it has been repeated. If you see Qwen 4 Max described as a multi-trillion-parameter model this week, the number has been borrowed from a different slide.

What can be said honestly is that the predecessor flagship is a 2.4-trillion-parameter mixture-of-experts with 95 billion parameters active per token, confirmed on the official model card for Qwen3.8-Max and on the open-weight release that followed it. Whatever Qwen 4 Max turns out to be, that is the baseline it has to beat, and it is a published, checkable number rather than a roadmap aspiration.

A two-column scoreboard titled "Qwen 4 Max vs the model you can call today". Left column Qwen 4 Max: Release date not given, Input price not given, Output price not given, Context window not given, Open weights not given, Benchmarks not given. Right column Qwen3.8-Max: Release date August 3 2026, Input price $2.00 per 1M tokens, Output price $6.00 per 1M tokens, Context window 1M tokens, Open weights published August 12 2026, Benchmarks vendor and independent. Footer reads "Qwen 4 Max status per Alibaba's September 22 2026 conference; Qwen3.8-Max from its published model card.", with the OrcaRouter logo bottom-right.

The architecture is not a rumour: a preview of it already shipped

The most useful thing about the Qwen 4 announcement is that part of the architecture was already released under a different name. Qwen3.8-Flash-Next was open-sourced in late August 2026, and its model card labels it plainly as a preview of the Qwen 4 architecture. That makes it the only concrete evidence anyone outside Alibaba has about how the Qwen 4 family is built.

It is a sparse model with 125 billion main parameters plus 51 billion parameters of N-gram embeddings, activating roughly 6 billion per inference step. It carries a 262,000-token native context that the card says extends toward 1 million, and Alibaba reports it trained at roughly 90 percent lower cost than Qwen3.7-Plus. Three techniques are named on the card:

• QSA, a Qwen-specific sparse attention scheme, which is the mechanism behind the cheap long-context training claim

• Gated residual connections, a stability measure for deep sparse networks

• N-gram embeddings, a separate 51-billion-parameter parameter store that is looked up rather than densely computed

If the Qwen 4 tiers inherit that design, the interesting consequence is economic rather than architectural. A family designed to reach frontier-adjacent quality at roughly a tenth of the training cost is a family that can be priced aggressively, and aggressive pricing from Alibaba has been the single most reliable force in open-weight model economics for two years. That is a prediction, not a fact, and it should be labelled as one — but it is the reason to track this roadmap rather than dismiss it.

What to run while you wait

Nothing in the Qwen 4 announcement changes what is available this week, and the honest answer for anyone building today is the Qwen 3.8 generation. Qwen3.8-Max has been generally available since August 3, 2026 at $2.00 per million input tokens and $6.00 per million output tokens, flat across the full 1-million-token context with no long-prompt surcharge, and its weights were published on August 12, 2026 as Qwen3.8-2.4T-A95B in both BF16 and FP8 under a custom Qwen licence. It is the model the Qwen 4 tiers will be measured against, and it is serving production traffic today.

If you want to compare the shipping Qwen generation against the rest of the frontier without managing a separate account, key and bill for each lab, OrcaRouter carries the Qwen 3.8 family — Qwen3.8-Max, Qwen3.8-Flash and Qwen3.8-27B — alongside Claude Opus 5, DeepSeek V4 Pro, Gemini 3.1 Pro Preview and GLM 5.3 on one OpenAI-compatible endpoint. That is one API for close to 200 models at 0% markup, meaning the provider's list price is passed through untouched, so a vendor price cut lands on your key the same day rather than at the next billing cycle. When a Qwen 4 tier does ship, the same catalogue is where it would appear; today, none of them do.

A screenshot of the OrcaRouter model page for Qwen3.8 Max (qwen/qwen3.8-max, English UI), showing the on-page nav, the 1M token context badge, the model ID qwen/qwen3.8-max, multi-modality shown as text plus image plus video input with text output, the Vision, Tools, JSON, Reasoning and Thinking capability tags, the listing date 2026-08-03, a p50 time-to-first-token of 3.29s, the code-samples panel with the OpenAI-compatible Python import, the Public benchmarks and Community buzz blocks, and the opening line of the model description calling Qwen3.8-Max Alibaba's newest flagship and highest-capability tier to date.

The bottom line

Qwen 4 Max is a real announcement and not a real product, and both halves of that sentence matter. The roster is confirmed: four tiers, a new-generation architecture, an agentic focus, and an open-weight 27B tier that will matter more to more people than the flagship will. Everything a purchasing decision needs — price, context, modality, weights, scores, a date — is missing.

Treat this as a roadmap signal with an unusually good paper trail, because the Qwen3.8-Flash-Next release gives you the architecture preview for free and Qwen3.8-Max gives you the incumbent baseline with published weights and a published price. Track the 27B tier for the ecosystem it will create and the Flash tier for the throughput economics. Do not republish the 5-to-10-trillion-parameter figure as a Qwen 4 fact, and do not plan a migration around a model that has no release date — plan it around Qwen3.8-Max, and move when there is an endpoint to move to.

A screenshot of the OrcaRouter model catalogue at www.orcarouter.ai/models (English UI), showing the header reading "199 models, 15 providers, one API key, one bill", the sort and filter rail with Newest selected alongside controls for input modalities and context length, the "How to call any model" panel showing a POST to api.orcarouter.ai/v1/chat/completions with a model and messages JSON body, and the first catalogue cards including openai/gpt-5-mini with a 200 Status badge.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily