
Qwen3.8-Flash-Next Is Out: Alibaba Ships the Qwen4 Architecture Before Qwen4 Exists
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 83 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 122 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 354 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
Qwen3.8-Flash-Next finally has numbers that Alibaba did not produce — and they do not say what the loudest version of the story says. On the Artificial Analysis Intelligence Index, restated to v4.3 on September 7, Alibaba's open-weight preview scores 40 and ranks fifth among the 113 models in its comparison class. DeepSeek V4 Flash 0731 (Reasoning, Max Effort) — the model it keeps being measured against — scores 35 and ranks ninth. That is not parity; it is a five-point lead for the smaller model. It is also the first broad, independent scoreboard to reach a model whose only published numbers, until this month, were the vendor's own. The model itself is not new: the weights went live on Hugging Face on August 24 and the formal ModelScope release landed on August 26. What changed this month is that a reader can now check the claims instead of taking them.
The prompt to revisit this post came from @teortaxesTex on X, who asked whether the preview is quietly underhyped: supposedly on the same Artificial Analysis tier as DeepSeek V4 Flash 0731, "almost 4x smaller," with "almost 3x fewer active parameters." Two of those three assertions do not survive contact with the published data — and the correction runs in the model's favour, because on the numbers that matter the preview is ahead of its comparison, not level with it.
Start with size, where the numbers are unambiguous. Qwen3.8-Flash-Next stores roughly 180B parameters — a 125B mixture-of-experts body, a separate 51B n-gram embedding table, and about 4B of multi-token-prediction layers — and activates 6B of them per token. DeepSeek V4 Flash 0731 stores 284B and activates 13B. Those are the figures on Qwen's model card and DeepSeek's documentation, and they are the figures Artificial Analysis carries on both model pages. The ratios they produce are 1.6x on stored parameters and 2.2x on active parameters. That is a genuine efficiency story, and it is the reason this architecture matters — but it is not 4x and not 3x. We could not find a published figure anywhere that supports the larger multipliers. If the intent was the served memory footprint rather than the parameter count, that number is not published either.
What was announced
Alibaba published the Qwen4 architecture before it published a Qwen4 model. Qwen3.8-Flash-Next — an open-weight, multimodal mixture-of-experts model built on the next-generation architecture that will power the Qwen4 family — came out in two steps: the official Qwen/Qwen3.8-Flash-Next repository went live on Hugging Face on August 24, two days before the ModelScope release the teaser timer had been pointing at. The formal ModelScope drop landed August 26 at 23:00 UTC+08:00, with the served Qwen3.8-Flash variant priced at the same time. The card's headline claim, which has been circulating ever since, is that a model activating 6B parameters per token beats Claude Opus 4.6 Max on 8 of 9 comparable benchmarks on Alibaba's own table. The full Qwen4 family, Alibaba keeps saying, comes later.
If you have not been tracking Alibaba's cadence this month, the anchor is Qwen3.8-Max — the 2.4-trillion-parameter flagship released on August 3 and open-sourced as Qwen3.8-2.4T-A95B less than two weeks later. Qwen3.8-Flash-Next's weights are out less than a month after that. The same lab that put out a 2.4-trillion-parameter open-weight model is now handing out the architecture for the generation after it, before the flagship of that generation is even named.

The part worth re-reading is Alibaba's own framing. The model is described as built on the next-generation architecture "that will power the upcoming Qwen4 family" — and explicitly not as Qwen4 itself. Alibaba's stated reason is that the architectural changes should be in the community's hands before the family arrives, so that runtimes, quantizations, and applications are ready when the real thing ships. That is a marketing position, but it is also a technical commitment: preview the hard part first, take the community along.
What the model card settles, and what it does not:
• Confirmed, per the official model card: Qwen3.8-Flash-Next is open-weight, multimodal, and a mixture-of-experts model on the Qwen4 architecture — the card calls it "an experimental preview of the architecture that will underpin Qwen4."
• Confirmed, per the model card and config: two headline architectural changes — Gated Delta Network (GDN) and Qwen Sparse Attention (QSA) — inside a 48-layer hybrid that runs 36 linear-attention layers against 12 full-attention layers.
• Confirmed from the repository: 125B total parameters with 6B activated per token (512 experts, 10 routed plus one shared) — Artificial Analysis lists 180B once the separate 51B n-gram embedding table and the MTP layers are counted — with a 262,144-token native context extensible to 1,000,000 under the Qwen Community License 1.0.
• No longer unconfirmed: the model has now been independently scored, twice over. Alibaba's table is still the vendor's own and still unreproduced, but it is no longer the only evidence in the room.
The first independent scores are in — and they say "ahead", not "level"
Artificial Analysis shipped Intelligence Index v4.3 on September 7, and the re-scoring is the reason any of this is comparable at all. The v4.3 update replaced Terminal-Bench v2.1 with the much harder Terminal-Bench v4.0 and swapped τ³-Banking for AutomationBench-AA, an agentic workflow benchmark built with Zapier on a private held-out test set of 657 tasks. Private held-out weighting rose to 45%. The stated purpose was to stop the frontier from gaming the index, and the effect on scores was large: GPT-6 Astra and Claude Fable 5.1, the two leaders, both landed at 53 having been scored in the sixties on the old scale.
That matters when you read the community numbers. The "56 versus 52" figures that circulated for Qwen3.8-Flash-Next and DeepSeek V4 Flash 0731 in early September were computed on the pre-v4.3 index; under v4.3 the pair sits at 40 and 35. Quoting either set against the other is comparing two different exams.
On the current index, here is how the two models actually compare on Artificial Analysis's own evaluations:
• Intelligence Index — Qwen3.8-Flash-Next 40 (5th of 113 in class) vs DeepSeek V4 Flash 0731 (Reasoning, Max Effort) 35 (9th of 113).
• Agentic knowledge work — AA-Briefcase 1587 vs 1259; GDPval-AA v2 1647 vs 1468; AutomationBench-AA 56% vs 54%.
• Terminal and coding — Terminal-Bench v4.0 25% vs 12%; SciCode 51% vs 50%.
• General and scientific reasoning — Humanity's Last Exam 38% vs 39%; CritPt 11% vs 17%; GDP.pdf 16% vs 11%; AA-LCR v1.1 80% vs 80%.
• Factual recall versus confabulation — AA-Omniscience −10 vs −14, a four-point edge to the Qwen preview.
• Price — $0.15 per million input / $0.47 per million output vs $0.44 / $1.32; blended $0.0882 per million tokens vs $0.2298. Cost per Intelligence Index task: $0.37 vs $0.22.
• Speed and latency — 53 output tokens per second vs 210; time to first token 2.80s vs 0.98s; end-to-end response 49.76s vs 12.88s; time per task 1,572.60s vs 221.65s.
• Token use — 108k output tokens per task vs 62k, of which 72k are reasoning tokens vs 45k. Artificial Analysis calls the preview "very verbose" and "slower than average."
• Context and size — 256k tokens vs 1,000k; 180B total / 6B active vs 284B / 13B.
Read together, that is a sharper answer than "same tier." Qwen3.8-Flash-Next wins the accuracy-and-efficiency argument on almost every axis Artificial Analysis measures — it is the higher-scoring model, it is smaller, and it costs about a third as much per token. DeepSeek V4 Flash 0731 wins the production argument outright: four times the throughput, a quarter of the end-to-end latency, seven times less wall-clock per completed task, and four times the context window. And on cost per completed task — the number that survives token verbosity — DeepSeek is the cheaper model at $0.22 against $0.37, despite the higher sticker price. If "underhyped" means "better than its reputation on quality per parameter," the label is fair. If it means "you should be running this instead," the speed and latency columns say otherwise.

There is independent evidence outside Artificial Analysis too. A third-party harness, smf-bench, published an "Official A" run on September 5: Qwen3.8-Flash-Next in NVIDIA's NVFP4 quantization on a single DGX Spark, thinking off, scored 137 of 157 (87.3%) with a clean 30 of 30 on its coding section, against 117 of 157 for a two-Spark DeepSeek V4 Flash Vision-Exp run on the same 157-test harness. That is one harness on one machine class, and it is not a substitute for a broad evaluation — but it is a second data point pointing the same direction, and it arrives from someone with no stake in either vendor.
What the Qwen4 architecture actually is
Two mechanisms carry the story. The first, GDN — Gated Delta Network — is Alibaba's linear-attention layer. Where a standard transformer keeps a key-value cache that grows with sequence length, a GDN layer maintains a fixed-size recurrent state and updates it with a learned gating rule; the lineage runs through DeltaNet and Mamba-2. The payoff is the one that has been quietly powering the Qwen line since Qwen3-Next: long contexts stay cheap because the state does not grow, while a minority of full-attention layers stays in the mix to do the precise retrieval that linear attention is bad at. In the current generation that mix runs at roughly three GDN layers for every full-attention layer — community analysis of the Qwen3.8-2.4T-A95B checkpoint counts 69 GDN layers against 23 attention layers. Qwen3.8-Flash-Next shows the same three-to-one split: 36 linear-attention layers against 12 full-attention layers.
The second, QSA — Qwen Sparse Attention — is the genuinely new piece, and the public description of it is still thin: the model card adds that it operates at micro-block level with a 512-block / 2048-token budget, but no full technical write-up exists yet. That scarcity of detail is itself part of the pattern. Qwen3-Next's preview in late 2025 introduced Gated DeltaNet with the same thin documentation, and the architecture only got fully specified once the Qwen3.5 series adopted it. For a reader the practical takeaway is the same both times: the architecture canary shows up first, the documentation and the production models show up after.
The playbook that already worked once
Alibaba has run this exact play before, and it worked. In late 2025, Qwen3-Next previewed Gated DeltaNet and a hybrid of linear and full attention. When the Qwen3.5 series shipped, it adopted that blueprint at scale — and the hybrid has carried through the Qwen3.8 generation since, including the 2.4T flagship. The reason the precedent matters here is the timing it implies. The full Qwen4 family should be expected on the far side of this preview, not inside it; if the Qwen3-Next gap is any guide, the distance between "here is the architecture" and "here is the family" is measured in months. Nothing in the model card confirms that — it is an inference from a pattern Alibaba has now run twice.
What we still don't know
The unknowns have shrunk, but they have not disappeared:
• The vendor benchmark table is still the vendor's. Artificial Analysis has now scored the model independently, which resolves the biggest gap in the launch coverage — but Alibaba's own headline claim, that the model beats Claude Opus 4.6 Max on 8 of 9 comparable benchmarks, rests on Alibaba's own in-house evals (CoWorkBench, JobBench, AndroidWorld) alongside public ones, and none of those has been reproduced by a third party.
• Whether the preview is the same model as the served one. The card positions Qwen3.8-Flash-Next as the open-weight experimental preview, while the production served variant — Qwen Cloud's Qwen3.8-Flash — adds a default 1,000,000-token context and built-in tools. Whether that served model is the same architecture under a bigger context window is still not stated, and the independent scores above are for the open-weight preview, not the served API model.
• The license is known and is a real license: Qwen Community License 1.0 permits commercial use, including selling derivative works, but requires a separate commercial agreement with Alibaba if you run a Model-as-a-Service or AI-work-assistant business past a defined revenue and monthly-active-user threshold. It is not the same as the revenue-gated Qwen3.8-Max license, but it is worth reading before building on the weights.
• One field report worth flagging: a September 6 write-up of a community NVFP4 build logs a grammar-constrained tool-calling failure when thinking and multi-token prediction are both enabled, traced to the xgrammar constrained-decoding path. That is a runtime bug in a quantized community build rather than a property of the model — but if your evaluation harness leans on grammar-constrained tool calls, it is the first thing to test.
A family iterating on a two-week clock
The surrounding cadence is its own story. Qwen3.8-Max, the 2.4-trillion-parameter flagship, released August 3; the open-weight Qwen3.8-2.4T-A95B followed on August 12–13; the free Apache-2.0 Qwen3.8-27B landed days later; and now, before the month is out, the next generation's architecture is being previewed — and independently benchmarked a week later. No other lab is currently iterating on this clock. For a reader picking models, the practical consequence is that "the best Qwen" is a moving target — and the delta between generations is arriving faster than the surrounding ecosystem usually adapts. Buying a decision rather than a model is increasingly the defensible move.
What a developer should do now
The efficiency numbers change the local-serving calculus more than the launch itself did. A 6B-active model that scores 40 on the current index is compressible: NVIDIA's official NVFP4 quantization is the one the September 5 smf-bench run used, and community NVFP4 and FP8 builds beyond it are already in circulation, which is what makes a single-GPU-class deployment of a 180B-stored model plausible at all. The caveat is unchanged — this is an unproven preview, the vendor table is unreproduced, and the shape of the independent scores (strong on knowledge work and terminal tasks, weaker on long-horizon repo generation and one-shot agent pass rates) means the model's weaknesses show up exactly where production code changes live. Run a low-stakes copy of a real workload at it before it touches anything that matters, and let automatic failover absorb the moments when a preview misbehaves.
On price, the picture that was murky at launch is now concrete. Artificial Analysis lists the preview's served rate at $0.15 per million input tokens and $0.47 per million output, against $0.44 and $1.32 for DeepSeek V4 Flash 0731 — the same order of magnitude as the served Qwen3.8-Flash variant, which Qwen Cloud lists at ¥1 and ¥3 (roughly $0.16 and $0.47). The production variant is the one you can actually call today: it is routed here as qwen/qwen3.8-flash, and OrcaRouter passes the vendor's list price through at 0% markup, so when Alibaba moves that number the endpoint moves with it the same day. The same is true of the comparison model in this piece — DeepSeek V4 Flash 0731 is routed as deepseek/deepseek-v4-flash-0731 at the provider's list price, which makes the cost-per-token half of the matchup above testable against your own traffic rather than against a benchmark's token counts. What is not routed here is the open-weight Flash-Next preview itself: to use it today you pull the weights and serve them yourself, which is precisely why the quantization story above matters more than the list price. The ceiling this generation has to clear is still visible on the same endpoint: Qwen3.8-Max (qwen/qwen3.8-max) sits at $2.00 per million input tokens and $6.00 per million output tokens, also passed through at the vendor's list price.

The practical sequencing has not changed much, but the confidence behind it has. If you want the served production variant without pulling 100GB of weights, Qwen3.8-Flash is already callable at list price. If you want the architecture — GDN, QSA, the n-gram table, the offload tricks — the weights are downloadable and the runtimes are ready: vLLM serves it from day one with an offload path that keeps the 51B-parameter n-gram embedding table in host memory rather than permanent VRAM, SGLang and TensorRT-LLM are in NVIDIA's Day 0 list, and Ollama's library carries an MLX build you can pull on Apple Silicon. The one thing to stop doing is treating the model as an unknown on quality. It has been measured now, and it measured well.
What to watch next
• Whether the independent numbers hold as the index matures. Qwen3.8-Flash-Next's 40 comes with an unusually high token bill — it burned 245M tokens on the index run against a 140M median — and Artificial Analysis's own note calls it "very verbose." A score produced by reasoning longer is still a score, but it is the kind of thing that moves when the evaluation changes. The next index revision is the real test of whether 40 was a level or a local maximum.
• Whether the served Qwen3.8-Flash gets scored separately. Every independent number in this piece is for the open-weight preview. The production variant with the 1M context and built-in tools has no independent score of its own, and if it is the same architecture under a longer context, that is a cheap evaluation someone should publish.
• How the day-0 serving support holds up in practice — the vendor-stated GB300 throughput figures are still vendor-stated, and the TP4 expert-parallel and KV-offload configurations are still new enough to be fragile.
• How long Alibaba waits before the full Qwen4 family follows the preview.
The what-we-know-so-far part of this story has now largely closed. The spec sheet is public, the weights are downloadable, the architecture is confirmed, the served Qwen3.8-Flash variant is live on hosted APIs at list price, and the model has been independently scored by two separate parties — with the smaller model winning the quality argument and the larger one winning the throughput argument. What remains open is whether a 6B-active preview generalises past the benchmarks it was tuned against, and how long Alibaba waits before the full Qwen4 family follows. That last question is still the one the release leaves unanswered, and it is still the one that will matter most.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
