
Qwen 4 QSA gets decode context parallelism: inside vLLM's draft PR for Qwen3.8-Flash-Next
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 397 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1136 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 51 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 106 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
On 2026-09-29, a draft pull request appeared in the vLLM repository titled “[Model][DCP] Support Qwen4Exp QSA,” and for the model it describes, it carries the most concrete serving numbers anyone has published all month: paired runs of Qwen3.8-Flash-Next on four GPUs showing KV token capacity going from 9,759,529 to 17,603,636, maximum concurrency from 37.23× to 67.15×, and time to first token falling from 1,869 ms to 767 ms. Qwen3.8-Flash-Next is the open-weight, 125-billion-parameter mixture-of-experts preview whose Hugging Face card describes it as “A Preview of the Qwen4 Architecture”; the pull request adds decode context parallelism to the sparse-attention path that architecture is built around. Qwen4 itself — the Qwen4 Max, Flash, Plus and 27B tiers the vendor named at its Apsara conference on 2026-09-22 — is still unreleased, with no weights, no identifier, no price and no date. So read this as what it is: not a launch, not a benchmark, but an engineering artifact that tells you how the Qwen4 serving envelope is being widened before the family exists.
This is a what-we-know-so-far piece, and the sourcing matters more than usual. The pull request is a draft, open, and unmerged — vllm-project/vllm#59279, opened on 2026-09-29 by Sungsoo Ha, an NVIDIA software engineer, and still sitting in draft. Every number below is the author’s own paired measurement, reported in the PR body, taken on an earlier revision of the same work. Nothing here is independently audited, nothing here has landed in a release, and the caveat the author attaches is material enough that it gets its own section below.
What the pull request actually changes
Decode context parallelism — DCP — is a serving technique, not a model change. Instead of one GPU group holding an entire KV cache, DCP splits that cache across ranks, so each rank reads only its slice of the context while the attention results are combined across ranks at the end. The point is capacity: with the cache partitioned, a deployment can hold far more concurrent long-context traffic on the same hardware, which is exactly the constraint that bites when every request is carrying a quarter of a million tokens.
The complication is that Qwen Sparse Attention — QSA — is not a plain attention layer. As the Qwen3.8-Flash-Next model card specifies it, a lightweight indexer compresses keys into micro-blocks at a compression ratio of 4, scores them, and keeps the best 512 blocks, roughly 2,048 token positions, while the final softmax and value aggregation still run on the uncompressed K and V. That means QSA carries more state than a KV cache: there is the main cache, and there are the selector and side caches the indexer maintains. The generic DCP implementation in vLLM does not know about any of that.
What #59279 does, per its description, is teach DCP about the QSA-specific parts:
• Each rank reads its own part of the main KV cache, while QSA’s selector and side caches stay replicated across ranks rather than sharded.
• Attention results are combined across ranks after the split read.
• The selector and the main KV cache are kept in one cache group, so they cannot drift apart.
• Synthetic V2 batches are prevented from writing QSA side caches.
![A capture of the vLLM GitHub pull request #59279, titled '[Model][DCP] Support Qwen4Exp QSA', showing an Open state with a Draft badge, the head branch sungsooha:n4/qsa-dcp-clean-20260929, the description of how decode context parallelism is enabled for Qwen4Exp QSA, and the labels kv-cache-manager, mrv2, speculative-decoding, dflash, nvidia, qwen and ci/build.](https://cms.orcarouter.ai/api/media/file/2-1437.png)
That last pair of details is the interesting part if you care about correctness rather than throughput. A sharded attention cache that quietly disagrees with a replicated selector is the kind of bug that shows up as a slow accuracy decay at long context rather than a crash, and the change is explicit about keeping the two in step. The author also states that AI assistance was used and Codex is credited as a co-author — worth saying plainly, because in a draft PR of this shape it is a fair question to ask who wrote what.
The paired numbers, and how they were taken
The test plan is specific enough to be checkable, which is why the results are worth quoting. Both arms serve Qwen/Qwen3.8-Flash-Next-FP8 on four GPUs with tensor parallelism 4 and expert parallelism enabled, at --gpu-memory-utilization 0.90 with prefix caching on. The only difference between the two arms is --decode-context-parallel-size: omitted for DCP=1, set to 2 for DCP=2, with a restart between arms so the benchmark begins from a cold cache. Load is an AgentX 256k trace at 128 users for 900 seconds; accuracy is EvalScope for GSM8K plus the checked-in MRCR evaluator, run six times per arm with the first run after restart discarded.
The reported throughput deltas, DCP=2 against DCP=1:
• KV tokens — 9,759,529 vs 17,603,636, a 1.80× increase in cache capacity.
• Maximum concurrency — 37.23× vs 67.15×, also 1.80×.
• Requests per second — 1.69 vs 2.30, 1.36×.
• Input tokens per second — 128,730 vs 179,702, 1.40×.
• Time to first token — 1,869 ms vs 767 ms, 2.44× lower.
• Inter-token latency — 43.48 ms vs 26.27 ms, 1.66× lower.
• Prefix cache hit rate at steady state — 67.85% vs 88.98%, a gain of 21.1 percentage points.

Accuracy, reported as mean ± sample standard deviation across post-warmup runs, was essentially flat: MRCR aggregate 0.8630 ± 0.0005 at DCP=1 against 0.8697 ± 0.0153 at DCP=2, and GSM8K 0.9788 ± 0.0020 against 0.9790 ± 0.0016. The 2-needle and 4-needle MRCR samples were fixed at 0.9960 and 0.9906 on both arms, so all of the run-to-run movement came from the 8-needle samples — and one DCP=2 aggregate run scored 0.8970 while the other four sat between 0.8620 and 0.8632. That is a real spread, not noise you can wave away, and it is stated in the PR rather than smoothed over.
What these numbers do not establish
The caveat is in the PR text and it is not small. The paired AgentX and accuracy results were measured on an earlier QSA DCP revision, using a vLLM nightly based on commit 3df4ae153eb. The final clean commit in the pull request includes a subsequent QSA localization-kernel fix and has passed focused B200 validation — but the full AgentX and accuracy evaluations have not been repeated on that exact source. In other words: the throughput story and the shipped diff are not the same artifact, and the author says so.
Beyond that, the usual discipline applies, and it applies hard here. These are single-configuration numbers from a single contributor on a single four-GPU setup. They are vendor-adjacent rather than neutral: a framework contributor measuring a framework change is a normal and useful thing, but it is not an independent audit, and no third party has reproduced the run. There is no released vLLM version you can install today that contains this change, because the change has not been merged. And DCP=2 is a two-way split of one specific shape — the deltas are not a promise about what DCP=4 or DCP=8 would do, and nothing in the PR claims they are.
Why a serving PR about an unreleased architecture is still worth your time
The obvious objection: the model in the title does not exist, so why care? Because the thing being tuned is not Qwen 4. It is Qwen3.8-Flash-Next, and that model does exist — Alibaba published it on 2026-08-24 as a 125B-parameter MoE with 6B activated, a 51-billion-parameter n-gram embedding table, a 4B MTP head for speculative decoding, 48 layers arranged as twelve repeats of three Gated DeltaNet blocks followed by one QSA block, 512 experts with 10 routed and 1 shared active, and a native context of 262,144 tokens that the card says is extensible to 1,000,000. It is the reference implementation of the Qwen4 architecture in open weights, and QSA — the micro-block sparse attention that this pull request is teaching DCP to shard — is the single most distinctive part of it.
What the numbers describe is what happens when you stop treating that 262K context as something one GPU group has to hold whole. The 1.80× jump in KV token capacity and concurrency is the arithmetic of splitting a cache in two, which is the least surprising result in the list. The more interesting figures are the latency ones: 2.44× lower time to first token and 1.66× lower inter-token latency at the same offered load, plus a 21-point improvement in steady-state prefix cache hit rate. Those say the DCP path is not merely buying capacity at a latency cost — in this paired run it bought both. That is the shape of change that matters to anyone serving agent traffic with very long system prompts, because prefix-cache behaviour at long context is usually where long-context throughput quietly dies.
And this is not an isolated patch. The same week produced a cluster of Qwen4Exp engine work: #59214 adds SM100 low-latency decode GEMM plans for B200 shapes, #59010 adds an SM90 native sparse prefill kernel for the QSA path on Hopper, #58977 covers BF16 INC PLE embeddings, and #58961 — the one that has actually merged, on 2026-09-28 — fixed a profiling KV cache that QSA key views were keeping alive. Read together, they are the serving envelope of the Qwen4 architecture being built in public, in the runtimes, months before the family ships. If you are planning for Qwen 4, the useful signal is not a launch date — there isn't one — it is what the kernels and cache layouts already assume about how you will have to serve it.
What you can call today
If you want to test long-context behaviour on the architecture this PR is about, the model to reach for is the Flash tier Alibaba actually serves. Qwen3.8-Flash — the production deployment built on Qwen3.8-Flash-Next, with a 1,000,000-token context and 131,072-token maximum output, taking text, image and video input — is live, and it is one endpoint for the model that actually runs the Qwen4Exp architecture today, listed as qwen/qwen3.8-flash at $0.15 per million input tokens and $0.47 per million output tokens, with cache reads at $0.0184. Because those are provider list prices passed through with no markup on our side, a vendor price or limit change on it reaches you the same day it is announced.

Two honest qualifications. First, Qwen3.8-Flash-Next itself — the FP8 weights in the pull request’s test plan, the ones you would need to reproduce any of these measurements locally — is not on our catalogue; the served Flash tier is the QwenCloud production line, not the raw preview checkpoint. If you want to run the exact configuration in the PR you are self-hosting on four GPUs. Second, the DCP change is unmerged, so nothing you can call anywhere today is running it. What the served tier gives you is a way to find out whether your workload is even shaped for the problem DCP solves: if your prompts are long, agentic and prefix-heavy, then the 1.80× capacity and the prefix-cache delta are the numbers to watch for in your own traces.
And if the interesting part for you is not one model but the switching question — which tier to build on while the Qwen 4 lineup is still nameless — that is a routing problem rather than a serving one, and one API for 200+ models is how you keep the option open without a second contract or a code change when the family finally lands.
Questions worth answering directly
Does #59279 mean Qwen 4 is out, or about to be?
No. The pull request is about the Qwen4Exp architecture as implemented in Qwen3.8-Flash-Next, which Alibaba shipped on 2026-08-24. The Qwen 4 family — Max, Flash, Plus and 27B — was named on a stage at Apsara on 2026-09-22 and placed on a company roadmap with a successor line projected at 5 to 10 trillion parameters, and it still has no model card, no weights, no API identifier, no context window, no price and no date. A framework PR adding a parallelism mode to the preview architecture is a step toward serving Qwen 4 well. It is not a step toward Qwen 4 existing.
How is decode context parallelism different from tensor parallelism?
They split different things and fail in different ways. Tensor parallelism partitions the weights and the computation of each layer across GPUs, so every rank participates in every token but sees the whole sequence. Decode context parallelism partitions the KV cache itself, so each rank holds and reads only a slice of the context and the partial attention results are merged afterward. TP is about fitting the model; DCP is about fitting the context and the concurrent traffic that rides on it. That distinction is exactly why this PR is nontrivial: QSA’s selector and side caches cannot simply be sharded the way the main KV cache can, so the change has to shard one and replicate the others, and then prove the two stay consistent.
If I call Qwen3.8-Flash-Next today through a hosted API, do I already get these numbers?
No, and the gap has three parts. The change is unmerged, so no released vLLM build contains it. Even once merged, the provider has to adopt that build and choose to run with a DCP size above one — it is a serving configuration, not a default. And the measured deltas are from an earlier revision of the patch rather than the final commit, which the author states has only had focused B200 validation so far. Treat the reported deltas as a well-documented upper bound on what the approach buys in one configuration, not as a specification of any endpoint you can rent this week.
The open question
The thing to watch is not whether this specific draft merges — it probably will in some form, since the QSA-specific cache handling it adds is a genuine gap rather than a preference. The thing to watch is whether the final commit gets the same paired evaluation the intermediate revision got. A serving change whose throughput claims come from one build and whose correctness claims come from another is, for now, a well-argued proposal rather than a measured result, and the accuracy spread at the 8-needle MRCR samples is wide enough that repeating the run on the shipped source would be the single most useful thing anyone could publish about it. Until then: the direction is legible, the ledger is not closed, and the only Qwen4-architecture model in open weights remains the one from August.
