
Qwen4Exp batch-sharded sampling: inside vLLM PR #61018, and what it says about Qwen 4
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 82 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 122 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 358 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
The most informative number in vLLM pull request #61018 is a loss: 1.5%. That is the upper bound the author puts on his own change — roughly 0.6 to 0.8 milliseconds saved out of a 42-millisecond average step, measured on a machine he does not own, in a patch he cannot run at all. The pull request, titled "[Model] Qwen4Exp: support batch-sharded sampling (compute_logits_local)" and opened on 2026-10-10 by contributor kimseunghyun-kr, adds three lines of model code to two files so that the Qwen4Exp architecture can take a sampling path vLLM shipped in August. It is a draft. It is 14 lines of model code plus 57 lines of test. And it is worth reading closely anyway, because of what those three lines are patching over: Qwen4Exp is the architecture inside Qwen3.8-Flash-Next, the 125-billion-parameter open-weight model the vendor published on 2026-08-24 as "an experimental preview of the architecture that will underpin Qwen4" — and a two-path bug in the text-only half of that architecture is exactly the kind of detail you only learn by watching the serving layer rather than the launch post.
To be unambiguous about the framing, because it matters here: Qwen 4 itself is unreleased. The vendor named four Qwen 4 tiers — Qwen 4 Max, Flash, Plus and 27B — at its Apsara conference on 2026-09-22, and has published no weights, no identifier, no price, no context length and no benchmark for any of them. Nothing below is a release. This is a what-we-know-so-far readout on a single draft pull request, and everything in it that carries a number is either a timestamp you can verify, a figure the contributor typed into his own PR body, or a value read out of a public model configuration file.
What batch-sharded sampling is, in one paragraph
Tensor parallelism splits a model's weights across GPUs; the vocabulary projection is the widest single tensor in the stack, so each rank normally computes only its own slice of the vocabulary — and then every rank all-gathers, so that each one ends up holding the full logits for every request in the batch. Sharded sampling inverts that exchange. Instead of replicating the vocabulary across ranks, it shards the batch: each rank samples one slice of the requests, and the ranks trade vocabulary slices with each other over an all-to-all. vLLM's own CLI documentation describes the flag plainly — "Each rank samples a slice of the batch instead of every rank sampling all of it" — and states the constraints: --enable-batch-sharded-sampling defaults to False, requires tensor_parallel_size greater than 1, at least tensor_parallel_size maximum sequences, and a non-negative max_logprobs. The last line of that documentation is the hook this pull request hangs on: "Models opt in by implementing compute_logits_local."
The feature itself is not new. vLLM merged it as PR #50465 — "[Model Runner V2] batch-sharded sample," by Giancarlo Delfin — on 2026-08-24, the same day Qwen3.8-Flash-Next's weights went up. The motivation there was memory and latency: materialising full target logits costs on the order of batch size × (speculative tokens + 1) × vocabulary size, and sharding cuts that allocation by a factor of the tensor-parallel degree while letting the sampler's top-k and top-p work run in parallel. It is the enabling step, not the destination: the PR's own text flags sharded draft-logits as future work.
Why three lines was the whole job
![GitHub page for vllm-project/vllm pull request 61018, titled "[Model] Qwen4Exp: support batch-sharded sampling (compute_logits_local)", showing the Draft badge, a three-file diff with 71 additions, the qwen label and the PR body.](https://cms.orcarouter.ai/api/media/file/2-1829.png)
Here is the actual defect, and it is a good one because it is invisible from outside the code. vLLM's Qwen4Exp implementation ships as two classes. Qwen4ExpForConditionalGeneration is the vision-language wrapper — a Qwen3-VL vision tower bolted onto the language model — and Qwen4ExpForCausalLM is the text-only path. The wrapper inherits compute_logits_local from Qwen3_5ForConditionalGeneration, which forwards the call down to language_model.compute_logits_local. But the language model class the wrapper was pointing at never defined that method.
So the two paths were in different states. A vision-language deployment of Qwen3.8-Flash-Next could take the sharded path; a text-only one could not, because the method the wrapper delegated to did not exist. The fix is one method, identical in the NVIDIA and AMD copies of the model file:
• The method returns self.logits_processor(self.lm_head, hidden_states, skip_gather=True) — the rank's own vocabulary shard, with no gather and no full-vocabulary materialisation on any rank.
• It follows what the PR calls "the same 3-line pattern as Qwen3.5 and MiniMax M3," which is worth pausing on: two other model families in the same repository had already opted in. Qwen4Exp was simply the one that had not.
The test file is where the honesty of the patch is easiest to check, and where its limits are easiest to see. It is 57 lines, it is parametrised over both the NVIDIA and AMD modules, and it does not download a model. The helper builds the class with object.__new__, substitutes nn.Identity for the language-model head, and installs a fake logits processor that returns its input plus one while recording how it was called. The first test asserts that a 4.0 going in comes out as 5.0 and that the recorded call carried skip_gather=True. The second wraps the language model in the conditional-generation class and asserts the delegation lands. That is a real test of the wiring and a test of nothing else — no kernel is exercised, no rank boundary is crossed, and the number of GPUs involved is zero.
Both of those facts are stated in the PR rather than buried. The test file's own docstring calls Qwen4Exp "a tiny CPU-only test double." The author's validation section notes that he could not import the model on his macOS setup at all, because the transformers build he had did not ship a Qwen4ExpConfig. The CUDA run was still pending. The end-to-end benchmark — MLPerf agentic dataset, 20 concurrent sessions, three 60-minute arms — was still pending. The MTP drafter's own logits path is flagged as unchecked. LoRA is already rejected by the flag upstream, so it is out of scope rather than a regression. There is also a line disclosing that the draft was written with AI assistance, and the commit trailer names Claude Opus 5.5 as a co-author, which is the sort of disclosure that should be normal in a stack this size and mostly is not.
The 1.5% estimate, and the measurements it sits next to
The contributor's estimate is an nsys trace at 16 concurrent sessions on Qwen3.8-Flash-Next-FP8 with tensor parallelism 8 and expert parallelism across eight A100-SXM4-40GB cards: about 1.6 milliseconds of a 42-millisecond step, trimmed to roughly one third of that, for an end-to-end gain of 1.5 to 2%. His headline is the arithmetic — 0.6 to 0.8 ms over 42 ms. Note what is attached to that: 42 ms is an inter-token latency figure, so a 1.7% cut is a 1.7% cut on token generation time, not a throughput claim in the abstract.
The honest comparison is against the parent feature's own merged numbers, because those were measured end to end by someone with the hardware. In the Speed-Bench 2K/2K runs published in PR #50465, batch-sharded sampling moved DeepSeek V4 with DSpark at 7 speculative tokens and concurrency 64 from 2.62 to 2.66 requests per second — a 1.53% throughput gain — with median inter-token latency falling 3.09% and time to first token rising 1.45%. On MiniMax M3 with DSpark at 8 speculative tokens, request throughput rose 5.38% and median TPOT fell 8.33% from 15.61 to 14.31 milliseconds. And the same plan documents where it does not help: at concurrency 4 to 16 the runs came out flat to slightly negative, with the concurrency-16 arm down 0.54% on throughput and 1.89% on acceptance length.
That pattern is the thing to carry away. Sharded sampling pays when sampling is heavy and the batch is wide — high concurrency, many speculative tokens, top-k and top-p doing real work — and it costs a little when the batch is small enough that the all-to-all is pure overhead. A 1.5% estimate for one model opting in is consistent with that, not in tension with it: the 5.38% and 8.33% figures belong to different models with different speculative budgets, and the model in this PR has its own configuration, its own 248,320-token vocabulary and its own speculative-decoding path.
None of it is audited. The parent PR's numbers are one contributor's paired runs on one node; the smoke-test numbers here are a trace on hardware the author does not have, from a revision of the patch he says was measured at 16 sessions only. A single-contributor, single-configuration measurement is a useful signal about direction and a poor basis for a capacity plan. The reason the topic is worth writing about at all is the direction, not the decimal.
What the surrounding patch cluster says about Qwen4Exp
One draft PR would not be a story. The story is that Qwen4Exp has become a sustained serving target in the week this landed, and the smallest patch in that cluster is the one that makes the pattern legible. In the seven days to 2026-10-10, vLLM carried, from the same contributor and others: an FP8 main KV cache path for sparse attention on Ampere, with KV capacity up 1.83× on eight A100s and 1.87× on four RTX 3090s and roughly 2.7× the requests at concurrency 16, at the cost of about 6% higher single-stream decode TPOT; a fix for W4A4 MoE padding at tensor-parallel 1 and expert parallelism; an AMD path that falls back from unsupported AITER FP8 MoE operations and serves in fp16; a fix carrying the PLE short-convolution state through align mode; and a decode kernel that unpacks e4m3 bytes four at a time per register on sm_80. One of them, a retune of the H200 M=4 merged QSA LL-GEMM plan, was merged into main on 2026-10-09.
Read that list as a whole and it says something concrete. The Qwen4Exp architecture — the one the vendor has not released — is being tuned for Ampere, Hopper, ROCm and fp16 fallback simultaneously, in a project that ships day-zero support for models people can actually download. The reason is not mysterious: Qwen3.8-Flash-Next is a real, downloadable, heavily used model, and it is built on the architecture Qwen 4 will use. Whether the Qwen 4 weights ever arrive looking like this is unknown and the vendor has said nothing, but the serving envelope is being widened in public, and that is checkable information in a way that a rumoured October-to-November window is not.
Some of the architectural shape is public too, for anyone who reads the configuration rather than the announcement. Qwen3.8-Flash-Next's config lists a 248,320-token vocabulary over 48 layers, 512 experts with 10 routed plus one shared active per token at an expert intermediate width of 640, a hidden size of 2,560, and a native context of 262,144 tokens described as extensible to one million. It is a hybrid: the layer-type list alternates three linear-attention blocks with one full-attention block, and the full-attention blocks use the sparse-attention path with a one-head indexer that compresses keys by a factor of four and keeps a 2,048-position budget. There is a per-layer embedding table at layer 2, an n-gram embedding with a 20-million-entry vocabulary — that is the 47.7 GiB table other patches in this cluster are busy host-staging — and a one-layer MTP head for speculative decoding. The model card's own summary is "125B with 6B activated, plus 51B n-gram embedding and 4B MTP." A 248,320-entry vocabulary is why the logits projection is worth sharding in the first place.
What this does not change for you
It is worth being exact, because a patch cluster this dense can read like a launch. None of the above is merged, and one piece of it — the Ampere FP8 KV cache work — is explicitly unvalidated in CI because vLLM's CI has no A100. There is no released vLLM version you can install today that carries the Qwen4Exp sharded-sampling opt-in. There is no independent benchmark of Qwen3.8-Flash-Next's serving behaviour under any of these changes; every number quoted above comes from the PR bodies, which makes them contributor-reported and unaudited in the specific sense that no third party has reproduced the run. And the headline number in the PR that started this piece is 1.5%, which is a real gain in a serving stack and not a reason to change a model choice.
What would change a decision is a full, audited serving run, and that does not exist yet. The pacing measurements that do exist for Qwen3.8-Flash-Next are outside these patches entirely: the model rates an Intelligence Index of 40 on Artificial Analysis, well above the median of 18 for open-weight models of similar size, and that figure is independent of everything discussed here.
What you can call today, and the routing angle

This is where the picture gets practical for anyone who has read this far and wants to use a hosted model rather than instrument one. Qwen3.8-Flash-Next is not on OrcaRouter's catalogue — it is not one of the 205 models we route, and there is no hosted endpoint for it here. But the production sibling it previews is: qwen/qwen3.8-flash, the official Qwen3.8-Flash release that the vendor's own model card describes as carrying more features than the preview, including a one-million-token context by default and built-in tools, is available at $0.15 per million input tokens and $0.47 per million output tokens, passed through at provider list price with no markup added. For scale within the same family, qwen/qwen3.8-27b runs $0.33 and $2.40, and qwen/qwen3.8-max $2.00 and $6.00 — all reachable on one key.
There are two reasons that matters more than usual for a piece about an unreleased architecture. The first is the switching cost. If you want to calibrate what a sparse-attention model feels like on your traffic before Qwen 4 exists, the comparison you want to run is against qwen/qwen3.8-flash at list price — and doing it through a router means the model you are measuring and the model you might fall back to are behind the same endpoint, the same SDK and the same key, with no second contract to sign. The second is that every serving change in this patch cluster targets self-hosted vLLM. If you are not running eight A100s, the 1.83× KV capacity and the 1.5% sampling win are things you read about, not things you get. A routed endpoint is the version of this that arrives without a build step: automatic failover if a provider degrades, and a routing DSL that lets you place a hosted call beside a self-hosted one in a single endpoint when you do have hardware of your own to measure.
The one thing not to do is read this article as a reason to wait for Qwen 4. There is no date. There is no price. There is no weight count for the real thing — the numbers above describe the preview build, not the product. What exists is a serving stack being prepared in public for an architecture that is, for now, downloadable only in preview form.
The short version

Three lines closing a gap between the text-only and vision-language paths of one model file is not, on its own, news. It is the kind of patch that would be buried in a merge commit if anyone had time to review it, and it may well be folded into a larger change or closed outright — the vLLM bot's own agent guidelines, quoted on the PR, instruct AI-assisted contributors to close their work if it lacks significant benefit, and 1.5% is a number that invites that question. What makes it worth your attention is the thing it documents: a 20-million-entry n-gram table, a 512-expert MoE with ten experts live per token, a hybrid of linear attention and compressed sparse attention, a speculative head — all of it being tuned across NVIDIA and AMD, Ampere through Hopper, before the product that carries it exists. If you care about the Qwen4 serving envelope, the PR bodies are where its real specifications currently live. If you want to call a model today, Qwen3.8-Flash is the one that is actually there.
One API for 200+ models, automatic failover, routing DSL. Explore the OrcaRouter model catalogue
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
