
Qwen 4 Leak: SGLang Shards Qwen4Exp Prefill for 18% Faster TTFT and a GiB Back Per GPU
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 64 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
A pull request opened in the SGLang repository at 03:47 UTC this morning, 2026-10-08, promises something that no Qwen 4 announcement has yet produced: a number. It is titled feat(qwen4-exp): enable LayerNorm sequence parallelism for GR and PLE, and on four H20 GPUs running the open-weight Qwen3.8-Flash-Next checkpoint in FP8, its author reports time-to-first-token gains of 17.7–18.4% at 32K input and 14.6–14.7% at 235K input, roughly a gigabyte of peak memory returned per GPU, and up to 22% more input throughput. Qwen 4 itself — the family the vendor named but did not ship at the Apsara conference in Hangzhou on 2026-09-22 — is still unreleased, with no weights, no identifier and no price. The only model that instantiates the Qwen4Exp architecture today is Qwen3.8-Flash-Next, published on 2026-08-26, and its managed sibling Qwen3.8-Flash is the version an API caller can actually reach. So read what follows for exactly what it is: one engineer's paired measurements, attached to an open, draft, unmerged pull request, about the serving envelope of a model family that does not exist yet.
Sourcing first, because this is a leak piece and the distinction does real work. The signal is sgl-project/sglang#43048, opened 2026-10-08 at 03:47 UTC by the GitHub account shiyang814-cpu, last touched at 03:55 UTC, and still marked draft, with no approving review recorded and no merge. It changes six files — two test files, the Qwen4Exp model file, the LayerNorm-SP module, a layer-boundary factory and an argument-group hook — for +345 and −50 lines. Every performance figure below comes from the PR description, is the author's own paired OFF/ON measurement, and has not been reproduced by anyone. Three CI runs on the revision at the tip of the branch are marked failing. Nothing here is a shipped capability.

What the pull request actually changes
Sequence parallelism is not a model change and it is not a new capability. It is a rewiring of where a few layers do their arithmetic. Its lineage runs back to Megatron-style sequence parallelism — the trick from arXiv:2205.05198 — and SGLang already ships it: the module's own docstring explains the mechanism it reuses, that under pure tensor parallelism a row-parallel all_reduce is algebraically a reduce_scatter followed by an all_gather. Because those two collectives move exactly the same bytes as the single all_reduce they replace, splitting the operation that way costs no extra communication volume at all. What it buys is the freedom to let the normalization and residual regions run on sequence-sharded activations — each tensor-parallel rank holding one-1/tpth of the token rows — which cuts the transient activation memory that long-context prefill needs to keep alive.
What this particular pull request does is extend that existing path from the architecture it was validated on to the Qwen4Exp architecture. During prefill, the Gated Residual and Per-Layer Embedding activations stay sharded along the token dimension across the TP group. Before attention, before GDN, before QSA, and before the Mixture-of-Experts block run, the full token rows are all-gathered back, the existing full-row tensor-parallel computation runs unchanged behind a shared fallback, and a reduce-scatter then sums the partial contributions and restores each rank's shard. Decode never touches the new path at all. The feature is reached through the option that already exists — --enable-layernorm-sp — with no Qwen4Exp-specific flag, and with the flag absent the code behaves exactly as it did before.
Why Qwen4Exp is the architecture that needs this
The reason this matters for Qwen4Exp specifically, and not equally for every model, is in the architecture's own design. Qwen3.8-Flash-Next applies Gated Residual projections to all token rows in every decoder layer — the configuration declares four residual streams and a bottleneck rank of 320 across 48 layers — and Per-Layer Embedding adds a second replicated, token-row-by-token-row projection on top of that. Under tensor parallelism, both of those operations are duplicated identically on every rank, because they carry no TP-sharded weight matrix of their own to force a split. Sharding their token dimension removes the replicated work directly, and as the PR's motivation section puts it, that happens while preserving the existing tensor-parallel layout and the reduction semantics of attention, GDN/QSA and MoE — which is the part that makes the change safe rather than clever.
It is worth saying plainly what that means for the reader. The interesting thing about this PR is not that SGLang is getting faster. It is that the Qwen4 architecture carries per-layer costs that scale with token count rather than with parameter count, and those costs are the ones that bite on long prefills. That is a design fingerprint, and it is the kind of thing a spec sheet never mentions.
The measured deltas
The author's benchmark fixes one configuration and toggles the flag: four NVIDIA H20 GPUs, Qwen3.8-Flash-Next-FP8, tensor parallel 4 and expert parallel 4, chunked prefill size 8192, FlashInfer linear-attention prefill and decode backends, the same server configuration for OFF and ON, alternating OFF → ON → OFF → ON service restarts, and fixed token inputs with warmup requests. Every figure below is from that setup and is unaudited:
• 32K input, batch size 1 — TTFT improved 17.72–18.36%, end-to-end latency about 16%, input throughput about 20%
• 235K input, batch size 1 — TTFT improved 14.63–14.71%, end-to-end latency about 14%, input throughput about 17%
• 32K input, batch size 4 — TTFT improved 18.74%, end-to-end latency 18.14%, input throughput 22.14%
• Peak memory — down by approximately 1.0 GiB per GPU
• Decode, batch size 1 — time per output token effectively unchanged
The last line is the one to read twice, and the author is candid about why: this optimization is enabled only for prefill, so single-stream decode gets nothing from it. The batch-4 time-per-token improvement, where it appears, reflects reduced scheduling delays from concurrent long prefills rather than any faster decode kernel. If you were hoping this was a throughput story, it is not — it is a latency-to-first-token and memory story, and those are the two constraints that decide whether a 235K-token request is servable at all.

The layout bug they had to fix first
The most informative part of the pull request is not the speedup table. It is the section on PLE physical-row layout, because it shows what the Qwen4Exp serving stack still gets wrong.
Per-Layer Embedding operates on a fixed physical CUDA-graph bucket, while only a prefix of the rows in that bucket may hold real tokens. The padding therefore has to be applied before the sequence is sharded, not after. In the final chunk of a 235K-token request, the author's numbers are 5,624 processed tokens inside an 8,192-row physical bucket at TP 4, and the only correct layout is rank 0 holding 2,048 valid rows, rank 1 holding 2,048 valid rows, rank 2 holding 1,528 valid rows plus 520 padding rows, and rank 3 holding 2,048 rows of padding. Sharding the 5,624 processed rows first — the obvious implementation — inserts padding between globally contiguous valid ranges and corrupts the result. The author records that a real 235K OFF/ON test only produced identical 16-token greedy output after this layout was fixed.
That is a small detail with a large implication. The PLE path in SGLang was added recently enough that a token-ordering bug of this shape was still reachable, and the person who found it was writing the sequence-parallel extension. Day-zero serving support for this architecture is not finished; it is actively being built, in public, by contributors, one layout at a time.
What it costs you: the constraints
A flag that only helps some deployments is only useful if you know which. The PR states its requirements explicitly, and configurations outside them fail during argument validation rather than silently degrading:
• Tensor parallel size must be greater than 1 — a single-GPU deployment gains nothing, because there is no rank to shard across
• Expert parallel size must equal tensor parallel size
• Pipeline parallel size must equal 1
• Data-parallel attention must be disabled
• Speculative decoding must be disabled
The last constraint is the one with a real decision behind it. For a sparse model that activates around 6B parameters per token, speculative decoding is one of the few levers that accelerates decode, and this feature explicitly turns that lever off in exchange for a prefill win. If your workload is long-prompt and short-output — document and codebase analysis, video summarisation, a large context read once — the trade is straightforwardly good. If your workload is a short prompt and a long generation, you are giving up the thing that was helping you and buying a number that does not apply to you. The expert-parallel-equals-tensor-parallel requirement is the other one to notice: it means the moe sharding geometry has to line up with the TP geometry exactly, which rules out several otherwise reasonable multi-node layouts.
What this says about the Qwen 4 timeline
Read the diff another way and you get a calendar. SGLang's LayerNorm-SP module on the main branch today carries an explicit allowlist of architectures for which the feature has been validated, and as of this writing that allowlist contains exactly one entry, Qwen3ForCausalLM — every other architecture is rejected at construction if you pass the flag. Adding Qwen4Exp to that path is therefore not a tweak to a mature abstraction; it is the first time the Qwen4 architecture has been brought onto an optimization that predates it by generations.
Set that against the public record and the picture is coherent. The vendor announced on 2026-09-22 that Qwen 4 is in training and previewed four tier names — Qwen 4 Max, Qwen 4 Flash, Qwen 4 Plus and Qwen 4 27B — with no specifications attached to any of them. The open-weight preview that shares the architecture, Qwen3.8-Flash-Next, has been downloadable since 2026-08-26. What is happening in the three weeks since is exactly what you would expect between "in training" and "launch": engine authors shaking down the runtime so that day-zero support is real rather than nominal. A pull request that makes a serving optimization work on the architecture, opened on the morning of 2026-10-08 and still in draft, is a better signal about how close Qwen 4 is to being servable than any date anyone has floated. It is also, emphatically, not a release date — the flag is off by default, the change is unmerged, and the model it benchmarks is the preview, not Qwen 4.
What you can call today
None of which changes what is actually available this afternoon. Qwen3.8-Flash-Next is real, its weights are on Hugging Face, and you can self-host it — but it is not on our catalogue, and we will not pretend otherwise. The tier we do serve is qwen/qwen3.8-flash, the managed sibling running the same Qwen4-preview architecture with a 1M-token context window and text, image and video input, priced at $0.15 per million input tokens, $0.47 per million output and $0.0184 per million cache reads. That is a list price passed through at 0% markup, so when the vendor moves it, the number in your invoice moves the same day rather than whenever a middleman republishes a table.

There is a second, less obvious reason to care about a routing layer here. Everything in this article is about an unproven preview plus a draft patch — the sort of thing you want to test without betting a production path on it. That is what failover is for: put the preview behind the same key as the model you already trust, watch how it behaves on your traffic, and let the request fall through to the known-good route when a provider wobbles or the endpoint is not there. One API for 200+ models, one set of credentials, no second contract signed to find out whether a new architecture is worth your attention.
Two things to watch from here, neither of which we can predict. The first is whether this patch merges at all: it is a draft with three failing CI runs on a six-file change from a contributor account with no prior history in the repository, and the fluidity of the PLE row-layout work suggests the author is still iterating. The second is whether the allowlist grows — if Qwen4Exp joins Qwen3ForCausalLM as a validated architecture, then this stops being a leak and becomes the default way a Qwen4-family model is served on long context. Until one of those happens, treat the 18% as a promise about where the runtime is going, not a number you can rent.
