Hero title card reading 'Qwen4-Exp QSA comes to Huawei's Ascend' under an 'UNVERIFIED — DRAFT PR, UNMERGED' badge, with the subtitle 'Inside SGLang PR #41855 — the opt-in CANN prefill path', three chips reading 'Opened 2026-09-30', 'Flag default off' and 'Ascend 910C / CANN 9.0', and a footer line reading 'Contributor-reported figures; not independently reproduced.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Qwen4-Exp QSA comes to Huawei's Ascend: inside SGLang's opt-in CANN prefill PR

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

On 2026-09-30, a contributor opened SGLang pull request #41855, titled “[NPU] Add opt-in CANN sparse attention for Qwen4-Exp QSA prefill,” and the interesting part is not the arithmetic. It is the hardware. The Qwen4Exp architecture now has a hand-written sparse-attention path for Huawei’s Ascend 910C accelerator, behind a flag that defaults to off, in a draft pull request that has not merged — while the model that architecture name belongs to, Qwen4-Exp, has never been published in any form. The one checkpoint that carries this architecture in open weights is still Qwen3.8-Flash-Next, the 125-billion-parameter mixture-of-experts preview the vendor released on Hugging Face on 2026-08-24, and whose card actually declares its architecture as qwen4_exp. Qwen 4 itself — the Max, Flash, Plus and 27B tiers the vendor named at its Apsara conference on 2026-09-22 — still has no weights, no identifier, no price and no date. So this is a what-we-know-so-far piece about an engineering artifact, not a launch: one more vendor’s serving stack quietly deciding that an unreleased architecture is worth supporting early.

What the pull request actually adds

The change is deliberately small and deliberately narrow. Five files, one commit, +355 lines against a main branch at commit b87a241, carrying the SGLang npu label. The author, w1ida, states the intent up front: an opt-in CANN main-attention path for Qwen4-Exp QSA eager prefill, built on torch_npu.npu_sparse_flash_attention, with the indexer, the Top-K selection, the token budget and the KV-cache contents all left exactly as they were.

The trick it uses to get there is worth one paragraph, because it explains why this is a layout adapter rather than a new attention kernel. For already-rotated Q and K, the path packs the cache as C = [K, V] and the query as Q' = [Q, 0]. The product Q' @ C.T then equals Q @ K.T, and because the padded query contributes nothing, softmax(scale * Q' @ C.T) @ C returns [P @ K, P @ V] stacked — so the V half can be sliced out. In the author’s own words, this is “an attention-layout embedding, not a change to the model’s attention or a low-rank KV compression.” Each KV head becomes an independent batch in the native MLA layout, the original D256 scale is preserved, and auxiliary RoPE is zeroed.

The operational details matter as much as the math:

• Enablement — SGLANG_NPU_QSA_NATIVE_PREFILL=1, default off. Only ordinary ForwardMode.EXTEND opts in; decode, speculative modes, mixed forward and graph capture all stay on the existing paths, and graph capture bypasses the adapter entirely.

• Hardware and dtype — BF16 with head dimension 256, Ascend 910C (Ascend910_93*), tested against CANN 9.0 and torch-npu 2.10. Any unsupported dtype or shape falls back silently to the reference path.

• Head shapes — supported local (Q heads, KV heads) pairs are (16,2), (24,2), (12,1), (6,1) and (3,1). CANN rejects a query/KV ratio of 12 outright — its tiler only accepts powers of two — so heads are padded 12→16, 6→8 or 3→4 and the added outputs discarded. This is the clearest sign in the whole PR that the hardware was not designed with sparse-attention head ratios in mind, and the adapter is absorbing that mismatch rather than the model changing shape for it.

• Sizing bounds — only the referenced physical cache extent is packed, capped at 262,144 tokens, which the author computes as at most 512 MiB for the packed BF16 K/V tensor at two KV heads. Interior -1 padding is detected and routed to the fallback, because CANN requires contiguous valid slots; fully masked rows keep the existing zero-output convention.

• Why prefill only — the extent and layout check copies two scalars to the host, and the temporary copies plus native workspace cost memory. That synchronisation is the reason the path is restricted to eager prefill and kept off during capture.

A capture of the SGLang GitHub pull request #41855, titled '[NPU] Add opt-in CANN sparse attention for Qwen4-Exp QSA prefill', showing an Open state with no merged marker, the line 'w1ida wants to merge 1 commit into main from npu/qsa-cann-prefill', the npu label, and the start of the Motivation section describing the Q' = [Q, 0] and C = [K, V] packing and stating that the approach avoids modifying the indexer, Top-K selection, or KV-cache layout.

Nine tests that passed, and a speed number that did not come from this branch

The correctness evidence is specific and reproducible, which is more than most kernel PRs offer. The author reports 9 tests passing in 40.772 seconds on an Ascend 910C (Ascend910_9362) with CANN 9.0 and torch-npu 2.10.0, needing no checkpoint: nonzero random BF16 Q/K/V against an FP32 CPU reference computed from the same BF16 inputs, unordered physical slots at widths 1/63/64/65/2051, default and explicit scales, fully masked rows, zero rows, zero selection width, non-contiguous tensors, compression-ratio-4 causal tail remainders of 0 through 3, two-request physical mapping with a shared prefix, cache-content reuse, and the Flash-Next local head shapes at 1 and 257 query rows.

The headline case is a long prefill: 7,810 query tokens against a 65,536-token cache with 2,051 selected slots per query, all outputs finite, with eight sampled rows compared against the FP32 reference. Observed relative L2 error reached 0.209% on the small head-shape cases and 0.231% on the sampled long-prefill rows, against test gates of atol=0.025, rtol=0.025 and relative L2 under 0.008, with empty rows required to be exactly zero. Peak allocated NPU memory for that run is quoted at 1,042.7 MiB — and the author labels it as the PyTorch allocator metric, not board HBM and not full-model memory, which is exactly the right caveat to attach.

Then there is the number that will get quoted and should not be. The PR body carries a speed table showing the existing local attention path at 2,270.79 new tokens per second and the native packed main attention at 3,890.06 — a 1.713× / +71.3% gain, with mean time to first token falling from 3.109 s to 1.812 s. The author is explicit that these are historical prototype measurements taken on 2026-09-29 on an adapted Whittle-Next-26B-A3B checkpoint, run at TP1 with W8A8 weights and BF16 attention on 910C under CANN 9.0, using the official sglang.bench_serving harness at concurrency 1 with six requests, one output token each, and 36,096 cached prefix tokens excluded from the new-token throughput. They are not a benchmark of the upstream branch in the PR, the original serving JSON artifacts are not present in the checkout, and later local results around 5,000 tokens per second used additional native indexer and block4 work that is explicitly not attributed to this change. The author also notes that the adapted checkpoint’s indexer weights are inert and its budget differs from the original, so none of it is evidence about indexer correctness or full-budget generation quality.

One naming clarification, since it will trip up anyone searching for the checkpoint: the benchmark model is the contributor’s own adapted artifact. Separately, “Whittle-Next” is also the name of a public series of Qwen3.8-derived MoE fine-tunes published by a third-party Hugging Face account, including a 26B-A3B variant uploaded in September. Those are not the Qwen4Exp model this PR targets, and they should not be read as the benchmark configuration behind that 1.713× figure.

Two further caveats come from the author rather than from me. Full serving integration, distributed tensor parallelism and the complete Qwen3.8-Flash-Next model have not been validated on this branch; the contributor says keeping it a draft while the integration and dependency question is discussed is deliberate, and even asks in the PR body whether the adapter belongs in SGLang or in the separate sgl-kernel-npu repository. CI is not clean either — the PR-body state block shows failures on PR Test (Base), PR Test (Extra) and the AMD ROCm 10 run. The correctness described above is operator-level; nothing in the PR claims an end-to-end accuracy or throughput result on the integrated stack.

Why QSA is the awkward part, in numbers

Qwen Sparse Attention is not a conventional attention layer, and the published configuration shows why an accelerator vendor has to write a bespoke path for it. From the Qwen3.8-Flash-Next config: 48 layers laid out as twelve repeats of three Gated DeltaNet blocks followed by one full-attention block, full_attention_interval 4, hidden size 2,560, attention head dimension 256, 24 query heads against 2 KV heads, RoPE dimension 64. The indexer that makes the attention sparse is a multi-query structure with 4 query heads sharing 1 key head, head dimension 128, a compression ratio of 4 and a budget of 2,048 selected micro-blocks per query.

That budget is what the PR keeps fixed. The sparse block size stays 1, the sparse mode stays 0, the attention mode stays 2, and the selected-token interface is untouched; native indexer and block4 optimisations are explicitly out of scope. So this is an adapter bolted under an existing selection mechanism, not a reimplementation of QSA — which is also why the author can credibly claim the KV-cache contents are unchanged.

A two-column scoreboard titled 'Qwen4-Exp QSA on Ascend — what was measured where'. The left column, headed 'Correctness (this branch)', reads: Suite 9 unit tests, Runtime 40.772 s, Reference FP32 CPU, Long prefill 7,810 query tokens, Relative L2 error up to 0.231%, Model needed none. The right column, headed 'Speed (historical prototype)', reads: Suite sglang.bench_serving, Tokens/s 2,270.79 to 3,890.06, Reported gain 1.713x, Mean TTFT 3.109 s to 1.812 s, Checkpoint adapted 26B-A3B, Model card not on this branch. A footer line reads that both columns are contributor-reported in SGLang PR #41855 and unaudited, and that the speed arm is a 2026-09-29 prototype, not this branch. The OrcaRouter logo is composited in the bottom-right corner.

The model card frames the design intent plainly: rather than selecting individual tokens, QSA works at the micro-block level to cut long-context latency, and that micro-block granularity, plus the replicated selector state that goes with it, is precisely what does not map cleanly onto a generic paged-attention kernel on either vendor’s silicon.

Where this sits in the Qwen4Exp serving build-out

Seen alone, one draft PR on one accelerator is a curiosity. Seen against the rest of September, it is the fourth or fifth plank of a platform that is being assembled in public before the family it serves exists:

• The architecture in open weights — Qwen3.8-Flash-Next, 2026-08-24, a 125B-parameter MoE with 6B activated, a 51-billion-parameter n-gram embedding table and a 4B MTP head, carrying model_type: qwen4_exp and architectures Qwen4ExpForConditionalGeneration.

• The vLLM side — #53909, the “qwen4 fuse op” PR adding HyperConnection, QSA and PLE kernels, still open and unmerged since 2026-08-26; #59279, which adds decode context parallelism to the same QSA path, a draft opened 2026-09-29; and the PLE-offload work that landed across September.

• The SGLang side — #38642 for DFlash hidden-state capture, #39548 for Qwen4-Exp PLE CPU offload on Ascend, #40235 adding host staging for the file-backed PLE table, and now #41855 for the Ascend attention path.

• The NPU enablement line — sglang #37570, adding Qwen3.8-Flash-Next to SGLang on NPU with graph replay, MTP and Triton kernels (opened 2026-09-02, still open, +2,590 lines across 20 files), and sgl-kernel-npu #807 for the companion Triton kernels (opened 2026-09-17, +4,643 lines). Both come from the same contributor. #41855 is the attention layer that sits inside that larger enablement effort.

Two observations a reader can act on. First, the whole Ascend Qwen4Exp story rests on a very small number of contributors — the enablement PRs and the kernel repository share an author, and the attention adapter is a different one. That concentration is a fair estimate of how far Ascend Qwen4Exp serving is from being a supported product path rather than an experiment. Second, the kernel bottleneck is not vendor-specific: two SGLang issues filed on 2026-08-28 document Qwen4Exp decode on an NVIDIA DGX Spark where QSA, PLE and Gated DeltaNet kernel time dominates, and where an NVFP4 KV cache was measured regressing decode by roughly 29% against fp8_e4m3. The attention and embedding layers of this architecture are the hard part everywhere.

What this does not mean

It does not mean Qwen 4 is out, or close. The Qwen 4 family that Alibaba named on 2026-09-22 — Max, Flash, Plus and a 27B tier — remains a roadmap with no model card, no weights, no API identifier, no context window, no licence and no price. A framework adapter that targets the internal architecture name is a step toward serving that family well one day; it is not a step toward the family existing.

It does not mean you can run this today. The PR is a draft with failing CI and no merge date. Even merged, the path needs an Ascend 910C, BF16, CANN 9.0 with torch-npu 2.10, and one of five specific local head shapes, and it is opt-in — meaning a deployment has to choose it. The author also declined to claim server-level validation, which is the part that would actually tell you whether it holds up under real batching.

And it does not mean Qwen3.8-Flash-Next is a supported product on Ascend, or anywhere else in a shipped engine build. The Qwen4Exp paths in both major open runtimes are unmerged pull requests. There is no released SGLang or vLLM version you can install that serves this architecture natively — the FastAPI-style convenience of a hosted endpoint is a different thing from a kernel you can run yourself, and the gap between them is exactly what PRs like this one are for.

What you can actually call while you wait

If the reason you care about Qwen4Exp is that you want to test the architecture’s long-context behaviour rather than its kernel internals, the model to reach for is the tier Alibaba actually serves. Qwen3.8-Flash — the production line built on Qwen3.8-Flash-Next, with official built-in tools and a 1,000,000-token context — is live on OrcaRouter as qwen/qwen3.8-flash: text, image and video input, 131,072-token maximum output, $0.15 per million input tokens and $0.47 per million output, with cache reads at $0.0184. Those are provider list prices passed through with 0% markup on our side, so a vendor price or limit change on it reaches you the same day it is announced. Over the trailing seven-day window the live card shows a p50 first-token latency of 4,416 ms, about 106 output tokens per second and a 2.68% error rate — the profile of a high-volume text tier rather than a lab preview.

Two honest qualifications, and they are the same two the sibling write-ups on this architecture carry. Qwen3.8-Flash-Next itself — the FP8 preview checkpoint, the thing you would need to reproduce any of these kernel measurements locally — is not on our catalogue; the served Flash tier is the production deployment built from it, not the raw preview artifact. And none of the Ascend or DCP work described above exists in anything you can call, because none of it has merged. What the served tier does give you is a cheap way to find out whether your workload is shaped for the problem these kernels solve — long, prefix-heavy, agentic prompts against a very long context. If it is, the throughput and cache behaviour you observe there is the same behaviour the Qwen 4 serving stack will be tuned to protect.

There is also a plumbing argument for not waiting on a family that has no date. Whichever tier ends up winning the Qwen 4 line-up, the switching cost is a routing question rather than an integration project, and one API for 200+ models is how you keep that option open without a second contract or a code change when the weights land. Failover matters for the same reason here in a specific way: if you want to build against an unproven tier, you want the request to fall over to something steady rather than fail when the path you bet on is having a bad minute.

A capture of the OrcaRouter model page for Qwen: Qwen3.8 Flash (qwen/qwen3.8-flash), showing the Qwen provider, a release date of 2026-08-26, the Quick Facts tags General Chat and High Volume, an independent-provider throughput of 106.2 output tok/s and 4,416 ms first-token latency, pricing of $0.23 cache write, $0.0184 cache read, $0.15 input and $0.47 output per 1M tokens, a PERFORMANCE panel for the seven-day window Sep 23 to Sep 30 2026 reading throughput 106.2 tok/s, first-token latency 4.4 s and error rate 3.9%, a traffic chart reading 2.57B tokens last 7 days with +6.9% versus the earlier half, a 0% markup note, and a vendor rate card cross-check block crediting Qwen Cloud, published 2026-08-26 and last checked 2026-09-30 12:08 UTC.

Three questions worth answering directly

Does SGLang #41855 mean Qwen 4 is out, or previewable?

No, on both counts. The pull request targets the Qwen4Exp architecture as implemented in Qwen3.8-Flash-Next, which Alibaba shipped on 2026-08-24. It does not touch Qwen 4 weights, and no Qwen 4 tier has weights to touch. The signal to read here is about serving capacity for a preview architecture, not about availability of the family.

If a model card says qwen4_exp, is that Qwen 4?

No — and this is the naming trap in the whole story. qwen4_exp is the internal architecture identifier, and it is what you will find in config.json for Qwen3.8-Flash-Next and its FP8 sibling. “Experimental architecture” is the operative word: the weights are published, the architecture is real, and the model is a preview of what the Qwen 4 family is expected to be built on. Searching the identifier and finding an SGLang or vLLM PR with Qwen4Exp in the title tells you about engine work, not about a release.

Is this an Ascend-versus-NVIDIA story?

Not really. The same attention path needed a bespoke adapter on the NVIDIA side too — decode context parallelism for QSA in vLLM, and a native sparse prefill kernel for Hopper — and the DGX Spark issues show QSA, PLE and Gated DeltaNet kernel time dominating decode there as well. QSA’s micro-block indexer and its replicated selector state are simply not what generic paged-attention kernels assume. Ascend’s contribution to the pattern is the sharper constraint: a head-ratio tiler that only accepts powers of two, which forces the padding the adapter has to hide.

The thing to watch

Not whether this merges. The adapter is honest about being an adapter, the correctness tests are reproducible without a checkpoint, and the author has flagged the integration question rather than pretending it is settled. The thing to watch is what happens after the NPU enablement line and this attention path are combined — whether the integrated branch gets the end-to-end run that neither of them has had, on real batching with the full model rather than local head-shape tensors. The 1.713× figure is the one that will circulate, and it is the one computed on a different build, on an adapted checkpoint, with an inert indexer. A measured number on the finished stack would be worth considerably more than a historical one.

Until then, the honest summary is the one the PR itself keeps to: the arithmetic checks out, the flag is off by default, the CI is red, and the model in the title still does not exist.