Hero title card for Qwen3.8-27B, the 27-billion-parameter dense vision-language model in Alibaba's Qwen3.8 line-up, with the subtitle '27B dense - native vision-language - Apache 2.0' and three feature chips reading '262,144-token context', 'text + image + video in' and 'text out', with the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Qwen3.8-27B: The Compact Multimodal Model in Alibaba's Qwen3.8 Line-Up

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Qwen3.8-27B is the dense, single-node member of Aliba​ba's Qwe​n3.8 generation: a 27-billion-parameter native vision-language model, published on Hugging Face under Apache 2.0, that accepts text, images and video and returns text with a 262,144-token context window. This is the reference page for that model — the canonical write-up of what it is, what it costs to call, and what it can do — and it is worth saying up front why a page exists for a model whose weights first appeared in August 2026 rather than in the last seven days. We are not reporting a launch. We are answering a standing question: readers reach us for the name "Qwen3.8-27B" thousands of times a month, and until now no page of ours owned that answer. The model's own release, dated on Qwe​n's card and recorded as 2026-08-14 by Artificial Analysis, is a fact this page uses to date the model, not an event it reports.

Read that as the exception it is: on a strict seven-day reading, a model from mid-August 2026 is not recent, and this item would be skipped as news. The warrant here is different. It is a reference page whose numbers were verified against Qwe​n's own model card, the repository's licence file, the Qwe​n Cloud rate card and Artificial Analysis on 2026-09-28, and whose reason to exist is search demand we measured first-party on the same day. It is not a second announcement, and nobody should read it as one.

Where 27B sits in the Qwen3.8 line-up

The Qwe​n3.8 generation is not one model, and the size suffix is the whole point of the name. Qwe​n's own repository notes across the family make the split explicit:

• Qwen3.8-27B — 27B dense parameters, native image and video input, Apache 2.0. The deployment-friendly member: a model sized to run on one GPU or a small node, and the subject of this page.

• Qwen3.8-Max, built on Qwen3.8-2.4T-A95B — 2.4 trillion total parameters with 95B active, the Qwen-Max-class model that Aliba​ba opened for the first time in this generation. Its repository carries a bespoke qwen3.8-max licence rather than Apache 2.0.

• Qwen3.8-Flash-Next — the experimental architecture preview Qwe​n says will underpin Qwen4, released under the qwen-community-1.0 licence. The hosted Qwen3.8-Flash is built on it.

So "27B" is not a trim level of the Max model; it is the size class that a single team can actually serve. What Aliba​ba claims it inherits is the training recipe: the card describes Qwen3.8-27B as taking the generation's advances in coding, professional work, research and long-horizon agentic tasks and compressing them into a dense checkpoint.

Scoreboard infographic titled 'Qwen3.8-27B - the scoreboard' with six rows: Parameters 27B dense (27.8B with vision), Context 262,144 native and 1M with YaRN, Modalities text + image + video in and text out, Licence Apache 2.0 with commercial use, Terminal Bench 2.1 of 73.0 marked vendor-reported, and AA Intelligence Index 33.7 marked independent, with a footer reading 'Qwen figures vendor-reported and unreproduced; AA figures per Artificial Analysis.' and the OrcaRouter logo in the bottom-right corner.

The architecture Qwen publishes

Every figure below is from the model card at Qwe​n/Qwen3.8-27B, which is the vendor's own documentation:

• Parameters — 27B as marketed. The repository's own safetensors metadata totals 27,781,427,952 parameters in BF16 across 18 shards, so roughly 27.8B once the vision tower is counted. There is no mixture-of-experts here: all parameters are active on every token.

• Shape — 64 layers, hidden dimension 5,120, feed-forward intermediate dimension 17,408, token embedding and LM output 248,320 (padded).

• Hybrid attention — the layout Qwe​n prints is 16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN)): three linear-attention blocks for every full-attention block, a 3:1 ratio. Gated DeltaNet runs 48 linear-attention heads for V and 16 for QK at head dimension 128; Gated Attention runs 24 query heads against 4 KV heads at head dimension 256, with a rotary dimension of 64. That ratio is what makes a 262K window affordable on a 27B model, and it is the same lineage Qwen3.8-Flash-Next extends with sparse attention.

• Multi-token prediction — the card states the model is trained with MTP over multiple steps, the drafting trick that speeds up generation in MTP-aware serving stacks.

• Thinking control — thinking mode is on by default and can be switched off per request. reasoning_effort takes xhigh (the default), medium or low, and preserve_thinking is on by default so reasoning traces carry across turns instead of being regenerated.

Context window and output ceiling

The card states 262,144 tokens natively, extensible to 1,000,000 with RoPE scaling (YaRN). Qwe​n's own serving guidance for long agentic runs, inside that 1M envelope, is to allow up to 262,144 tokens for reasoning content and up to 131,072 tokens for the final response — the ceiling is a serving configuration, not a fixed property of the checkpoint.

Two windows are worth keeping apart, because they are easy to conflate. The open weights run at 262,144 natively; the hosted version on Qwe​n Cloud, which the model card describes as still to come ("the service is coming soon"), is listed on the Qwe​n Cloud model page with a 1M context, 991K max input, 131K max output and a 262K max-reasoning budget. The hosted product's spec sheet is not the checkpoint's spec sheet.

Modalities: what goes in and what comes out

Input is text, image and video; output is text only. The card calls Qwen3.8-27B "a native vision-language model that understands images and videos" and its example code passes all three input types. There is no audio input, no image generation and no speech output. Artificial Analysis's model record for the same checkpoint agrees on the envelope — image, video and text in, text out — which is a useful second reading because it is not Aliba​ba's own page.

What the Apache 2.0 release actually permits

The licence is not a summary on a blog post; the repository carries the Apache License, Version 2.0 text itself, and the card's metadata declares license: apache-2.0. What that grants, read from the licence text: a perpetual, worldwide, royalty-free copyright licence to reproduce, prepare derivative works of, publicly display, sublicense and distribute the work and its derivatives, and a patent licence from contributors — with commercial use among the permitted purposes and no field-of-use restriction, no user-count threshold and no separate commercial agreement. Apache 2.0 is also permissive in the sense that matters for shipping products: derivative weights may be redistributed under different terms, provided you keep the licence and attribution notices and state what you changed (Sections 4a–4c). Section 6 grants no rights to the Qwe​n name or trademarks, so an Apache-2.0 fork cannot market itself as Qwe​n.

The comparison inside the family is instructive, and it is visible from the repositories rather than from coverage: the 2.4T-A95B checkpoint calls itself other with a qwen3.8-max licence, and Qwen3.8-Flash-Next uses qwen-community-1.0. The 27B is the one of the three that arrives under plain Apache 2.0.

The independent benchmark position

Artificial Analysis has run the model, and its record is worth separating from Aliba​ba's own table because the two answer different questions. The independent index, measured on the xhigh configuration:

• Intelligence Index 33.7 — Artificial Analysis's stored value, which its model page prints rounded to 34. Two ranks are published and they measure different pools: that page's own summary tile puts the model 1st of the 142 models in its size class, on the same-class comparison the page's tooltip describes, while our catalogue's Artificial Analysis–sourced row records 45th of 145, about the 67th percentile. Neither is a rank against the same field as the other, so do not read them as one number in two formats.

• Coding Index 68.1 — carried in our catalogue with artificialanalysis.ai as its named source, ranked 35th of 138 in that index's pool. The model sits higher on coding than on general intelligence, which is consistent with the shape of the vendor's own table.

• Effort ladder, same harness — dropping reasoning effort moves the index a long way: 33.7 at xhigh, 27.6 at medium, 26.2 at low, and 20.2 with reasoning off entirely. That is a measured spread of 13.5 points, and it is the most concrete argument for tuning effort per workload rather than per model.

• Where the two sources agree and disagree — on GPQA Diamond, Aliba​ba's card reports 89.2 and Artificial Analysis measures 90.5. On Humanity's Last Exam the card reports 30.8 and Artificial Analysis 33.9. Close, but not the same number, and that is normal: different harnesses, different runs. One caution belongs in the article rather than a footnote — no third party has published a head-to-head between Qwen3.8-27B and any of the models in Aliba​ba's comparison table run at matched effort on a matched harness, so every cross-model comparison on this page is a comparison of separate measurements, not a result.

Screenshot of the Artificial Analysis model page for Qwen3.8 27B in its xhigh configuration showing a same-class comparison rank of 1 of 142 next to an Intelligence tile, 46.6 output tokens per second at rank 55 of 142, $0.50 per million input tokens and $3.00 per million output tokens, a 256k-token context window, 27B total parameters, an Apache 2.0 licence, an open-weights marker and a release month of August 2026, with the summary text reporting a score of 34 on the Artificial Analysis Intelligence Index.

What Qwen reports, and how to read it

The model card carries roughly two dozen scores with comparison columns for Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B and Opus4.6 Max. All of it is vendor-reported and unreproduced — Aliba​ba's own evaluation, printed by Aliba​ba — and several rows use a Claude Code harness judged by a GPT-5.4 build, which the card's footnotes state. The headline vendor figures, on that basis:

• Agentic terminal coding — Terminal Bench 2.1 (Terminus) 73.0, against 63.4 for the Qwen3.6-27B it replaces, with Opus4.6 Max leading the row at 78.2.

• Agentic coding — SWE-bench Pro 61.7, QwenSWEBench 79.0, DeepSWE 1.1 42.2, NL2Repo-Bench 42.3, LiveCodeBench v6 90.3.

• Agents and office work — CoWorkBench 70.7, JobBench 33.4, Agents' Last Exam 42.9 by score (Pass@1 20.4).

• General reasoning — GPQA Diamond 89.2, HLE 30.8, IFBench 79.5. Opus4.6 Max leads both reasoning rows in the vendor's own table, at 91.3 and 40.0.

• Vision and computer use — OSWorld-Verified 84.3, AndroidWorld 81.9, WebArena-Verified 64.8, OmniDocBench 1.5 91.1, RealWorldQA 85.9.

The honest summary of that table is narrower than the table looks. Against its direct predecessor the gains are large and consistent — QwenSWEBench 79.0 versus 49.3, DeepSWE 42.2 versus 13.3 — and that is a claim about one generation of recipe change. Against the frontier models in the neighbouring columns it wins some rows and loses others, on Aliba​ba's own numbers, by Aliba​ba's own choice of harness.

Running it

The weights are 18 BF16 shards in the Transformers format, and the card lists Hugging Face Transformers, vLLM, SGLang and TokenSpeed as supported stacks. We have written the local-deployment and quantisation paths separately and they are the right places for that detail: Qwen3.8-27B on Ollama for a local daemon, and the benchmarks walkthrough for the full scorecard including what changed from Qwen3.6-27B. This page stays on the model itself.

Qwen3.8-27B on our catalogue

Two cards are live on OrcaRouter today, side by side, and both were re-read on 2026-09-28 before this paragraph was written. qwen/qwen3.8-27b lists at $0.33 per million input tokens and $2.40 per million output tokens with the full 262,144-token window, text, image and video input, and the reasoning, tools, structured-output and JSON surfaces exposed as parameters. Its sibling obsidian/Qwen3.8-27B — the same weights served in block-FP8 with the vision tower kept at full precision, and, in the card's own words, "designed to provide direct, complete responses across a wide range of prompts" — lists at $0.40 in and $4.21 out. Both answer to chat completions and the Responses API, so a document-parsing or screenshot job can be pointed at either model under one key and one endpoint, with no second contract and no code change.

For scale, our own playground has moved 105.1 million tokens through qwen/qwen3.8-27b in the last seven days, with a 3.8-second median time to first byte and a 3.3% error rate over the same window. Those are our serving numbers, not a verdict on the model.

Screenshot of the OrcaRouter model page for qwen/qwen3.8-27b showing a 262K-token context window, text, image and video input, text output, the vision, tools, JSON and reasoning badges, a 3.80-second p50 time to first token, 105.1M tokens served in seven days, and list pricing of $0.33 per million input tokens and $2.40 per million output tokens.

The searches that lead here

The demand behind this page is spread thin and sits just outside the click. In the 28 days to 2026-09-25 our property took 8,931 impressions and 273 clicks across 69 distinct exact spellings of the model's name — "qwen3.8 27b", "qwen3.8-27b", "qwen 3.8 27b" and dozens more ways of writing the same three tokens. Our best position on the largest spellings ran 9.0 to 10.6. Those are positions we held by accident of other pages, which is exactly the problem a canonical page fixes: at ninth place you are on the page, and nobody clicks. The one place the traffic converts is non-English — a Russian-language spelling of the name sits at position 1.8 for us and converts at 14.4%.

None of that is evidence about the model. It is evidence about what people type, and it is why the page exists.

What it adds up to: a 27B dense checkpoint with a genuinely unusual attention layout, a permissive licence, native video understanding, Apache-2.0 terms that let you ship a derivative, an independent coding rank in the top quarter of what Artificial Analysis measures, and a vendor table that is strong against its predecessor and mixed against the frontier. The open question is the one nobody outside Aliba​ba can answer yet — how the 1M-token hosted path behaves in production, and whether the Flash-Next architecture lands in a model of this size.

The 2.4T-A95B core behind Qwen3.8-Max is live on the same catalogue: qwen/qwen3.8-max at $2.00 / $6.00 per million tokens, if the 27B is not the size you need.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily