
Qwen3.8-27B Open Weights: The Max Shipped, the 27B Didn't — What We Know Now
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 36 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 181 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1277 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 110 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
On August 13, 2026, the open-weights week Alibaba promised for its Qwen3.8 generation split in two. The Max-class flagship's weights — Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter sparse mixture-of-experts model and the first Max-level Qwen release ever made downloadable — actually shipped, on Hugging Face and ModelScope, announced late the previous night through the ModelScope community. The smaller model this page is about, Qwen3.8-27B, did not: committed for the same week of August 10, it still has no official repository, model card, license file, or benchmark of its own, and third-party trackers now describe its release as delayed with no new date given. The question this what-we-know-so-far page was built to answer — "when does the 27B actually drop?" — has become two: why did the Max's weights ship first, and what happened to the 27B?
This remains a what-we-know-so-far piece, not a review. Qwen3.8-Max, the API flagship, has been live since August 3 at $2/$6 per million tokens; its underlying open-weights base, Qwen3.8-2.4T-A95B, became downloadable on August 13; and the Qwen3.8-27B weights are still unpublished. Every claim below is labeled confirmed, vendor-reported, or community-estimated accordingly.
The signal, and what the week actually delivered
The early signal was a post on X by @kimmonismus, who flagged "next week" as the one to watch: "Qwen-3.8 27b - Grok 4.6," with Astra conspicuously absent from the list. The Grok half of that week shipped on August 12. The Qwen half split in two. What the post captured was the mood among people who watch model releases for a living — the Qwen3.8 open-weights drop was the event on the calendar, and the week arrived with the specifics still not nailed down. A week later, the specifics have partly resolved: one of the two promised open-weight drops has landed, and it was not the one the post named.
What's confirmed
• The family is real and officially announced. Alibaba unveiled the Qwen3.8 base-model family on August 3, 2026, after a preview iteration in mid-July.
• Qwen3.8-Max shipped the same day. The API flagship — a 2.4-trillion-parameter mixture-of-experts model with roughly 95B active parameters, a 1M-token context window, and text, image, and video input — went live on the Qwen API at $2/$6 per million tokens. It is also available through OrcaRouter at that same list price, passed through at 0% markup.
• The Max's open weights shipped on August 13. Qwen3.8-2.4T-A95B — the base model the Qwen3.8-Max API is built on — became downloadable from Hugging Face and ModelScope, with an FP8 quantized version alongside it. This is the first Max-level Qwen open-weights release ever. The downloadable model is text-only with thinking mode forced on, carries a native 256K-token context expandable to about 1M, and deploys on SGLang, vLLM, and TokenSpeed.
• Qwen3.8-27B is the family's smaller open-weight model — still not shipped. Alibaba positioned it as the realistic path for local and on-premise deployment and committed its weights to Hugging Face and ModelScope alongside the Max's. As of August 13 there is still no official Qwen3.8-27B repository. The commitment was real; the delivery has not happened yet.
• The first independent score for the generation exists. Artificial Analysis puts Qwen3.8-Max at an Intelligence Index of 56 at roughly $1.14 per task — genuinely good, with the caveats any days-old flagship carries.
What we actually know about the 27B
The honest summary is that "Qwen3.8-27B" is still mostly a name and a parameter count. The confirmed facts are that it is a roughly 27-billion-parameter model, that it is the open-weights member of the Qwen3.8 generation, and that it is sized for single-GPU and on-premise work rather than a multi-hundred-GPU cluster. Unsloth's Daniel Han reports it should fit in roughly 17GB of VRAM on release — a single prosumer card.
Everything else is unconfirmed. Alibaba has not published whether the 27B is a dense model or a mixture-of-experts, its context window, which modalities it accepts, or any benchmark scores of its own. What changed this week is the timing: the Max and the 27B were promised for the same week, and only the Max arrived. Third-party trackers now report the 27B as delayed with no new date — a community read, not an official Alibaba statement, and it is possible the repo still lands before the promised week is fully out. But as of this writing the official Qwen organization on Hugging Face lists no Qwen3.8-27B, and the handful of "Qwen3.8-27B-FP8", "-GGUF", and "-MLX" repositories that surface in search are community placeholder cards with single-digit download counts, not the official release.
The useful reference point is still the previous generation's 27B: Qwen3.5-27B, an open-weight dense model with a 32K context window. It is a reasonable lower bound for what the Qwen3.8-27B has to beat, and a reminder that Qwen 27B-class models have historically been built to run on hardware a developer actually owns.
The Max is the reference point for what the 27B might inherit

The 27B will not match the Max's raw ceiling, but the Max's numbers — now published in a real model card — set expectations for what the generation's improvements look like. On Alibaba's own evaluations of the open-weights model, Qwen3.8-2.4T-A95B scores 93.0 on PaperBench, 86.1 on OSworld-Verified, 91.5 on parametric CAD, 67.7 on SWE-bench Pro, and 86.6 on Terminal-Bench 2.1, with the vendor claiming parity with Claude Opus 4.8, Claude Fable 5, GPT-5.6 Sol, and Gemini 3.1 Pro. These are vendor-reported figures, published in the release materials and not yet reproduced by an independent lab. The agentic gain the API launch claimed — FrontierSWE jumping from 40.7 to 73.5 in a single generation — carries into the same release materials. If even part of that improvement flows down to a 27B, it would reset what "good local coding model" means in its class.
One nuance to hold on to: the downloadable Max is not identical to the API Max. Qwen3.8-2.4T-A95B is text-only with thinking mode forced on, where the Qwen3.8-Max API adds vision input, a non-thinking mode, and a default 1M-token context. That gap is exactly the kind of thing a 27B is likely to inherit — and worth remembering when you read "Qwen3.8" benchmark headlines without checking which variant they describe.
Price is where the Max is most defensible: $2/$6 per million tokens across the whole 1M-token context on the API, with no long-prompt surcharge. That is aggressive for a frontier API, and it matters for the routing math — at those numbers the 3.8 generation becomes an option for workloads that used to default to more expensive flagships.
What can actually run it

For the 27B, Alibaba has still not published hardware requirements, so the numbers in circulation remain community projections based on the Qwen 3.6-27B quant table — estimates, not specs. The working ranges most people are planning around are unchanged:
• Q4_K_M quant, ~16GB VRAM — an RTX 4090 or equivalent, the sweet spot most local developers will target.
• Q3_K_M quant, ~13GB VRAM — fits a 12GB card like the RTX 3060 at reduced quality.
• Q6_K quant, ~21GB VRAM — near-lossless on 24GB cards.
• FP8 serving, ~27GB VRAM — a single L40S, per the earlier projection, for production-grade inference.
• Full precision, H100 80GB — the benchmark-quality path, and not something most teams will run.
The new wrinkle is that "you can self-host the Qwen3.8 generation now" is true, but not for the model people were planning around. The Max's open weights are a different hardware class entirely: the full-precision BF16 model is roughly 4.9TB, and Unsloth's 1-bit layered selective quantization brings it to about 397GB — runnable, but only with 410GB or more of combined RAM and VRAM. That is serious infrastructure, not a prosumer card. It is exactly the gap the 27B is meant to fill, which makes the 27B's continued absence the reason the single-GPU local path for this generation does not exist yet.
The license question: the Max just set the precedent
The license question for the 27B is still unanswered, but "what would it look like" is no longer a guess. The Max's open weights shipped under a custom license labeled "qwen3.8-max" — not Apache 2.0. As reported from the model card, it is generally free for commercial use and hosting, but it triggers additional requirements at scale: products with more than 100 million monthly active users or more than $20 million in monthly revenue must display the model name, and companies with more than $50 million in annual revenue that offer model-as-a-service or AI work-assistant services need separate licensing. A claim circulating that the license bans use in the US, EU, UK, and Korea is false — the license contains no territorial restrictions.
For the 27B specifically, no license has been named. But the Max shipped under a custom Qwen license rather than Apache 2.0, which makes Apache 2.0 a poor assumption for the sibling model too. The old Tongyi Qianwen license — with its 100-million-MAU clause — is not what the Max got; the Max's is a new, scale-tiered license. Read the LICENSE file in the actual repository before you plan around it, and expect the 27B's to be decided only when its repo ships.
Toolchain timing: day-one support just got proven
The serving engines moved as fast as expected — for the model that shipped. Qwen3.8-2.4T-A95B launched with working support in SGLang, vLLM, and TokenSpeed, which is what "API-servable almost immediately" looks like for a large open-weights release. For the 27B, the same engines are the most likely first movers whenever its repo appears, and the Max's day-one coverage makes the 27B inheriting that support the reasonable expectation. Community GGUF and AWQ quantizations still typically lag by one to two weeks, so the Ollama-style "pull and run" experience probably still will not exist on the 27B's release day — a manual vLLM or llama.cpp setup is the likely local path at first.
What's still unknown
An honest leak piece stops at what is actually public, and the boundary is wider than it was a week ago in one direction and narrower in another: the Max's questions got answered, the 27B's did not.
• The new drop date. "The week of August 10" was the commitment; the week is nearly over and the repo has not appeared. Third-party trackers report a delay with no new date, but Alibaba has made no official statement.
• Architecture. Dense or MoE for the 27B, and if MoE, the active-parameter count.
• The spec sheet. Context window, modalities, output limits, reasoning mode.
• Benchmarks. The 27B has no official scores and no independent evaluation.
• The license. Undecided until the 27B's repository ships; the Max's custom license is the precedent to watch against.
• Mirror risk. The placeholder problem is now visible in the wild — community "Qwen3.8-27B-FP8", "-GGUF", and "-MLX" cards exist with zero official files. Until the official Qwen organization publishes the repo, download only from the official org page.
What it means for builders
The practical split for the Qwen3.8 generation is now three-way instead of two. If you want the frontier API today, Qwen3.8-Max is live through OrcaRouter at the provider's $2/$6 list price, passed through at 0% markup — the price you see is the price Alibaba set. If you want to self-host the Max-class model, the open weights are downloadable now, but you need the hardware for a roughly 4.9TB full-precision model or a roughly 397GB quantized one. And if you want the single-GPU local deployment this generation was expected to provide, you are still waiting on Qwen3.8-27B, which has not shipped. One OpenAI-compatible key covers the API side today, and when the 27B's hosting settles, the same key can route to it.

That last part is the more general point, and it is exactly what an unproven, delayed open-weight release is the test case for. A model that has not even shipped yet has no track record, no independent benchmark suite, and no incident history — and when it does arrive, it will be brand-new. Routing is how you try it without betting a customer-facing path on it: send it the traffic that can tolerate surprise, fail over automatically to a proven model when it misbehaves, and switch the whole workload once it has earned trust. The routing layer is also what makes the price math honest — when Alibaba cuts the price of a model you call through one endpoint, the change is live the same day, because the list price passes straight through.
What to watch
• Whether the 27B's repo appears at all. A listing under the official Qwen organization on Hugging Face or ModelScope is the event that ends the delay — and the week of August 10 has nearly closed without one.
• The 27B's license file. The Max shipped with a custom scale-tiered license, not Apache 2.0. Whether the 27B follows that precedent is the single most consequential line of its eventual announcement.
• Architecture and context length. Dense-vs-MoE and the context window decide what the 27B is for.
• The first independent benchmark. The Max's 56 on the Artificial Analysis Intelligence Index set the bar; the 27B's first outside evaluation will tell you how much of the generation's gains survive at consumer size.
• Toolchain coverage. Whether the 27B inherits the Max's day-one SGLang/vLLM/TokenSpeed support, and how quickly community quantizations follow.
FAQ
Has Qwen3.8-27B been released yet?
No. Alibaba announced it on August 3, 2026, and committed the weights to Hugging Face and ModelScope for the week of August 10, but as of August 13 there is no official Qwen3.8-27B repository, model card, or license. The Max-class open weights — Qwen3.8-2.4T-A95B — shipped on August 13; the 27B did not, and third-party trackers report it as delayed with no new date.
What hardware do I need to run Qwen3.8-27B?
Unconfirmed, but community projections based on the Qwen 3.6-27B quant table put a Q4_K_M quant at roughly 16GB of VRAM (an RTX 4090), a Q3_K_M quant at roughly 13GB (12GB cards), a Q6_K quant at roughly 21GB (24GB cards), and FP8 serving at roughly 27GB (a single L40S). Unsloth's Daniel Han reports about 17GB on release. Treat these as estimates until the official model card publishes hardware guidance.
Is Qwen3.8-27B available through OrcaRouter?
Not yet. Qwen3.8-27B is an open-weight, self-host model, and it is not on the platform today. The flagship of the same family — Qwen3.8-Max — is available now through OrcaRouter at the provider's $2/$6 per million tokens list price with 0% markup, and when the 27B's hosting settles it can be added to the same OpenAI-compatible key.
Should I wait for Qwen3.8-27B or use Qwen3.8-Max today?
It depends on where the workload runs. If you need a frontier API now, Qwen3.8-Max is live and aggressively priced at $2/$6 per million tokens. If you want to self-host the Max class today, Qwen3.8-2.4T-A95B is downloadable now, but it needs serious hardware — roughly a 4.9TB full-precision model, or about 397GB quantized, with 410GB+ of combined RAM and VRAM. If you need a local or on-premise model on a single consumer card, Qwen3.8-27B is still the one to wait for — it is the realistic self-host path in this generation, and its drop is now the open question rather than a fixed date.
The promised week has nearly played out, and half of it did not arrive. Qwen3.8-2.4T-A95B is downloadable, licensed, and servable; Qwen3.8-27B is confirmed to exist, still committed to open weights, and still without a repository, a license, a benchmark, or a new date. Watch whether the 27B's repo lands before the promised week is fully out, what its license file says next to the Max's, and how the first independent evaluations score it. Those three signals will tell you whether this is a routine two-step open-weight release with the smaller model still to come, or a delay worth planning around.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
