
Qwen3.8-27B Open Weights: Everything We Know So Far Before the Drop
- metaNOWOŚĆMeta: Muse Spark 1.22026-08-0557Inteligencja72Kod
- qwenNOWOŚĆQwen: Qwen3.8 Max2026-08-0358Inteligencja72Kod
- deepseekNOWOŚĆDeepSeek: DeepSeek V4 Flash 07312026-07-3152Inteligencja69Kod
- qwenNOWOŚĆQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 za 1 mln tokenów · 2076 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Inteligencja78Kod
- googleGoogle: Gemini 3.6 Flash2026-07-2152Inteligencja69Kod
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Inteligencja49Kod
- metaMeta: Muse Spark 1.12026-07-1653Inteligencja71Kod
- kimiMoonshotAI: Kimi K32026-07-1560Inteligencja76Kod
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Inteligencja71Kod
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Inteligencja77Kod
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Inteligencja77Kod
- grokxAI: Grok 4.52026-07-0856Inteligencja72Kod
- tencentTencent: Hy32026-07-0642Inteligencja59Kod
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Inteligencja42Kod
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Inteligencja39Kod
- anthropicAnthropic: Claude Sonnet 52026-06-3055Inteligencja72Kod
- klingKling: Kling 3.0 Turbo2026-06-1757Inteligencja52Kod57Matematyka
- z-aiZ.ai: GLM 5.22026-06-1653Inteligencja69Kod60Matematyka
Qwen3.8-27B is the open-weight release everyone has had on their calendar this week — and, as of this writing, it is still a promise rather than a download. Alibaba announced the 27B model on August 3 alongside its flagship Qwen3.8-Max, saying both sets of weights would be published "within the week" on Hugging Face and ModelScope. That week is now here, and a search of the Hugging Face model hub for "qwen3.8" still turns up no official repository. The gap between the announcement and the payload is the whole story right now.
This is a what-we-know-so-far piece, not a review. The Qwen3.8 family is officially announced and the Qwen3.8-Max flagship is already shipping on Alibaba's own API, but the Qwen3.8-27B weights are not yet published, no license has been named, and most of the 27B's spec sheet is still empty. Every claim below is labeled as confirmed, vendor-reported, or community-estimated accordingly.
The signal: a busy week, with Qwen at the front of it
The early signal is a post on X by @kimmonismus, who flagged "next week" as the one to watch: "Qwen-3.8 27b - Grok 4.6," with Astra conspicuously absent from the list. The Grok half of that week has its own story; this article is about the Qwen half. What the post captures is the mood among people who watch model releases for a living: the Qwen3.8 open-weights drop is the event on the calendar, and the week has arrived with the specifics still not nailed down.
What's confirmed
• The family is real and officially announced. Alibaba unveiled the Qwen3.8 base-model family on August 3, 2026, after a preview iteration in mid-July.
• Qwen3.8-Max shipped the same day. The flagship — a 2.4-trillion-parameter mixture-of-experts model with roughly 95B active parameters, a 1M-token context window, and text, image, and video input — went live on the Qwen API at $2/$6 per million tokens. It is also available now through OrcaRouter at that same list price.
• Qwen3.8-27B is the family's smaller, open-weight model. Alibaba positioned it as the realistic path for local and on-premise deployment, and committed its weights to Hugging Face and ModelScope alongside the Max's — the first time a Max-generation Qwen family has carried an open-weights promise.
• The first independent score for the generation exists. Artificial Analysis puts Qwen3.8-Max at an Intelligence Index of 56 at roughly $1.14 per task — genuinely good, though it comes with the caveats any days-old flagship carries.
What we actually know about the 27B
The honest summary is that "Qwen3.8-27B" is mostly a name and a parameter count so far. The confirmed facts are that it is a roughly 27-billion-parameter model, that it is the open-weights member of the Qwen3.8 generation, and that it is sized for single-GPU and on-premise work rather than a multi-hundred-GPU cluster. Unsloth's Daniel Han reports it should fit in roughly 17GB of VRAM on release — a single prosumer card.
Everything else is unconfirmed. Alibaba has not published whether the 27B is a dense model or a mixture-of-experts, its context window, which modalities it accepts, or any benchmark scores of its own. Until the model card appears, treat every one of those cells as empty.
The useful reference point is the previous generation's 27B: Qwen3.5-27B, an open-weight dense model with a 32K context window. It is a reasonable lower bound for what the Qwen3.8-27B has to beat, and a reminder that Qwen 27B-class models have historically been built to run on hardware a developer actually owns.
The Max is the reference point for what the 27B might inherit

The 27B will not match the Max's raw ceiling, but the Max's numbers set expectations for what the generation's improvements look like. On Alibaba's own evaluations, Qwen3.8-Max scores 92.6 on GPQA Diamond, 86.6 on Terminal-Bench 2.1, and 67.7 on SWE-bench Pro — all vendor-reported and none independently confirmed yet. The agentic gain is the headline: FrontierSWE jumped from 40.7 to 73.5 in a single generation. If even part of that agentic improvement flows down to the 27B, it would reset what "good local coding model" means in its class.
Price is where the Max is most defensible: $2/$6 per million tokens across the whole 1M-token context, with no long-prompt surcharge. That is aggressive for a frontier API, and it matters for the routing math — at those numbers the 3.8 generation becomes an option for workloads that used to default to more expensive flagships.
What can actually run it

Because Alibaba has not published hardware requirements, the numbers in circulation are community projections based on the Qwen 3.6-27B quant table — treat them as estimates, not specs. The working ranges most people are planning around:
• Q4_K_M quant, ~16GB VRAM — an RTX 4090 or equivalent, the sweet spot most local developers will target.
• Q3_K_M quant, ~13GB VRAM — fits a 12GB card like the RTX 3060 at reduced quality.
• Q6_K quant, ~21GB VRAM — near-lossless on 24GB cards.
• FP8 serving, ~27GB VRAM — a single L40S, per the earlier projection, for production-grade inference.
• Full precision, H100 80GB — the benchmark-quality path, and not something most teams will run.
For context on what runs today: Qwen 3.6's 27B dense model serves from 8GB VRAM up, and the 35B MoE with 3B active is the current coding pick at 24GB. The Qwen3.8-27B inherits a well-worn local-deployment path.
The license question no one can answer yet
Apache 2.0 is not a safe assumption. Alibaba has named no license for the Qwen3.8 weights, and several prior Qwen releases shipped under the Tongyi Qianwen license, which includes a 100-million-monthly-active-users threshold that triggers a commercial conversation. That clause is precisely the kind of detail that decides whether "open weights" means "safe to build a product on" or "fine for personal use, legally complicated at scale." Read the LICENSE file in the actual repository before you plan around it.
Toolchain timing: what day-one support will look like
The serving engines move fastest. vLLM and SGLang are the most likely to have working support within days of the weight drop, which means the open-weights model should be API-servable almost immediately for teams that run their own stack. Community GGUF and AWQ quantizations typically lag by one to two weeks, so the Ollama-style "pull and run" experience is probably not available on release day — the local path will be a manual vLLM or llama.cpp setup at first.
What's still unknown
An honest leak piece stops at what is actually public, and the boundary here is wide.
• The exact drop date. "This week" is the commitment; there is no confirmed day, and the repo has not appeared yet.
• Architecture. Dense or MoE, and if MoE, the active-parameter count.
• The spec sheet. Context window, modalities, output limits, reasoning mode.
• Benchmarks. The 27B has no official scores and no independent evaluation.
• The license. Apache 2.0, Tongyi Qianwen, or something else — undecided until the repository ships.
• Mirror risk. Until the official Qwen organization publishes the repo, any "Qwen3.8-27B" you can download is either a placeholder or a third-party upload — download only from the official org page.
Co to oznacza dla budowniczych
If you are planning around the Qwen3.8 generation, the practical split is simple: the Qwen3.8-Max API is available now, and the Qwen3.8-27B is the self-host play arriving this week. You can evaluate the generation's behavior today without waiting — Qwen3.8-Max is live through OrcaRouter at the provider's list price of $2/$6 per million tokens, passed through at 0% markup, so the price you see is the price Alibaba set, with no platform spread on top. One OpenAI-compatible key covers it, and when the 27B's hosting settles, the same key can route to it.

That last part is the more general point, and it is exactly what a brand-new, unproven open-weight release is the test case for. A model that dropped this week has no track record in production, no independent benchmark suite, and no incident history. Routing is how you try it without betting a customer-facing path on it: send it the traffic that can tolerate surprise, fail over automatically to a proven model when it misbehaves, and switch the whole workload once it has earned trust. The routing layer is also what makes the price math honest — when Alibaba cuts the price of a model you call through one endpoint, the change is live the same day, because the list price passes straight through.
Co obejrzeć
• Whether the repo appears this week. A Hugging Face and ModelScope listing under the official Qwen organization is the event that turns this from a leak into a release.
• The license file. The single most consequential line of the announcement.
• Architecture and context length. Dense-vs-MoE and the context window decide what the model is for.
• The first independent benchmark. The Max's 56 on the Artificial Analysis Intelligence Index set the bar; the 27B's first outside evaluation will tell you how much of the generation's gains survive at consumer size.
• Toolchain coverage. Whether vLLM and SGLang land day-one support, and how quickly community quantizations follow.
Często zadawane pytania
Has Qwen3.8-27B been released yet?
No. Alibaba announced it on August 3, 2026, and committed the weights to Hugging Face and ModelScope "this week" — but as of this writing there is no official Qwen3.8-27B repository, no model card, and no license. It is announced and imminent, not released.
What hardware do I need to run Qwen3.8-27B?
Unconfirmed, but community projections based on the Qwen 3.6-27B quant table put a Q4_K_M quant at roughly 16GB of VRAM (an RTX 4090), a Q3_K_M quant at roughly 13GB (12GB cards), a Q6_K quant at roughly 21GB (24GB cards), and FP8 serving at roughly 27GB (a single L40S). Unsloth's Daniel Han reports about 17GB on release. Treat these as estimates until the official model card publishes hardware guidance.
Is Qwen3.8-27B available through OrcaRouter?
Not yet. Qwen3.8-27B is an open-weight, self-host model, and it is not on the platform today. The flagship of the same family — Qwen3.8-Max — is available now through OrcaRouter at the provider's $2/$6 per million tokens list price with 0% markup, and when the 27B's hosting settles it can be added to the same OpenAI-compatible key.
Should I wait for Qwen3.8-27B or use Qwen3.8-Max today?
It depends on where the workload runs. If you need a frontier API now, Qwen3.8-Max is live and aggressively priced at $2/$6 per million tokens. If you need local or on-premise deployment, the Qwen3.8-27B is the model to wait for — it is the realistic self-host path in this generation, and the drop is imminent.
The week has not played out yet, and that is the point of a what-we-know-so-far piece. Qwen3.8-27B is confirmed to exist, committed to open weights, and hours-to-days away from an official repository — with the architecture, license, context length, and benchmarks still unwritten. Watch whether the repo lands this week, what the license file says, and how the first independent evaluations score it. Those three signals will tell you whether this is a routine open-weight release or a genuine reset of what a 27B can do.
Porównane w tym artykule1
Wykryto na podstawie tego artykułu · Benchmarki: Artificial Analysis · aktualizowane codziennie
