Title card for the Qwen3.8-27B benchmark roundup, showing a benchmark report card with a seal reading OPEN WEIGHTS and icons for coding, throughput, vision, long context and local hardware, with chips reading 27B DENSE, 262K CONTEXT and APACHE 2.0.
Guides & Insights

Qwen3.8-27B Benchmarks: The Full Table, With the Caveats Attached

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Qwen/Qwen3.8-27B scores Terminal-Bench 2.1 at 73.0, SWE-bench Pro at 61.7, LiveCodeBench v6 at 90.3 and OSWorld-Verified at 84.3 — and every one of those numbers is Alibaba's own. The 27-billion-parameter open-weights model hit Hugging Face on August 14, 2026 under Apache 2.0, and within its first two weeks it picked up three independent third-party results: the #1 open-weight spot on Harvey's Legal Agent benchmark, in a study run by legal-AI firm Harvey and memory company Engram; a #9 overall ranking on Code Arena's WebDev leaderboard, the only model in its size class inside the top 10, per Alibaba's announcement on August 25; and, announced on August 26, the #1 open-weight spot on Arena.ai's Image-to-WebDev leaderboard, ranked #7 overall with a reported score of 1,574. This page is the complete, sourced rundown: all 25 official benchmarks from the model card, what changed versus its predecessor Qwen3.6-27B, what an independent AMD throughput test measured, those legal-agent and webdev-arena results, and which claims you should still treat as provisional.

A reading guide before the numbers: "official" here means Alibaba's evaluation printed on the model card at Qwen/Qwen3.8-27B. It is vendor-reported and unreproduced. "Independent" means someone outside the lab measured it — so far that covers a hardware-throughput test, a legal-domain agent study, and placements on two of Arena.ai's web-development leaderboards, Code Arena's WebDev and the Image-to-WebDev arena. Benchmarks whose harness matters are flagged inline.

The model, one line per spec

• 27B dense parameters (28B counting the vision encoder), 64 layers, hidden size 5120.

• Hybrid attention — 48 Gated DeltaNet linear-attention layers plus 16 full-attention layers, a 3:1 ratio — which is what makes a 262K-token window affordable to run.

• 262,144 tokens native context, extensible to 1,000,000 with YaRN scaling.

• Native image and video input (text output), Apache 2.0 weights, roughly 55.6GB in BF16.

The short version: this is the model meant to take what Alibaba learned building the 2.4-trillion-parameter Qwen3.8-Max and compress it into something a single GPU can run.

The scorecard: every benchmark on the model card

All scores below are Alibaba-reported from the August 14 model card, with the Qwen3.6-27B score the model replaces in parentheses.

Coding. Terminal-Bench 2.1 (agentic terminal coding) 73.0 (63.4); SWE-bench Pro 61.7 (53.5); QwenSWEBench 79.0 (49.3); DeepSWE 1.1 42.2 (13.3); NL2Repo-Bench 42.3 (36.2); LiveCodeBench v6 90.3 (83.9). The SWE-bench Pro and QwenSWEBench runs used the Claude Code harness, per a model-card footnote — worth keeping in mind before comparing them to a model scored on a different harness.

Long-horizon agents and office work. CoWorkBench 70.7 (61.0); JobBench 33.4 (21.8); Agents' Last Exam Pass@1 20.4 / score 42.9 (10.6 / 27.3).

Reasoning and knowledge. GPQA Diamond 89.2 (87.8); Humanity's Last Exam 30.8 (24.0); IFBench instruction-following 79.5 (69.1). HLE is GPT-4o-judged, per the card.

Vision and computer use. OSWorld-Verified 84.3 (63.9); WebArena-Verified 64.8 (48.8); AndroidWorld 81.9 (70.3); MathVision 94.6 with CI / 90.0 without (85.1); BabyVision 85.6 with CI / 65.7 without (28.9); OmniDocBench 1.5 91.1 (89.4); RealWorldQA 85.9 (84.1). "CI" is Alibaba's inference-time enhancement; the gap between its on and off states is the single most informative number on the card about how much score you can buy with extra inference compute.

A scoreboard card for Qwen3.8-27B, listing Terminal-Bench 2.1 73.0 (was 63.4), SWE-bench Pro 61.7 (was 53.5), LiveCodeBench v6 90.3 (was 83.9), GPQA Diamond 89.2 (was 87.8), OSWorld-Verified 84.3 (was 63.9) and native context 262K (1M via YaRN), with a footer reading Alibaba-reported, Aug 14 2026 model card — no independent scores yet.

What actually changed versus Qwen3.6-27B

Every benchmark on the card went up, but the deltas are not uniform — and the shape of the improvement is the story. The gains cluster in agentic coding and computer use, not in knowledge recall:

• DeepSWE 1.1 (agentic coding) 13.3 → 42.2, roughly a 3x jump

• QwenSWEBench 49.3 → 79.0

• Agents' Last Exam score 27.3 → 42.9

• OSWorld-Verified 63.9 → 84.3 and AndroidWorld 70.3 → 81.9

• BabyVision (with CI) 28.9 → 85.6 — the largest relative gain on the card

Meanwhile GPQA Diamond moved 87.8 → 89.2 and RealWorldQA 84.1 → 85.9: real but small. This is not a "smarter at everything" release. It is a release aimed squarely at agents that read screens, use tools, and write code across many steps.

A card showing the biggest Qwen3.8-27B benchmark gains over Qwen3.6-27B: DeepSWE 1.1 13.3 to 42.2, QwenSWEBench 49.3 to 79.0, Agents' Last Exam score 27.3 to 42.9, OSWorld-Verified 63.9 to 84.3, BabyVision with CI 28.9 to 85.6, WebArena-Verified 48.8 to 64.8, with a footer reading Delta vs Qwen3.6-27B, Alibaba-reported.

Since this page went up, the most important development is legal, not technical. Alibaba announced that Qw​en3.8-27B is the #1 open-weight model on Harvey's Legal Agent benchmark — a ranking claim from the vendor's own account, so treat it as Alibaba-stated. The result behind it is a study that is not Alibaba's: Harvey, the legal-AI company, and Engram, an AI-memory startup, built a synthetic law firm called Calderwood & Harkness out of 100M+ tokens across roughly 10,000 documents and 266 matters, then adapted Qw​en3.8-27B to it with parametric memory, compressed notes, and a retrieval tool.

On 250 legal tasks the adapted 27B averaged 67%, ahead of every model the study tested — including Claude Opus 4.8 — and on strict all-or-nothing correctness it led Opus 4.8 30% to 25%, at roughly $0.13 per query versus $1.32 for Opus 4.8. Those figures are Harvey/Engram-reported, not yet independently reproduced. Before training on the firm's own documents the same model scored just 4.7% on firm-specific questions; after studying them, 72.6%.

Read the #1 with one big caveat: the model that topped the benchmark had studied the test corpus beforehand, adapted on the firm's 100M tokens before it was evaluated. That makes this a demonstration of what the 27B can absorb with domain-specific tuning, not a score for the stock Apache-2.0 weights on a neutral harness — and it covers one domain, law, not general quality. What it does support is the pitch behind the tweet: a 27B small enough to run on a local machine is strong enough for professional-grade work once it has the right knowledge.

The second independent result: #9 overall on Code Arena's WebDev

The second independent result is a coding one, and it is why this page changed on August 25. Code Arena is Arena.ai's web-development leaderboard — the project formerly known as WebDev Arena, on the site formerly known as LMArena — where users build real apps with anonymized pairs of models and vote on which output they would rather ship. On that public leaderboard Qw​en3.8-27B sits at #9 overall with a score of 1595, the only model in its size class inside the top 10.

The placement is the independent part — it is on a third-party leaderboard anyone can open. The headline around it is Alibaba-stated: the "only small model in the top 10" framing, and the companion placements, come from Qw​en's August 25 announcement. They are worth repeating with that label. The 27B ranks six spots below the far larger Qwen3.8-Max (which the announcement places at #3 overall), and well ahead of the similarly sized Gemma 4-31B (which the same announcement places at #80). The point is not that a 27B beats the frontier at building web apps; it is that a model small enough to run on one GPU is inside the top ten of a blind coding arena whose other occupants are much larger models.

The third independent result: #1 open model on Arena.ai's Image-to-WebDev

The third independent result arrived on August 26, and it is the strongest yet for a model this size. Image-to-WebDev is the sibling board of Code Arena's WebDev on Arena.ai — the arena where a model is given a screenshot or other visual reference and has to turn it into a working web interface, with blind human voters picking the output they would rather ship. On that public leaderboard Qw​en3.8-27B ranks #1 among open models and #7 overall, with a reported arena score of 1,574.

The placement is the independent part — it is on a third-party leaderboard anyone can open. The headline around it is Alibaba-stated: the Qw​en team's August 26 announcement bills the 27B as "on par with models 100x its size" on this test, a comparison the leaderboard does not document. And a preference ranking is not a code-quality verdict — Image-to-WebDev votes measure which output a human would rather ship, not whether the HTML is secure, accessible, or maintainable. What the result does establish is that a 27B open-weights model now leads its entire open field at turning a picture into a page, which is the skill a growing "screenshot to site" category of products is built on.

The honest gaps: where it still loses, and the methodology caveats

Against the frontier the picture is flattering but not one-sided. Where Qw​en3.8-27B and Claude Opus 4.6 Max overlap, Qw​en wins the majority — SWE-bench Pro, QwenSWEBench, CoWorkBench, LiveCodeBench v6, OSWorld-Verified, AndroidWorld, IFBench, and the whole vision block. Against Meta's Muse Glimmer-30B, the natural local-class rival, it wins all eight overlapping benchmarks outright. But it still loses Terminal-Bench 2.1 (73.0 vs 78.2), GPQA Diamond (89.2 vs 91.3), Humanity's Last Exam (30.8 vs 40.0) and NL2Repo-Bench (42.3 vs 47.6) to Claude Opus 4.6 Max. If your workload is the hardest frontier reasoning — frontier math, massively multidisciplinary questions — the 27B is not there yet.

Three caveats before you quote any of this. First, the model-card table above is still entirely Alibaba-reported — there is no Artificial Analysis index for the 27B yet. It now has three independent third-party results, but each is narrow: the Code Arena ranking measures web-app building from text prompts, the Image-to-WebDev ranking measures turning a screenshot into a working page, both judged by blind votes on anonymized outputs, and the Harvey legal-agent study covers a single domain with a model trained for the test. None is the stock weights run across a general-purpose harness. Second, the model card uses the Claude Code harness for SWE-bench Pro and QwenSWEBench and a GPT-4o judge for Humanity's Last Exam — comparisons with models scored on other setups will move the numbers. Third, Alibaba's MathVision run used a fixed prompt while competitors got the higher of two variants. None of this means the results are wrong; it means they are a ceiling, not a floor, until independent labs run the model on neutral general-purpose harnesses.

What it takes to run it

Qw​en3.8-27B is a local model, which changes what "performance" means: throughput on your own hardware instead of a price card. The BF16 weights are roughly 55.6GB; quantized to 4-bit it fits a 24GB card such as an RTX 3090 or 4090. AMD shipped Day-0 support on launch day, and its own llama.cpp/Vulkan measurements — August 2026, three or more runs averaged — put it at 24.5 tokens/s on a Ryzen AI Max+ 395 and 51.8 tokens/s on a Radeon AI PRO R9700 with 32GB, on systems with more than roughly 24GB of variable graphics memory or VRAM. Community tests of the FP8 build on an NVIDIA GH200 held 10 concurrent 262K-context requests with sub-10ms time-to-first-token. Your own numbers will differ — those are someone else's GPUs — but the shape is real: this is the first Qw​en release in this class that runs, not just in a review, on a single workstation.

Free weights, or a paid API?

The decision this model forces is the one that divides local from hosted: the weights cost nothing, but your GPUs are not free. Self-hosting means owning the silicon — one 24GB card if you accept 4-bit quantization and slower generation, something bigger if you want the full 262K context at full quality. The hosted alternative on OrcaRouter is $0.33 per million input tokens and $2.40 per million output, with the same 262K context and image-plus-video input, and a p50 time-to-first-token of 225ms measured over the last seven days — no GPU to buy, no VRAM to babysit. The arithmetic is the one you always do with open weights: the token throughput you actually need times your cost per token, versus the flat price of a card. At heavy sustained volume, self-hosting wins; at bursty or modest volume, paying $0.33/$2.40 beats buying silicon you mostly idle.

A decision card titled Free weights, your own compute — or pay per token. The self-host side lists $0 per token, Apache 2.0 weights, roughly 55.6GB BF16 with 4-bit fitting a 24GB card, 24.5 tokens/s on Ryzen AI Max+ 395 and 51.8 tokens/s on Radeon AI PRO R9700. The hosted-API side lists $0.33 input and $2.40 output per 1M tokens, 262K context with image and video input, and a 225ms p50 first token. A footer reads API price per OrcaRouter Aug 15 2026; throughput per AMD llama.cpp/Vulkan tests.

Bottom line

Qw​en3.8-27B is the strongest open-weights model Alibaba has put in this size class, on Alibaba's own numbers: best-in-table at agentic coding, computer use and vision reasoning, with a 262K context window and Apache 2.0 weights. The correct reading of the scorecard is conditional — every model-card figure is vendor-reported, harness-specific, and as yet unreproduced, and the frontier still wins the hardest reasoning benchmarks. If you are deciding whether to self-host it or call it as an API, treat the benchmark table as the promise, the AMD throughput figures as the first independent hardware measurement, the Harvey legal-agent result as the first independent sign that the model's capability survives outside Alibaba's own harnesses, and the two webdev-arena placements — #9 overall on Code Arena's WebDev and #1 open model on Arena.ai's Image-to-WebDev — as the first independent coding-leaderboard confirmations of that capability. We will update this page as more independent labs publish scores.