
Claude Haiku 5.5 vs Qwen3.8 27B: A Clean Sweep on the API, and Why the Open Weights Still Exist
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 127 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 58 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 350 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Most of the comparisons in this series come down to a trade. This one does not, and the absence of a trade is itself the finding. Claude Haiku 5.5, shipped on October 7, 2026, beats Qwen3.8 27B, the open-weight 27B dense model from August 14, 2026, on the composite intelligence score, on cost per token, on cost per finished task, on context window, on response ceiling, on agentic coding, and on the science-reasoning rows — and it does so from a single proprietary API that requires no hardware. The one row Qwen3.8 27B wins outright is video input, and the other is a licence.
That is not a reason to write the model off, because an Apache-2.0 licence and a video modality are not consolation prizes. It is a reason to be precise about why anyone would choose the open-weights side, and to stop pretending the argument is about benchmarks.
The numbers both models were actually run on
Per million tokens in US dollars. The vendor's rates are from its own pricing documentation; the shared measurements are Artificial Analysis Intelligence Index v4.3.2, read 2026-10-08, both in their highest published effort configuration.

• Input — Claude Haiku 5.5 $0.10 up to 100,000 tokens and $0.50 above; Qwen3.8 27B $0.50 on the board's recorded third-party rate.
• Output — Claude Haiku 5.5 $0.50 up to 100,000 tokens and $2.50 above; Qwen3.8 27B $3.00.
• Cached input — Claude Haiku 5.5 $0.01 and $0.05, a 90% discount; Qwen3.8 27B $0.10, an 80% discount.
• Independent score — 43.40 for Claude Haiku 5.5 (Max) against 33.70 for Qwen3.8 27B (Xhigh). A 9.70-point gap, the widest in this batch.
• Cost per finished task — $0.21 against $1.01, on the same evaluation run. A 4.7x gap, also the widest here.
• Context and output — 1M tokens and a 128,000-token response ceiling for Claude Haiku 5.5; a 256K-token context for Qwen3.8 27B.
• Modality — both take text and images and return text. Qwen3.8 27B adds video input; Claude Haiku 5.5 does not.
• Weights — Qwen3.8 27B is 27B dense parameters under Apache 2.0, downloadable. Claude Haiku 5.5 is proprietary and API-only.

Why the per-task gap is four times the per-token gap
The rate card already shows a 5x input gap and a 6x output gap in the vendor's favour, so the direction is not in question. What is worth reading is that the per-task figure is smaller than the per-token figure, which is the opposite of what happened in this batch's other matchups, and the reason is instructive.
Claude Haiku 5.5 emits 162,164 output tokens per Intelligence Index task — 129,047 reasoning and 33,118 answer. Qwen3.8 27B emits 66,797, split 47,711 reasoning and 19,087 answer. The vendor's model spends 2.4 times as many tokens per task, so its 6x output-rate advantage compresses to a 4.7x per-task advantage. It is still enormous. It is just not the number the rate card advertises.
The blend math is a cleaner way to see the shape of it. On an input-heavy 7:2:1 blend — the weight most production traffic actually has — Claude Haiku 5.5 comes to $0.077 per million tokens against $0.470 for Qwen3.8 27B. On a balanced 1:1 blend it is $0.30 against $1.75. The vendor's 100,000-token cliff is the only thing that ever closes this: past it, Claude Haiku 5.5's meters become $0.50 and $2.50 on the whole request, and the two models meet at roughly the same price — at which point you are paying a 9.7-point score premium for nothing.
Where the vendor card and the independent board disagree
This is the row set worth slowing down on, because Qwen3.8 27B's own model card reads considerably stronger than the independent measurements do, and the difference is the whole reason a "27B model that rivals much larger ones" claim needs a second source.
The vendor's own published figures, as recorded on our catalogue page for the model, include 90.3 on LiveCodeBench v6, 89.2 on GPQA Diamond, 84.3 on OSWorld-Verified, and 79.0 on QwenSWEBench. Those are vendor-reported and unreproduced, and they describe a model that is competitive well above its weight class.
On Artificial Analysis's independent run, Qwen3.8 27B scores 0.0556 on Terminal-Bench 4.0, 0.4664 on SciCode, 0.3392 on Humanity's Last Exam and 1423.21 on GDPval-AA v2.1. Set the 0.0556 next to the 84.3 the vendor card reports on OSWorld and the shape of the disagreement is clear: on computer-and-terminal agentic work, the independent board puts this model roughly where a 27B dense model belongs, and the vendor card does not.
Two rows do cut the other way, and they are worth stating for the same reason. Qwen3.8 27B scores 0.4824 on AutomationBench against Claude Haiku 5.5's 0.3541 — a 12.8-point independent win on the closest thing either board has to a business-automation benchmark. That is a real number from the same run, on the row most adjacent to what people actually deploy small models for.
The honest read is not "the vendor inflated everything." It is that a model card measures the tasks its authors chose, the independent board measures a fixed set, and Qwen3.8 27B's strengths are on the vendor's list and not on the board's.
The economics change when you own the hardware
Every rate quoted above is an API price, and for Qwen3.8 27B that is the wrong frame.
This is a 27B dense model under Apache 2.0. It fits on a single modern accelerator at the quantisation most teams would accept, which means the marginal cost of a token stops being a vendor's meter and becomes your own utilisation. That is why the price to call it differs by host in a way neither of the proprietary models in this comparison ever could: the board records a third-party rate of $0.50 in and $3.00 out, and the OrcaRouter catalogue carries it at $0.33 in and $2.40 out with no cache-read line at all, because we run the weights on our own infrastructure rather than reselling someone's endpoint.
For a team already paying for GPUs, the comparison is not $0.21 versus $1.01 per task. It is the vendor's billed rate against the incremental cost of a request on hardware you own, and those are different numbers. That is the argument for the open-weights side, and it is the only one that survives contact with the benchmark table.
There is a second, less obvious consequence of the licence being Apache 2.0: the model can be shipped inside a product, fine-tuned on your own traffic, pinned to a specific revision, and audited down to the weights. None of that is available at any price from a proprietary API-only model, and for regulated workloads it is frequently the deciding constraint rather than a preference.

Calling either of them
Qwen3.8 27B is on the OrcaRouter catalogue as qwen/qwen3.8-27b at $0.33 and $2.40 per million tokens, self-hosted on our own infrastructure, with the OpenAI-compatible endpoint at https://api.orcarouter.ai/v1. Run against our own traffic over the last week it returns a median 1,837 ms to first token at 319 output tokens per second — a fast model, measured on our playground rather than a standardised harness, so treat that as a serving figure and not a benchmark. Video, image and text input, 262K-token context, native tool calling, structured outputs and a reasoning mode are all supported on the same key as everything else on the catalogue.
Claude Haiku 5.5 is not on the catalogue. It is reachable through the vendor's API directly and through AWS, Azure and the major cloud platforms, and that is the accurate statement. What the catalogue does carry from the vendor's line-up is Claude Sonnet 5.5 and Claude Opus 5.5, which makes the escalation path out of a self-hosted small model and into a hosted mid-tier agentic one a routing change rather than a second integration — automatic failover sits underneath it either way.
The verdict, which is unusually lopsided
As an API call on text or images, Claude Haiku 5.5 wins this comparison on every row that is not about licensing. Nine point seven index points, 4.7x cheaper per finished task, four times the context window, twice the response ceiling, and a five-to-six-times lower sticker inside the short-prompt regime. There is no workload-shape argument that rescues Qwen3.8 27B on price, because the vendor's disadvantage — heavy output token use — is worth far less than its rate advantage.
What rescues it is everything the API price does not capture: a licence that permits redistribution, weights that can be held and pinned, video input, and a self-hosted path where the per-token cost is your own hardware rather than a vendor's card. If you are shopping for the cheapest way to classify, extract, summarise and route text, Claude Haiku 5.5 is the answer and the gap is not close. If you are shopping for a model you can put inside your own product, on your own metal, under your own audit, then the benchmark gap stopped being the question several paragraphs ago.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
