A generated title card reading 'Claude Haiku 5.5 vs Qwen3.8-Flash-Next' with the subtitle 'The open-weight model is not the cheap one', a bar labelled 'Cost per finished Intelligence Index task' marked $0.21 for Claude Haiku 5.5 and $0.37 for Qwen3.8-Flash-Next, and a footer line reading 'Independent figures per Artificial Analysis; pricing per Anthropic and Alibaba.' The real OrcaRouter logo is composited bottom-right.
Guides & Insights

Claude Haiku 5.5 vs Qwen3.8-Flash-Next: The Open-Weight Model Is Not the Cheap One

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The instinct is that an open-weight model from Alibaba's flash tier will undercut a closed Anthropic small model on price, and on the two sticker numbers it does: Qwen3.8-Flash-Next is served at $0.15 per million input tokens and $0.47 per million output on Alibaba's own API, against $0.10 and $0.50 for Claude Haiku 5.5. Run both against the same benchmark suite, though, and the ordering flips. Artificial Analysis measures Claude Haiku 5.5 at Max effort at $0.21 per completed Intelligence Index task; Qwen3.8-Flash-Next costs $0.37 on the same measurement, for a score three points lower. The cheaper sticker belongs to the model with the larger bill, and the reason is the same one that decides most of these comparisons: the rate card is not the price, and the open-weight model earns its keep somewhere other than a managed API bill. That "somewhere else" is the actual subject of this matchup.

What each one is

The two landed six weeks apart and they are not the same kind of artifact.

• Claude Haiku 5.5 — a closed-weight model released October 7, 2026, on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. 1M-token context, up to 128,000 output tokens on the synchronous Messages API, adaptive thinking with an effort dial from Low to Max and a default of medium, a June 2026 knowledge cutoff, and a rate card that steps at 100,000 tokens.

• Qwen3.8-Flash-Next — an open-weight mixture-of-experts model released August 26, 2026 under the qwen-community-1.0 license, described by Alibaba as an experimental preview of the architecture that will underpin Qwen4. It stores 125B main-model parameters plus a separate 51B n-gram embedding table — roughly 176B before a 4B multi-token-prediction module — and activates 6B per token, with 512 experts, top-10 routing and 48 layers. It is natively 262,144 tokens of context, extensible to 1M with YaRN, and takes text, image and video input.

• Qwen3.8-Flash — the served commercial counterpart, also live since August 26, 2026 on Alibaba's QwenCloud at $0.16 per million input and $0.47 per million output with a default 1M-token context. Worth keeping separate from Flash-Next: one is a checkpoint you can download and inspect, the other is a priced managed endpoint.

The distinction between the last two is the whole reason this comparison is interesting. Claude Haiku 5.5 and Qwen3.8-Flash-Next compete on very little, because only one of them can be run on your own hardware.

The measured bill, which points the other way

All the following independent figures are from Artificial Analysis on Intelligence Index v4.3.2. The Anthropic and Alibaba rate cards are the vendors' own published prices.

• Intelligence Index — Claude Haiku 5.5 at Max effort 43, second of 182 models. Qwen3.8-Flash-Next 40, sixth of 117 in its class. Three points apart, in Anthropic's favour.

• Cost per task — $0.21 against $0.37. The more expensive model per token is 43% cheaper per finished task.

• Output tokens on the Index run — 440 million for Claude Haiku 5.5 and 240 million for Qwen3.8-Flash-Next, both against a 100 million median. The Qwen model is the more efficient writer and still the pricier task, which tells you the gap is not token volume alone but the mix of input and valuation the evaluator applies.

• Output speed — 241.9 tokens per second against 56.0. Claude Haiku 5.5 is more than four times faster on the same measurement, and its 295-second time to first token on that run is an artefact of thinking at Max effort rather than a network figure; Qwen3.8-Flash-Next's time to first token is 2.53 seconds.

• Rate card shape — Claude Haiku 5.5 steps from $0.10 and $0.50 to $0.50 and $2.50 at 100,000 tokens. Qwen3.8-Flash is flat at $0.16 and $0.47 regardless of length, which is a genuine advantage on long prompts and the one place the sticker argument survives contact with a long document.

• Tokeniser — Claude Haiku 5.5 uses the newer tokenizer shared with Claude 4.7 and later, so the same text counts as roughly 30% more tokens than on the previous Haiku. That inflates prompts toward the 100,000-token step; it is Anthropic's own documentation, and it cuts against the model in exactly the long-prompt case where the Qwen flat rate looks best.

So the shape of the trade is: pay more per token, get a faster, slightly smarter answer that costs less per completed task, with a rate cliff to watch on long inputs. Or pay less per token and get a slower model that costs more per completed task, with a flat rate and a 262,144-token native window.

A screenshot of OrcaRouter's model page for qwen/qwen3.8-flash, showing the model card and its list pricing of $0.15 per million input tokens and $0.47 per million output tokens, captured in English.A screenshot of the Artificial Analysis model page for Claude Haiku 5.5 at maximum effort, showing an Intelligence Index of 43 on v4.3.2 at rank 2 of 182 models, a 1,000,000-token context window, 241.9 output tokens per second at rank 8 of 182, a cost of $0.21 per Intelligence Index task, and 440 million output tokens against a 100 million median.

Where the open weights actually change the arithmetic

Everything above compares two managed APIs, and that is the comparison Qwen3.8-Flash-Next was not designed to win. Its case is that it does not have to be rented.

• Day-zero serving support. SGLang shipped a cookbook page and a dedicated image; vLLM has a build for it while the architecture-support pull request is still open. NVIDIA validated both plus TensorRT LLM on a GB300 NVL72 rack, and NeMo covers fine-tuning.

• The hardware floor is low for a 176B checkpoint. The 51B n-gram embedding table can live in host memory rather than VRAM, because its lookup addresses are known in advance and can be prefetched. That is what lets the model run on DGX Spark clusters and 4×RTX PRO 6000 workstations, and same-day community quantizations — 1-bit builds that fit in 75GB of RAM at roughly 80% of full accuracy — put it on hardware a small team owns.

• Fine-tuning is available. Not in the sense of a hosted tuning API: the weights are yours, under a community license with the usual conditions, and the architecture is documented in a technical report that publishes the ablations and the negative results.

• The window is smaller than it looks. 262,144 tokens natively, extensible to 1M with YaRN — versus 1M native on Claude Haiku 5.5, with no rope-scaling caveat. On the long-document workloads where the flat Qwen rate looks most attractive, the model that can actually hold the document is the other one.

• The benchmarks are vendor-reported. Alibaba's own table puts Flash-Next ahead of DeepSeek-V4-Flash-0731 on DeepSWE 1.1 (58.7 against 54.4), SWE-bench Pro (62.5 against 56.0) and GPQA Diamond (91.7 against 90.8), and behind it on NL2Repo-Bench (48.1 against 54.2) — repository-level code generation, the one agentic dimension where the smaller model does not keep up. None of it has been reproduced by a third party, and the model is six weeks old.

That is the honest boundary between the two. Claude Haiku 5.5 is a metered service with a fixed quality ceiling and no operational surface. Qwen3.8-Flash-Next is an artifact whose cost curve bends down toward the price of electricity and whose quality ceiling is whatever you can serve, plus the risk of an unreproduced benchmark table and an architecture still being absorbed by the runtimes.

The realistic way to decide

Three questions settle it, and only the first is about benchmarks.

Do you have, or want, GPUs? If not, the self-host case is theoretical and the managed comparison above is the whole story — and it goes to Claude Haiku 5.5 on cost per task, speed and score, with the caveat that its 100,000-token step punishes long prompts. If you do have accelerators sitting under a utilisation target, Qwen3.8-Flash-Next is one of the few open models where a 6B active-parameter decode is genuinely cheap to run at volume, and the $0.37 per task figure stops being the relevant number.

How long is the average prompt? This is the one place the sticker argument holds. Under 100,000 tokens — after a tokenizer that inflates by roughly 30% — Claude Haiku 5.5 is cheaper on every meter. Above it, the flat $0.16 and $0.47 is the better card, and its advantage widens with length right up to the 262,144-token native window, past which YaRN extension is a different engineering problem.

What is the workload's tolerance for latency variance? 241.9 output tokens per second against 56.0 is a large gap, and a self-hosted deployment adds queueing to it. A live-support or browser-driving agent — the workloads Anthropic names for this tier — will feel the difference.

For anyone who wants to answer the second and third questions rather than reason about them, the cheapest instrument is a shared endpoint. Qwen3.8-Flash-Next is not on OrcaRouter's catalogue yet, and neither is Claude Haiku 5.5; the Qwen sibling we do route is Qwen3.8-Flash, at the provider's list price with 0% markup passed through, which is the closest served equivalent to the model in this comparison and the right baseline for a long-prompt cost test. On the Anthropic side, Claude Haiku 4.5, Claude Sonnet 5.5 and Claude Opus 5.5 are all callable from the same key today, so a two-generation price experiment is a configuration rather than a second contract — and because the pricing is pass-through rather than blended, a vendor cut on either leg shows up on our side the same day it is announced. Automatic failover is what makes the experiment safe to run on real traffic: a preview model or an unreproduced benchmark table is exactly the case for pointing a route at something new while a trusted model sits behind it.

What would change the answer

Two things, both checkable. The first is independent evaluation of Qwen3.8-Flash-Next on a harness that also runs Claude Haiku 5.5 — the vendor table is against DeepSeek, and the independent board's 40 is a single placement. A repository-level coding score is the row to watch, because NL2Repo is where the model has already been shown to trail. The second is whether the runtime work lands: the vLLM support is still being merged through open pull requests, and until it does, self-hosting this architecture is an exercise for teams who enjoy that sort of thing rather than a supported path.

Until then the summary is short. Qwen3.8-Flash-Next is the more interesting model to own and the more expensive one to rent. Claude Haiku 5.5 is the cheaper, faster, higher-scoring managed endpoint, with a rate card that punishes anything long. Pick by whether you are buying tokens or buying a checkpoint — and if you are buying tokens, note that the open-weight model is not the discount you assumed.

A generated two-column scoreboard titled 'Claude Haiku 5.5 vs Qwen3.8-Flash-Next — the scoreboard'. Left column Claude Haiku 5.5: weights, closed and API-only; input price $0.10 per million up to 100,000 tokens; output price $0.50 per million up to 100,000 tokens; native context 1,000,000 tokens; independent score 43 on Intelligence Index v4.3.2 at Max effort, rank 2 of 182; cost per task $0.21; measured output speed 241.9 tokens per second. Right column Qwen3.8-Flash-Next: weights, open under the qwen-community-1.0 license; served price $0.16 per million input and $0.47 per million output, flat; native context 262,144 tokens, extensible to 1M with YaRN; independent score 40, rank 6 of 117; cost per task $0.37; measured output speed 56.0 tokens per second. A footer line reads 'Vendor pricing per Anthropic and Alibaba; independent figures per Artificial Analysis.' The real OrcaRouter logo is composited bottom-right.