
Claude Haiku 5.5 vs Qwen3.8-Flash-Next: Three Cents Apart Per Token, Seventy-Eight Apart Per Finished Job
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 127 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 58 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 350 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Put the two rate cards next to each other and they look like they came out of the same meeting. Claude Haiku 5.5 is $0.10 in and $0.50 out. Qwen3.8-Flash-Next is $0.15 in and $0.47 out. The vendor's model is a third cheaper on input and six percent dearer on output. Neither vendor built that shape by accident, and neither published a comparison note explaining what they were aiming at — but the arithmetic of a mixed 3:1 workload makes the intent legible, and it makes the pair the closest pricing match in this whole tier.
That is where the resemblance ends. Claude Haiku 5.5 is a closed API model launched on 2026-10-07 with a one-million-token window and a 128,000-token maximum output. Qwen3.8-Flash-Next is a 180-billion-parameter open-weights model with 6 billion active parameters, weights on Hugging Face under the Qwen Community License 1.0, and a published context window that its own checkpoint config and the independent board do not fully agree on. Released 2026-08-26, it is neither new nor exotic to anyone building in this band.
The two are priced together on purpose. What the rate cards cannot tell you is that they are built for different jobs, and that the cheaper-looking one on a spreadsheet is the more expensive one to run.
The arithmetic that gives the pricing away
Blend the two rates at a common 3:1 input-to-output ratio — the ratio most chat and code-assist traffic settles near — and you get $0.20 per million tokens for Claude Haiku 5.5 and $0.23 per million for Qwen3.8-Flash-Next. Anthropic's model is about 13% cheaper on that blend, which is a smaller gap than the input column implies and a bigger one than the output column does.
Swap the ratio and the position moves. At 1:1, the two are $0.30 and $0.31 per million — a tie inside a rounding error. At 1:3, output-dominant, Claude Haiku 5.5 comes in at $0.40 and Qwen3.8-Flash-Next at $0.39, a one-cent win for Qwen and the only ratio at which Anthropic's model is not the cheaper of the two.
So there is no single honest answer to "which is cheaper," and any comparison that gives one is choosing a ratio on the reader's behalf. What is defensible is this: Claude Haiku 5.5's card is weighted for input-heavy traffic, Qwen3.8-Flash-Next's is weighted for output-heavy traffic, and the three cents per million tokens that separate them at 3:1 is the closest anyone has come to matching Anthropic's number in this tier. That is a deliberate positioning, whatever the launch materials say.
Our catalogue lists Qwen3.8-Flash-Next at $0.15 input and $0.47 output per million tokens with a $0.0184 cached-input rate. The $0.15 and $0.47 are the same figures Artificial Analysis records as the provider's list price, which is what the pass-through arrangement is supposed to produce.
What a day of real volume actually costs
Take a pipeline pushing 10 million input tokens and 2 million output tokens a day. That is a small production workload, comfortably inside the first pricing tier at typical prompt lengths.
• Input cost — $1.00 a day on Claude Haiku 5.5, $1.50 a day on Qwen3.8-Flash-Next.
• Output cost — $1.00 a day on Claude Haiku 5.5, $0.94 a day on Qwen3.8-Flash-Next.
• Daily total — $2.00 against $2.44. About $160 a year apart at that scale, or about 3.7 cents per million tokens.
Now apply the two adjustments that the rate card leaves out, and the direction reverses.
First, Claude Haiku 5.5 uses the newer tokenizer shared with Claude 4.7 and later, which per Anthropic's own documentation counts the same text as approximately 30% more tokens than the model it replaces. If your input is English prose of the kind those token counts were measured on, the effective input rate is not $0.10 — it is closer to $0.13 per million words-worth of text, which puts it within two cents of Qwen3.8-Flash-Next rather than a third below it.
Second, and larger: Claude Haiku 5.5 is verbose. Artificial Analysis recorded 162,164 output tokens for it while running the Intelligence Index, against 107,884 for Qwen3.8-Flash-Next. If that ratio holds on your workload rather than the abstract 3:1, the output line above stops being $1.00 against $0.94 and becomes something closer to $1.50 against $0.94 — and the daily total flips in Qwen's favour.
Neither adjustment is decisive on its own. Together they are the difference between the two models being a rounding error apart and Qwen3.8-Flash-Next being meaningfully cheaper to operate, and the only way to know which applies is to count tokens on your own traffic rather than reason from the rate card.

Where the flat rate stops looking flat
Claude Haiku 5.5's price has a step in it, and Qwen3.8-Flash-Next's does not.
• Price — Claude Haiku 5.5 $0.10 input / $0.50 output per million tokens up to 100,000 tokens of prompt, then $0.50 / $2.50 above; Qwen3.8-Flash-Next $0.15 / $0.47 flat.
• Cached input — $0.01 read and $0.125 write for Claude Haiku 5.5; $0.016 read and $0.23 write for Qwen3.8-Flash-Next.
• Context — one million tokens for Claude Haiku 5.5. For Qwen3.8-Flash-Next the number depends on who you ask: Qwen publishes a one-million-token window, the checkpoint's own config caps position embeddings at 262,144, and Artificial Analysis records the evaluated configuration at 256,000.
• Maximum output — 128,000 tokens for Claude Haiku 5.5; 131,072 on our catalogue entry for Qwen3.8-Flash-Next.
• Input modality — text and images for Claude Haiku 5.5; text for Qwen3.8-Flash-Next.
• Release — 2026-10-07 for Claude Haiku 5.5, 2026-08-26 for Qwen3.8-Flash-Next.
• Weights — proprietary, API-only for Claude Haiku 5.5; open weights under the Qwen Community License 1.0 for Qwen3.8-Flash-Next.
Three of those lines pull in opposite directions from the price headline, and the prompt-length step is the sharpest. Below 100,000 tokens of prompt Claude Haiku 5.5 is the cheaper model on both rate lines. Above it, Anthropic's input rate quintuples and Qwen3.8-Flash-Next retains a 3.3× advantage on input and a 5.3× advantage on output. The threshold roughly coincides with where the context windows diverge anyway — one million tokens for one model, 262,144 for the other — so a pipeline that needs the long window is also the pipeline that pays the second-tier rate to get it.
That is not a criticism of the pricing; tiered rates keyed to prompt length are a normal way to charge for a long-context model. It is a warning about the comparison. A head-to-head that runs at 8,000-token prompts is describing one pair of prices. The same two models at 400,000-token prompts are a different pair entirely, and the cheaper one in the second case was the more expensive one in the first.
What the vendor tables say underneath
Neither vendor has run the other's evaluation suite, so the shared picture is limited to the rows Artificial Analysis measured on both.
• Intelligence Index v4.3.2 — 43.40 for Claude Haiku 5.5 (Max) against 39.82 for Qwen3.8-Flash-Next. A 3.57-point lead to Anthropic, the largest composite gap in this series.
• Cost per finished Index task — $0.21 against $0.37 on the same board.
• Task time — 424.6 seconds against 1,524.3 seconds.
• Terminal Bench 4.0 — 0.3283 against 0.2525. Claude Haiku 5.5 by 7.6 points.
• SciCode — 0.5498 against 0.5058. Claude Haiku 5.5 by 4.4 points.
• Humanity's Last Exam — 0.4439 against 0.3804. Claude Haiku 5.5 by 6.4 points.
• AutomationBench, partial score — 0.3541 against 0.5591. Qwen3.8-Flash-Next by 20.5 points.
Six rows to Anthropic, some of them wide, and one row to Qwen with a 20-point margin. The shape is the same one that runs through this entire price band: Anthropic's small model is stronger at reasoning, science and the cost-and-latency of getting an answer, and weaker at driving a multi-step task to completion. Qwen3.8-Flash-Next is the opposite.
The spread between the composite and the agentic row is unusually wide here — 3.57 points one way, 20.5 the other — which makes this the least ambiguous pair in the band on that axis. If your work is a single hard question answered once, Claude Haiku 5.5 is ahead on every measure you will notice. If your work is an agent that has to finish, Qwen3.8-Flash-Next is ahead on the only row that predicts it, and by more than the composite gap by a factor of five.

The 180 billion parameters you can hold
The line that settles this comparison for a lot of teams is not in the benchmark table at all.
Qwen3.8-Flash-Next is a sparse mixture-of-experts model, 180 billion parameters total and 6 billion active — a ratio that keeps the compute spent per token modest relative to the model's total capacity. It is published under the Qwen Community License 1.0, which Artificial Analysis records as a commercial licence rather than a permissive one; that distinction is worth reading before you plan around it, because a community licence is not MIT and carries its own terms on redistribution and scale. It can be downloaded, self-hosted, pinned to a checkpoint, quantised and fine-tuned, subject to those terms.
Claude Haiku 5.5 cannot be any of those things. It is proprietary, API-only, with no published parameter count, no licence and no repository. Anthropic commits to keeping it callable until not sooner than 2027-10-07, which is a real and specific guarantee — but a guarantee about availability is a different product from the weights themselves.
The practical consequence runs both ways, and it is worth saying plainly rather than as a pitch for either side. Teams that need inference inside their own tenancy, a fixed checkpoint they can diff against, or the ability to fine-tune on domain data are not choosing between these two models on index points. They are choosing Qwen3.8-Flash-Next, because the other option does not offer the thing they need at any price. Teams that need a managed endpoint with a published system card, a documented evaluation set and a retirement commitment are choosing the other direction, and the weights are not a feature they were going to use.
Running the flat-rate side
A note on the name. The independent board files this model as Qwen3.8-Flash-Next; our own catalogue lists the same release under the shorter vendor name Qwen3.8 Flash, at the identical $0.15 / $0.47 with a $0.0184 cached-input rate, and points its repository at the Hugging Face weights published on 2026-08-26. Where the figures below come from Artificial Analysis they belong to the entry it calls Qwen3.8-Flash-Next. The two are the same model; the board simply carries the longer of its names.
Qwen3.8-Flash-Next is on the OrcaRouter catalogue at the provider's list price with 0% markup, passed through rather than blended — so the $0.15 and $0.47 above are the rates on the bill, and a change on Qwen's side lands on ours the same day rather than at some later repricing. Claude Sonnet 5.5 and Claude Opus 5.5 are also on the catalogue, which makes the escalation path from a small model to an Anthropic mid tier a routing configuration rather than a second integration. One API for 200-plus models, one key, automatic failover underneath, and the routing DSL when the cost ceiling belongs in the route rather than in application code.
Claude Haiku 5.5 is not on the catalogue. It is reachable through Anthropic's own API, and the Anthropic leg of this comparison is not something we serve today — better to say that than to leave it ambiguous.

Which of the two, and on what grounds
If your traffic is input-heavy, your prompts sit under 100,000 tokens, and you are paying for real volume, Claude Haiku 5.5 is the cheaper model on both rate lines and the higher-scoring one on six of the seven shared measurements — with the caveat that the 30% tokenizer difference and Anthropic's verbosity both eat into a discount that looks larger on the card than it is on the invoice.
If your traffic is output-heavy, or your prompts run long enough to cross Anthropic's threshold, or you need the flat rate for forecasting, Qwen3.8-Flash-Next is cheaper — sometimes by a factor of five on output above the step — and it is the model that wins the agentic completion row by 20 points.
If you need the weights, the comparison is not close, and the price never enters into it.
The two cards are within three cents per million tokens of each other on a 3:1 blend, which is close enough that the rate should not be what decides it. What decides it is which of those three paragraphs you are in, and that is a question about your pipeline rather than about either model.
