
Claude Sonnet 5.5 vs Qwen 4 Max: One Model Has a Card, the Other Has a Name
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 931 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 192 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1177 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 70 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 107 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
The two halves of this title are not the same kind of object, and that is the first thing a reader needs. Claude Sonnet 5.5 shipped on September 28, 2026: it has an API identifier, a rate card, a context window, a published retirement floor and a row on the independent leaderboards. Qwen 4 Max was previewed at the Yunqi conference in Hangzhou on September 22, 2026 as the flagship of an announced four-model Qwen 4 line — and as of today it has no price, no context window, no published score, no model card, and no page on Artificial Analysis to open. That is not a scandal; it is how a tier gets announced. But it changes what a useful comparison looks like. The comparison worth making is between the newest released Sonnet and the Qwen model you can actually call today, which is Qwen3.8 Max — because that one has a rate card, and the rate card is instructive.
What Alibaba actually put on the table
Qwen 4 Max was shown as a name in a lineup, not as a product. The conference preview introduced the Qwen 4 family tier structure, and nothing about it has since been converted into a specification a developer could plan against: no per-million-token price, no window size, no maximum output, no benchmark table, no Hugging Face repository, no entry in the vendor's own model list. Whether it ships as a closed API tier, as open weights, or in both forms is not stated anywhere the vendor has published. When a tier is that early, the only honest thing an article can say about its specs is that there are none — and any page offering a Qwen 4 Max spec sheet right now is printing a projection.

The one thing the announcement does usefully establish is direction. Alibaba's most recent shipping flagship, Qwen3.8 Max, is positioned by the vendor's own migration guide directly against GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro, and it is the tier Qwen 4 Max is being built to replace. So the successor's target is legible even when its numbers are not, and the model it will have to beat on price is the one already on sale.
The Qwen you can call today, priced against the new Sonnet
Qwen3.8 Max has been generally available since August 3, 2026. It is Alibaba's highest-capability tier to date: natively multimodal with text, image and video input and text output, a 1,000,000-token window, thinking mode, function calling, built-in tools and structured outputs, served through OpenAI-compatible, Anthropic-compatible and native DashScope endpoints. Alibaba positions it for the same demanding multi-step analysis and agentic work that the vendor positions Claude Sonnet 5.5 for, and the two land on almost exactly the same headline rate — which is where the comparison gets interesting.
• Input price — Claude Sonnet 5.5 $2.00 per million vs Qwen3.8 Max $2.00 per million, identical
• Output price — Claude Sonnet 5.5 $10.00 per million vs Qwen3.8 Max $6.00 per million
• Cache read — Claude Sonnet 5.5 $0.20 per million vs Qwen3.8 Max $0.25 per million
• Cache write — Claude Sonnet 5.5 $2.50 per million vs Qwen3.8 Max $2.50 per million
• Context — Claude Sonnet 5.5 1,000,000 tokens vs Qwen3.8 Max 1,000,000 tokens on the vendor's spec sheet; Artificial Analysis records the snapshot it measured at 983,616
• Max output — Claude Sonnet 5.5 128,000 synchronous, 300,000 on Batches vs Qwen3.8 Max not published by the vendor
• Input modality — Claude Sonnet 5.5 text, image and file vs Qwen3.8 Max text, image and video
An identical input rate and a 40% cheaper output rate is a genuinely strong price position, and it is the reason Qwen3.8 Max keeps appearing in comparisons against frontier models. Then you look at the independent measurements and the advantage moves somewhere else entirely.
What the independent numbers say about the price advantage
Both models are on the same Artificial Analysis Intelligence Index, revision v4.3.2, and both have been measured rather than estimated.
• Intelligence Index — Claude Sonnet 5.5 56 at Max Effort vs Qwen3.8 Max 45 on the September 2 snapshot
• Cost per Index task — Claude Sonnet 5.5 $7.60 vs Qwen3.8 Max $5.41
• Output speed — Claude Sonnet 5.5 138.7 tokens/sec vs Qwen3.8 Max 38.2 tokens/sec
• Time to first chunk — Claude Sonnet 5.5 370.8 s at Max Effort vs Qwen3.8 Max 3.0 s
• Output tokens generated across the Index — Claude Sonnet 5.5 410M vs Qwen3.8 Max 190M, which the board flags as very verbose against a median of 88M
• GPQA Diamond — Claude Sonnet 5.5 not published on the current board vs Qwen3.8 Max 92.8%
• Humanity's Last Exam — Claude Sonnet 5.5 55.0% vs Qwen3.8 Max 43.1%
• Long-context recall — Claude Sonnet 5.5 82.7% vs Qwen3.8 Max 80.3%
Two things stand out, and neither is the price. The first is that a 40% cheaper output rate compresses to a 30% cheaper task once you account for token counts — Qwen3.8 Max emits 190M tokens across the index against Claude Sonnet 5.5's 410M, which is leaner, but not lean enough to preserve the full rate gap. The second is speed, and here the direction reverses hard: Claude Sonnet 5.5 generates output 3.6 times faster than Qwen3.8 Max, 138.7 tokens per second against 38.2. On a long generation that is the difference between a minute and several, and it is not a difference any per-token discount repairs.
The first-token figure cuts the other way and deserves its caveat stated plainly. 370.8 seconds is Claude Sonnet 5.5 at Max Effort with default fallback, a reasoning configuration that spends a very large thinking budget before answering; a caller who configures a lower effort gets a first token in seconds. Qwen3.8 Max's 3.0 seconds is measured on a model whose thinking behaviour differs. The comparison is real but it is a comparison of configurations, not of vendors.

Where the missing Sonnet migration cost lands
One operational difference does not show up in any of the rows above and matters more than the price gap to anyone with running code. Claude Sonnet 5.5 changed six behaviours relative to Claude Sonnet 5, and five of them break existing integrations: non-default temperature, top_p and top_k now return a 400; thinking: {"type": "disabled"} returns a 400 and must become between_tools with effort limited to low, medium or high; forced tool choice of "any" or a named tool is rejected, so loops must move to auto with strict tool use or structured outputs; the computer_20251124 computer-use tool is refused on the Claude API and Google Cloud; and thinking blocks are now bound to the producing model and account, so reasoning carries forward from Sonnet 5 but not out to another family. The sixth change fails nothing and is therefore the easy one to ship broken: text between tool calls now returns inside thinking blocks, so a streaming application goes quiet between calls until it sets a display value or disables up-front thinking.
Qwen3.8 Max asks for none of that from a Qwen caller — it is an OpenAI-compatible endpoint with the full sampling surface and the standard tool-calling shape — and its Anthropic-compatible endpoint exists precisely so that code written against Claude can point at it with a base-URL change. For a team already on Sonnet 5, moving to Sonnet 5.5 is a day of migration work with a list of five items to check; moving to Qwen3.8 Max is a configuration change with a different set of risks, mostly around behaviour drift rather than API shape.
Getting the reachable one onto the same key
A realistic pipeline in late September 2026 does not pick one of these. It runs a Claude model for the agentic path and a Qwen model where the video input or the price matters, and the cost of that arrangement is normally two contracts, two keys, two rate-limit surfaces and two places a request can fail. OrcaRouter collapses it into one OpenAI-compatible endpoint across 200-plus models, with provider list price passed through and no markup added, so a vendor rate change on either side takes effect here the same day rather than at the next renewal, plus automatic failover between upstream providers and a routing DSL for composing calls out of several models.
Qwen3.8 Max is routable here today as qwen/qwen3.8-max at Alibaba's $2.00 input and $6.00 output across the full 1,000,000-token window, and the September 2 snapshot that Artificial Analysis measured is addressable as qwen/qwen3.8-max-0902 when you need the exact revision a score was taken on. Neither Qwen 4 Max nor any other Qwen 4 tier is in our catalogue, and no Qwen 4 tier will be until Alibaba publishes something to route. Claude Sonnet 5.5 is not in the catalogue either — it is a day old, and the route to it is Anthropic's own API, with Claude Sonnet 5 available here at Anthropic's $2.00 and $10.00 as the failover target in the meantime. Qwen3.8 Max carries 52.0 million tokens of traffic a week through this catalogue and Claude Sonnet 5 carries 7.9 million, which is the practical version of the argument above: the Qwen tier is not a fallback in name only.

The rule for an announced tier
Treat an announced tier as a scheduling input, not a procurement option. Qwen 4 Max is worth tracking because it tells you where Alibaba's flagship is heading and because it will eventually displace Qwen3.8 Max at the top of the line. It is not worth planning against, because there is nothing to plan against: no identifier, no price, no window, no weights, no measured score. The evidence that will make it real is specific and easy to recognise — an entry in Alibaba's own model list, a price row, a published context window, a Hugging Face repository if the weights are open, and a page on Artificial Analysis, because a model nobody has measured is not a model you can compare.
Until then the decision is between two models that both exist. If your work is agentic, the seventeen Index points and the 3.6× output speed are what you are buying, and Claude Sonnet 5.5's higher per-task cost is the price of them. If your work takes video input, needs a native Anthropic-compatible endpoint on a Qwen bill, or simply cannot justify a $10 output rate, Qwen3.8 Max is on sale right now at the same input price with a 40% cheaper output and a first token that arrives in seconds — and the honest caveat is that it will feel slow once the generation starts, which is the trade the cheaper rate buys.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
