
Claude Sonnet 5.5 vs Qwen3.8 Max: Six Weeks, Two Tiers, One Verdict on Cost
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 984 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Run the arithmetic on three prompts and the comparison mostly resolves before you reach the benchmarks. A short support-reply call: 2,000 input tokens, 500 output. On Qwen3.8 Max at $2 per million in and $6 per million out that is four-tenths of a cent; on Claude Sonnet 5.5 at $2 and $10 it is nine-tenths. A repository refactor: 400,000 input, 30,000 output — forty-three cents against eighty-three. A long agentic session that actually spends its budget: two million tokens in, 200,000 out, so $5.20 against $6.30 once the numbers get big enough for the input side to dominate.
The gap that starts as a 2.25× difference on output price narrows to roughly 1.2× at agentic scale, and that scaling is not a curiosity. Qwen3.8 Max and Claude Sonnet 5.5 are both 1M-token models with heavy output ceilings, both priced at $2 per million on input, and both released within eight weeks of each other — the vendor's on August 3, 2026 with a dated refresh on September 2, Anthropic's on September 28. The difference between them is not the rate card. It is what each one spends to finish.
What shipped, and when
Qwen3.8 Max is Alibaba's highest-capability tier to date. Alibaba positions it directly against GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro in its own migration guide, and recommends it whenever a task needs the strongest available reasoning. It carries a 1M-token context window, accepts text, image and video input, and emits text — the video input surface is the one capability here that Claude Sonnet 5.5 does not have. Pricing is $2 per million input and $6 per million output, with cache reads at $0.25 and cache writes at $2.50. The August 3 entry in our catalogue is joined by a September 2 refresh listed as Qwen3.8 Max (0902), which carries the same headline rates and a 131K output ceiling.

Claude Sonnet 5.5 is the second model in Anthropic's 5.5 generation, released on September 28, 2026, six days after Claude Opus 5.5. It ships as claude-sonnet-5-5 on the Claude API and across Amazon Bedrock, Google Cloud and Microsoft Foundry, with a 1M-token window, a 128K synchronous output ceiling that extends to 300K on the Message Batches API behind the output-300k-2026-03-24 header, and Claude Sonnet 5's rates held exactly: $2 in, $10 out, $0.20 cache reads, $2.50 cache writes. Zero data retention from launch, retirement not sooner than September 28, 2027.
So the input rate is identical, the context windows are the same size, and the two vendors refreshed within a month of each other. Everything that separates them is on the output side of the bill or in the benchmarks.
Benchmarks: vendor table against independent index
Two kinds of evidence exist here and they are not interchangeable.
Anthropic's launch table is vendor-reported, unreproduced by any third party, and covers eight evaluations. The rows that matter
• Terminal-Bench 4.0 — Claude Sonnet 5.5 70.6% against Claude Sonnet 5's 10.3%
• CursorBench 4.0 — 55.5% against Claude Sonnet 5's 34.1%
• GDPval-AA v2.1 — 1844 against 1449 for Claude Sonnet 5, 1846 for Claude Opus 5.5 and 1487 for GPT-6 Sol
• Humanity's Last Exam, with tools — 64.5% against 54.9%; Claude Opus 5.5 sits at 67.7%
• OSWorld 2.1, partial — 80.1% against 57.0%
• Chartography, no tools — 61.6% against 15.6%
Alibaba does not publish a directly equivalent four-column table, so the only figures that can be placed on the same axis are the independent ones. Artificial Analysis, v4.3.2 revision
• Composite Intelligence Index — Claude Sonnet 5.5 56 at rank 3 of 216; Qwen3.8 Max 45 at rank 27
• Cost per Intelligence Index task — $7.60 against $5.41
• Output speed — Anthropic's model runs against a Claude Sonnet 5 baseline measured at 78.7 tokens per second; Qwen3.8 Max is measured at 39.1
• Verbosity — 410M output tokens on the Index against 190M
Qwen3.8 Max's own per-evaluation scores are respectable rather than marginal: 92.8% on GPQA Diamond, 88.8% on Terminal-Bench 2.1, 80.3% on long-context recall, 47.8% on tau-banking. It loses ground on the composite because it wins nothing decisively and the newer Anthropic tier wins several things outright. Note also that the independent page for Qwen3.8 Max carries a September 2, 2026 release date and a 984K context reading — it is describing the (0902) refresh, not the August build, and mixing figures between the two would be a mistake.

What is missing from both columns is as important as what is in them. Nobody has independently reproduced Anthropic's OSWorld or Terminal-Bench numbers. Nobody has published a Chartography or CursorBench figure for Qwen3.8 Max. The composite is the only number measured the same way on both models, which is exactly why the cost-per-task figure beneath it does so much work.
The number that decides it, and it is not on either rate card
Artificial Analysis charges both models the same way: run the ten evaluations in the Index, count the tokens, multiply by the list price. Claude Sonnet 5.5 comes out at $7.60 per task and Qwen3.8 Max at $5.41, a 40% premium for the Anthropic model despite a higher composite score. The verbosity readings explain the direction — 410M tokens against 190M — and the speed readings explain the rest: a model that emits fewer tokens and takes longer per token is not the one paying for output at twice the rate.
Anthropic's own efficiency claim fits the same story rather than contradicting it. The company says Claude Sonnet 5.5 generates output 30%+ faster than Claude Sonnet 5 and costs "up to 30% less per task" at identical prices — a claim about tokens per task, not about rates. Whether it holds against a model that spends less than half as many tokens is precisely what the $7.60 figure is asking, and one independent run is one data point.
The practical consequence: the $6-versus-$10 output rate is the number procurement will look at, and it is the least predictive number in the article. The predictive one is tokens per task on your own prompts, and neither vendor will hand it to you.
Inputs, video, and the migration trap
The capability difference that shows up immediately is modality. Qwen3.8 Max accepts video alongside text and images; Claude Sonnet 5.5 accepts text and images. For screen-recording analysis, meeting-footage pipelines or anything that treats a video as a first-class input, that is not a benchmark gap — it is a binary eligibility question, and only one of these models passes.

The difference that shows up a week later is behavioural. Claude Sonnet 5.5 makes five breaking changes to code running on Claude Sonnet 5. A non-default temperature, top_p or top_k now returns a 400. Sending thinking: {"type": "disabled"} returns a 400 pointing at between_tools, the new lowest thinking setting, accepted at low, medium and high effort and rejected at xhigh or max. Forced tool use — tool_choice of "any" or a named tool — returns a 400 outright. The older computer_20251124 computer-use tool is refused on the Claude API and Google Cloud. And thinking blocks are bound to the model and the account that produced them, so reasoning can be carried forward from Claude Sonnet 5 but not out to another family, and an edited conversation prefix can turn a replay into an error.
Qwen3.8 Max asks for none of that. It is an OpenAI-format model with a conventional request surface, and moving to it from another Qwen tier is a model-string change.
Weigh the two costs honestly. Choosing the Qwen model costs you the eleven-point composite gap and the Anthropic vendor numbers you cannot independently verify anyway. Choosing the Anthropic model costs you video input and a checklist of four API changes that will each fail loudly with a 400 the first time they are hit — which is the good kind of failure, but a day of work regardless.
Where OrcaRouter fits in
The reason to route rather than choose is that the deciding number — cost per task on your traffic — cannot be computed from either vendor's documentation.
OrcaRouter is a single OpenAI-compatible endpoint over 200-plus models with provider list price passed through at no markup, so a rate cut or a tier change on either side is live the same day it is announced. Both Qwen3.8 Max builds are in the catalogue at Alibaba's own $2 / $6 — the August entry and the (0902) refresh — alongside Claude Sonnet 5, the model most Claude API traffic still runs on, routable since June 30 at Anthropic's $2 / $10 with a measured 150 output tokens per second. Claude Sonnet 5.5 itself is not on our catalogue yet; while that is true, the honest route to it is the vendor's own API and the three clouds Anthropic lists, and the way to try it without moving a workload is a slice of traffic behind automatic failover to Claude Sonnet 5 or Qwen3.8 Max.
Which one, for which job
Qwen3.8 Max wins on video input, on output rate, on measured cost per index task, and on integration effort. It is the right default for high-volume, well-specified work where a twelve-point composite gap does not show up in output quality.
Claude Sonnet 5.5 wins on the independent composite, on agentic coding evaluations where Anthropic's vendor numbers showed a 60-point jump over its own predecessor on Terminal-Bench 4.0, on the 300K batch output ceiling, and on a job-shaped decision that a rate card cannot make: it is the model to pick when the task is hard enough that being right matters more than being cheap.
What would change the answer for either is the same thing, and it is not a new benchmark release. It is a tokens-per-task count on your own prompts, at your own effort settings, for both models. Anthropic's effort parameter and Qwen's output ceiling are both levers on that number, and both are set by you rather than by the vendor's price list.
The one thing this comparison does settle: at $2 per million on input, the input side of the bill is no longer a differentiator between a Chinese flagship and Anthropic's mid-tier. That race moved to the output side and to the token count months ago.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
