A hero title card for "Qwen 4 Max vs Claude Opus 5" with the subtitle "The benchmark that moved under both of them", three pill badges reading "Index revision v4.3.2", "51 vs 53" and "$5.00 and $25.00", and the OrcaRouter logo in the bottom-right corner.
Engineering & Research

Qwen 4 Max vs Claude Opus 5: the benchmark that moved under both of them

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Qwen 4 Max was previewed at the Yunqi conference in Hangzhou on September 22, 2026 — the flagship tier of a four-model Qwen 4 line that is announced and not yet released, with no price, no context window and no published score. Claude Opus 5 has been generally available since July 24, 2026 at $5.00 per million input tokens and $25.00 per million output tokens, and it has been the target every flagship since the Qwen3.8-Max release has been aimed at. The obvious way to write this comparison is a spec sheet with one column blank. The useful way is to notice that on September 19, 2026, three days before the announcement, Artificial Analysis rescored its Intelligence Index and Claude Opus 5 stopped being the model at the top of it. That single change reframes what Qwen 4 Max has to beat.

What changed on September 19

Artificial Analysis published index revision v4.3.2 on September 19, 2026. Under the current revision Claude Opus 5 scores 51 on the Intelligence Index at its maximum reasoning effort — 50 at xhigh, 48 at high, 45 at medium and 39 at low — and it sits behind Claude Fable 5.1 and GPT-6 Astra, both at 53. The per-task cost on that index is $5.86.

If that reads like a downgrade, it is not one. The model did not change. The index did. Our own coverage through July and August cited Claude Opus 5 at 61 to 63 and at the top of the table, and those numbers were correct under the revisions in force at the time. Anyone still quoting them without a revision date is quoting a retired figure, and this is the single most common error in comparisons written this month. The practical consequence is that a claim like "Claude Opus 5 is the highest-scoring model available" is no longer true, and a comparison built on it collapses.

For the Qwen side the consequence is sharper still. A flagship tier announced in September 2026 is being aimed at a target that was redefined in September 2026. Whatever Alibaba eventually publishes for Qwen 4 Max will be read against a leaderboard where 53 is the number to beat, not 63.

A two-column scoreboard titled "Qwen 4 Max vs Claude Opus 5 - the scoreboard". Left column Qwen 4 Max: Status announced, no date; Input price not published; Output price not published; Context window not published; Independent score none exists; Open weights not published. Right column Claude Opus 5: Status generally available; Input price $5.00 per 1M tokens; Output price $25.00 per 1M tokens; Context window 1M tokens; Independent score 51 on index v4.3.2; Open weights none. Footer reads "Claude Opus 5 from Anthropic published pricing and Artificial Analysis index v4.3.2, September 19 2026.", with the OrcaRouter logo bottom-right.

The comparison, honestly stated

One side of this table is complete and the other side is empty, and pretending otherwise would be the whole failure mode of this article. Here is what is actually known:

• Status — Claude Opus 5 generally available since July 24, 2026, versus Qwen 4 Max announced September 22, 2026 with no release date

• Input price — Claude Opus 5 $5.00 per 1M tokens, versus Qwen 4 Max not published

• Output price — Claude Opus 5 $25.00 per 1M tokens, versus Qwen 4 Max not published

• Cached input — Claude Opus 5 $0.50 per 1M tokens on a cache read, versus Qwen 4 Max not published

• Context — Claude Opus 5 1M tokens at the default and maximum, versus Qwen 4 Max not published

• Output ceiling — Claude Opus 5 128K tokens synchronously and 300K on the Batches API with a beta header, versus Qwen 4 Max not published

• Modality — Claude Opus 5 text, image and file input with text output, versus Qwen 4 Max not published

• Long-context pricing — Claude Opus 5 has no surcharge, so a 900,000-token request bills at the same per-token rate as a 9,000-token one, versus Qwen 4 Max not published

• Independent score — Claude Opus 5 51 on Artificial Analysis Intelligence Index v4.3.2 at maximum effort, versus Qwen 4 Max no independent evaluation exists

• Vendor-reported scores — Claude Opus 5 SWE-bench Verified 96.0%, SWE-bench Pro 79.2%, OSWorld 2.0 70.6%, Terminal-Bench 2.1 in the mid-80s depending on harness, versus Qwen 4 Max nothing reported

• Reasoning control — Claude Opus 5 adaptive thinking on by default with a five-step effort ladder, versus Qwen 4 Max not published

Ten rows of "not published" is not a rhetorical device. It is the state of the evidence, and the honest conclusion from it is that no capability comparison between these two models can be made yet by anyone.

The part of the gap that will not close

Pricing and context will be published eventually. One difference will survive the release, and it is worth deciding about now rather than in the week Qwen 4 Max ships: the two models are built to be bought in different ways.

Claude Opus 5 is a single closed model at a single price from a single vendor, with a published retirement commitment of not sooner than July 24, 2027. That is a stable target for capacity planning. Qwen 4 Max is the flagship of a family that Alibaba has repeatedly released across a much wider price range — the Qwen 3.8 generation spans a frontier flagship at $2.00 input and a Flash tier well below it — and the flagship tier of the Qwen line has historically been the one with the shortest useful life before a cheaper tier closes the gap on it.

The other difference is open weights, and it points the other way. The Qwen 3.8 generation published weights for its largest model under a custom Qwen licence, which is what makes fine-tuning, quantisation and self-hosting possible at all. Claude Opus 5 has no weight release and will not have one. If your workload needs a model you can run inside your own perimeter, or modify, that is a structural advantage no benchmark revision can take away from Alibaba — and it is the strongest reason to care about this particular comparison rather than treating it as one flagship against another.

A screenshot of the OrcaRouter model page for Claude Opus 5 (anthropic/claude-opus-5, English UI), showing the 1M token context badge, the model ID anthropic/claude-opus-5, a 128K maximum output, text plus image plus file input with text output, the Vision, Tools, JSON and Reasoning tags, the listing date 2026-07-24 credited to Anthropic, a p50 time-to-first-token of 5.75s, the OpenAI-compatible code sample, the Public benchmarks block, and the opening of the model description calling Claude Opus 5 Anthropic flagship for demanding reasoning, coding and long-horizon agentic work.

How to run this comparison today

Claude Opus 5 is available now, and OrcaRouter routes it as anthropic/claude-opus-5 at Anthropic's list price with 0% markup — the provider's rate passed through untouched, so the $5.00 and $25.00 figures above are what you are billed rather than a marked-up equivalent. That matters more than usual for a model in this price band: a 10 percent platform margin on a $25.00 output rate is $2.50 per million tokens, which is more than the entire input cost of most open-weight alternatives. Alongside it the catalogue carries close to 200 models on one OpenAI-compatible endpoint, with automatic failover when a provider path degrades and a routing DSL for expressing the policy explicitly instead of hard-coding a model name into your application.

What OrcaRouter does not carry is Qwen 4 Max, or any Qwen 4 tier. The catalogue has zero entries for the family, which is what you would expect for a model announced and not released. When the tiers do ship they would appear in the same place, and until then the Qwen model worth comparing against Claude Opus 5 is the one you can actually call: Qwen3.8-Max, generally available since August 3, 2026 at $2.00 input and $6.00 output per million tokens, with published weights and a recorded independent intelligence score of 56 at $1.14 per task — a figure banked under an earlier index revision, so compare it to Claude Opus 5's 51 only with that caveat attached.

What to do with this before a model card exists

There is no verdict to give here, because one of the two models does not exist as a product. What can be said is that the comparison changed shape on September 19, 2026 without either vendor doing anything: Claude Opus 5 is no longer the top-scoring model on the independent index that most comparisons lean on, and it now sits two points behind two other models. A Qwen 4 Max launch will land into that table, not the one from July.

So the decision for the next few months is not Qwen 4 Max versus Claude Opus 5. It is Claude Opus 5 versus the Qwen model that already shipped — a flagship at a fifth of the input price with published weights — and that comparison is available now, with real numbers on both sides. Revisit it when Alibaba publishes a model card, and treat any Qwen 4 Max benchmark table you see before then as unverified.

A screenshot of the OrcaRouter model page for Qwen3.8 Max (qwen/qwen3.8-max, English UI), showing the 1M token context badge, the model ID qwen/qwen3.8-max, text plus image plus video input with text output, the Vision, Tools, JSON, Reasoning and Thinking capability tags, the listing date 2026-08-03, a p50 time-to-first-token of 3.29s, the OpenAI-compatible Python code sample, and the model description naming Qwen3.8-Max Alibaba newest flagship and highest-capability tier to date.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily