
Qwen3.8-Max vs Qwen3.7-Max: Two Max Tiers, 16 Index Points, and a Promo That Inverts the Price
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 495 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 186 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1306 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 113 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 224 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
Four days ago, the vendor told us that Qwen3.8-Max is not the model it shipped in August. In a press release dated September 22, 2026, from the Apsara Conference in Hangzhou, the company disclosed that its flagship completed 33 iterative cycles of fully automated training and post-training work over more than a month, and that the resulting build lifted its Artificial Analysis Intelligence Index from 40 to 45. The model ID never changed. The number behind it did.
That matters most in one specific place: the comparison between Qwen3.8-Max and Qwen3.7-Max, the Max tier it replaced. Qwen3.8-Max reached general availability on August 3, 2026; Qwen3.7-Max shipped in May. On paper this looks like an easy generational call — newer flagship, higher score, move on. The numbers say something more awkward. The older tier is three to four times faster, roughly five times cheaper per completed task, and identical to the newer one on pure knowledge benchmarks. The newer tier is multimodal, dramatically better at agentic work, and cheaper on the international rate card. Which one you should be calling depends on a question most comparisons skip: whether you are buying knowledge or buying tool calls.
What the Apsara disclosure actually claims
The disclosure is a vendor statement, published by Alibaba on its own press site, and it should be read as one. What it says is specific: over a month of fully automated runs spanning pipeline design, data validation, iterative experimentation and error diagnosis, Qwen3.8-Max completed 33 cycles, and autonomous training optimisation plus post-training techniques produced the score movement. Alibaba also described a second experiment in chip design, where the model ran more than 60 hours of self-improvement across a design lifecycle, made over 10,000 EDA tool calls, and produced production-grade chip bus modules while reducing chip area by 42% with no stated performance compromise. All of that is Alibaba's account of its own model, unreproduced by anyone outside the company.
The score movement itself, though, is not a vendor claim. The Index is Artificial Analysis's, and the 45 is on that third party's Qwen3.8 Max (0902) page. The 40 it moved from is on the same organisation's page for the August 3 build, which is now marked deprecated and points forward to the current one. So the delta is a vendor-reported cause attached to an independently-scored effect, which is a stronger position than either half alone. It is also, as of this writing, the most concrete public evidence anyone has that a frontier lab moved its flagship's measured capability without shipping a new version number.
Two other things about the release are easy to miss and both are load-bearing. First, Alibaba confirmed that Qwen 4 is in training, with the Qwen 4.5 and Qwen 5 series projected to scale to 5–10 trillion parameters. Second, there is a dated snapshot of the refreshed model. Alibaba pinned the post-trained build as qwen3.8-max-0902 on September 2, and it is separately routable. That pin is the practical answer to everything the RSI disclosure implies, and we will come back to it.
The scoreboard
Six dimensions, both sides, with the sourcing discipline kept explicit because the two columns are not equally verified.
• Context window — Qwen3.8-Max 1,000,000 tokens vs Qwen3.7-Max 1,000,000 tokens (identical)
• Max output — 131,072 tokens vs 131,072 tokens per Alibaba's own model pages (identical; note that our catalogue page for Qwen3.7-Max carries 64,000, a disagreement we flag rather than paper over)
• Input modality — text, image and video vs text only
• International list price per 1M tokens — $2.00 in / $6.00 out vs $2.50 in / $7.50 out; the same rate card already charges both models $1.65 / $4.951 in Beijing, Frankfurt and Virginia
• Artificial Analysis Intelligence Index — 45 vs 29
• Independent output speed — 39.7 tokens/sec vs 211.3 tokens/sec, measured against the same 211-model comparison class, where Artificial Analysis grades the newer tier notably slow and the older one notably fast
The last two rows are the whole article. On the same index family, the newer tier is 16 points ahead. On speed, it is more than five times behind.

Where the 16 points come from — and where they don't
Split the Index into its parts and the shape of the upgrade becomes clear, and it is not the shape the marketing suggests.
On knowledge and reasoning, the two models are effectively tied. GPQA Diamond: 92.8 vs 92.3. Humanity's Last Exam: 43.1 vs 40.5. Long-context recall: 80.33 vs 79. SciCode: 52.1 vs 49.5. If your workload is question answering over a large corpus, the generational jump is a rounding error. Paying a multiple for it would be hard to justify.
On agentic and coding work, the gap opens hard. Terminal-Bench 2.1: 88.76 vs 74.53. Artificial Analysis's coding index: 76.2 vs 66. And the widest divergence in the set is the banking tool-use task, where Qwen3.8-Max records 47.84 against Qwen3.7-Max's 11.75 — a four-fold gap in exactly the capability that the RSI description says the automated cycles were optimising: pipeline design, data validation, iterative experimentation, error diagnosis. A model that spends a month improving its own training loop is a model that got better at loops.
One honest caveat on those individual rows. Both model pages sit on the same Intelligence Index revision family (v4.3 and v4.3.2), which makes the headline 45-versus-29 comparison a fair one. The per-benchmark rows are not same-cohort: Qwen3.7-Max's task-level figures date from the May evaluation window, while Qwen3.8-Max's come from the September series. Read the direction of travel, not the third significant figure.
The price question is two questions
On Alibaba's international rate card, the newer Max is cheaper than the one it replaced: $2.00 / $6.00 for Qwen3.8-Max against $2.50 / $7.50 for Qwen3.7-Max, per million tokens. A 20% cut on both directions as the generation advances is unusual and worth noting on its own. Cache economics favour the newer tier too — implicit cache at $0.25 against $0.50, explicit cache read at $0.17 against $0.25.
That is the first question, and it has a regional answer that most coverage flattens. Alibaba charges both models $1.65 input / $4.951 output in Beijing, Frankfurt and Virginia. The 20% gap exists only on the Singapore international card. If your traffic lands in the other three regions, input and output tokens cost the same for both Max tiers today, and the decision collapses to pure capability — at which point the newer model wins on everything except throughput, and there is no arithmetic left to do.
The second question is the trap. Qwen3.7-Max does not currently cost $2.50 / $7.50. It is served at $1.25 / $3.75 — half list, a promotional rate. That makes the previous-generation tier the cheaper option today, by a wide margin. It also means the price you can measure this week is not the price you would be committing to. Alibaba's own model page states plainly that the published table shows original pricing "excluding any limited-time promotions" and directs you to the console for current offers. A promotion is, by construction, a rate that can end. If it does, Qwen3.7-Max does not merely become more expensive than it is today; it becomes 25% more expensive than the model that beats it.
We pass the provider rate through at 0% markup, so both the list price and the promotion are live on our side the same day they change — which is convenient for a comparison and unhelpful if you are trying to budget. Build the model against the list card, and treat the promotional rate as upside.

The number that reverses the cost argument
Token price is not what a task costs. Artificial Analysis publishes a cost-per-task figure that folds in cache economics, reasoning volume and answer length, and on that measure the ranking inverts completely: $5.41 per Intelligence Index task for Qwen3.8-Max against $1.15 for Qwen3.7-Max. A 4.7× premium, against a token rate that is either 20% cheaper or identical depending on your region.
The gap is mostly verbosity. On the same evaluation the newer model generated 190M output tokens against the older one's 120M, and it reasons at length to reach its answers. If you are paying per output token for a high-volume, moderately-difficult workload, the newer Max is not the cheaper model at any token price on either rate card. It is the more expensive one, and the token list price is the wrong instrument for deciding it.
Speed, and what our own traffic shows
The throughput gap is large enough to change architectures. Artificial Analysis places Qwen3.8-Max at 39.7 tokens per second — its summary calls the model notably slow — against 211.3 for Qwen3.7-Max, which it calls notably fast. That is the difference between a model you put behind an interactive surface and one you put behind a queue — and on Qwen3.8-Max it is the first thing our own stat panel shows you.
Our own playground telemetry over the seven days to September 25 tells a slightly more nuanced story on latency, and it comes with a caveat: these are our measurements on our routing, not a controlled benchmark. Qwen3.8-Max showed a median 2,611 ms to first response at 54.7 tokens per second with a 2.45% error rate; the pinned 0902 build showed a 3,360 ms median at 46.2 tokens per second with a notably cleaner 0.66% error rate. Qwen3.7-Max sat at a 4,509 ms median but 169.8 tokens per second with a 3.80% error rate. So the older tier answers more slowly at the front and then streams three times faster, while failing more often. If your workload is bursty, that error-rate difference is worth more attention than the median.
Both tiers are live on one OrcaRouter key alongside the pinned snapshot, with automatic failover across providers. For a comparison like this that is the practical point: measuring your own prompts against both Max generations is a configuration change, not a second contract.

Reproducibility is now the deciding factor
Here is the part of the Apsara disclosure that nobody has priced. A live model ID whose measured capability moved from 40 to 45 over a month, with no version change and no announcement at the time, is a reproducibility hazard. If your product behaviour is a function of a moving ID, you have a model that can improve — or shift — underneath a frozen evaluation suite, a cached prompt library, or a signed-off regression set.
Both generations offer dated snapshots, and that is the mitigation. Qwen3.7-Max is documented by Alibaba as functionally equivalent to qwen3.7-max-2026-05-20, with further dated builds at 2026-05-17 and 2026-06-08. Qwen3.8-Max has qwen3.8-max-0902, the post-trained September build, which is what the 45 refers to. If you are shipping anything that needs to reproduce, pin the dated ID and treat the floating one as a staging target. The RSI cycles are a feature of the floating ID; the pin is how you opt out of them.
One naming detail worth knowing so you do not misread the catalogue. Alibaba lists a third dated form, qwen3.8-max-2026-09-02, alongside the 0902 alias; they refer to the same build. The original August 3 release is no longer routable on our side — we serve Qwen3.8-Max, its 0902 pin, and Qwen3.7-Max.
Who should pick which
The decision does not split on which model is better. It splits on what you are buying.
Stay on Qwen3.7-Max if your workload is text-only, knowledge-heavy rather than tool-heavy, and throughput-sensitive — high-volume classification, extraction, long-context retrieval, batch summarisation. On GPQA Diamond and long-context recall it is within a point of the newer model, it streams three to four times faster, and it costs $1.15 per task against $5.41. It is also the right answer for anyone who wants a cheap second opinion in a fusion or routing configuration, because at its promotional rate it is a different price class entirely. Two caveats: the promotion is a promotion, and Artificial Analysis has already flagged the model as deprecated, superseded by Qwen3.8-Max. Do not start a multi-year commitment on it.
Move to Qwen3.8-Max if any of these is true: you need image or video input, which Qwen3.7-Max does not accept at all; your workload is agentic, with long tool-call chains or multi-step coding loops, where the 88.76 versus 74.53 on Terminal-Bench 2.1 and the four-fold tool-use gap are the difference between shipping and not; or you write against the Singapore international rate card, where the newer model is simply 20% cheaper and there is no trade to weigh. Pin qwen3.8-max-0902 if you need the behaviour to hold still.
If you are in Beijing, Frankfurt or Virginia, the choice is cleaner than any comparison suggests: both Max tiers cost $1.65 / $4.951 per million tokens. Identical. At those rates, paying nothing extra for multimodal input, a 16-point higher Index and a four-fold better tool-use score is not a decision, it is free — and the only reason to stay on Qwen3.7-Max becomes raw throughput.
Two questions worth answering directly
Does the 33-cycle training run mean the model changed under existing API traffic? That is what the disclosure implies, and it is why the dated pin exists. Alibaba describes the improved model as "the updated Qwen3.8-Max," which is the floating ID, not a new one. If you had workloads on the floating ID through September, they were served by a model whose measured capability moved during that window. Alibaba has not published a changelog for the intervening builds, so the only reliable way to know which behaviour you have is to pin.
Is the RSI claim independently verified? The score it produced is — the Index is Artificial Analysis's, and the 45 sits on their page. The mechanism behind it is Alibaba's own description of its own training process, with no external replication, no published evaluation protocol, and no third party observing the 33 cycles. Treat "33 iterative cycles" as a vendor claim and "45 after 40" as a measured pair. They are not the same kind of statement, and the difference is the reason the headline is worth reading carefully rather than repeating.
The next thing to watch is not Qwen 4. It is whether Qwen3.7-Max's half-price promotion survives the arrival of the model that is meant to replace it. If it does not, a great many teams will discover they were comparing a list price against a discount, and that the older of two siblings was the more expensive one all along.
