
GPT-6 Sol vs Qwen 4 Max: The Input Price Is a Tie and Almost Everything Else Isn't Comparable
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 58 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 347 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Here is the number that makes the GPT-6 Sol and Qwen 4 Max pairing worth its own page: both are listed at $2.00 per million input tokens. That is where the usable part of the comparison starts and, for Qwen 4 Max, very nearly where it ends. The company announced the Qwen 4 Max tier on 22 September 2026 at its Apsara conference — the same day GPT-6 Sol was released — and as of this writing there is no published model card for it: no price, no context window, no weights statement, and no score from any evaluator. The Qwen flagship that does have measurements, and that you can call today, is Qwen3.8 Max, released 3 August 2026 at $2.00 / $6.00. So the practical question is not how Qwen 4 Max performs against GPT-6 Sol. It is what the announced tier tells you about where the company's flagship is going, and what the Qwen you can actually route does against the GPT-6 model today.
That distinction is not pedantry. A tier that has been announced but not specified can absorb any assumption you want to make about it, and an article that quietly assumes Qwen 4 Max beats GPT-6 Sol — because the version number is higher — is making an argument the vendor has not made and the evidence does not support.
Three things that are comparable right now
• Input price — GPT-6 Sol $2.00 per million vs Qwen3.8 Max $2.00 per million, an exact tie
• Output price — GPT-6 Sol $10.00 per million vs Qwen3.8 Max $6.00 per million
• Cached input — GPT-6 Sol $0.20 per million vs Qwen3.8 Max $0.25 per million
• Context window — GPT-6 Sol 1,050,000 tokens vs Qwen3.8 Max 1,000,000 tokens
• Input modalities — both take text, image and video or file; GPT-6 Sol text, image and file, Qwen3.8 Max text, image and video
• Long-request step — GPT-6 Sol reprices the whole request above 272,000 input tokens to $4.00 / $15.00 vs Qwen3.8 Max no step on its card
• Intelligence Index — GPT-6 Sol 47.6 vs Qwen3.8 Max 45.4
• Coding index — GPT-6 Sol not published on this revision vs Qwen3.8 Max 76.2
• GPQA Diamond — GPT-6 Sol 96.1 vs Qwen3.8 Max 92.8
• Terminal-Bench 4.0 — GPT-6 Sol 43.9 vs Qwen3.8 Max 38.9
Two of those rows are the ones to sit with. The cached-input line runs against the pattern — Qwen3.8 Max charges five cents more per million for cache reads than GPT-6 Sol does, which on a caching-heavy agent loop is the meter that moves first. And the long-request clause is where the tie on input price breaks for good: at exactly $2.00 in, the two look interchangeable until a prompt crosses 272,000 tokens, at which point GPT-6 Sol's whole request reprices to $4.00 / $15.00 while Qwen3.8 Max's card carries no equivalent step at all. On a full-context prompt the company's model is not merely cheaper on output. It is the only one of the two whose published rate you can still predict from the headline.
Two things that are not comparable, and why the difference matters
The first is Qwen 4 Max's capability. The company presented it as a tier in a lineup rather than as a product: no benchmark table, no context figure, no statement about weights, no rate. Everything written about how it will perform is a projection from the Qwen line's trajectory, including this sentence. The one thing that is legible is the target — the tier it replaces and the model the successor is being built to overtake — and that tells you the direction of travel, not the destination.
The second is the open-weight question, and it is the one most worth flagging before you plan around a Qwen 4 Max deployment. The company publishes weights for parts of the Qwen family and not for others. Qwen3.8 Max, the flagship, is listed as a closed API model — no licence file, no downloadable checkpoint — even though smaller siblings in the same generation are published. Whether the 4-series flagship follows the smaller siblings or stays closed is unanswered, and a team that builds a self-hosting plan around the expectation of open weights would be betting on a precedent that the current flagship does not set.
The comparison that actually runs is GPT-6 Sol against Qwen3.8 Max, and it is closer than the generation numbers suggest: 47.6 against 45.4 on the Intelligence Index, both figures third-party on the same revision, and GPQA Diamond a genuine 96.1 against 92.8. A 2.2-point index gap at three-quarters of the output price is the tightest version of this matchup available, and it is the one to reason from until the company publishes something for the 4-series.

What the announcement itself is evidence of
An announced-but-unspecified tier is not an empty signal. It tells you that the company intends its next flagship to sit at the top of the line and that the company is willing to name it publicly before it is ready, which is a change in communication habit rather than a capability claim. It also sets a clock: a tier announced in September creates an expectation of a shipping window, and the gap between announcement and specification is where competitive positioning usually happens.
For anyone currently running Qwen3.8 Max, the practical consequence is that the migration will be a capability question rather than a cost question. The tier Qwen 4 Max is being built to replace is already at $2.00 / $6.00, and the successor has to beat a model that scores 45.4 while costing a third less on output than GPT-6 Sol. Whether it also closes the gap to 47.6 — and on what date — is the only thing worth tracking, and there is nothing in the announcement that answers it.

Routing a model that does not exist yet
Only one of the two Qwen models in this conversation is on our catalogue, and it is not the 4-series. Qwen3.8 Max is routed through OrcaRouter under the model id qwen/qwen3.8-max, on the same key as GPT-6 Sol and 200+ other models, and the pairing is worth exactly that: a route change rather than an integration project when the company publishes the successor. Qwen 4 Max will be added when there is something to route — a card, a rate, an endpoint — and not before, because there is no responsible way to sell a slot against a specification that has not been written.
In the meantime, the reason to keep both columns open is that the two models disagree about where they are cheap. Qwen3.8 Max gives up nothing on input price and takes a third off GPT-6 Sol's output rate and has no long-context cliff; GPT-6 Sol carries a marginally better index and a much richer cached-input discount, at the cost of repricing the whole request above 272,000 tokens. A workload that is long and output-heavy and a workload that is short and cache-friendly want different answers, and both are one routing rule apart rather than two procurement conversations.
Failover matters here for a reason specific to this matchup. Qwen3.8 Max is a high-volume model on our routes — more than 110 million tokens in a seven-day window — and volume is what exposes endpoint variance. A second configured route behind it means a degraded primary is a retry, and a rate change on the company's card, which is exactly the kind of event a flagship handover produces, is live on our side the same day because our catalogue passes provider list prices through with no markup.

What to do with this, this month
If you need a Qwen flagship now, the one to call is Qwen3.8 Max, and it is competitive with GPT-6 Sol on reasoning at a meaningfully lower output rate — that is the honest state of the matchup today. If you need the strongest published result on the reasoning benchmarks in this pair, GPT-6 Sol's 47.6 is the number, and its cached-input discount is the cheapest in the comparison. If you are waiting for Qwen 4 Max, wait for a card rather than for a date, because the announcement gave you the tier and nothing else.
And whichever you pick, check the GPT-6 side before you commit to it. GPT-6 Sol has already been superseded by GPT-6.1 Sol at the same $2.00 / $10.00 with cached input halved to $0.10, and its own model page now redirects readers to it. A comparison against GPT-6 Sol is a comparison against a model whose successor is on sale, and the same caution applies to any Qwen figure derived from the 4 Max announcement: both are shapes of news, and only one of them is a purchasable product right now.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
