A hero title card for an article comparing GPT-6 Sol with Qwen3.8-Max, reading 'GPT-6 Sol vs Qwen3.8-Max' with the subtitle 'The one line where the closed model cannot compete', showing GPT-6 Sol as text and image input and Qwen3.8-Max as text, image and video input.
Guides & Insights

GPT-6 Sol vs Qwen3.8-Max: The One Line Where the Closed Model Cannot Compete

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Almost every comparison between GPT-6 Sol and Qwen3.8-Max will be decided on cost, and almost every one of them will get the cost wrong in the same direction. Qwen3.8-Max, the vendor's hosted flagship, has been priced at $2.00 per million input tokens and $6.00 per million output since its August 3, 2026 general availability, with the dated Qwen3.8-Max-0902 snapshot landing on September 2. GPT-6 Sol reached general availability on September 22, 2026 at $2.00 in and $10.00 out. Identical input price, and 40% cheaper output on the Qwen3.8-Max rate card — and, on the independent harness, roughly five times the cost to finish a task.

But there is a second, cleaner difference that the cost argument keeps burying, and it is the one that actually decides some workloads. Qwen3.8-Max accepts video. GPT-6 Sol does not. That is not a price line or a benchmark line; it is a capability one of these models has and the other cannot approximate, and it is the only part of this matchup that no amount of token-efficiency work will fix.

A six-row scoreboard card titled 'GPT-6 Sol vs Qwen3.8-Max - the scoreboard', comparing the two models on output price, Intelligence Index, cost per task, input modality, maximum output and long-context handling. The left column lists GPT-6 Sol at $10.00 output, an Index of 48, $1.06 per task, text and image input, a 128,000-token maximum output and whole-request repricing above 272K; the right column lists Qwen3.8-Max at $6.00 output, an Index of 45, $5.41 per task, text, image and video input, a 131,072-token maximum output and a flat tier above 272K. A footer reads 'Vendor list rates; Index per Artificial Analysis.'

Two flagships, two different shapes

Qwen3.8-Max is a 2.4-trillion-parameter sparse mixture-of-experts model with roughly 95 billion parameters active per query — the second-largest hosted model in production behind Kimi K3. Its context window is one million tokens with a 131,072-token output cap, its input modalities are text, image and video, and it supports reasoning effort across low, medium and extra-high with structured outputs and tool calling. The hosted flagship is not open weights; the downloadable artifact of the same generation is a separate release with its own limitations.

GPT-6 Sol is text and image in, text out. Its window is 1,050,000 tokens with a 922,000-token maximum input and a 128,000-token output ceiling, priced at a flat $2.00 and $10.00 with a $0.20 cached read and $2.50 cache writes, repricing the whole request above 272,000 input tokens to $4.00 and $15.00. It is closed, single-identifier, and OpenAI has described the rates as permanent rather than promotional.

The window figures are within about 7% of each other, and the output ceilings are in the same band. The modality list is the only structural gap — and it runs one way.

Line by line

• Input — GPT-6 Sol $2.00 per million tokens vs Qwen3.8-Max $2.00 per million tokens; the headline number is identical

• Output — GPT-6 Sol $10.00 per million tokens vs Qwen3.8-Max $6.00 per million tokens; Alibaba's rate is 40% lower

• Cached input — GPT-6 Sol $0.20 per million tokens vs Qwen3.8-Max $0.25 per million tokens for implicit cache reads, with an explicit cache read at $0.17 and an explicit cache write at $2.50

• Context window — GPT-6 Sol 1,050,000 tokens vs Qwen3.8-Max 1,000,000 tokens

• Maximum output — GPT-6 Sol 128,000 tokens vs Qwen3.8-Max 131,072 tokens

• Input modalities — GPT-6 Sol text and image vs Qwen3.8-Max text, image and video

• Long-context handling — GPT-6 Sol reprices the whole request above 272,000 input tokens at 2× input and cache rates and 1.5× output vs Qwen3.8-Max a flat tier across the full window, which is the one place the Qwen rate card is structurally better rather than marginally cheaper

• Reasoning effort — GPT-6 Sol none, low, medium (default), high, xhigh, max vs Qwen3.8-Max low, medium and extra-high

• Knowledge cutoff — GPT-6 Sol April 20, 2026 vs Qwen3.8-Max not published as a single date

• Dated snapshot — GPT-6 Sol a single rolling identifier vs Qwen3.8-Max-0902 available as a pinned September 2, 2026 revision at the same price

The long-context line deserves more attention than it usually gets. GPT-6 Sol's surcharge is a step applied to the entire request, so a 900,000-token agent run bills at $4.00 and $15.00 on every token. Qwen3.8-Max charges one tier across the whole window. For a workload that lives above 272,000 tokens, the two rate cards stop being comparable at all — the cheaper model on the headline is the more expensive one by a wide margin, and the direction of the gap depends entirely on where your context sits.

Screenshot of the Artificial Analysis model page for Qwen3.8 Max (0902), captured 23 September 2026, showing an Intelligence Index score of 45, list pricing of $2.00 per million input tokens and $6.00 per million output tokens, a cost of $5.41 per Intelligence Index task, 190 million output tokens generated during the index run, a 984,000-token context window and 39.2 output tokens per second.

The rate card says 40% cheaper. The harness says five times the bill.

Artificial Analysis measured Qwen3.8-Max-0902 at an Intelligence Index score of 45, costing $5.41 per completed index task, having generated 190 million output tokens across the run. Its summary describes the model as "notably slow and very verbose." GPT-6 Sol scored 48 at $1.06 per task on 77 million output tokens.

Read those three pairs together and the shape of the matchup appears. Qwen is three index points behind, generated two and a half times the tokens, and cost about five times as much per finished task — on a rate card where its output is 40% cheaper. The mechanism is not subtle. At $6.00 per million output tokens, 190 million tokens costs $1.14 per index run in output alone before input is counted; at $10.00 per million, Sol's 77 million costs $0.77. The cheaper rate is buying more tokens, not a smaller bill.

This is the same pattern that shows up across the September field, and it is worth stating as a rule rather than a fact about these two models: on a per-token rate card, a 40% output discount is worth nothing if the model emits two and a half times as many tokens. Cost per completed task is the only figure that survives.

What Alibaba published, and the effort-setting problem

Alibaba's own benchmark table is the longest in this comparison and the least comparable. Vendor-reported figures put Qwen3.8-Max ahead on PaperBench and IFBench while trailing on SWE-bench Pro and Humanity's Last Exam, with a reasoning effort scale — low, medium, extra-high — that has no direct mapping to OpenAI's six-level scale. A score published at extra-high is not the same product as a score published at medium, and neither vendor labels which one produced a given number on a comparison chart.

That is the honest reason not to lean on either vendor's table here. Alibaba's numbers are run on Alibaba's harness at Alibaba's effort setting; OpenAI's are run on OpenAI's. The Artificial Analysis figures exist precisely because both of those are unusable as a head-to-head, and they are the ones used above.

Reproducibility is a real feature, and it is priced at zero

The dated snapshot is the most underrated line in this comparison. Qwen3.8-Max-0902 is a pinned September 2, 2026 revision that Alibaba says carries the same capabilities and pricing as the base model. Pinning it means the model behind your endpoint does not change under you between a Tuesday and a Thursday.

GPT-6 Sol has no equivalent. It is a single rolling identifier — the same model page carries a default snapshot and no dated variants — which means the thing you benchmarked in September is not guaranteed to be the thing you are calling in November. For most teams that is fine. For a regulated pipeline that has to demonstrate what produced an output, it is a real difference, and it is one Alibaba gives away for free while OpenAI does not offer it at all.

The video line is the one that decides

Everything above is a comparison. This part is not. If your input includes video — screen recordings, inspection footage, lecture capture, anything where the information is in motion rather than in a still frame — Qwen3.8-Max accepts it natively and GPT-6 Sol has no path to it. You cannot prompt your way around a modality. You cannot pre-process a video into images and get the same thing, because temporal information is the point, and no amount of index points on a text benchmark substitutes for the input your task actually requires.

The reverse case is worth naming too, because it is where the comparison stops being lopsided. If your input is text and images, Sol is cheaper per task, four points ahead on the neutral board, and cheaper on cached reads. The video capability is not a tiebreaker in your favour; it is simply unused.

Routing a decision that has two separate answers

Screenshot of OrcaRouter's model page for Qwen3.8 Max (0902), captured 23 September 2026, showing the model id qwen/qwen3.8-max-0902 from provider Qwen, a 1M-token context window with a 131k-token maximum output, text, image and video input with text output, list pricing of $2.00 per million input tokens and $6.00 per million output tokens, and a p50 time to first token of 3.50 seconds.

This is a matchup where the right architecture is genuinely two models, not one, and the reason is that the split is by input modality rather than by difficulty. A pipeline that ingests video and reasons over text needs both, and the routing rule is a property of the request rather than a judgement call.

Qwen3.8-Max is on OrcaRouter's catalogue — the dated qwen/qwen3.8-max-0902 snapshot is routable at Alibaba's list price under the pass-through model, which means an Alibaba rate change is live on our side the same day rather than after a reseller renegotiates. GPT-6 Sol is not in our catalogue as of this writing and is reachable through OpenAI's own API. The routing DSL is the piece that fits this specific case: a rule that sends video-bearing requests one way and everything else the other is exactly what it is for, and automatic failover means the modality branch does not become the single point of failure.

Where each one wins

• Video in, text out — Qwen3.8-Max, and there is no contest to describe.

• Text and image in, cost per completed task as the metric — GPT-6 Sol, by roughly five times on the independent harness.

• Cached-prefix work — GPT-6 Sol. Its cached read is marginally cheaper and its token volume is less than half.

• Requests above 272,000 input tokens — Qwen3.8-Max. Sol's whole-request repricing is the single largest cost difference in this comparison, and it is invisible on the headline rates.

• Reproducible pinning — Qwen3.8-Max-0902, which costs nothing extra and is the only pinned revision of the two.

• If you were shown a chart with Qwen ahead on SWE-bench Pro — check whose harness produced it and at which effort setting. The independent board has Qwen three points behind on the composite, and the vendor tables are not on the same scale as each other.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily