
GPT-6.1 Sol vs Qwen3.8-Max: The Invoice, Not the Price List
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 161 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 79 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 53 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 301 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Take one nightly repository-maintenance job, run it once a night for a month, and put GPT-6.1 Sol and Qwen3.8-Max on it in turn. On Artificial Analysis's v4.3.2 per-task figures, GPT-6.1 Sol costs $21.72 for those thirty nights and Qwen3.8-Max costs $162.27 — a gap of $140.55 a month, about $1,690 a year, for a job whose per-token price list reads $2.00 in and $10.00 out against $2.00 in and $6.00 out. The per-million rates are within a rounding error of each other on input and the Chinese model is 40% cheaper on output, yet the invoice differs by a factor of 7.5. The vendor released GPT-6.1 Sol on 29 September 2026; Qwen3.8-Max has been live since 3 August 2026.
That gap is not a pricing trick and it is not a benchmark artefact — it is the whole of what this comparison has to say, and it comes from one number that no rate card shows you.
What each model is
GPT-6.1 Sol is OpenAI's current reasoning and coding flagship, shipped a week after the GPT-6 Sol it supersedes, at the same $2.00 / $10.00 list price, with a $0.10 cached-input rate and a 1,050,000-token context window. It is sold for agentic work: its best-for tags on our own model card read reasoning, coding and agentic, and it carries the top-tier quality score.
Qwen3.8-Max is Alibaba's agentic flagship — a closed, API-only model that takes text, image and video input and holds a 1,000,000-token context. Unlike the smaller Qwen releases, this one has no published weights. On Artificial Analysis it sits at 45.42 on the Intelligence Index, 17th of 147 models, with a coding index of 76.2 that ranks 9th in that field. It is not a weak model. It is simply an expensive one to let think.
The number that is not on the rate card

Artificial Analysis records 107,730 output tokens per task for Qwen3.8-Max, of which 70,830 are reasoning tokens. GPT-6.1 Sol produces 38,128 output tokens, 25,190 of them reasoning. The Chinese model writes 2.8 times as many tokens to answer the same question, and it writes them at 40% less per token.
• Output tokens per task — 38,128 vs 107,730; Qwen writes 2.8× more
• Of which reasoning — 25,190 vs 70,830
• Cost per task — $0.724 vs $5.409; Qwen costs 7.5× more
• Time per task — 640 s vs 2,152 s; about 11 minutes vs about 36
• Time to first answer token — 332.0 s vs 55.1 s
• Intelligence Index — 51.83 vs 45.42
• Context window — 1,050,000 vs 1,000,000 tokens
• Inputs accepted — text, image and file vs text, image and video
Do the multiplication and the discount disappears. A model that emits close to three times the tokens at a 40% lower rate does not cost less — it costs about 1.7 times as much on output alone, and the measured per-task figure puts the true multiple at 7.5 once input, cached reads and the length of the productive answer are accounted for.
Note the fourth and fifth lines, because they point in opposite directions and together they explain what Qwen3.8-Max is. It returns its first token in 55 seconds where GPT-6.1 Sol takes five and a half minutes, then spends thirty-six minutes completing the task where the OpenAI model is done in eleven. That is a model that starts talking immediately and thinks at very great length. For a chat interface with a streaming answer that reads as responsiveness. For an unattended nightly job, it is wall-clock time you are also paying for.
Three places Qwen3.8-Max is genuinely ahead
The index gap is 6.4 points across the whole suite, which is the sort of number that gets quoted as if it were uniform. It is not, and the disagreement is instructive:
• GPQA Diamond — Qwen3.8-Max 0.9283 vs GPT-6.1 Sol not published in this run
• GDPval — 1,671.4 vs 1,575.09
• Terminal-Bench 2.1 — 0.8876 vs not published in this run
• AutomationBench — 0.5619 vs 0.6487
• SciCode — not published in this run vs 0.5417 for GPT-6.1 Sol
• CritPT — 0.1771 vs 0.3171
Qwen3.8-Max wins the graduate-level science questions and the occupational-task sample, which is a real result for a model of its generation and the reason its coding index sits 9th out of 138. It loses the harder research-adjacent evaluations — CritPT by 14 points — and it loses AutomationBench by 8.7, which is the one that most resembles unattended computer work. Whichever of those two you believe matters more for your workload is the actual decision; the 6.4-point headline is an average over a disagreement, not a verdict.
What the extra tokens buy, and what they do not
Long reasoning traces are not waste by definition. A model that spends 70,830 tokens of deliberation on a hard task is doing something the shorter model may be skipping, and on the evaluations where Qwen3.8-Max is ahead, that deliberation is plausibly why. The failure mode is that you pay for it on every task, not only the hard ones. Reasoning-token counts are a property of the model's calibration, not of the question's difficulty, and a model that over-thinks a one-line extraction bills you for it.
The way to find out which of your tasks is which is to stop treating the model choice as global. Which brings us to the part that is a config change.
Our own model card for GPT-6.1 Sol carries the subject's served numbers: the $2.00 / $10.00 rate, the 1,050,000-token context window, a 4.54-second p50 to first token, and a stored Artificial Analysis index block on the same page as the pricing an application would actually route against.

One key, one meter, two models
Both models are reachable through a single OrcaRouter endpoint — one API for 200+ models, provider list price passed through at 0% markup. The OrcaRouter card for Qwen3.8-Max carries the served rate alongside the 1,000,000-token context window, the text/image/video input modalities and the 101.4-million-token seven-day volume — the numbers an application routes against, next to the ones the article argues about. That pass-through matters specifically for Qwen3.8-Max, because a model whose economics are dominated by output-token volume is one whose vendor can reprice it at any time, and a pass-through meter means the change reaches your invoice the day Alibaba publishes it rather than on a reseller's schedule.

The more interesting use of the routing layer here is the routing DSL, which lets you compose models into one call rather than choosing one for the whole application. Split the nightly job: a cheap model classifies and extracts, GPT-6.1 Sol handles the reasoning and verification steps, and Qwen3.8-Max is reserved for the tasks where its long deliberation and its GPQA-class science strength actually change the answer. Automatic failover covers the rest — if a provider degrades mid-run, the request moves rather than failing the batch.
Neither of these is a reason to pick a model. They are the reason you do not have to make the pick once and live with it.
The call, and the thing to watch
Pay the 7.5 times only where the deliberation earns it. On a general-purpose nightly job, GPT-6.1 Sol at $0.724 a task is the defensible default and the $1,690 a year is real. Where the task is hard science or occupational reasoning with a checkable answer, Qwen3.8-Max's 0.9283 on GPQA Diamond is worth the tokens — that is the specific thing it does better.
The thing to watch is whether Alibaba reprices. A 2.8×-token model at $2.00/$6.00 is priced as if tokens were the scarce thing; the model's own behaviour makes tokens abundant. If the output rate moves materially, the arithmetic in this article moves with it, and the pass-through meter will show you that on the same day it happens.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
