A generated hero card titled 'GPT-6.1 Sol vs Qwen3.8-Max' with two rounded panels: left 'GPT-6.1 Sol' listing 'OpenAI - released 29 Sep 2026', '$2.00 in / $10.00 out per 1M', '$0.72 per task' and 'AA Index 51.8'; right 'Qwen3.8-Max' listing 'Alibaba - released 3 Aug 2026', '$2.00 in / $6.00 out per 1M', '$5.41 per task' and 'AA Index 45.4'; below them the strip 'Same input price. 7.5x the invoice.'; footer 'Figures: Artificial Analysis v4.3.2.' and the OrcaRouter logo bottom-right.
Guides & Insights

GPT-6.1 Sol vs Qwen3.8-Max: The Invoice, Not the Price List

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Take one nightly repository-maintenance job, run it once a night for a month, and put GPT-6.1 Sol and Qwen3.8-Max on it in turn. On Artificial Analysis's v4.3.2 per-task figures, GPT-6.1 Sol costs $21.72 for those thirty nights and Qwen3.8-Max costs $162.27 — a gap of $140.55 a month, about $1,690 a year, for a job whose per-token price list reads $2.00 in and $10.00 out against $2.00 in and $6.00 out. The per-million rates are within a rounding error of each other on input and the Chinese model is 40% cheaper on output, yet the invoice differs by a factor of 7.5. The vendor released GPT-6.1 Sol on 29 September 2026; Qwen3.8-Max has been live since 3 August 2026.

That gap is not a pricing trick and it is not a benchmark artefact — it is the whole of what this comparison has to say, and it comes from one number that no rate card shows you.

What each model is

GPT-6.1 Sol is OpenAI's current reasoning and coding flagship, shipped a week after the GPT-6 Sol it supersedes, at the same $2.00 / $10.00 list price, with a $0.10 cached-input rate and a 1,050,000-token context window. It is sold for agentic work: its best-for tags on our own model card read reasoning, coding and agentic, and it carries the top-tier quality score.

Qwen3.8-Max is Alibaba's agentic flagship — a closed, API-only model that takes text, image and video input and holds a 1,000,000-token context. Unlike the smaller Qwen releases, this one has no published weights. On Artificial Analysis it sits at 45.42 on the Intelligence Index, 17th of 147 models, with a coding index of 76.2 that ranks 9th in that field. It is not a weak model. It is simply an expensive one to let think.

The number that is not on the rate card

A generated two-column scoreboard titled 'GPT-6.1 Sol vs Qwen3.8-Max - the scoreboard'. Left column 'GPT-6.1 Sol': Output tokens 38,128, Reasoning tokens 25,190, Cost per task $0.72, AA Index 51.8, Time per task 640 s, Context 1.05M. Right column 'Qwen3.8-Max': Output tokens 107,730, Reasoning tokens 70,830, Cost per task $5.41, AA Index 45.4, Time per task 2,152 s, Context 1M. Footer: 'Both columns Artificial Analysis v4.3.2.' OrcaRouter logo bottom-right.

Artificial Analysis records 107,730 output tokens per task for Qwen3.8-Max, of which 70,830 are reasoning tokens. GPT-6.1 Sol produces 38,128 output tokens, 25,190 of them reasoning. The Chinese model writes 2.8 times as many tokens to answer the same question, and it writes them at 40% less per token.

• Output tokens per task — 38,128 vs 107,730; Qwen writes 2.8× more

• Of which reasoning — 25,190 vs 70,830

• Cost per task — $0.724 vs $5.409; Qwen costs 7.5× more

• Time per task — 640 s vs 2,152 s; about 11 minutes vs about 36

• Time to first answer token — 332.0 s vs 55.1 s

• Intelligence Index — 51.83 vs 45.42

• Context window — 1,050,000 vs 1,000,000 tokens

• Inputs accepted — text, image and file vs text, image and video

Do the multiplication and the discount disappears. A model that emits close to three times the tokens at a 40% lower rate does not cost less — it costs about 1.7 times as much on output alone, and the measured per-task figure puts the true multiple at 7.5 once input, cached reads and the length of the productive answer are accounted for.

Note the fourth and fifth lines, because they point in opposite directions and together they explain what Qwen3.8-Max is. It returns its first token in 55 seconds where GPT-6.1 Sol takes five and a half minutes, then spends thirty-six minutes completing the task where the OpenAI model is done in eleven. That is a model that starts talking immediately and thinks at very great length. For a chat interface with a streaming answer that reads as responsiveness. For an unattended nightly job, it is wall-clock time you are also paying for.

Three places Qwen3.8-Max is genuinely ahead

The index gap is 6.4 points across the whole suite, which is the sort of number that gets quoted as if it were uniform. It is not, and the disagreement is instructive:

• GPQA Diamond — Qwen3.8-Max 0.9283 vs GPT-6.1 Sol not published in this run

• GDPval — 1,671.4 vs 1,575.09

• Terminal-Bench 2.1 — 0.8876 vs not published in this run

• AutomationBench — 0.5619 vs 0.6487

• SciCode — not published in this run vs 0.5417 for GPT-6.1 Sol

• CritPT — 0.1771 vs 0.3171

Qwen3.8-Max wins the graduate-level science questions and the occupational-task sample, which is a real result for a model of its generation and the reason its coding index sits 9th out of 138. It loses the harder research-adjacent evaluations — CritPT by 14 points — and it loses AutomationBench by 8.7, which is the one that most resembles unattended computer work. Whichever of those two you believe matters more for your workload is the actual decision; the 6.4-point headline is an average over a disagreement, not a verdict.

What the extra tokens buy, and what they do not

Long reasoning traces are not waste by definition. A model that spends 70,830 tokens of deliberation on a hard task is doing something the shorter model may be skipping, and on the evaluations where Qwen3.8-Max is ahead, that deliberation is plausibly why. The failure mode is that you pay for it on every task, not only the hard ones. Reasoning-token counts are a property of the model's calibration, not of the question's difficulty, and a model that over-thinks a one-line extraction bills you for it.

The way to find out which of your tasks is which is to stop treating the model choice as global. Which brings us to the part that is a config change.

Our own model card for GPT-6.1 Sol carries the subject's served numbers: the $2.00 / $10.00 rate, the 1,050,000-token context window, a 4.54-second p50 to first token, and a stored Artificial Analysis index block on the same page as the pricing an application would actually route against.

A screenshot of the OrcaRouter model page for openai/gpt-6.1-sol, showing a 1,050,000-token context window, 128K max output, text/image/file input with text output, the $2.00 input and $10.00 output per-million prices, a 4.54 s p50 time to first token and 284.2M tokens routed over seven days, above the OpenAI-compatible code sample and the supported-parameters list.

One key, one meter, two models

Both models are reachable through a single OrcaRouter endpoint — one API for 200+ models, provider list price passed through at 0% markup. The OrcaRouter card for Qwen3.8-Max carries the served rate alongside the 1,000,000-token context window, the text/image/video input modalities and the 101.4-million-token seven-day volume — the numbers an application routes against, next to the ones the article argues about. That pass-through matters specifically for Qwen3.8-Max, because a model whose economics are dominated by output-token volume is one whose vendor can reprice it at any time, and a pass-through meter means the change reaches your invoice the day Alibaba publishes it rather than on a reseller's schedule.

A screenshot of the OrcaRouter model page for qwen/qwen3.8-max, showing the $2.00 input and $6.00 output per-million prices, a 5.66 s p50 time to first token, 101.4M tokens over seven days, a 1,000,000-token context window and text/image/video input, above the OpenAI- and Anthropic-compatible code samples.

The more interesting use of the routing layer here is the routing DSL, which lets you compose models into one call rather than choosing one for the whole application. Split the nightly job: a cheap model classifies and extracts, GPT-6.1 Sol handles the reasoning and verification steps, and Qwen3.8-Max is reserved for the tasks where its long deliberation and its GPQA-class science strength actually change the answer. Automatic failover covers the rest — if a provider degrades mid-run, the request moves rather than failing the batch.

Neither of these is a reason to pick a model. They are the reason you do not have to make the pick once and live with it.

The call, and the thing to watch

Pay the 7.5 times only where the deliberation earns it. On a general-purpose nightly job, GPT-6.1 Sol at $0.724 a task is the defensible default and the $1,690 a year is real. Where the task is hard science or occupational reasoning with a checkable answer, Qwen3.8-Max's 0.9283 on GPQA Diamond is worth the tokens — that is the specific thing it does better.

The thing to watch is whether Alibaba reprices. A 2.8×-token model at $2.00/$6.00 is priced as if tokens were the scarce thing; the model's own behaviour makes tokens abundant. If the output rate moves materially, the arithmetic in this article moves with it, and the pass-through meter will show you that on the same day it happens.

Compared in this article3

Detected from this article · Benchmarks: Artificial Analysis · updated daily