
Solar Mini 4 vs Qwen3.8-Max: What the Last Twenty Index Points Actually Cost
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-29$2.00 / $10.00 per 1M tokens
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-28$2.00 / $10.00 per 1M tokens · 159 tok/s
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 220 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 118 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 102 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 217 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Solar Mini 4 costs $0.10 per million input tokens and $0.40 per million output tokens. Qwen3.8-Max, the model most people mean when they type "Qwen 3.8" and ask for the API, costs $2.00 and $6.00. On a rate card that reads as a twenty-fold difference in input price. On the independent evaluation both models have now been through, the honest figure is that Qwen3.8-Max scores 45.4 on the Artificial Analysis Intelligence Index against Solar Mini 4's 24.1 — roughly twice the measured capability — while costing about fifteen times as much per completed task. Everything worth knowing about this pairing is in the gap between those two ratios, and in the fact that Upstage's model runs at 204 tokens per second while Alibaba's flagship runs at 39.
One naming note before the numbers, because it changes what you are buying. Qwen 3.8 is a generation, not a model: the hosted flagship is Qwen3.8-Max, refreshed as a dated 0902 snapshot, and the downloadable member of the same generation is Qwen3.8-2.4T-A95B, the 2.4-trillion-parameter weights Alibaba published on 12 August 2026. Solar Mini 4, released 22 September 2026, is Upstage's 35-billion-parameter sparse mixture-of-experts with about 3 billion parameters active per token and no open weights at any price.
The two columns, on the dimensions that decide procurement
• Intelligence Index — 45.4 for Qwen3.8-Max (0902) against 24.1 for Solar Mini 4, both Intelligence Index v4.3.2.
• Cost per completed task — $5.41 for Qwen3.8-Max against $0.36 for Solar Mini 4. Roughly fifteen to one, and the ratio nobody quotes.
• Rate card — $2.00 / $6.00 per million tokens, cached input $0.25, for Qwen3.8-Max against $0.10 / $0.40, cached input $0.01, for Solar Mini 4.
• Context — 1,000,000 tokens on both, which is the one dimension where the cheap model gives nothing away. Artificial Analysis lists 983,616 for Qwen3.8-Max and 1,048,576 for Solar Mini 4; Upstage's own documentation says 512K, so the small model's ceiling is the less certain of the two.
• Weights — Qwen3.8-Max is proprietary hosted; its open sibling Qwen3.8-2.4T-A95B is downloadable under a custom licence and scores 39.9 on the same index at 262,000 tokens of context. Solar Mini 4 is proprietary with no open sibling.
• Throughput — 204 output tokens per second for Solar Mini 4 against 39.5 for Qwen3.8-Max, as measured by the same lab. The model that costs a fifteenth as much per task is five times faster doing it.
What twenty-one index points buy

They buy the hard end of the distribution. On Humanity's Last Exam, Qwen3.8-Max answers correctly on 43.1% of items against Solar Mini 4's 25.8%. On scientific coding it is 52.1% against 47.6%. On agentic knowledge work the gap widens into a different category: Qwen3.8-Max reaches 1663 Elo on GDPval-AA, where Solar Mini 4 manages 1072 — nearly six hundred Elo points, which is not a rounding difference but a change in what the model can be trusted to do unattended.
The one evaluation where the small model does not just hold its ground but wins is long-context reasoning. Solar Mini 4 scores 83.3% on AA-LCR v1.1 and Qwen3.8-Max scores 80.3%. That is the strongest argument for Solar Mini 4 in this entire comparison: on the benchmark built to test "read a very long document and reason across it," fifteen times the price per task buys a result that is three points worse.
Where neither model is the right instrument is autonomous terminal work. Solar Mini 4 scores 1% on Terminal-Bench 4.0; Qwen3.8-Max scores 38.9%. A model that clears eight of twenty hard terminal tasks is genuinely usable there and the other one is not, which is the clearest example in this comparison of a gap in kind rather than in degree.
Why the flagship's per-task bill is so much worse than its token price suggests
Qwen3.8-Max emits more output per task than Solar Mini 4 does — 107,700 tokens against 88,300 — and it emits them at $6.00 per million rather than $0.40. That is most of the gap right there, before cache behaviour is considered. Alibaba does offer a $0.25 cached-input rate, which is a real discount against its $2.00 input price, and the family's open sibling, Qwen3.8-2.4T-A95B, is a 2.4-trillion-parameter sparse model with about 95 billion parameters active, so every token it produces is being produced by a substantially larger machine. None of that is a criticism. It is the reason a fifteen-fold per-task multiplier exists, and the reason a per-token price comparison between these two models is misleading in both directions.
The practical consequence: Solar Mini 4 is not "the budget Qwen 3.8." It is a specialist that happens to be cheap, fast, and strong on exactly one axis — long documents — and weak on the axis that costs the most to get wrong, which is agentic reliability. Qwen3.8-Max is the generalist whose advantage shows up precisely where a wrong answer is expensive.
Running the pair instead of choosing between them
This is the unusual case in the Solar Mini 4 series where both sides of the comparison are routes we can actually sell you. Qwen3.8-Max and its dated 0902 snapshot sit on OrcaRouter at the vendor's own $2.00 / $6.00 list, alongside Qwen3.8-Flash at $0.15 / $0.47 and Qwen3.8-27B at $0.33 / $2.40 — three rungs of the same generation, one key, no second contract and no code change beyond a model string. As of 1 October 2026 our own playground measurement puts Qwen3.8-Max at a 2.5-second p50 with 87.7 output tokens per second over the trailing seven days; that window rolls, so read the live model card rather than this sentence. Upstage's models are not available through us, and Solar Mini 4 has to be reached through Upstage's own API.

What that makes possible is the split this comparison argues for: send the long-document, Korean-language and structured-extraction traffic somewhere cheap, route the reasoning-heavy and tool-using requests to the flagship, and let failover cover the provider that has a bad afternoon. It is also where the routing DSL earns its place — composing models into a single call means a cheap first pass and an expensive verification pass can be one request against one endpoint instead of two integrations that must be kept in sync.
The short version

If your workload is long context, a fixed output shape and a tolerable latency budget, Solar Mini 4 gives you 83% of the long-context result at a fifteenth of the task cost and five times the throughput, and the twenty-one missing index points will never appear on your dashboard. If your workload involves tools, multi-step agents or questions where a confident wrong answer is worse than no answer, the missing points are the entire product and Qwen3.8-Max is worth fifteen times the price. Every figure attributed to Artificial Analysis above is that organisation's independent measurement; the Qwen3.8 naming and release dates are Alibaba's, and the OrcaRouter serving figures are our own playground measurement on a rolling seven-day window.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
