
MiMo-V2.6-Pro vs Qwen3.8-Max: One Index Point, Six Times the Price
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Qwen3.8-Max is a 2.4-trillion-parameter model that activates 95 billion of them per token and runs on a single provider. Xiaomi MiMo-V2.6-Pro is a 1.02-trillion-parameter model that activates 42 billion per token, scores one point higher on the same independent index, and costs about a sixth as much to call. Those facts sit together awkwardly, and how you resolve them is the whole decision between the two. The short version: on Artificial Analysis Intelligence Index v4.3.2, Xiaomi MiMo-V2.6-Pro scores 46 and Qwen3.8-Max scores 45 — a tie inside the noise — while the price sheet reads $0.43 and $0.87 per million tokens against $2.00 and $6.00.
What each model actually is
Qwen3.8-Max is Alibaba's flagship, generally available since 3 August 2026, and it was the largest model the Qwen line had shipped when it landed. It is a sparse mixture-of-experts built on the Qwen 3.5 architecture at 2.4 trillion total parameters with roughly 95 billion active per token, a 1M-token context window, up to 131,000 tokens of output, and multimodal input across text, image and video. Alibaba cut the international rate to $2.00 per million input tokens and $6.00 per million output, with cached input at $0.25, and pitched that as roughly 40% of Claude Opus 5's input price. A point update, Qwen3.8-Max-0902, followed on 2 September 2026 and is what Artificial Analysis currently benchmarks at 45.
Xiaomi MiMo-V2.6-Pro reached general availability on 22 September 2026, three days before this comparison was written, with its reinforcement-learning checkpoint repositories appearing the day before. It is a sparse mixture-of-experts at 1.02 trillion total parameters and 42 billion active per token, with the same 1M-token context window, native multimodality across text, image, video and audio, and weights published under the MIT licence. Xiaomi held the previous generation's pricing rather than repricing the new checkpoint upward, which is why it lists at $0.43 and $0.87 with a 99% cache discount and a blended $0.18.
• Active compute per token — Qwen3.8-Max 95B; Xiaomi MiMo-V2.6-Pro 42B, less than half
• Total parameters — Qwen3.8-Max 2.4T; Xiaomi MiMo-V2.6-Pro 1.02T
• Index score — Xiaomi MiMo-V2.6-Pro 46; Qwen3.8-Max 45, both on Artificial Analysis v4.3.2
• Output speed — Xiaomi MiMo-V2.6-Pro 134.3 tokens/second; Qwen3.8-Max 39.2, against a tier median of 72.5
• Time to first token — Xiaomi MiMo-V2.6-Pro 2.15 seconds; Qwen3.8-Max 2.93
• List price — Xiaomi MiMo-V2.6-Pro $0.43 / $0.87 per 1M; Qwen3.8-Max $2.00 / $6.00
• Blended rate — Xiaomi MiMo-V2.6-Pro $0.18 per 1M at 7:2:1 cache-hit/input/output; Qwen3.8-Max $1.18
• Weights — Xiaomi MiMo-V2.6-Pro MIT-licensed and downloadable; Qwen3.8-Max served as a proprietary endpoint, with a text-only open-weight release that drops the 1M context and vision input
• Reachability — one API provider counted for each on Artificial Analysis
The score is a tie. The speed is not.
One point on a ten-evaluation composite is not a difference anyone should act on, and the more useful signal is what happened underneath it. The model with less than half the active parameters per token scored marginally higher and generated output three and a half times faster — 134.3 tokens per second against 39.2, where the median for comparable reasoning models is 72.5. Qwen3.8-Max is the slowest of the frontier-class models measured on that page, and that is a design consequence of activating 95 billion parameters per token rather than a serving accident.
Latency tells the opposite story from the one the parameter counts suggest. Qwen3.8-Max reaches its first token in 2.93 seconds against Xiaomi's 2.15, and the gap is much smaller than the throughput gap — Qwen3.8-Max is slow once it starts writing, not slow to start. For a chat interface where the answer is short, that difference barely surfaces. For an agent loop generating tens of thousands of output tokens, throughput is the whole wall clock, and a 3.4× gap compounds across every turn of the run.

Where the vendor tables disagree, and why that is not a contradiction
Alibaba published a broad table for Qwen3.8-Max and it is genuinely strong: Terminal-Bench 2.1 at 86.6, SWE-bench Pro at 67.7, PaperBench at 93.0, GPQA Diamond at 92.6, IFBench at 82.8, OSWorld-Verified at 86.1 and Humanity's Last Exam at 43.6. Those are vendor-run figures on Alibaba's own harnesses and some of them on Alibaba's own benchmarks, so they are claims rather than measurements — but they are claims on widely used benchmarks, which makes them more comparable across labs than a bespoke harness result.
Xiaomi's headline number is 72.57 on DeepSWE v1.1, up from a 58.4 baseline, measured with mini-swe-agent at average-of-three on Xiaomi's own grader, over a reinforcement-learning run Xiaomi documented as thirty steps and roughly 750,000 trajectories at a published spend of about $2.62 million. It has not been submitted to the public DeepSWE leaderboard and no third party has reproduced it. Putting 72.57 against Qwen's 67.7 SWE-bench Pro and declaring a five-point Xiaomi win would be reading across two different benchmarks, two different harnesses and two different scoring conventions. There is exactly one ruler both models were run against by someone other than their maker, and it says 46 to 45.
Two of the more consequential Qwen3.8-Max numbers are not scores at all. Alibaba announced that weights would be open-sourced alongside Qwen3.8-27B, and the checkpoint that landed is text-only and does not carry the full 1M-token context or the vision input that the served Max endpoint has. So "Qwen3.8-Max is open weights" and "Qwen3.8-Max is a multimodal 1M-context endpoint" are both true statements about different artefacts, and a deployment plan built on the first while expecting the second will not survive contact. The 0902 point release, meanwhile, moved the model to the top of one public web-development arena without a version-number change — a reminder that the model you benchmarked in August is not necessarily the model serving your requests in September.
Does open weights mean you can self-host either one?
For Xiaomi MiMo-V2.6-Pro, in principle yes: MIT licence, no commercial agreement needed. In practice a 1.02T-parameter checkpoint at 42B active needs a multi-node serving cluster, which is a different order of commitment from calling an API. Publishing weights makes self-hosting legal, not cheap.
For Qwen3.8-Max the answer is narrower than the reputation suggests. The open release is text-only and without the full context window, so self-hosting it gives you a materially smaller model than the one the 45 was measured on. If what you want is the benchmarked configuration, the served endpoint is the product, and the endpoint is proprietary.
Calling either one, and the arithmetic that decides it
Both models are counted at a single API provider by Artificial Analysis, which is unusual for two flagship models and makes failover the practical question rather than a nicety. Qwen3.8-Max is on OrcaRouter as qwen/qwen3.8-max — one OpenAI-compatible endpoint at Alibaba's own list price with 0% markup, so the $2.00 and $6.00 above are passed through unchanged and a price change on Alibaba's side is live on ours the same day. Our own measurements put the endpoint's p50 at around 3.3 seconds, which is the number to plan against rather than the theoretical best case, and a fallback chain means a stalled request retries onto another upstream before the response begins. Xiaomi MiMo-V2.6-Pro is not on OrcaRouter; no Xiaomi model is, and the route to it today is Xiaomi's own platform and the third-party catalogues that list it.
The cost arithmetic is where this matchup is actually settled, and it is not the arithmetic the headline rates suggest. At a 7:2:1 cache-hit/input/output mix the blended rates are $0.18 and $1.18 per million — a 6.6× gap, wider than the 4.6× input and 6.9× output gaps, because Xiaomi's 99% cache discount is deeper than Alibaba's 88%. Artificial Analysis also records a cost of $0.13 per Intelligence Index task for Xiaomi MiMo-V2.6-Pro, measured identically across every model on the board. Run the same workload through both and the cheaper model is the one that scored higher, which is an unusual thing to be able to write and the reason this comparison is not a formality.

Who should switch, and who should not
If you are on Qwen3.8-Max today and it works, there is no urgent case to move. The index gap is a point, the vision and long-context behaviour is proven in your workload, and Alibaba is still developing the line — the 0902 update shows that. The case for moving is throughput and blended cost, and it is a strong case only if your traffic is output-heavy or long-horizon: agents, code generation, anything that writes far more than it reads.
If you are starting fresh, the ordering is difficult to argue with on the published evidence. The Xiaomi model is faster, cheaper on every axis, MIT-licensed, and marginally ahead on the only independently run composite covering both. The counterweights are real and specific: three days of production history, a single serving route, and no public benchmark entry to inspect. That combination argues for evaluating it in a lane beside an incumbent rather than replacing one — which is what a routing configuration is for, and why the useful move is to put the candidate and the incumbent behind one key rather than to pick a winner from a scoreboard.

What would change this answer
Three things. A second and third provider for Xiaomi MiMo-V2.6-Pro would remove the only structural objection to it and turn the price gap into a live decision rather than a tempting one. An independent run of the DeepSWE configuration would either confirm the 72.57 or retire it, and either outcome is worth more than the number itself. And per-evaluation breakdowns on v4.3.2 would show which of the ten evaluations each model wins — the composite is a composite, and a model that leads on agentic coding and one that leads on retrieval-heavy question answering can post the same index score while being wrong for each other's jobs.
