OrcaRouter model radar hero card headed 'MiMo-V2.6-Flash vs Qwen 3.8', subtitled 'A download against a meter, and only one side has a third-party score', with a left panel for MiMo-V2.6-Flash reading 309B total / 15B active, MIT licence, ungated and No independent score, a right panel for Qwen3.8-Max reading 2.4T total / 95B active, $2.00 in / $6.00 out per 1M and AA Index 45, current scale, and three badges below reading Flash: published 21 Sept 2026, Qwen: metered since 3 Aug 2026 and Routable: Qwen3.8-Max only.
Guides & Insights

MiMo-V2.6-Flash vs Qwen3.8-Max: A Download Against a Meter, and Only One of Them Has a Score

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

One number decides how this comparison is usually written, and it is not a benchmark. Qwen3.8-Max is a hosted model that Artificial Analysis measures continuously, so it carries a current Intelligence Index of 45 on the recalibrated v4.3 scale, a 39.2-token-per-second output speed, and a list price of $2.00 per million input tokens and $6.00 per million output. MiMo-V2.6-Flash, which Xiaomi published as downloadable weights on 21 September 2026, carries none of those things: no provider rate card, no model identifier Xiaomi has announced, and no Artificial Analysis page at all. Everything published about MiMo-V2.6-Flash's capability comes from Xiaomi's own evaluation tables in its model card. Not one figure has been rerun by somebody who does not work for Xiaomi.

That asymmetry is not a mark against the model. It is a fact about what kind of purchase each one is, and it changes what you can actually conclude from a head-to-head. The rest of this piece is about the three things you genuinely can compare — the shape of the claim, the cost of a token, and the licence — plus one number that is real, independently measured, and belongs to neither model in the headline.

First, which Qwen 3.8, and which MiMo

Both names in that headline are families pretending to be checkpoints, and the substitutions are not harmless.

• Qwen3.8-Max — Alibaba's hosted flagship, unveiled 3 August 2026. A 2.4-trillion-parameter sparse mixture-of-experts model with about 95 billion parameters active per token, 1M-token context, and text, image and video input. This is the model the rest of this article treats as the opponent.

• Qwen3.8-Max (0902) — the dated September 2 snapshot of the same model, served as its own entry with the same $2.00 and $6.00 pricing and a 131,072-token output ceiling. The current Artificial Analysis page is for this build, which matters when you quote an index score.

• Qwen3.8-2.4T-A95B — the open-weight checkpoint of the same core, downloadable since 12 August 2026 under the Qwen3.8-Max License. It is text-only, thinking is always on, and its native context is 262K rather than a million. It is not the model in the headline.

• MiMo-V2.6-Flash — Xiaomi's 309B-total, 15B-active sparse mixture-of-experts checkpoint, ungated on Hugging Face under the MIT licence, native omnimodal, with a 1M-token context claim. It is the efficiency member of a two-checkpoint release whose flagship is MiMo-V2.6-Pro at 1.02T total and 42B active.

• MiMo-V2-Flash — the December 2025 model. Same 309B total, same 15B active, same 128-token sliding window. Different training run, different benchmark table, and a 256K context. Any page ranking "MiMo-V2-Flash" against a Qwen model is comparing the wrong generation, and the spec sheet alone will not tell you so.

The MiMo collision is the sharper of the two traps, because the numbers that identify the model also happen to be the numbers that are identical between the 2025 and 2026 checkpoints. The Qwen collision is milder but has a price attached: the open 2.4T weights and the hosted API are different artifacts with different terms, and a comparison that swaps them will get both the capability and the licence wrong.

The index was rebuilt, and the number most pages still print is gone

Here is the failure mode to watch for before any score appears.

Artificial Analysis recalibrated its Intelligence Index to v4.3 on 7 September 2026. Scores from before that date are not comparable to scores after it, and the live Qwen3.8-Max page carries an Updated badge marking a restated result. The widely circulated figure of 58 for Qwen3.8-Max belongs to the old scale. Current Qwen3.8-Max (0902) reads 45, ranked 19th of the 202 models Artificial Analysis places in that class.

This is not pedantry. A 58 next to a 46 reads as a twelve-point gap. The same two models on the same current scale are a single point apart — and the 46 is not even the model this article is about:

• Qwen3.8-Max (0902), independently measured — Intelligence Index 45 on v4.3. Output speed 39.2 tokens per second, ranked 151st of 202, which the page itself notes is slow. Cost per Intelligence Index task $5.41. Output tokens generated to complete the index: 190M.

• MiMo-V2.6-Pro, independently measured — Intelligence Index 46, ranked first of the 114 models in the open-weights class. Output speed 134.3 tokens per second. Cost per Intelligence Index task $0.13. Output tokens to complete the index: 140M.

• MiMo-V2.6-Flash, the actual subject — no Artificial Analysis page exists. Nothing to quote.

Read those three together and the honest summary is uncomfortable for both vendors. The point of comparison everybody wants — Flash against Qwen3.8-Max on a neutral scale — does not exist and cannot be constructed from vendor tables, because the two vendors ran different harnesses on different checkpoints at different times and then compared themselves to other companies' models they also ran. What does exist is a one-point gap between the flagship Qwen and Xiaomi's larger checkpoint, at roughly one-fortieth the cost per index task. That is the most interesting real number in this matchup, and it is about a model neither side of the headline names.

Scoreboard card headed 'MiMo-V2.6-Flash vs Qwen3.8-Max', with paired rows for parameters (309B total / 15B active against 2.4T total / 95B active), licence (MIT, ungated, no revenue trigger, against a hosted API whose open 2.4T sibling uses the Qwen3.8-Max License), price per 1M (no vendor rate card against $2.00 in and $6.00 out with an 88% cache discount), independent score (None published, no Artificial Analysis page, against AA Index 45 on the v4.3 scale, 19th of 202 in class), DeepSWE v1.1 (67.9 on Xiaomi's own harness against Not published) and routability (No - nowhere, including OrcaRouter, against Yes - on OrcaRouter, base tier and the 0902 snapshot).

What the vendor tables will and will not tell you

Both vendors publish numbers. They are not the same kind of number, and lining them up column by column produces a chart rather than a finding — but the shape of the claims is still worth reading, because it tells you what each lab optimised for.

• Code agent — Xiaomi reports MiMo-V2.6-Flash at DeepSWE v1.1 67.9, Terminal Bench 2.1 87.6, ProgramBench 26.0 and MiMo Code Bench 61.2. Alibaba reports Qwen3.8-Max at SWE-bench Pro 67.7 and Terminal-Bench 2.1 86.6. The Terminal-Bench figures are close enough that the harness, not the model, is plausibly deciding the order.

• General agent — Xiaomi reports AutomationBench v1.0.6 52.3, Toolathlon-Verified 73.6, OSWorld-Verified 80.8 and Agents' Last Exam 27.6 for Flash. Alibaba reports OSWorld-Verified 86.1, WideSearch 81.9 and Agent's Last Exam 52.4 for Qwen3.8-Max. On the two benchmarks both labs ran, Qwen's reported figures are ahead — on Alibaba's own harness.

• Knowledge and security — Xiaomi runs CyberGym (Flash 95.1), MiMo Cyber Bench (77.2), ExploitGym (6.0), ExploitBench (25.3) and SEC Bench Pro (47.5). Alibaba runs GPQA Diamond (92.6), Humanity's Last Exam (43.6) and PaperBench (93.0). Almost no overlap, so almost nothing to compare.

The one genuinely useful thing a vendor table can show is an internal inconsistency, and Xiaomi's has one. Its public RL dashboard, which streamed the training run through September, published MiMo-V2.6-Flash at 65.68 on DeepSWE v1.1 using a mini-swe-agent harness at average-of-three. The finished model card prints 67.9 for the released weights. Those describe a training snapshot and a released artifact respectively, and the two-point difference is the same size as the gap between Xiaomi's two checkpoints. If you have been quoting the dashboard number since last week, it is stale.

The bill, which is where the two models stop being comparable

This is the section that decides deployments, and it is not close.

• Qwen3.8-Max is metered. $2.00 per million input tokens and $6.00 per million output on Alibaba's international rate card, with a cache discount Artificial Analysis measures at 88% and a cost of $5.41 per Intelligence Index task. Alibaba's China-region documentation lists CNY 12 and CNY 36 per million for the same pair. You can call it this afternoon and get a bill.

• MiMo-V2.6-Flash is a download. Xiaomi publishes no per-token rate for the MiMo-V2.6 generation on its own site and has announced no callable identifier for either checkpoint, so there is nothing to meter. What there is instead is 172.9 GB of FP8 weight data across 65 shards, before KV cache for a 1M-token context, and a deployment section recommending SGLang at tensor parallel 16 with a data-parallel factor of 2, or a vLLM recipe at tensor parallel 8 — a multi-node serving job.

• The third-party listings are not a rate card. Catalogue pages do carry figures for the MiMo-V2.6 series, and Artificial Analysis counts a single API provider serving MiMo-V2.6-Pro at $0.435 and $0.87 per million. That is a provider's listing for the larger checkpoint, not a Xiaomi price, and it does not exist for Flash at all.

So the arithmetic is a metered line item against a capital decision. At a few hundred million tokens a month, Qwen3.8-Max is a bill you can forecast and Flash is a node you would be buying to serve a model nobody outside Xiaomi has measured. At billions of tokens a month the open path starts winning on cost, and this is the point at which the licence stops being a legal footnote and becomes a procurement input.

The licences are not the same permission

MIT is about as permissive as open weights get. The Qwen3.8-Max License is not MIT, and the difference has teeth.

MiMo-V2.6-Flash ships under MIT with no revenue threshold, no user-count trigger, no separate commercial agreement and no research clause. Commercial deployment, modification, redistribution and further training are all permitted, and the repository is ungated.

The Qwen3.8-Max License grants use, modification, distribution, sublicensing, sale, deployment, hosting and fine-tuning, on two conditions. A product or service exceeding 100 million monthly active users or US$20 million in monthly revenue must display the model name prominently in its interface. And a model-as-a-service or AI-work-assistant business with aggregate revenue above US$50 million across any twelve consecutive months needs a separate licence from Alibaba before commercial use. Internal use is carved out, provided the model or its capabilities are not exposed to third parties.

For most teams neither condition binds. For a platform business that succeeds, one does — and it binds at precisely the moment the model has become load-bearing. Note also that those conditions attach to the downloadable 2.4T weights, not to the hosted API, where the provider's terms apply instead. The two paths have different obligations, which is a further reason not to compare the open checkpoint against the hosted one as if they were the same product.

Where OrcaRouter fits, and where it does not

Qwen3.8-Max is routable on OrcaRouter, on both the base entry and the dated 0902 snapshot, through one API key at provider list price with 0% markup passed through. That matters specifically here for two reasons. First, the 0902 build is a separate entry rather than a silent replacement, so the routing DSL can pin reproducibility-sensitive traffic to the dated snapshot while everything else follows the base tier — no second integration, no code change. Second, because list price is passed through rather than marked up, a change to Alibaba's rate card lands on your side the same day instead of at the next contract renewal, which is what keeps the metered-versus-node arithmetic above honest as prices move. Cache-heavy workloads also keep the provider's cache discount rather than losing part of it to a reseller's margin.

Screenshot of the OrcaRouter model page for Qwen3.8 Max, marked FEATURED, giving the model id qwen/qwen3.8-max, a 1M-token context, text, image and video input with text output, an input price of $2.00 and output price of $6.00 per 1M tokens, a p50 time to first token of 3.33 seconds, and seven-day traffic of 30.8M tokens.

MiMo-V2.6-Flash is not routable anywhere, ours included. We route no Xiaomi model, and the availability line in Xiaomi's own card points at Xiaomi's own channels: the MiMo API platform, AI Studio, MiMo Code and the desktop app. If you want this checkpoint, you are downloading 172.9 GB and finding a node.

Screenshot of the Artificial Analysis model page for Qwen3.8 Max (0902), showing an Intelligence Index of 45 on the updated v4.3 scale ranked 19th of 202, output speed of 39.2 tokens per second ranked 151st of 202, a price of $2.00 per million input tokens and $6.00 per million output with an 88% cache discount and a $5.41 cost per Intelligence Index task, 190M output tokens generated on the index, and a 984k-token context window.

Which one, and when

Take Qwen3.8-Max if you want capability without hardware, if your volume is uncertain, or if you want the option of an independent number before you commit. Accept that 45 on the current index is a middle-of-the-frontier position rather than a top one, that 39.2 tokens per second is slow enough to matter inside an agent loop, and that 190M output tokens to finish the Intelligence Index is roughly twice the median — verbosity costs money on the output-priced side of the meter.

Take MiMo-V2.6-Flash if the constraint is the licence, the data path, or a long-run bill you can amortise — and if you have the node. MIT terms with no revenue gate are genuinely a different permission from a licence with a MAU threshold and a revenue-share trigger, and for a company that expects to be large, that is the reason to pick it. Budget for multi-node serving, size from the published weight data rather than the repository page, and treat every number in Xiaomi's card as vendor-reported until somebody outside the lab reruns it.

And if you are choosing today rather than next quarter, run the metered option first. A month of Qwen3.8-Max tells you what your workload actually needs. A node tells you only what you hoped it would.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily