
MiMo-V2.6-Pro vs Gemini 3.1 Pro: A 16-Point Gap That Only Exists on One Index Version
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Put Xiaomi MiMo-V2.6-Pro and Gemini 3.1 Pro Preview side by side on the Artificial Analysis Intelligence Index and you get 46 against 30 — a sixteen-point lead for a model that costs a fifth as much on input. Put them side by side on the numbers both vendors published at their own launches and Gemini 3.1 Pro wins by eleven. Both readings come from the same evaluator. The reason they disagree is the single most useful thing to understand about comparing models in September 2026: the index has been rebuilt underneath both of them, and a score is only meaningful next to the version number that produced it.
The v4.3.2 composite is the current one. It runs ten evaluations — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1 — and on it, read on 22 September 2026, Xiaomi's model scores 46 and Google's scores 30. At launch on 19 February 2026, Gemini 3.1 Pro Preview was reported at 57 on the then-current scale, which topped the board. Those two Gemini figures are not a decline. They are two different rulers.
What each model actually is
Gemini 3.1 Pro Preview is Google's frontier reasoning model, released 19 February 2026 and still carrying the preview label seven months later. It has a 1M-token context window, caps output at 65,536 tokens — listed as 64K by Artificial Analysis and 65K on our own model page — and takes text, images, PDFs, audio and video in, text out. It is proprietary, with an undisclosed parameter count and a knowledge cutoff of January 2025. Google's own API charges $2.00 per million input tokens and $12.00 per million output below 200K of input, moving to $4.00 and $18.00 above it, with cached input at $0.20.
Xiaomi MiMo-V2.6-Pro went generally available on 22 September 2026, with the reinforcement-learning checkpoint repositories appearing the day before. It is a sparse mixture-of-experts model at 1.02 trillion total parameters and 42 billion activated per token, with a 1M-token context window, native multimodality across text, image, video and audio, and MIT-licensed downloadable weights. Xiaomi kept the pricing of the previous MiMo-V2.5 generation rather than repricing the new checkpoint upward: $0.43 per million input tokens and $0.87 per million output, with a 99% cache discount and a blended rate of $0.18 at a 7:2:1 cache-hit/input/output mix.
• Weights — Xiaomi MiMo-V2.6-Pro MIT-licensed and downloadable; Gemini 3.1 Pro Preview proprietary, API only
• Parameters — MiMo-V2.6-Pro 1.02T total / 42B active; Gemini 3.1 Pro Preview undisclosed
• Context — 1M tokens both; Gemini 3.1 Pro Preview caps output at 65,536 tokens, MiMo-V2.6-Pro publishes no equivalent cap
• Modality — MiMo-V2.6-Pro takes text, image, video and audio; Gemini 3.1 Pro Preview takes text, image, PDF, audio and video
• List price — MiMo-V2.6-Pro $0.43 / $0.87 per 1M tokens; Gemini 3.1 Pro Preview $2.00 / $12.00 below 200K input, $4.00 / $18.00 above
• Cache — MiMo-V2.6-Pro 99% discount; Gemini 3.1 Pro Preview $0.20 per 1M cached reads
• Speed — MiMo-V2.6-Pro 134.3 output tokens/second; Gemini 3.1 Pro Preview 121.8
• Time to first token — MiMo-V2.6-Pro 2.15 seconds; Gemini 3.1 Pro Preview 63.62 seconds
• Measured cost to run the index — MiMo-V2.6-Pro $206.66, or $0.13 per task; Gemini 3.1 Pro Preview $1,310.21, or $0.67 per task

The preview label is the real difference
Everything above is a specification. The one line that changes how you should read the whole comparison is the word "preview". A preview model is a candidate that Google has not committed to keeping at that endpoint. That matters for a comparison page in a way that a sixteen-point index gap does not, because a model you cannot rely on being there in six months is not a model you standardise on — and Google's own family has already moved past it. Gemini 3.6 Flash and Gemini 3.8 Flash sit above Gemini 3.1 Pro on coding and agentic work at lower cost, which is why one technical review of the 3.1 Pro release described it as decidedly previous-generation. The AA page itself lists it at rank 71 of 202 on intelligence.
Xiaomi's model is in the opposite position. It shipped this week, the weights are out under MIT, and the serving characteristics are the strongest thing about it. 134.3 output tokens per second puts it eleventh of 114 models measured for speed, and 2.15 seconds to first token is roughly thirty times faster than Gemini 3.1 Pro Preview's 63.62 seconds. For an interactive application, that is not a tie-breaker. It is the difference between a product and a demo, and it is the number most likely to be invisible on a benchmark table.
Where the money actually goes
Per-token arithmetic understates the gap between these two, because the models do not burn the same number of tokens to finish the same job. The cleaner figure is what Artificial Analysis measured to run its full Intelligence Index against each: $206.66 for MiMo-V2.6-Pro and $1,310.21 for Gemini 3.1 Pro Preview. Same suite, same harness, measured the same way — a 6.3x difference on an identical task set, and one that already accounts for the fact that reasoning models emit a lot of tokens thinking. Gemini 3.1 Pro Preview generated 67 million output tokens across the index; the median for models in its price tier is 90 million, so it is actually more economical per point than its peers, and it still costs six times more to evaluate than the Xiaomi model does.
Then put a production loop against it. A 30-step agent run reading 200,000 tokens of context per step and writing 2,000 tokens per step is 6 million input and 60,000 output tokens per run. On Gemini 3.1 Pro Preview below the 200K tier that is $12.00 of input and $0.72 of output — $12.72 per run. On MiMo-V2.6-Pro it is $2.58 and $0.05 — $2.63. Cross 200,000 tokens of input on the Google model and the input rate doubles to $4.00, at which point the same loop costs $24.00 before output. Caching moves both numbers, but a 99% discount on a cheap model against a $0.20 cache read on an expensive one leaves the expensive model an order of magnitude out in front on a context-heavy loop.

The routing question, stated honestly
Here is the part that is easy to get wrong. Gemini 3.1 Pro Preview is on OrcaRouter as google/gemini-3.1-pro-preview, reachable through one OpenAI-compatible key at provider list price with 0% markup — so a Google price change is live on our side the same day rather than at the next billing cycle. Xiaomi MiMo-V2.6-Pro is not on OrcaRouter. No Xiaomi model is, and saying so plainly is more useful than letting a model page imply otherwise. Reaching it today means Xiaomi's own platform or one of the third-party catalogues listing it.
That asymmetry is exactly why a router earns its place in this comparison rather than being a footnote to it. The two models are not competing for the same requests. Gemini 3.1 Pro Preview is the incumbent you may already be paying for, and the question this page answers is whether the new cheap model can take a share of its traffic. Running that test properly means sending your own prompts to both and comparing the answers, not reading an index — and doing it by opening a second contract, a second key and a second billing relationship is the friction that stops the test from ever happening. One key, both lanes where we route them, and the comparison becomes a config change. Automatic failover covers the other half: a preview model can be withdrawn or rate-limited with little notice, and a second lane behind the same endpoint is what turns that from an outage into a routing decision.

Which one to reach for
Pick Gemini 3.1 Pro Preview when the task needs PDF and audio input natively, when you are already inside Google's ecosystem, or when you specifically need the behaviour of that model and can live with a preview endpoint's deprecation risk. Its 63-second time to first token means it is a batch and offline-workload model in practice, whatever the marketing says about interactivity.
Pick MiMo-V2.6-Pro when volume is the constraint, when the loop is context-heavy and repetitive, when you want the weights on your own hardware — MIT permits commercial deployment, modification and further training — or when first-token latency is part of the product. It is the rare case of a cheap model that is also the fast one, and that combination is worth more than the sixteen index points suggest.
What would change this picture is a non-preview successor to Gemini 3.1 Pro, or an independent reproduction of the DeepSWE figures Xiaomi published from its own harness — 72.57 for Pro, 65.68 for Flash, moving from 58.4 and 48.8 respectively across the reinforcement-learning run. Those numbers are vendor-reported, evaluated on Xiaomi's grader with mini-swe-agent at average-of-three, and not submitted to a public leaderboard. Until someone re-runs them, the Artificial Analysis 46 is the number to quote, and it is measured on the same ruler as the 30.
OrcaRouter puts 200+ models behind one OpenAI-compatible endpoint at provider list price with 0% markup, so a vendor price change is live the same day.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
