
MiMo-V2.6-Pro vs DeepSeek V4 Pro: Two Open Flagships, Ten Index Points Apart
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Xiaomi MiMo-V2.6-Pro and DeepSeek V4 Pro are the same kind of object — trillion-parameter Chinese mixture-of-experts flagships, MIT-licensed, 1M-token context, open weights you can download today — and they are ten points apart on the Artificial Analysis Intelligence Index v4.3.2, 46 against 36. They also disagree about what a flagship is for. DeepSeek V4 Pro, released August 13, 2026, is a text-only reasoning model: 1.6 trillion parameters with 49 billion active, and no vision, audio or video input anywhere in its published surface. Xiaomi MiMo-V2.6-Pro, released September 21, 2026, is 1.02 trillion parameters with 42 billion active and takes text, image, video and audio. That difference decides more deployments than the ten points do, and it is the first thing to settle before you compare anything else.
Same licence, different machines
Start with what is genuinely identical, because it is more than you would expect across two competing labs. Both ship under an MIT licence tag on public repositories, which means commercial deployment, modification and further training without a revenue gate or a field-of-use clause. Both claim a 1M-token context window. Both are sparse mixture-of-experts designs — the architecture has become the default at this scale, not a differentiator. And both were trained by labs that publish less than Western frontier labs do about method.
• Parameters — MiMo-V2.6-Pro 1.02T total / 42B active; DeepSeek V4 Pro 1.6T total / 49B active
• Input modality — MiMo-V2.6-Pro text, image, video, audio; DeepSeek V4 Pro text only
• Context — 1M tokens both; DeepSeek V4 Pro caps output at 384K tokens
• Index score — MiMo-V2.6-Pro 46 on Intelligence Index v4.3.2; DeepSeek V4 Pro 36 on the same version
• Cost per Index task — MiMo-V2.6-Pro $0.13; DeepSeek V4 Pro $0.67
• Listed rate — MiMo-V2.6-Pro $0.43 / $0.87 per 1M tokens; DeepSeek V4 Pro $1.32 / $3.96 as Artificial Analysis lists it, before DeepSeek's own peak/off-peak schedule
• Speed — MiMo-V2.6-Pro 134.3 output tokens/second; DeepSeek V4 Pro 97.4
• Time to first token — MiMo-V2.6-Pro 2.15s; DeepSeek V4 Pro 1.62s
Read that list twice, because two rows are counter-intuitive. The smaller model scores ten points higher — 42 billion active parameters beating 49 billion is not a rounding difference, it is a post-training result, and it is the single clearest signal in this comparison that reinforcement-learning quality has overtaken raw scale as the lever. And the cheaper model is also the faster one, on both throughput and cost per task, which is the opposite of the usual trade. Xiaomi is not selling you a discount in exchange for patience. DeepSeek's edge is time to first token, 1.62 seconds against 2.15 — real, but the only category where V4 Pro leads.

Why the same index version matters here
Comparing open-weights models across index revisions is how most published comparisons go wrong, so the version number is not a footnote. Artificial Analysis rebuilt the Intelligence Index as v4.3 in September 2026, and scores moved substantially across that boundary — Claude Opus 5 lost three points between v4.2 and v4.3, and the open-weights field re-sorted. Both figures above are v4.3.2, which means the 46 and the 36 are directly comparable to each other. They are not comparable to the older numbers you will find quoted elsewhere for either model.
That is not a hypothetical. DeepSeek V4 Pro was widely reported at 53 on an earlier index scale in August 2026, and our own coverage of that period carried it. On v4.3.2 the same model reads 36. Nothing about the model changed; the ruler did. If you have a spreadsheet with 53 in it next to a 46 for MiMo-V2.6-Pro, you have a comparison that will point you at the wrong model. The correct pairing is 46 against 36, and the correct conclusion is that the model released five weeks later, at two-thirds the parameter count, is materially ahead.
The reporting gap between the two labs
There is a second, subtler difference, and it is about what each lab is willing to have checked.
MiMo-V2.6-Pro's 46 is an Artificial Analysis measurement. The evaluation house ran its own suite — ten evaluations across reasoning, knowledge, mathematics and coding — on the released model and published the result, along with the operating numbers: 134.3 tokens/second, 2.15 seconds to first token, a 99% cache discount, $0.13 per index task, and a total evaluation cost of $206.66. That is a third party's number about Xiaomi's model.
DeepSeek V4 Pro's headline agentic figures are DeepSeek's own, and the gap between them and the independent ones has been documented. The company's launch material reports Terminal-Bench 2.1 at 87.9, DeepSWE at 62.7, CyberGym at 83.3 and SWE-bench Verified at 80.6. Independent runs put Terminal-Bench 2.1 at 78.7 — roughly nine points below the vendor figure — and the Artificial Analysis Coding Index at 68.8. None of those vendor numbers has been reproduced, and the size of the Terminal-Bench gap is the reason to treat the rest of the set as claims rather than measurements.
Xiaomi has its own version of the same problem. Its dashboard reports MiMo-V2.6-Pro moving from 58.4 to 72.57 on DeepSWE v1.1 across its reinforcement-learning run, evaluated on Xiaomi's own harness with mini-swe-agent at average-of-three. That is not a leaderboard submission and no outside party has rerun it. So neither model arrives with a clean bill of health on vendor benchmarks. The difference is that MiMo-V2.6-Pro also arrives with an independent composite score that DeepSeek's launch did not have at the equivalent moment — and if you are choosing between them on evidence rather than on marketing, that is the asymmetry that should carry weight.

What this looks like in production
The modality difference is the practical fork, and it cuts in DeepSeek's favour less often than the ten-point gap suggests. If your pipeline ingests screenshots, PDFs rendered as images, video frames or audio, MiMo-V2.6-Pro handles it natively in one model, and DeepSeek V4 Pro cannot handle it at all — you would need a separate vision model in front, which means a second contract, a second failure mode and a second bill. If your workload is text-only reasoning, the extra modality is dead weight you are paying nothing for.
Cost per index task is the other number to carry. $0.13 against $0.67 is a fivefold gap on a controlled, identical task set, measured the same way for both. DeepSeek runs a peak/off-peak billing schedule that can cut its effective rate substantially depending on when you call, so the real-world ratio narrows at off-peak hours and widens at peak — a detail that matters if your traffic is bursty and you cannot choose when it arrives.
This is where the DeepSeek side has a genuine advantage that has nothing to do with the model: it is on our platform. DeepSeek V4 Pro is reachable as deepseek/deepseek-v4-pro through an OrcaRouter key at provider list price with 0% markup, with automatic failover across providers and the routing DSL available on top — so the peak/off-peak arithmetic, the provider spread and the fallback path are all things you configure rather than things you build. MiMo-V2.6-Pro is not on our platform, and we will say so plainly: reaching it today means Xiaomi's own platform or one of the third-party catalogues that lists it, and Artificial Analysis counted a single API provider for it at capture time. One route is not a production route, which is the strongest argument for treating MiMo-V2.6-Pro as a model to evaluate now and to route to carefully, with something else behind it.
Which one, and when
Choose DeepSeek V4 Pro if your work is text-only, if you need the 384K output ceiling, if you want the fastest time to first token of the two, or if you are already on our platform and want a trillion-parameter open-weights lane with failover and no new integration. It is the incumbent for a reason, and being five weeks older has meant five weeks of tooling, quantisation and serving recipes that a September model has not accumulated yet.
Choose MiMo-V2.6-Pro if the task is multimodal, if cost per task dominates your unit economics, if throughput matters more than a half-second of latency, or if you want the current open-weights leader on an independently measured index. The MIT licence on both means the weights question is settled either way — you can run either on your own hardware, fine-tune it, and ship it commercially.
The configuration worth considering is not a choice at all. A router lets you keep the DeepSeek lane for text-only reasoning at off-peak rates and add a MiMo lane for the multimodal requests, with a fallback in front of the thinner route. That is not hedging; it is matching each request to the model that actually handles it, which is a cheaper and more durable answer than picking a winner and hoping the workload never changes shape.

The question this comparison cannot answer yet
Both models' most-quoted benchmarks are vendor-run and unreproduced. For DeepSeek V4 Pro that gap has been open since August and is the reason its independent scores read lower than its launch numbers. For MiMo-V2.6-Pro the independent composite exists but the agentic breakdown does not — Artificial Analysis publishes a 46 and not the per-evaluation spread underneath it, which is the detail that tells you whether a model belongs in an agent loop or a question-answering pipeline.
Until that spread is published and until someone reruns Xiaomi's DeepSWE harness, the defensible position is narrower than either lab's marketing: MiMo-V2.6-Pro leads on an independent composite and on measured cost per task, DeepSeek V4 Pro leads on maturity, output length, first-token latency and being reachable through a route that has redundancy behind it. Five weeks is a short head start, and it is the only one DeepSeek has.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
