OrcaRouter model radar hero card for the comparison between MiMo-V2.6-Pro and DeepSeek V4 Pro, subtitled 'Two open flagships, ten index points apart.', with one panel per model and a footer noting that both are MIT-licensed with a 1M-token context window and that the scores are v4.3.2 and not comparable to earlier index versions.
Guides & Insights

MiMo-V2.6-Pro vs DeepSeek V4 Pro: Two Open Flagships, Ten Index Points Apart

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Xiaomi MiMo-V2.6-Pro and DeepSeek V4 Pro are the same kind of object — trillion-parameter Chinese mixture-of-experts flagships, MIT-licensed, 1M-token context, open weights you can download today — and they are ten points apart on the Artificial Analysis Intelligence Index v4.3.2, 46 against 36. They also disagree about what a flagship is for. DeepSeek V4 Pro, released August 13, 2026, is a text-only reasoning model: 1.6 trillion parameters with 49 billion active, and no vision, audio or video input anywhere in its published surface. Xiaomi MiMo-V2.6-Pro, released September 21, 2026, is 1.02 trillion parameters with 42 billion active and takes text, image, video and audio. That difference decides more deployments than the ten points do, and it is the first thing to settle before you compare anything else.

Same licence, different machines

Start with what is genuinely identical, because it is more than you would expect across two competing labs. Both ship under an MIT licence tag on public repositories, which means commercial deployment, modification and further training without a revenue gate or a field-of-use clause. Both claim a 1M-token context window. Both are sparse mixture-of-experts designs — the architecture has become the default at this scale, not a differentiator. And both were trained by labs that publish less than Western frontier labs do about method.

• Parameters — MiMo-V2.6-Pro 1.02T total / 42B active; DeepSeek V4 Pro 1.6T total / 49B active

• Input modality — MiMo-V2.6-Pro text, image, video, audio; DeepSeek V4 Pro text only

• Context — 1M tokens both; DeepSeek V4 Pro caps output at 384K tokens

• Index score — MiMo-V2.6-Pro 46 on Intelligence Index v4.3.2; DeepSeek V4 Pro 36 on the same version

• Cost per Index task — MiMo-V2.6-Pro $0.13; DeepSeek V4 Pro $0.67

• Listed rate — MiMo-V2.6-Pro $0.43 / $0.87 per 1M tokens; DeepSeek V4 Pro $1.32 / $3.96 as Artificial Analysis lists it, before DeepSeek's own peak/off-peak schedule

• Speed — MiMo-V2.6-Pro 134.3 output tokens/second; DeepSeek V4 Pro 97.4

• Time to first token — MiMo-V2.6-Pro 2.15s; DeepSeek V4 Pro 1.62s

Read that list twice, because two rows are counter-intuitive. The smaller model scores ten points higher — 42 billion active parameters beating 49 billion is not a rounding difference, it is a post-training result, and it is the single clearest signal in this comparison that reinforcement-learning quality has overtaken raw scale as the lever. And the cheaper model is also the faster one, on both throughput and cost per task, which is the opposite of the usual trade. Xiaomi is not selling you a discount in exchange for patience. DeepSeek's edge is time to first token, 1.62 seconds against 2.15 — real, but the only category where V4 Pro leads.

Two-column scoreboard card headed 'MiMo-V2.6-Pro vs DeepSeek V4 Pro — the scoreboard'. The Xiaomi MiMo-V2.6-Pro column lists 1.02T total and 42B active parameters, text, image, video and audio input, a 1M-token context window, Intelligence Index v4.3.2 score 46, a listed rate of $0.43 / $0.87 per million tokens, $0.13 per index task, 134.3 tokens per second with 2.15 seconds to first token, and a release date of 21 September 2026. The DeepSeek V4 Pro column lists 1.6T total and 49B active parameters, text-only input, a 1M-token context window with output capped at 384K, index score 36, a listed rate of $1.32 / $3.96, $0.67 per index task, 97.4 tokens per second with 1.62 seconds to first token, and a release date of 13 August 2026.

Why the same index version matters here

Comparing open-weights models across index revisions is how most published comparisons go wrong, so the version number is not a footnote. Artificial Analysis rebuilt the Intelligence Index as v4.3 in September 2026, and scores moved substantially across that boundary — Claude Opus 5 lost three points between v4.2 and v4.3, and the open-weights field re-sorted. Both figures above are v4.3.2, which means the 46 and the 36 are directly comparable to each other. They are not comparable to the older numbers you will find quoted elsewhere for either model.

That is not a hypothetical. DeepSeek V4 Pro was widely reported at 53 on an earlier index scale in August 2026, and our own coverage of that period carried it. On v4.3.2 the same model reads 36. Nothing about the model changed; the ruler did. If you have a spreadsheet with 53 in it next to a 46 for MiMo-V2.6-Pro, you have a comparison that will point you at the wrong model. The correct pairing is 46 against 36, and the correct conclusion is that the model released five weeks later, at two-thirds the parameter count, is materially ahead.

The reporting gap between the two labs

There is a second, subtler difference, and it is about what each lab is willing to have checked.

MiMo-V2.6-Pro's 46 is an Artificial Analysis measurement. The evaluation house ran its own suite — ten evaluations across reasoning, knowledge, mathematics and coding — on the released model and published the result, along with the operating numbers: 134.3 tokens/second, 2.15 seconds to first token, a 99% cache discount, $0.13 per index task, and a total evaluation cost of $206.66. That is a third party's number about Xiaomi's model.

DeepSeek V4 Pro's headline agentic figures are DeepSeek's own, and the gap between them and the independent ones has been documented. The company's launch material reports Terminal-Bench 2.1 at 87.9, DeepSWE at 62.7, CyberGym at 83.3 and SWE-bench Verified at 80.6. Independent runs put Terminal-Bench 2.1 at 78.7 — roughly nine points below the vendor figure — and the Artificial Analysis Coding Index at 68.8. None of those vendor numbers has been reproduced, and the size of the Terminal-Bench gap is the reason to treat the rest of the set as claims rather than measurements.

Xiaomi has its own version of the same problem. Its dashboard reports MiMo-V2.6-Pro moving from 58.4 to 72.57 on DeepSWE v1.1 across its reinforcement-learning run, evaluated on Xiaomi's own harness with mini-swe-agent at average-of-three. That is not a leaderboard submission and no outside party has rerun it. So neither model arrives with a clean bill of health on vendor benchmarks. The difference is that MiMo-V2.6-Pro also arrives with an independent composite score that DeepSeek's launch did not have at the equivalent moment — and if you are choosing between them on evidence rather than on marketing, that is the asymmetry that should carry weight.

Screenshot of the Artificial Analysis model page for DeepSeek V4 Pro, showing the Intelligence Index score of 36 on version 4.3.2, an open-weights model under the MIT licence, text-only input, 1.6 trillion total and 49 billion active parameters, a 1M-token context window, a listed price of $1.32 per million input tokens and $3.96 per million output tokens, output speed of 97.4 tokens per second and time to first token of 1.62 seconds.

What this looks like in production

The modality difference is the practical fork, and it cuts in DeepSeek's favour less often than the ten-point gap suggests. If your pipeline ingests screenshots, PDFs rendered as images, video frames or audio, MiMo-V2.6-Pro handles it natively in one model, and DeepSeek V4 Pro cannot handle it at all — you would need a separate vision model in front, which means a second contract, a second failure mode and a second bill. If your workload is text-only reasoning, the extra modality is dead weight you are paying nothing for.

Cost per index task is the other number to carry. $0.13 against $0.67 is a fivefold gap on a controlled, identical task set, measured the same way for both. DeepSeek runs a peak/off-peak billing schedule that can cut its effective rate substantially depending on when you call, so the real-world ratio narrows at off-peak hours and widens at peak — a detail that matters if your traffic is bursty and you cannot choose when it arrives.

This is where the DeepSeek side has a genuine advantage that has nothing to do with the model: it is on our platform. DeepSeek V4 Pro is reachable as deepseek/deepseek-v4-pro through an OrcaRouter key at provider list price with 0% markup, with automatic failover across providers and the routing DSL available on top — so the peak/off-peak arithmetic, the provider spread and the fallback path are all things you configure rather than things you build. MiMo-V2.6-Pro is not on our platform, and we will say so plainly: reaching it today means Xiaomi's own platform or one of the third-party catalogues that lists it, and Artificial Analysis counted a single API provider for it at capture time. One route is not a production route, which is the strongest argument for treating MiMo-V2.6-Pro as a model to evaluate now and to route to carefully, with something else behind it.

Which one, and when

Choose DeepSeek V4 Pro if your work is text-only, if you need the 384K output ceiling, if you want the fastest time to first token of the two, or if you are already on our platform and want a trillion-parameter open-weights lane with failover and no new integration. It is the incumbent for a reason, and being five weeks older has meant five weeks of tooling, quantisation and serving recipes that a September model has not accumulated yet.

Choose MiMo-V2.6-Pro if the task is multimodal, if cost per task dominates your unit economics, if throughput matters more than a half-second of latency, or if you want the current open-weights leader on an independently measured index. The MIT licence on both means the weights question is settled either way — you can run either on your own hardware, fine-tune it, and ship it commercially.

The configuration worth considering is not a choice at all. A router lets you keep the DeepSeek lane for text-only reasoning at off-peak rates and add a MiMo lane for the multimodal requests, with a fallback in front of the thinner route. That is not hedging; it is matching each request to the model that actually handles it, which is a cheaper and more durable answer than picking a winner and hoping the workload never changes shape.

Decision card headed 'MiMo-V2.6-Pro vs DeepSeek V4 Pro — the decision', with two columns. 'Choose MiMo-V2.6-Pro when' lists a multimodal task with screenshots, rendered PDFs, video frames or audio in the same request; cost per task dominating unit economics at $0.13 against $0.67 on an identical task set; throughput mattering more than latency at 134.3 against 97.4 tokens per second; and wanting the current open-weights leader on an independently measured index at 46 against 36. 'Choose DeepSeek V4 Pro when' lists text-only reasoning work; needing the 384K output ceiling; first-token latency being the binding constraint at 1.62s against 2.15s; and wanting a route with failover behind it.

The question this comparison cannot answer yet

Both models' most-quoted benchmarks are vendor-run and unreproduced. For DeepSeek V4 Pro that gap has been open since August and is the reason its independent scores read lower than its launch numbers. For MiMo-V2.6-Pro the independent composite exists but the agentic breakdown does not — Artificial Analysis publishes a 46 and not the per-evaluation spread underneath it, which is the detail that tells you whether a model belongs in an agent loop or a question-answering pipeline.

Until that spread is published and until someone reruns Xiaomi's DeepSWE harness, the defensible position is narrower than either lab's marketing: MiMo-V2.6-Pro leads on an independent composite and on measured cost per task, DeepSeek V4 Pro leads on maturity, output length, first-token latency and being reachable through a route that has redundancy behind it. Five weeks is a short head start, and it is the only one DeepSeek has.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily