OrcaRouter model radar hero card for the comparison between MiMo-V2.6-Pro and Claude Opus 5, subtitled 'Five index points. Twenty-one times the price.', with one panel per model.
Guides & Insights

MiMo-V2.6-Pro vs Claude Opus 5: 21x Cheaper, Five Index Points Behind

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Xiaomi MiMo-V2.6-Pro and Claude Opus 5 are the two ends of a single trade, and the trade has a price. On the Artificial Analysis Intelligence Index v4.3.2, Claude Opus 5 (max effort) scores 51 and Xiaomi MiMo-V2.6-Pro scores 46. On the rate card, Claude Opus 5 costs $5.00 per million input tokens and $25.00 per million output; MiMo-V2.6-Pro costs $0.43 and $0.87. Run the division and the answer is 11.6x on input and 28.7x on output — and 21x on the blended rate Artificial Analysis publishes for each. Five index points cost twenty-one times the money. Whether that is a bargain or a waste depends entirely on what your workload does with the fifth point, and most teams never work that out before they commit.

The two models, stated plainly

Claude Opus 5 is Anthropic's proprietary flagship, released July 24, 2026, with a 1M-token context window and undisclosed parameter count. Xiaomi MiMo-V2.6-Pro is a sparse mixture-of-experts model at 1.02 trillion total parameters and 42 billion activated, with a 1M-token context window, native multimodality across text, image, video and audio, and MIT-licensed published weights.

• Weights — MiMo-V2.6-Pro MIT-licensed and downloadable; Claude Opus 5 proprietary, API only

• Parameters — MiMo-V2.6-Pro 1.02T total / 42B active; Claude Opus 5 undisclosed

• Context — 1M tokens both; Claude Opus 5 caps output at 128K on the Messages API and 300K on the Batch API

• Modality — MiMo-V2.6-Pro takes text, image, video and audio; Claude Opus 5's published input surface is text and image

• List price — MiMo-V2.6-Pro $0.43 / $0.87 per 1M tokens; Claude Opus 5 $5.00 / $25.00

• Cache — MiMo-V2.6-Pro lists a 99% cache discount; Claude Opus 5 reads cache at $0.50 per 1M with $6.25 writes at a 5-minute TTL

• Measured cost per Intelligence Index task — MiMo-V2.6-Pro $0.13; Claude Opus 5 $5.86

• Serving — MiMo-V2.6-Pro 134.3 output tokens/second and 2.15s to first token; Claude Opus 5 57.9 tokens/second

That last pair is the one that gets lost. The cheaper model is also the faster one — by a factor of 2.3 on output throughput — because it is served on hardware Xiaomi and its partners control and can batch aggressively. Claude Opus 5 at max effort is a reasoning model doing more work per token; the speed gap is not a defect, it is a different product. But it does mean the cheap lane is not paying for its discount in latency. It is paying for it in five index points and in the tail of tasks where that fifth point lives.

Two-column scoreboard card headed 'MiMo-V2.6-Pro vs Claude Opus 5 — the scoreboard'. The Xiaomi MiMo-V2.6-Pro column lists Intelligence Index v4.3.2 score 46, list price $0.43 / $0.87 per million tokens, measured cost per index task $0.13, serving speed 134.3 tokens per second, size 1.02T total and 42B active parameters, MIT-licensed downloadable weights, and text, image, video and audio input. The Claude Opus 5 column lists index score 51, list price $5.00 / $25.00, measured cost per index task $5.86, serving speed 57.9 tokens per second, undisclosed size, proprietary API-only weights, and text and image input.

Where the five points actually are

A five-point gap on a composite of ten evaluations is not evenly spread, and the aggregate hides the thing that matters. Both models were run on the same v4.3.2 suite, which includes AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. The published pages do not break out per-evaluation scores for either model, which is a real limitation on both sides of this comparison and worth saying rather than papering over.

What is on the record is directional. Claude Opus 5 at max effort scored 49.0% on Terminal-Bench 4.0 within the v4.3 suite, behind GPT-6 Astra at 59.1% and Claude Fable 5.1 at 52.0%. On AutomationBench-AA, Opus 5 completed every objective without a guardrail violation on 28.3% of workflows, against 41.6% for GPT-6 Astra and 32.1% for Fable 5.1. Those are hard agentic numbers, and they are exactly the categories where a five-point composite gap is least likely to be evenly distributed — MiMo-V2.6-Pro's own index entry does not publish a comparable agentic breakout.

Meanwhile the vendor numbers both companies publish are not comparable to each other, and it is worth being blunt about why. Anthropic's launch material reports SWE-bench Verified at 96.0% and SWE-bench Pro at 79.2%; one tracker notes that Opus 5 shipped without a standalone SWE-bench entry, and there is no independently audited SWE-bench Verified figure for it. Xiaomi's material reports MiMo-V2.6-Pro moving from 58.4 to 72.57 on DeepSWE v1.1 across its RL run, on Xiaomi's own harness with mini-swe-agent at average-of-three, not submitted to the public DeepSWE board. Two vendor harnesses, two different tasks, no overlap. Placing 96.0% next to 72.57 and drawing a conclusion would be inventing a comparison that does not exist.

Screenshot of the OrcaRouter model page for Claude Opus 5, showing the 1M-token context window, a 128K maximum output, input pricing of $5.00 per million tokens and output pricing of $25.00 per million, cache reads at $0.50 per million and cache writes at $6.25 per million on a five-minute TTL, and the provider uptime table listing the routes behind the model.

The cost comparison, done properly

Per-token arithmetic understates the gap, because the two models do not consume the same number of tokens to finish the same job. Artificial Analysis measured what each actually cost to run the full Intelligence Index: $0.13 for MiMo-V2.6-Pro, $5.86 for Claude Opus 5. That is a 45x gap on a controlled, identical task set, and it is the fairest single number available because both were measured the same way.

Now put a plausible production loop against it. A 30-step agent run that reads 200,000 tokens of context per step and writes 2,000 tokens per step is 6 million input tokens and 60,000 output tokens per run. On Claude Opus 5 at list that is $30.00 of input and $1.50 of output — $31.50 per run before any caching. On MiMo-V2.6-Pro it is $2.58 of input and $0.05 of output — $2.63. With the cache discounts both vendors offer the input side collapses on either model, and the ratio moves rather than disappearing: a 99% cache discount on a cheap model versus a $0.50 cache read on an expensive one still leaves the expensive model an order of magnitude out in front on a context-heavy loop.

Which is the whole argument for a cheap lane and against a wholesale switch. If your workload is a long-horizon agent loop that re-reads the same context hundreds of times, the per-run cost difference is not a rounding error — it is the difference between a feature you can afford to run per user and one you cannot. That is the case for putting MiMo-V2.6-Pro in front of the bulk of your traffic. It is not a case for deleting Claude Opus 5, because the tasks where the fifth index point pays for itself are, by construction, the tasks where a cheap model's failure is expensive.

Cost card headed 'The same agent run, priced twice', listing a stated assumption of a 30-step agent run reading 200,000 tokens of context per step and writing 2,000 tokens per step for 6M input and 60K output tokens per run; MiMo-V2.6-Pro at $0.43 / $0.87 costing $2.58 input plus $0.05 output for $2.63 per run before caching; Claude Opus 5 at $5.00 / $25.00 costing $30.00 input plus $1.50 output for $31.50 per run before caching; a per-token ratio of 11.6x on input, 28.7x on output and 21x on the blended rate; a measured $0.13 against $5.86 per Intelligence Index task, a 45x gap; and the caching comparison of a 99% cache discount against a $0.50 cache read and $6.25 cache write.

What a routing decision looks like here

Claude Opus 5 is reachable through a single OrcaRouter key at provider list price with 0% markup, as anthropic/claude-opus-5, with the frontier models you would compare it against sitting alongside it on the same key. Because we pass the provider's list price through without a markup, a vendor price change on either side is live on our side the same day rather than waiting for a re-quote. The routing DSL is where the interesting part sits: you can compose the two models into one call rather than choosing between them, send the bulk of a workload to the cheap lane and route the tail — the tasks that fail a self-check, or that match a complexity predicate — to the expensive one. Failover is the other half. Claude Opus 5 has broad provider coverage; MiMo-V2.6-Pro was listed by a single API provider when Artificial Analysis captured it, which is a thin route for a model you would put in a production path. A router that can fall back is what makes the cheap lane safe to adopt.

Worth stating plainly, because it is easy to assume otherwise: we do not host MiMo-V2.6-Pro. Reaching it today means Xiaomi's own platform or one of the third-party catalogues listing it. The routing argument here is about how you use a model like this alongside the ones we do carry, not a claim that it is one of ours.

Which one to pick

Pick Claude Opus 5 when the task's failure cost exceeds the token cost by enough that five index points is cheap insurance — long-horizon agentic work with real side effects, tool use where a wrong call is worse than a slow one, anything where you are already paying a human to check the output. The $5.86-per-index-task figure is not a reason to avoid it; it is the price of the tier, and the tier exists for a reason.

Pick MiMo-V2.6-Pro when volume is the constraint, when the workload is context-heavy and repetitive, when you want the weights in your own infrastructure — the MIT licence permits commercial deployment, modification and further training — or when you are building something whose unit economics only work at a fifth of a cent per thousand tokens. Its 134.3 tokens/second makes it viable for interactive use in a way that most trillion-parameter open models are not.

Pick both, if you have a router, because the honest answer to "which is better" is that they are not competing for the same requests. The failure mode to avoid is the one that costs the most: standardising on the cheap model because the headline ratio is 21x, discovering in month three that a specific class of request needs the expensive one, and having no path to route it there. The second failure mode is the mirror image — paying Opus 5 rates for a summarisation job that a 46-index model finishes identically. Neither is a model problem. Both are configuration problems, and configuration is the part you can change on a Tuesday afternoon.