
Xiaomi MiMo-V2.6-Pro vs Grok 4.7: 30× Cheaper Per Task, and Neither One Is Fast
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 36 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 181 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1277 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 110 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
Xiaomi MiMo-V2.6-Pro and Grok 4.7 both went out on 21 September 2026, both land on 46 on the Artificial Analysis Intelligence Index v4.3.2, and both get called slow by the page that measured them — 43.6 output tokens per second for MiMo-V2.6-Pro against a 67.1 median, 40 for Grok 4.7 at xhigh effort. What separates them is what an answer costs. The evaluator’s per-task figure is $0.13 for the Chinese open checkpoint and $3.74 for the Grok 4.7 model, a factor of roughly 29. That is not a rate difference, or not only one: Grok 4.7’s output rate is $6.00 per million tokens against $0.87, which is seven-fold, and the rest of the multiple comes from how many tokens each model burns getting to an answer.
Two models in the same month, on the same index, at the same score, priced like they belong to different decades of the market. The interesting question is not which one wins. It is which part of that 29× you can actually route around — because in this pairing exactly one of the two is a route you can buy today, and it is the expensive one.
Same day, same score, different animals
Grok 4.7 shipped on 21 September 2026 at $2.00 per million input tokens and $6.00 per million output with a 500,000-token context window, a May 2026 knowledge cutoff, a 75% cache discount, and text and image input against text output. It measured 46 on the index at xhigh effort, 56 on the Coding Agent Index in the Grok Build harness, and 1,657 Elo on AA-Briefcase — fourth on that board, behind three Anthropic entries, at roughly half Claude Opus 5’s cost per task.
Xiaomi MiMo-V2.6-Pro is a sparse mixture-of-experts checkpoint, 1.02 trillion total parameters with 42 billion active, MIT-licensed, 1M-token context, text, image, speech and video input against text output, $0.43 and $0.87 per million tokens with a 99% cache discount that collapses the effective input rate to near nothing on a warm prompt. Artificial Analysis places it first of 114 open-weights models on the index and twelfth of 114 on cost. It is reachable through three API providers. It is not on our routes, and we do not list a model before a provider has onboarded it.
The only line that decides it
• Cost per index task — MiMo-V2.6-Pro $0.13; Grok 4.7 $3.74 — 29× • Output rate — MiMo-V2.6-Pro $0.87 / 1M; Grok 4.7 $6.00 / 1M — 6.9× • Output tokens on the index run — MiMo-V2.6-Pro 140M; Grok 4.7 240M, against an 88M median • Output speed — MiMo-V2.6-Pro 43.6 t/s; Grok 4.7 40 t/s — both under the median • Context window — MiMo-V2.6-Pro 1M; Grok 4.7 500K • Weights — MiMo-V2.6-Pro MIT-licensed and downloadable; Grok 4.7 proprietary, API only
Read the first two bullets together and the shape of the gap becomes clear. At list price the two are 6.9× apart on output. On the index run they are 29× apart on what an answer costs. The multiplying factor in between is verbosity: Grok 4.7 generated 240 million output tokens against an 88 million median for its class, and MiMo-V2.6-Pro generated 140 million. That is a 1.71× token penalty sitting on top of the 6.9× rate penalty, and 1.71 × 6.9 is 11.8 — the rest of the distance to 29 comes from the input side, where MiMo-V2.6-Pro’s 99% cache discount and $0.43 input rate do the heavy lifting on a workload with any repeated prefix at all.

Verbosity is the number people leave out of the estimate
Cost models are usually built from a rate card and an assumed token count. Both halves of that are wrong here. The rate card is wrong because the cache discount is not a footnote — 99% against 75% is the difference between a system prompt costing you money every call and costing you almost nothing. The token count is wrong because the two models do not emit the same number of tokens for the same work, and the evaluator measured the difference: 240 million against 140 million, a 1.71× spread on an identical benchmark.
Neither of those is a proxy you can read off a spec sheet. Grok 4.7’s verbosity rank is #94 of 210 models in its class; MiMo-V2.6-Pro’s cost rank is #12 of 114 in its. A team that quotes the Grok 4.7 rate card and assumes a token-neutral workload will forecast roughly a quarter of what the index actually spent per task. The correction is not to use the index cost as your estimate either — it is a measurement of one benchmark’s shape. It is to measure your own token counts on both sides before committing, because the gap between the two models is 6.9× on paper and 29× in practice, and only one of those numbers is real for your traffic.

Both are slow, and that is not a tie
Artificial Analysis measured Grok 4.7 at 40 output tokens per second and calls it “among the slowest”; MiMo-V2.6-Pro at 43.6 and calls it “notably slow”. Those two readings are close enough that speed should not be the tiebreaker between them. But the medians underneath are not: 67.1 for the open-weights field MiMo-V2.6-Pro sits in. Both models are well under it, and MiMo-V2.6-Pro also carries a 3.17-second time to first token against a 2.31-second median, which is the number that hurts in an interactive loop rather than a batch one.
So the honest summary of the latency side is that 29× cheaper does not buy you a faster product, it buys you a slower one that costs less — and if your users are watching a cursor, the 3.17-second wait may cost more than the tokens saved. For overnight batch work, offline evaluation, or any pipeline where the clock is measured in hours, the trade is straightforward and the 29× is nearly free.
What a router can and cannot do with this pairing
This is the pairing where our own catalogue is genuinely one-sided, and it is worth being plain about it. Grok 4.7 is live on OrcaRouter as grok/grok-4.7 at the provider’s own $2.00 / $6.00 with a 450K maximum output and an OpenAI-SDK-compatible endpoint at https://api.orcarouter.ai/v1 — so if the comparison makes you want the xAI model behind the same key as everything else, that route exists today. Xiaomi MiMo-V2.6-Pro is not on our routes; for that side, the vendor’s own API or whichever of its three listed providers you already use is the answer, and we will not pretend otherwise.
Where a router earns its place in a comparison like this is on the parts that are not the model. One key across 200+ models at each provider’s list price with 0% markup means a vendor price cut lands on your invoice the same day rather than at the next contract renewal. Automatic failover means the 40 t/s route degrading does not take your pipeline with it. The routing DSL lets the split be made on the published per-task cost rather than on brand loyalty — cheap model for the high-volume, hard model for the hard calls, decided per request. And for the workloads where neither model is quite right, model fusion can run several behind one call.

Questions this comparison usually raises
Does the 29× hold for my workload? Only if your token mix resembles the index run’s. If your outputs are short and your prompts are long and cached, the multiple collapses toward the rate gap of 6.9× on output and widens on input, where MiMo-V2.6-Pro’s 99% discount is the decisive term. If your outputs are long reasoning traces, it widens past 29×, because Grok 4.7’s verbosity penalty is proportional to output length.
Is the 46 on each side the same 46? It is the same index, the same version, and the same scale — the underlying sub-scores were 46.32 and 46.45, a 0.13-point spread, and the index rounds both to 46. Artificial Analysis does not publish per-evaluation breakouts for either model, so the composite is the finest grain available and a 0.13-point difference should not be read as a capability ranking.
Which one should a new project default to? If the workload is batch, latency-tolerant, and cost-sensitive, MiMo-V2.6-Pro at $0.13 a task is the default and the open weights mean you are not exposed to a single provider’s pricing. If you need the xAI model specifically — the Grok Build harness results, the AA-Briefcase placement, the 500K context with a May 2026 cutoff — Grok 4.7 is a live route on OrcaRouter today and the per-task premium is the price of that specific capability. There is no model here that wins both.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
