
MiMo-V2.6-Pro vs Grok 4.6: Both Vendors Published DeepSWE Numbers, and Neither Is on the Board
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Two vendors, one benchmark, two numbers that cannot be subtracted. SpaceXAI's launch material for Grok 4.6 reports 65.9% on DeepSWE v1.1. Xiaomi's material for MiMo-V2.6-Pro reports 72.57 on DeepSWE v1.1. Read together they suggest the cheaper open-weights model beats the proprietary frontier one by nearly seven points. Read carefully they say almost nothing, because one is a launch-deck percentage on the vendor's own harness and the other is an average-of-three from a reinforcement-learning run on Xiaomi's own grader, and neither has been submitted to the public DeepSWE board. The comparison between MiMo-V2.6-Pro and Grok 4.6 that survives scrutiny is on the independent index, and there the answer is the reverse of the vendor numbers: 46 against 44, in Xiaomi's favour, by two points.
Grok 4.6 was released on 12 August 2026 by SpaceXAI and is the incumbent here — a shipped, broadly served frontier model with a documented price list and months of independent measurement behind it. Xiaomi MiMo-V2.6-Pro went generally available on 22 September 2026, with its reinforcement-learning checkpoint repositories appearing the day before, and it is the newcomer. Both are reasoning-capable, both are built for long-running agentic work, and the price gap between them is large enough that the newcomer's inexperience is the only thing holding the comparison back.
What each model actually is
Grok 4.6 is proprietary, takes text and image input and returns text, and carries a knowledge cutoff of 1 February 2026. Its context window is 500,000 tokens, which is half the size of its main competitors' and the single most consequential spec on its sheet. Its list rates are $2.00 per million input tokens, $0.50 per million cached input and $6.00 per million output below 200,000 prompt tokens, with a cliff at that boundary: at or above 200K the rates become $4.00, $1.00 and $12.00, and the higher tier bills the whole request rather than the portion over the line. Priority processing doubles the standard rate. Artificial Analysis measured 63.0 output tokens per second and 42.79 seconds to first token at maximum effort, and $2,351.83 to run the full index.
Xiaomi MiMo-V2.6-Pro is a sparse mixture-of-experts model at 1.02 trillion total parameters and 42 billion activated per token, with a 1M-token context window and native multimodality across text, image, video and audio. It is MIT-licensed with downloadable weights, and Xiaomi kept the previous generation's pricing rather than repricing the new checkpoint upward — $0.43 per million input tokens and $0.87 per million output, a 99% cache discount, and a blended $0.18 at a 7:2:1 cache-hit/input/output mix. Artificial Analysis measured 134.3 output tokens per second, 2.15 seconds to first token, and $206.66 to run the index — $0.13 per task.
• Weights — MiMo-V2.6-Pro MIT-licensed and downloadable; Grok 4.6 proprietary, API only
• Parameters — MiMo-V2.6-Pro 1.02T total / 42B active; Grok 4.6 undisclosed
• Context — MiMo-V2.6-Pro 1M tokens; Grok 4.6 500K, with a price cliff at 200K of input
• Modality — MiMo-V2.6-Pro takes text, image, video and audio; Grok 4.6's published input surface is text and image
• Index score — MiMo-V2.6-Pro 46; Grok 4.6 44, both Artificial Analysis v4.3.2
• List price — MiMo-V2.6-Pro $0.43 / $0.87 per 1M; Grok 4.6 $2.00 / $6.00, doubling above 200K input
• Blended rate — MiMo-V2.6-Pro $0.18 per 1M; Grok 4.6 $1.35
• Cost to run the index — MiMo-V2.6-Pro $206.66; Grok 4.6 $2,351.83
• Speed — MiMo-V2.6-Pro 134.3 output tokens/second; Grok 4.6 63.0
• Time to first token — MiMo-V2.6-Pro 2.15 seconds; Grok 4.6 42.79 seconds
• Knowledge cutoff — MiMo-V2.6-Pro not published; Grok 4.6 1 February 2026

The benchmark both vendors ran, and why it does not compare
DeepSWE v1.1 is a real, public, third-party coding-agent evaluation run by Datacurve, and it is the one task family both companies chose to publish against. That makes it tempting. It should not be used as a head-to-head, and the reason is not that either vendor is lying — it is that the two figures describe different measurements of the same task name.
Xiaomi's 72.57 is an average-of-three on its own harness with mini-swe-agent, produced during a reinforcement-learning run rather than from a frozen release candidate, and reported alongside a starting figure of 58.4. The run is documented as 30 steps and roughly 750,000 trajectories per model, completed in under six days at a published spend of about $2.62 million for the Pro model. It has not been submitted to Datacurve's board. SpaceXAI's 65.9% for Grok 4.6 is a launch-deck figure from the vendor's own evaluation set, reported next to gains over Grok 4.5 of 11.9 points on the same benchmark — and the public board carries a submitted Grok 4.6 configuration at around 67%, which is close enough to the vendor number to suggest the harness is at least roughly aligned. That is not true of the Xiaomi figure, which has no public counterpart at all.
So the honest statement is this: on the public board, Grok 4.6 is present and MiMo-V2.6-Pro is absent. On the independent composite both were actually run against by the same evaluator, Xiaomi's model leads by two points. Those two facts are not in tension. They simply measure different things, and the second one is the one to quote.
The 500K context is the spec that decides real workloads
Half a million tokens is a large context window by any normal standard and a small one for 2026. The models MiMo-V2.6-Pro is grouped with all sit at 1M, and the practical consequence shows up in exactly the workloads Grok 4.6 is marketed for. A long-running coding agent holding a repository, a tool transcript and an accumulated plan will cross 200,000 tokens of prompt in a session — and at that line, Grok 4.6's rates double and apply retroactively to the whole request. The same session on MiMo-V2.6-Pro runs to a million tokens without a tier change.
That is a specification difference and a pricing-structure difference at the same time, and the second one is easy to miss when comparing headline rates. A blended $1.35 for Grok 4.6 against $0.18 for MiMo-V2.6-Pro is a 7.5x gap. On a prompt that crosses 200K it widens to something closer to 15x, because the Grok rate doubles and the Xiaomi one does not move.
The counterargument is on the record too, and it deserves to be. Grok 4.6's launch thesis was turn efficiency rather than raw context: in the agentic evaluation where it was measured against Claude Opus 5 on identical tasks, it finished in roughly 53 turns and about 500 million input tokens against roughly 103 turns and two billion tokens for the other model, at an estimated $0.84 per task — the lowest figure among the frontier models in that comparison. A model that goes around the loop half as many times does not need as much context, and it does not burn as many tokens. That is a genuine architectural advantage and it is the strongest argument for Grok 4.6 against any cheaper per-token alternative, including this one.

Getting to each of them
Grok 4.6 is on OrcaRouter as grok/grok-4.6, reachable through one OpenAI-compatible endpoint at provider list price with 0% markup — so a SpaceXAI price change, including a change to that 200K cliff, is live on our side the same day. Xiaomi MiMo-V2.6-Pro is not on OrcaRouter. No Xiaomi model is, and the route to it today is Xiaomi's own platform or one of the third-party catalogues listing it. Saying that plainly is more useful than letting a model page imply otherwise.
What that asymmetry is worth in practice is a cheap experiment. Grok 4.6 is the model you may already be paying for. MiMo-V2.6-Pro is a two-point index improvement at a fraction of the rate, launched this week, with a single serving route behind it. The question worth answering is not which is better in the abstract but how much of your Grok traffic the Xiaomi model can take on your own prompts — and answering it by opening a second contract, a second key and a second billing relationship is the friction that stops the test from happening. With one endpoint and automatic failover across providers, the comparison is a config change, and the failover half matters more than usual here: a model that shipped this week and was listed by one API provider when Artificial Analysis captured it is not something to put in a production path without a lane to fall back to.

Which one to reach for
Pick Grok 4.6 when turn efficiency is the thing you are optimising — long agent loops where the model's ability to finish in half the steps is worth more than its per-token rate — or when you want a model with months of independent measurement behind it rather than a week. Its 63 tokens per second and 42.79-second time to first token make it a batch and offline-workload model in interactive terms, and its 500K context with a 200K price cliff argues for keeping prompts short.
Pick MiMo-V2.6-Pro when you are running context-heavy, repetitive loops that would cross Grok's 200K cliff, when you need first-token latency low enough for a human to wait, when you want the weights in your own infrastructure, or when volume makes the 7.5x blended-rate gap the deciding number. Its two-point index lead is small enough that it should not be the reason on its own; the $206.66 against $2,351.83 cost to run the same evaluation suite is a better one.
What would settle the open question is a submitted DeepSWE configuration for MiMo-V2.6-Pro. Grok 4.6 already has one. The moment Xiaomi's model appears on the public board with a stated harness, a confidence interval and a cost per task, the 72.57 stops being a vendor number and the comparison above becomes a real one — and given the two-point gap on the index, it could go either way.
OrcaRouter puts 200+ models behind one OpenAI-compatible endpoint at provider list price with 0% markup, so a vendor price change is live the same day.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
