
MiMo-V2.6-Pro vs GPT-5.6 Sol: One Index Point, Twenty-Two Times the Blended Rate
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
On the Artificial Analysis Intelligence Index v4.3.2, read on 22 September 2026, the vendor's GPT-5.6 Sol at maximum effort scores 47 and Xiaomi's MiMo-V2.6-Pro scores 46. One point. The blended rates Artificial Analysis publishes for the two — a 7:2:1 cache-hit, input, output mix — are $3.08 for GPT-5.6 Sol and $0.18 for MiMo-V2.6-Pro. Twenty-two times the money for one index point is the entire shape of this matchup, and the interesting question is not whether the point is worth it. It is which specific requests the point lives in, because the aggregate does not tell you.
What makes the pair worth comparing at all is that they are not obviously in the same category. GPT-5.6 Sol is a proprietary frontier model with a multi-agent mode and a price list to match. Xiaomi MiMo-V2.6-Pro is an MIT-licensed open-weights checkpoint that shipped on 22 September 2026 with the reinforcement-learning repositories landing the day before, and it is priced as if it were a mid-tier model. That a trillion-parameter open model lands within one point of OpenAI's flagship is the news. What you do with that fact is a separate decision.
The two models, stated plainly
GPT-5.6 Sol was released on 9 July 2026 as the flagship of the GPT-5.6 family, alongside Terra and Luna, and OpenAI routes the bare gpt-5.6 alias to it. It is proprietary, takes text and image input and returns text, supports reasoning effort from none through xhigh plus an "ultra" multi-agent mode exclusive to Sol, and carries a February 2026 knowledge cutoff. Its list price is $4.00 per million input tokens and $20.00 per million output on OpenAI's first-party API, with cached input at $0.50 and a 1M-token context window. Artificial Analysis measured 77.3 output tokens per second and 127.21 seconds to first token, at a cost of $3,464.84 to run the full index.
Xiaomi MiMo-V2.6-Pro is a sparse mixture-of-experts model at 1.02 trillion total parameters and 42 billion activated per token, with a 1M-token context window and native multimodality across text, image, video and audio. Xiaomi kept the previous generation's pricing rather than repricing the new checkpoint upward, which is why a model this size lists at $0.43 and $0.87 per million tokens with a 99% cache discount. Artificial Analysis measured 134.3 output tokens per second, 2.15 seconds to first token, and $206.66 to run the index — $0.13 per task.
• Weights — MiMo-V2.6-Pro MIT-licensed and downloadable; GPT-5.6 Sol proprietary, API only
• Parameters — MiMo-V2.6-Pro 1.02T total / 42B active; GPT-5.6 Sol undisclosed
• Context — 1M tokens both; GPT-5.6 Sol caps output at 128K
• Modality — MiMo-V2.6-Pro takes text, image, video and audio; GPT-5.6 Sol's published input surface is text and image
• List price — MiMo-V2.6-Pro $0.43 / $0.87 per 1M; GPT-5.6 Sol $4.00 / $20.00, cached input $0.50
• Blended rate — MiMo-V2.6-Pro $0.18 per 1M; GPT-5.6 Sol $3.08
• Measured cost per index task — MiMo-V2.6-Pro $0.13; GPT-5.6 Sol $3,464.84 for the full index
• Speed — MiMo-V2.6-Pro 134.3 output tokens/second; GPT-5.6 Sol 77.3
• Time to first token — MiMo-V2.6-Pro 2.15 seconds; GPT-5.6 Sol 127.21 seconds
• Reasoning control — GPT-5.6 Sol exposes none/low/medium/high/xhigh/max plus an ultra multi-agent mode; MiMo-V2.6-Pro publishes no equivalent ladder

Where the one point actually sits
A one-point gap on a ten-evaluation composite is a rounding error at the aggregate and a chasm at the component. The suite covers AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1 — reasoning, knowledge, mathematics and coding — and neither vendor publishes a per-evaluation breakout for these two models. That absence is a real limitation on both sides of this page and worth saying rather than papering over with a composite.
What is on the record is directional, and it points at agentic work as the place the gap concentrates. GPT-5.6 Sol's launch material reports 53.6% on Agents' Last Exam, a long-running professional-workflow evaluation spanning 55 fields, which OpenAI described as a new high; 64.6% on SWE-Bench Pro; 88.8% on Terminal-Bench 2.1; and 62.6% on OSWorld 2.0 for computer use. Its BrowseComp score of 90.4% rises to 92.2% in the ultra mode that runs four parallel sub-agents at roughly four times the token cost. Those are vendor figures on vendor harnesses, and the usual caveat applies to every one — but they are the categories where a one-point composite gap is least likely to be evenly distributed, and Xiaomi publishes no comparable agentic breakout for MiMo-V2.6-Pro.
The counterweight is that MiMo-V2.6-Pro's own vendor numbers are on a different task family entirely. Xiaomi reports the model moving from 58.4 to 72.57 on DeepSWE v1.1 across its reinforcement-learning run, evaluated with mini-swe-agent at average-of-three on Xiaomi's own grader and never submitted to a public board. Placing 72.57 beside any Sol figure would be inventing a comparison that does not exist. The only number both models were actually run against by the same evaluator is the 46 and the 47.
The cost comparison, done properly
Per-token arithmetic understates the gap, because the two models do not consume the same number of tokens to finish the same job. The fair single number is what Artificial Analysis measured to run its full index against each: $206.66 for MiMo-V2.6-Pro and $3,464.84 for GPT-5.6 Sol. Same suite, same harness, measured identically — a 16.8x difference on a controlled task set, and one that already absorbs the fact that Sol is a reasoning model emitting tokens to think.
Now a production loop. A 30-step agent run reading 200,000 tokens of context per step and writing 2,000 tokens per step is 6 million input and 60,000 output tokens per run. On GPT-5.6 Sol at list that is $24.00 of input and $1.20 of output — $25.20 per run before caching. On MiMo-V2.6-Pro it is $2.58 and $0.05 — $2.63. Caching moves both: Sol's cached input at $0.50 against MiMo-V2.6-Pro's 99% discount still leaves the expensive model an order of magnitude out in front on a context-heavy loop, which is precisely the loop an agent runs.
The latency half of the comparison is starker than the price half and gets less attention. GPT-5.6 Sol measured 127.21 seconds to first token. MiMo-V2.6-Pro measured 2.15. That is a fifty-nine-fold difference, and it is the number that decides whether an interactive product is possible at all. Sol at maximum effort is a batch model in practice — you submit, you wait, you come back. If your application blocks on the first token, the choice has already been made for you, whatever the index says.

One key, both lanes
Both models are reachable through a single OrcaRouter key at provider list price with 0% markup — GPT-5.6 Sol as openai/gpt-5.6-sol, which is our own route to it. Because we pass the provider's list price through without a markup, a price change on either side is live on our side the same day rather than waiting for a re-quote.
The routing DSL is where this pairing gets interesting rather than merely convenient. Two models one index point apart, one costing sixteen times more per task, are not alternatives to choose between — they are a two-lane configuration. Send the bulk of a workload to the cheap lane and route the tail to the expensive one on a predicate you define: requests that fail a self-check, requests matching a complexity rule, requests whose failure cost exceeds the token cost by enough that the point is cheap insurance. Failover is the other half of it. Sol is served broadly; MiMo-V2.6-Pro launched this week and was listed by a single API provider when Artificial Analysis captured it, which is a thin route for a model you would put in a production path. A router that can fall back is what makes the cheap lane safe to adopt, and it is also what makes the expensive lane optional rather than mandatory.

Which one to pick
Pick GPT-5.6 Sol when the task's failure cost dwarfs the token cost — long-horizon agentic work with real side effects, computer-use automation, tool use where a wrong call is worse than a slow one, anything where you are already paying a human to check the output. The $3,464.84 index run is not a reason to avoid it; it is the price of the tier, and the tier exists for a reason.
Pick MiMo-V2.6-Pro when volume is the constraint, when the workload is context-heavy and repetitive, when you need first-token latency low enough for a human to wait, when you want the weights in your own infrastructure — MIT permits commercial deployment, modification and further training — or when the unit economics only work at a fifth of a cent per thousand tokens. Its 134.3 tokens per second makes it usable interactively in a way that most trillion-parameter open models are not, and that is a capability, not a discount.
Pick both if you have a router, because the honest answer to "which is better" is that a one-point composite gap is not a verdict — it is a measurement of how similar they are on average and how different they probably are at the tail. The failure mode to avoid is standardising on the cheap model because the ratio is 22x, discovering in month three that a class of request needs the expensive one, and having no path to route it there. The mirror failure is paying Sol rates for a summarisation job a 46-index model finishes identically. Neither is a model problem. Both are configuration, and configuration is the part you can change on a Tuesday afternoon.
OrcaRouter puts 200+ models behind one OpenAI-compatible endpoint at provider list price with 0% markup, so a vendor price change is live the same day.
