A generated hero card titled 'GPT-6.1 Sol vs MiniMax M3' with two rounded panels: left 'GPT-6.1 Sol' listing 'OpenAI - released 29 Sep 2026', '$2.00 in / $10.00 out per 1M', 'AA Index 51.8' and '$0.72 per task'; right 'MiniMax M3' listing 'MiniMax - released 31 May 2026', '$0.30 in / $1.20 out per 1M', 'AA Index 29.2' and '$0.51 per task'; between them the line 'Cheaper card, smaller bill - 22.6 index points apart'; footer 'Figures: Artificial Analysis v4.3.2.' and the OrcaRouter logo bottom-right.
Guides & Insights

GPT-6.1 Sol vs MiniMax M3: The Cheaper Rate Card Loses on the Bill

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

MiniMax M3 asks for a fraction of what GPT-6.1 Sol asks — $0.30 per million input tokens against $2.00, $1.20 per million output against $10.00 — and by the meter that actually runs, it is cheaper: $0.508 per task against $0.724 on Artificial Analysis's v4.3.2 evaluation. That is a genuine win, and it is also the whole of the argument, because the same measurement that hands M3 the lower bill hands GPT-6.1 Sol a 22.6-point index lead, and the benchmark set that measures whether a model can finish agentic work rather than describe it turns the gap into a wall. The vendor shipped GPT-6.1 Sol on 29 September 2026. The company released M3, its 428-billion-parameter open-weight model, on 31 May 2026.

Both are live, both are priced, and both are one API call away. What follows is the arithmetic that decides between them, and the two places where the arithmetic is not the whole story.

What each model actually is

GPT-6.1 Sol is OpenAI's current flagship for reasoning and software engineering — a closed model, released a single week after the GPT-6 Sol it replaces, and sold on the same $2.00 / $10.00 list price as its predecessor. It carries a $0.10 cached-input rate, half of what GPT-6 Sol charged for cache hits, and it is the model OpenAI points at when it talks about agentic coding rather than chat.

MiniMax M3 is a mixture-of-experts model with 428 billion total parameters and 23 billion active per token, published under the MiniMax Community License — an open-weights release with a non-commercial-flavoured custom licence, not an OSI-approved one. It is served by a wide field of inference providers: Artificial Analysis records fifteen hosts for M3 against eight for GPT-6.1 Sol. That breadth is the practical advantage of open weights, and it is also why price competition on M3 has been faster than on the closed model.

The rate card, and the meter that overrules it

A generated two-column scoreboard titled 'GPT-6.1 Sol vs MiniMax M3 - the scoreboard'. Left column 'GPT-6.1 Sol': Input $2.00 / 1M, Output $10.00 / 1M, AA Index 51.8, Cost per task $0.72, Output tokens 38,128, Time per task 640 s. Right column 'MiniMax M3': Input $0.30 / 1M, Output $1.20 / 1M, AA Index 29.2, Cost per task $0.51, Output tokens 48,192, Time per task 486 s. Footer: 'Both columns Artificial Analysis v4.3.2.' OrcaRouter logo bottom-right.

Line the two published rate cards up and it is not close.

• Input — GPT-6.1 Sol $2.00 per 1M vs MiniMax M3 $0.30 per 1M; M3 is 85% cheaper
• Output — $10.00 per 1M vs $1.20 per 1M; M3 is 88% cheaper
• Cached input — $0.10 per 1M vs $0.06 per 1M
• Cost per AA task — $0.724 vs $0.508; M3 is 30% cheaper, not 85%
• Output tokens per task — 38,128 vs 48,192; M3 writes 26% more
• Time per task — 640 s vs 486 s
• Context window — 1,050,000 vs 1,048,576 tokens; effectively level
• Maximum output — 128,000 tokens vs 512,000; M3 will write one very long answer
• Inputs accepted — text, image and file vs text, image and video
• Index rank — GPT-6.1 Sol 10th of 147 models vs MiniMax M3 62nd

The two cheap lines at the top are the ones a procurement sheet sees. The third line is the one that pays the bill, and it says a 85–88% rate-card advantage becomes a 30% real-world advantage, because M3 emits more tokens per task and because a task is not a fixed quantity of work. Rate cards price tokens; work is priced in tasks. Any comparison that stops at the first two lines is comparing the wrong numbers, in the wrong unit.

Where the thirty percent stops mattering

The task cost is close enough that the index gap decides it for anything beyond bulk text. Artificial Analysis puts GPT-6.1 Sol at 51.83 and MiniMax M3 at 29.22 — a 22.6-point spread, and one that is not evenly distributed across the suite. On the evaluations that ask a model to operate a terminal, edit a repository and verify its own work, the two models are not in the same category:

• Terminal-Bench 4.0 — GPT-6.1 Sol 0.5606 vs MiniMax M3 0.0202
• Terminal-Bench Science — 0.5810 vs 0.0048
• AA-Omniscience — 41.53 vs 1.35
• MMLU-Pro — 0.8595 vs 0.7856
• AutomationBench — 0.6487 vs 0.2125
• τ²-bench — not published for GPT-6.1 Sol in this run vs 0.8889 for M3
• GPQA Diamond — not published for GPT-6.1 Sol in this run vs 0.9293 for M3
• CritPT — 0.3171 vs 0.0371

Read that carefully, because it is not a clean sweep and the exceptions are the interesting part. MiniMax M3 scores 0.8889 on τ²-bench, the tool-use and customer-service dialogue evaluation, and 0.9293 on GPQA Diamond — a graduate-level science score that beats every OpenAI figure published in this particular run. M3 is a competent reasoner and a good conversational tool user. On Terminal-Bench 4.0 it scores 0.0202. That is not "worse"; that is a model that cannot hold a multi-step repository task together at all, and no discount on tokens makes a task the model fails cheaper than the task the model finishes.

There is a second-order cost the per-task figure hides. M3 records 48,192 output tokens per task and a time-to-first-answer of 23.0 seconds against GPT-6.1 Sol's 332.0 seconds — the small model starts talking almost immediately and then reasons at length. If your workload is latency-sensitive but shallow, that is a point in M3's favour. If it is agentic, the token count is what you pay for and the terminal score is what you get.

Our own model card for GPT-6.1 Sol carries the same numbers in the form you would actually route against: the $2.00 / $10.00 rate, a 1,050,000-token context window, a 4.54-second p50 to first token, and the Artificial Analysis index and benchmark block sitting on the same page.

A screenshot of the OrcaRouter model page for openai/gpt-6.1-sol, showing a 1,050,000-token context window, 128K max output, text/image/file input with text output, the $2.00 input and $10.00 output per-million prices, a 4.54 s p50 time to first token and 284.2M tokens routed over seven days, above the OpenAI-compatible code sample and the supported-parameters list.

What the open weights are worth here

M3 being downloadable is the strongest argument for it, and it is worth being precise about what that buys. It buys licence terms that let you run the model inside your own boundary — useful if the data cannot leave — and it buys the price competition that comes from fifteen independent hosts. It does not buy you a cheap deployment: 428 billion parameters at 23 billion active still needs serious inference hardware, and the MiniMax Community License is a custom licence with conditions attached, not an unrestricted one. For most teams, "open weights" means a better price on hosted inference, not a box in the rack.

The OrcaRouter card for MiniMax M3 is where the served numbers live: the $0.30 / $1.20 rate, the 1,048,576-token context window, the 512,000-token maximum output, text/image/video input, and a 56.4-million-token seven-day volume across the providers routing it.

A screenshot of the OrcaRouter model page for minimax/minimax-m3, showing the $0.30 input and $1.20 output per-million prices, a 10.60 s p50 time to first token, 56.4M tokens over seven days, a 1,048,576-token context window, 512,000 max output and text/image/video input, above the OpenAI-compatible code sample.

Both of them behind one key

The practical reason this comparison is a config change rather than a migration is that both models are reachable from a single OrcaRouter endpoint — one API for 200+ models, with the provider's list price passed through at 0% markup, which means a MiniMax rate change lands on your bill the same day the vendor publishes it rather than whenever a reseller renegotiates. That matters more for M3 than for its rival: the whole reason to route to an open-weight model is that its price and its host list are moving, and a pass-through meter is the only kind that tracks a moving target.

Automatic failover is the other half. M3 is served by many providers of varying reliability, and the routing layer can move a request to another host when one degrades, without the caller knowing which host it landed on. That is the honest way to adopt a model whose Terminal-Bench score says "not for the critical path" — put it where a retry is cheap.

Which one, if you have to pick today

If the work is bulk text — extraction, summarisation, classification, dialogue, anything where the output is prose and the failure mode is a wrong sentence rather than a broken build — MiniMax M3 is the correct default and the thirty percent is real money. Use the GPQA and τ²-bench numbers as your evidence: this is a capable model that happens to be cheap.

If the work is agentic — repositories, terminals, tool chains, anything with a verification step — the 0.0202 on Terminal-Bench 4.0 is disqualifying, and GPT-6.1 Sol at 0.5606 for $0.724 a task is the better deal even at 85% more per token. The rate card was never the thing you were buying.

The useful answer is that you do not have to choose once. Route the shallow tier to M3, route the agentic tier to GPT-6.1 Sol, and keep both behind one credential — the routing DSL composes the two into a single call when a task starts as extraction and turns into a repair job.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily