Generated title card reading GPT-6 Sol vs MiniMax M3, what eight times the output price does not buy, with stat cards reading output price $10.00 vs $1.20 per million, Intelligence Index 47.6 vs 29.2, and weights closed vs open, with a footer reading Index per Artificial Analysis v4.3.2; MiniMax M3 rates per MiniMax and GPT-6 Sol rates per OpenAI's rate card.
Guides & Insights

GPT-6 Sol vs MiniMax M3: What Eight Times the Output Price Still Does Not Buy

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The first thing to say about GPT-6 Sol against MiniMax M3 is that the price line is not close and the score line is not close either, and they point in opposite directions. MiniMax M3 is an open-weight model you can download and run yourself, and it is listed at $0.30 per million input tokens and $1.20 per million output; GPT-6 Sol, the middle tier of the GP​T-6 generation released 22 September 2026, is $2.00 and $10.00 on the same units. That is a little over eight times the output rate. MiniMax M3 also accepts video as an input modality, which GPT-6 Sol does not — the Sol model takes text, image and file. What M3 does not have is a score anywhere near Sol's: 29.2 against 47.6 on Artificial Analysis's Intelligence Index, on the same revision v4.3.2 board.

That asymmetry is the whole story, and it is a more interesting one than a simple price-versus-quality trade, because the open-weight column gets to keep something the closed column cannot sell at any price: M3's checkpoints are on Hugging Face and the model can be run inside a network boundary. If you are choosing between these two, the useful question is not which is better. It is which of the four separate things you are buying — price, throughput, control, or capability — you can afford to lose.

Four dimensions, and they do not rank the same way

• Output price — GPT-6 Sol $10.00 per million vs MiniMax M3 $1.20 per million
• Input price — GPT-6 Sol $2.00 per million vs MiniMax M3 $0.30 per million
• Cached input — GPT-6 Sol $0.20 per million vs MiniMax M3 $0.06 per million
• Weights — GPT-6 Sol closed, API only vs MiniMax M3 open-weight and self-hostable
• Input modalities — GPT-6 Sol text, image and file vs MiniMax M3 text, image and video
• Context window — GPT-6 Sol 1,050,000 tokens vs MiniMax M3 1,048,576 tokens
• Max output — GPT-6 Sol 128,000 tokens vs MiniMax M3 512,000 tokens
• Intelligence Index — GPT-6 Sol 47.6 vs MiniMax M3 29.2
• Coding index — GPT-6 Sol not published on this revision vs MiniMax M3 58.6
• Long-request step — GPT-6 Sol reprices the whole request above 272,000 input tokens vs MiniMax M3 no step on its card

Two of those lines deserve more than a row. The max-output ceiling runs the opposite way to everything else: M3 will emit up to 512,000 tokens in one response, four times GPT-6 Sol's 128,000. For a workload that generates a whole file or a long report in a single call, that is a structural advantage, not a rounding error. And the long-request step is a real GPT-6 Sol cost that M3 does not have at all — above 272,000 input tokens, Sol reprices the entire request to $4.00 / $15.00, while M3's card carries no equivalent clause. On a 300,000-token prompt the output rate comparison is not eight times, it is twelve and a half.

So the honest ranking depends entirely on the workload. For high-volume classification or extraction where a wrong answer is cheap and a slow answer is not, M3 at $0.30 / $1.20 with no long-context penalty is the sane default and Sol is money burned. For a task where the first answer has to be right — a migration plan, a security review, a piece of reasoning that will be read by a person who will act on it — a 18.4-point index gap at eight times the output price is a trade most teams take, because the alternative is paying a human to fix it.

Generated comparison scoreboard titled GPT-6 Sol vs MiniMax M3, the scoreboard, with six matched rows. GPT-6 Sol: input $2.00 per million, output $10.00 per million, weights closed, Intelligence Index 47.6, context window 1,050,000 tokens, reprices above 272K input tokens. MiniMax M3: input $0.30 per million, output $1.20 per million, weights open and downloadable, Intelligence Index 29.2, context window 1,048,576 tokens, no long-request step. A footer reads MiniMax M3 figures per Artificial Analysis v4.3.2 and MiniMax; GPT-6 Sol figures per Artificial Analysis v4.3.2 and OpenAI.

Where M3's published numbers are unusually good, and where they are thin

MiniMax M3's strongest independent figures are not where you would expect from a model at this price. Long-context recall lands at 83.0, which is 0.7 below GPT-6 Sol's 83.7 and effectively a tie — on a million-token window, at a sixth of the input price. GPQA Diamond comes in at 92.9, close enough to the frontier band that the difference is within the range you would expect from reasoning-effort settings rather than from capability. Those two lines are why M3 is worth taking seriously rather than filing under "cheap".

The thin parts are equally clear and should be stated rather than glossed. Retail-style tool use is poor: on tau-banking M3 scores 15.3, and on the newer Terminal-Bench 4.0 revision it scores 2.0. Both of those are third-party measurements, and both suggest a model whose benchmark-average hides a specific, large weakness in agentic tool-calling against a simulated service. The vendor's own reported numbers paint a much better picture — SWE-Bench Pro 59, BrowseComp 83.5, MCP Atlas 74.2, Terminal-Bench 2.1 at 66 — and those are vendor figures on its own harness, which is a claim rather than a measurement. Any deployment that hinges on tool use should read the pair of numbers together and test it directly.

There is a second-order point about the benchmark comparison itself. GPT-6 Sol is measured on a smaller benchmark set in our catalogue than M3 is — five scores against M3's twenty-three — because the vendor publishes less and third parties have run fewer of the tests against it. A model with twenty-three published numbers will always look more thoroughly evaluated than one with five, and the difference is documentation, not capability.

What open weights actually get you, beyond a lower bill

Self-hosting M3 is the option GPT-6 Sol cannot match, and its value is not mainly financial. At $0.30 / $1.20 the API is cheaper than most teams' cost of running the hardware, so the reason to download it is a constraint rather than a spreadsheet: data that cannot leave a boundary, a workload with no acceptable external dependency, a fine-tune, or a latency budget that a shared endpoint cannot meet. Open weights are the only one of those four the closed model can answer, and it answers it by not being available. GPT-6 Sol's answer is that you do not need to run it — which is a different argument, and true for a different set of teams.

Throughput deserves its own line because it is the number a demo never shows. MiniMax M3 is measured at roughly 495 output tokens per second in our rolling seven-day window, and GPT-6 Sol at roughly 244. M3 is not merely the cheaper model on this comparison, it is the faster one on our routes by about a factor of two, which changes what a batch job costs in wall-clock as well as in dollars. The caveat is that these are rolling windows on a shared playground, not a controlled benchmark, and they move.

OrcaRouter model page for MiniMax M3 (minimax/minimax-m3) by MiniMax, dated 2026-05-31, tagged Vision, Tools, JSON and Reasoning, describing it as MiniMax's flagship open-weight foundation model with a 1M-token context window and text, image and video input, listing $0.30 per million input and $1.20 per million output, with a model-page note that it is available through OrcaRouter under the model id minimax/minimax-m3.

Running the pair as one system

Both models sit behind a single OrcaRouter key alongside 200+ others, which makes the interesting configuration here a cheap-first cascade rather than a choice. Route the volume to MiniMax M3, keep GPT-6 Sol behind it for the requests that fail a check or trip a difficulty signal, and the blended cost lands far closer to the open-weight rate than to Sol's while the hard calls still get the model that scores 47.6. OrcaRouter's routing DSL is where that composition is expressed — a call can name several models and the conditions between them instead of hard-coding one endpoint per task.

For teams that genuinely cannot decide, model fusion is the blunt version of the same idea: a panel of models answers together and the responses are combined. It is a poor fit for volume work and a reasonable fit for a rarely-run, high-stakes prompt where the difference between a 29.2 and a 47.6 answer is the only thing on the page.

Failover matters more with an open-weight model than with a closed one, and for a specific reason: M3 is served by more than one infrastructure provider, and endpoints differ. A second route configured in advance means a slow or degraded host is a retry rather than an outage. And because our catalogue passes provider list prices through with no markup, a vendor rate change is live on our side the same day — which matters on a $1.20 output line where a vendor moving a single cached-input rate visibly changes a bill.

Artificial Analysis model page for MiniMax-M3, described as an open-weights model released June 2026, showing an Intelligence rank of 18 of 117 with an Intelligence Index of 29 on a comparable-model median of 18, an output speed rating above average, a 1M-token context window, and support for text, image and video input with text output.

Which one you should actually call

Take MiniMax M3 for volume, for anything that wants a video input, for anything that needs to emit more than 128,000 tokens in one response, and for anything that has to run inside your own perimeter. Take GPT-6 Sol for the requests whose answer is the product, for tool-heavy agentic loops where M3's tau-banking and Terminal-Bench 4.0 scores are a warning rather than a footnote, and for anything where an 18-point index difference is worth eight times the output rate to you. Take both, and let a router decide, if neither of those describes your whole traffic.

One thing to check before either, and it is the same checkpoint this blog keeps hitting: M3 has been the flagship of the M-series line since May 2026 and remains the newest M-series model in our catalogue, so it is genuinely current — but GPT-6 Sol has already been superseded by GPT-6.1 Sol at the same short-context rates with cheaper cached input, and its own GPT-6 Sol page carries a pointer to it. If the reason you are looking at Sol here is capability rather than the eight-day-old integration, look at the successor first.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily