A generated two-column comparison scoreboard titled 'GPT-6.1 Sol vs Kimi K3 - the scoreboard'. Left column 'GPT-6.1 Sol': rows reading 'Intelligence Index: 51.8', 'Cost per index task: $0.72', 'Output tokens on index run: 67M', 'AA-LCR long-context: 0.83', 'Context window: 1,050,000', 'Weights: closed'. Right column 'Kimi K3': rows reading 'Intelligence Index: 43.6', 'Cost per index task: $2.00', 'Output tokens on index run: 160M', 'AA-LCR long-context: 0.89', 'Context window: 1,048,576', 'Weights: open on Hugging Face'. A footer reads 'Independent figures per Artificial Analysis v4.3.2 at maximum effort.'
Guides & Insights

GPT-6.1 Sol vs Kimi K3: The Download Exists, and It Is Still Not the Cheap Option

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Kimi K3 is the counter-example that usually settles this kind of argument. Moonsh​oot AI shipped the 2.8-trillion-parameter model in July 2026 and published the weights — moonshotai/Kimi-K3 is on Hugging Face, ungated, with more than a million downloads and a last revision in September. So there is no "open in principle, withheld for safety hardening" asterisk here, no two-week delay to wait out. GPT-6.1 Sol, which Open​AI released on 29 September 2026, is closed and API-only and always will be. And on the independent board the downloadable model scores 43.6 to the closed model's 51.8 while costing $2.00 per completed index task against $0.72. Open weights are supposed to be the cheap escape hatch; in this pairing they are the expensive side of the trade, and the reason is not the license.

What each model is

Kimi K3 is a sparse mixture-of-experts model with 2,800 billion total parameters and roughly 104 billion — about 3.7% — active per token. It carries a 1,048,576-token context window, takes text and images as input, and returns text. The published checkpoint is an FP8 artifact and it is a genuinely large download; a production deployment on your own hardware is a cluster decision rather than a workstation one. Moonsh​oot's architectural claim is a hybrid linear-attention scheme it says is substantially faster at long-context decoding, plus the residual and routing work that made a 2.8-trillion-parameter model servable at all.

GPT-6.1 Sol is Open​AI's mid-tier refresh — below GPT-6 Astra, above GPT-6 Luna — with a 1,050,000-token window, a 128,000-token output ceiling, text, image and file input, and an April 30, 2026 knowledge cutoff. Reasoning runs low through max with medium as the default; the none rung that the previous GPT-6 Sol accepted is rejected outright, and Open​AI's own migration note points developers at low instead. It is priced at $2.00 and $10.00 per million tokens with cached input at $0.10.

One line per dimension, both sides on each:

• Input — GPT-6.1 Sol $2.00 per 1M vs Kimi K3 $3.00 per 1M
• Output — GPT-6.1 Sol $10.00 per 1M vs Kimi K3 $15.00 per 1M
• Cached input — GPT-6.1 Sol $0.10 per 1M vs Kimi K3 $0.30 per 1M
• Context window — 1,050,000 tokens vs 1,048,576 tokens; a 0.1% difference, immaterial
• Maximum output — 128,000 tokens on both
• Total and active parameters — undisclosed vs 2,800 billion total, 104 billion active per token
• Modality — text, image and file in, text out vs text and image in, text out
• Weights — closed, API only vs open, ungated, 1M-plus downloads
• Long-context clause — reprices the whole request above 272,000 input tokens at 2× input and cache and 1.5× output vs no equivalent clause on the published card
• Intelligence Index, independent, max effort — 51.8 vs 43.6
• Cost per index task, independent — $0.72 vs $2.00
• Output tokens on the index run — 67 million vs 160 million, against a board median of 140 million for its comparison class

The one evaluation K3 wins, and why it matters

Before the arithmetic, the exception, because it is real. Scored evaluation by evaluation at maximum effort on Artificial Analysis v4.3.2, GPT-6.1 Sol leads seven of ten and ties nothing, but the model that wins the long-context reasoning evaluation is K3:

• AA-LCR v1.1 — Kimi K3 0.887 vs GPT-6.1 Sol 0.830. A six-point lead on the evaluation built specifically for reasoning over long inputs.
• SciCode — Kimi K3 0.595 vs GPT-6.1 Sol 0.542. K3 also leads on scientific coding.
• Terminal-Bench 4.0 — Kimi K3 0.126 vs GPT-6.1 Sol 0.561. A 43-point gap in the other direction, and by far the largest split in the pairing.
• Humanity's Last Exam — Kimi K3 0.469 vs GPT-6.1 Sol 0.529.
• AA-Omniscience — Kimi K3 19.7 vs GPT-6.1 Sol 41.5; on factual reliability the two are not the same class of model.
• GDPval-AA v2.1 — Kimi K3 Elo 1536.5 vs GPT-6.1 Sol Elo 1575.1.
• MMMU-Pro — Kimi K3 0.805 vs GPT-6.1 Sol 0.860.
• AutomationBench-AA, GDP.pdf, CritPt — K3 0.583, 0.22 and 0.234; Sol 0.649, 0.310 and 0.317.

The AA-LCR result is the one worth taking seriously because it is the single number that contradicts the cost-per-task story, and it points at a workload where the cheap-looking answer is wrong in both directions. If your traffic is very long inputs, a modest generated answer, and heavy reuse of a stable prefix — the shape where verbosity never gets to bite — K3's combination of a top-tier long-context score, an image input and a downloadable checkpoint is a better fit than Sol's, and the cost-per-task figure for the index run does not apply to you. That is the honest version of "open weights are worth a premium": sometimes the open model is simply the better model for the job, and here it is on one axis.

It is also worth naming what those two wins have in common. K3 is strong where a task can be answered in one long act of reading or reasoning, and weak where it has to be executed in many small steps and checked. The 43-point Terminal-Bench gap and the 22-point omniscience gap are the same finding seen from two angles: sustained agentic execution is not what this checkpoint was optimised for.

The arithmetic, which is what actually decides it

K3's rate card is 1.5× Sol's on every line and its measured cost per completed index task is 2.8×, and the difference between 1.5 and 2.8 is entirely explained by token volume. The board's own measurement is 160 million output tokens for K3 against 67 million for Sol on the same ten-evaluation suite. K3 produces 2.4 times the output at 1.5 times the price, and 2.4 × 1.5 is 3.6 — the board's $2.00-against-$0.72 figure is a little friendlier than the pure ratio because input and cache contributions offset it slightly, but the direction is not in dispute.

Moonsh​oot's own marketing runs against this in a way that is worth reading carefully. The company claims K3 emits fewer output tokens than its predecessor, and the architectural work — the attention scheme, the residual design — is aimed squarely at long-context decoding cost. Both of those can be true while the model is still twice as verbose as a benchmark median on a general suite. The two claims are about different things: efficiency relative to a previous Kimi, and efficiency relative to the field. Only the second one is what you are being billed against, and it is the one the independent board measures.

Caching softens it. Both cached-input rates are a tenth of their model's uncached input rate — $0.30 for K3, $0.10 for Sol — so the cache line does not change the ordering, but it does mean that a request-heavy, answer-light workload narrows the gap to roughly the rate-card ratio rather than the measured task ratio. And Sol's long-context clause cuts the other way: above 272,000 input tokens it reprices the whole request at $4.00 and $15.00 with cache at $0.20, which is the one band where K3's flat card wins outright. That is worth knowing for two reasons, because it is also the band where K3's long-context score is the best number in this comparison.

A generated hero card for a GPT-6.1 Sol versus Kimi K3 comparison, headed 'The download exists, and it is still the dearer option', with two badges reading 'Kimi K3: $2.00 per index task' and 'GPT-6.1 Sol: $0.72 per index task', and the OrcaRouter logo in the bottom-right corner.

What open weights are actually worth here

Three things, and it is worth separating them because "open weights" is usually used as one argument when it is three.

The first is that you can leave. If Moonsh​oot is acquired, changes the license for future versions, deprecates K3 or raises its API price, a checkpoint you have already downloaded keeps running. That is not a hypothetical for this lab: Moonsh​oot suspended new subscriptions within about 48 hours of an earlier release to protect service quality for existing paying customers, which is a live demonstration that API access is a policy decision rather than a property. None of this shows up in a cost-per-task table and all of it is real.

The second is that you can pin. A fixed revision, fine-tuned, running inside a boundary you control, is a thing an auditor can be shown and an API cannot provide. For regulated traffic that is worth more than a rate card.

The third — the one usually assumed — is that it will be cheaper. It is not. The checkpoint is a 2.8-trillion-parameter FP8 artifact, and the published deployment guidance is framed around multi-accelerator configurations rather than a single node. A team already operating that class of cluster has a genuine option here and the calculus changes entirely, because the marginal cost of a token becomes electricity. A team that does not is comparing an API bill against a capital purchase, and the API bill wins by a wide margin. The honest summary is that open weights here give you the ability to leave, not the ability to run cheaply.

Both on one key, which is the part that makes the trade useful

Kimi K3 is on OrcaRouter at Moonsh​oot's own list price — $3.00 and $15.00 per million tokens with a $0.30 cache read, unchanged from the vendor's card — and GPT-6.1 Sol landed on our routes at Open​AI's price on release day, $2.00 and $10.00 with cached input at $0.10 and the 272,000-token step passed through exactly as listed. Nothing is added to either. The platform passes provider list price through rather than marking it up, so a Moonsh​oot rate change and an Open​AI rate change both appear on our side the same day they appear on the vendor's, without a second contract or a reseller in between.

A screenshot of the OrcaRouter model page for kimi/kimi-k3, showing the listing dated 2026-07-15, the model described as Moonshot AI's 2.8-trillion-parameter mixture-of-experts flagship for long-horizon coding and end-to-end knowledge work, a 1M-token context with text and image input, and the list price of $3.00 input and $15.00 output per million tokens with an 8.48 s median time to first token.

The reason that matters more for this pair than for most is that the two models belong to different failure modes. GPT-6.1 Sol is the model you run for sustained execution, tool loops and factual reliability; Kimi K3 is the model you run for very long inputs where you want the best long-context score and you want a checkpoint as a hedge against someone else's business decision. Those are not competing choices, they are two routes on the same traffic. Holding both behind one endpoint, with automatic failover between them and a rule that moves long-input work to K3 and execution work to Sol, is what converts "we have an open-weight fallback" from a line in a design document into something that works at three in the morning on a call that is failing.

A screenshot of the OrcaRouter model page for openai/gpt-6.1-sol, showing the listing dated 2026-09-29, a 1M-token context with 128K maximum output, text, image and file input, a reasoning effort ladder running from low to max, and the pass-through list price of $2.00 input and $10.00 output per million tokens with cached input at $0.10.

Where each one belongs

• Terminal agents, repo-scale coding, multi-step execution — GPT-6.1 Sol, on a 43-point Terminal-Bench 4.0 gap. This is not close.
• Answers where a hallucination is expensive — GPT-6.1 Sol, on an AA-Omniscience gap of 22 points.
• Very long inputs with a short generated answer — Kimi K3. Best long-context reasoning score in this comparison, image input, and verbosity that never gets a chance to apply.
• Self-hosting or an inference stack you must be able to freeze — Kimi K3, and only if you already operate the hardware. The checkpoint is the deliverable; the economics of running it are separate.
• A hedge against a vendor changing its terms — Kimi K3, and this is the one argument GPT-6.1 Sol cannot answer at any price.
• Choosing on rate card — neither K3 nor Sol is the cheap option in this comparison, and neither is the cheap option in the GPT-6 line. If the invoice is the deciding input, the model you want is further down the price list, not across this table.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily