
GPT-6.1 Sol vs Kimi K3: The Download Exists, and It Is Still Not the Cheap Option
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 161 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 79 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 53 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 301 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Kimi K3 is the counter-example that usually settles this kind of argument. Moonshoot AI shipped the 2.8-trillion-parameter model in July 2026 and published the weights — moonshotai/Kimi-K3 is on Hugging Face, ungated, with more than a million downloads and a last revision in September. So there is no "open in principle, withheld for safety hardening" asterisk here, no two-week delay to wait out. GPT-6.1 Sol, which OpenAI released on 29 September 2026, is closed and API-only and always will be. And on the independent board the downloadable model scores 43.6 to the closed model's 51.8 while costing $2.00 per completed index task against $0.72. Open weights are supposed to be the cheap escape hatch; in this pairing they are the expensive side of the trade, and the reason is not the license.
What each model is
Kimi K3 is a sparse mixture-of-experts model with 2,800 billion total parameters and roughly 104 billion — about 3.7% — active per token. It carries a 1,048,576-token context window, takes text and images as input, and returns text. The published checkpoint is an FP8 artifact and it is a genuinely large download; a production deployment on your own hardware is a cluster decision rather than a workstation one. Moonshoot's architectural claim is a hybrid linear-attention scheme it says is substantially faster at long-context decoding, plus the residual and routing work that made a 2.8-trillion-parameter model servable at all.
GPT-6.1 Sol is OpenAI's mid-tier refresh — below GPT-6 Astra, above GPT-6 Luna — with a 1,050,000-token window, a 128,000-token output ceiling, text, image and file input, and an April 30, 2026 knowledge cutoff. Reasoning runs low through max with medium as the default; the none rung that the previous GPT-6 Sol accepted is rejected outright, and OpenAI's own migration note points developers at low instead. It is priced at $2.00 and $10.00 per million tokens with cached input at $0.10.
One line per dimension, both sides on each:
• Input — GPT-6.1 Sol $2.00 per 1M vs Kimi K3 $3.00 per 1M
• Output — GPT-6.1 Sol $10.00 per 1M vs Kimi K3 $15.00 per 1M
• Cached input — GPT-6.1 Sol $0.10 per 1M vs Kimi K3 $0.30 per 1M
• Context window — 1,050,000 tokens vs 1,048,576 tokens; a 0.1% difference, immaterial
• Maximum output — 128,000 tokens on both
• Total and active parameters — undisclosed vs 2,800 billion total, 104 billion active per token
• Modality — text, image and file in, text out vs text and image in, text out
• Weights — closed, API only vs open, ungated, 1M-plus downloads
• Long-context clause — reprices the whole request above 272,000 input tokens at 2× input and cache and 1.5× output vs no equivalent clause on the published card
• Intelligence Index, independent, max effort — 51.8 vs 43.6
• Cost per index task, independent — $0.72 vs $2.00
• Output tokens on the index run — 67 million vs 160 million, against a board median of 140 million for its comparison class
The one evaluation K3 wins, and why it matters
Before the arithmetic, the exception, because it is real. Scored evaluation by evaluation at maximum effort on Artificial Analysis v4.3.2, GPT-6.1 Sol leads seven of ten and ties nothing, but the model that wins the long-context reasoning evaluation is K3:
• AA-LCR v1.1 — Kimi K3 0.887 vs GPT-6.1 Sol 0.830. A six-point lead on the evaluation built specifically for reasoning over long inputs.
• SciCode — Kimi K3 0.595 vs GPT-6.1 Sol 0.542. K3 also leads on scientific coding.
• Terminal-Bench 4.0 — Kimi K3 0.126 vs GPT-6.1 Sol 0.561. A 43-point gap in the other direction, and by far the largest split in the pairing.
• Humanity's Last Exam — Kimi K3 0.469 vs GPT-6.1 Sol 0.529.
• AA-Omniscience — Kimi K3 19.7 vs GPT-6.1 Sol 41.5; on factual reliability the two are not the same class of model.
• GDPval-AA v2.1 — Kimi K3 Elo 1536.5 vs GPT-6.1 Sol Elo 1575.1.
• MMMU-Pro — Kimi K3 0.805 vs GPT-6.1 Sol 0.860.
• AutomationBench-AA, GDP.pdf, CritPt — K3 0.583, 0.22 and 0.234; Sol 0.649, 0.310 and 0.317.
The AA-LCR result is the one worth taking seriously because it is the single number that contradicts the cost-per-task story, and it points at a workload where the cheap-looking answer is wrong in both directions. If your traffic is very long inputs, a modest generated answer, and heavy reuse of a stable prefix — the shape where verbosity never gets to bite — K3's combination of a top-tier long-context score, an image input and a downloadable checkpoint is a better fit than Sol's, and the cost-per-task figure for the index run does not apply to you. That is the honest version of "open weights are worth a premium": sometimes the open model is simply the better model for the job, and here it is on one axis.
It is also worth naming what those two wins have in common. K3 is strong where a task can be answered in one long act of reading or reasoning, and weak where it has to be executed in many small steps and checked. The 43-point Terminal-Bench gap and the 22-point omniscience gap are the same finding seen from two angles: sustained agentic execution is not what this checkpoint was optimised for.
The arithmetic, which is what actually decides it
K3's rate card is 1.5× Sol's on every line and its measured cost per completed index task is 2.8×, and the difference between 1.5 and 2.8 is entirely explained by token volume. The board's own measurement is 160 million output tokens for K3 against 67 million for Sol on the same ten-evaluation suite. K3 produces 2.4 times the output at 1.5 times the price, and 2.4 × 1.5 is 3.6 — the board's $2.00-against-$0.72 figure is a little friendlier than the pure ratio because input and cache contributions offset it slightly, but the direction is not in dispute.
Moonshoot's own marketing runs against this in a way that is worth reading carefully. The company claims K3 emits fewer output tokens than its predecessor, and the architectural work — the attention scheme, the residual design — is aimed squarely at long-context decoding cost. Both of those can be true while the model is still twice as verbose as a benchmark median on a general suite. The two claims are about different things: efficiency relative to a previous Kimi, and efficiency relative to the field. Only the second one is what you are being billed against, and it is the one the independent board measures.
Caching softens it. Both cached-input rates are a tenth of their model's uncached input rate — $0.30 for K3, $0.10 for Sol — so the cache line does not change the ordering, but it does mean that a request-heavy, answer-light workload narrows the gap to roughly the rate-card ratio rather than the measured task ratio. And Sol's long-context clause cuts the other way: above 272,000 input tokens it reprices the whole request at $4.00 and $15.00 with cache at $0.20, which is the one band where K3's flat card wins outright. That is worth knowing for two reasons, because it is also the band where K3's long-context score is the best number in this comparison.

What open weights are actually worth here
Three things, and it is worth separating them because "open weights" is usually used as one argument when it is three.
The first is that you can leave. If Moonshoot is acquired, changes the license for future versions, deprecates K3 or raises its API price, a checkpoint you have already downloaded keeps running. That is not a hypothetical for this lab: Moonshoot suspended new subscriptions within about 48 hours of an earlier release to protect service quality for existing paying customers, which is a live demonstration that API access is a policy decision rather than a property. None of this shows up in a cost-per-task table and all of it is real.
The second is that you can pin. A fixed revision, fine-tuned, running inside a boundary you control, is a thing an auditor can be shown and an API cannot provide. For regulated traffic that is worth more than a rate card.
The third — the one usually assumed — is that it will be cheaper. It is not. The checkpoint is a 2.8-trillion-parameter FP8 artifact, and the published deployment guidance is framed around multi-accelerator configurations rather than a single node. A team already operating that class of cluster has a genuine option here and the calculus changes entirely, because the marginal cost of a token becomes electricity. A team that does not is comparing an API bill against a capital purchase, and the API bill wins by a wide margin. The honest summary is that open weights here give you the ability to leave, not the ability to run cheaply.
Both on one key, which is the part that makes the trade useful
Kimi K3 is on OrcaRouter at Moonshoot's own list price — $3.00 and $15.00 per million tokens with a $0.30 cache read, unchanged from the vendor's card — and GPT-6.1 Sol landed on our routes at OpenAI's price on release day, $2.00 and $10.00 with cached input at $0.10 and the 272,000-token step passed through exactly as listed. Nothing is added to either. The platform passes provider list price through rather than marking it up, so a Moonshoot rate change and an OpenAI rate change both appear on our side the same day they appear on the vendor's, without a second contract or a reseller in between.

The reason that matters more for this pair than for most is that the two models belong to different failure modes. GPT-6.1 Sol is the model you run for sustained execution, tool loops and factual reliability; Kimi K3 is the model you run for very long inputs where you want the best long-context score and you want a checkpoint as a hedge against someone else's business decision. Those are not competing choices, they are two routes on the same traffic. Holding both behind one endpoint, with automatic failover between them and a rule that moves long-input work to K3 and execution work to Sol, is what converts "we have an open-weight fallback" from a line in a design document into something that works at three in the morning on a call that is failing.

Where each one belongs
• Terminal agents, repo-scale coding, multi-step execution — GPT-6.1 Sol, on a 43-point Terminal-Bench 4.0 gap. This is not close.
• Answers where a hallucination is expensive — GPT-6.1 Sol, on an AA-Omniscience gap of 22 points.
• Very long inputs with a short generated answer — Kimi K3. Best long-context reasoning score in this comparison, image input, and verbosity that never gets a chance to apply.
• Self-hosting or an inference stack you must be able to freeze — Kimi K3, and only if you already operate the hardware. The checkpoint is the deliverable; the economics of running it are separate.
• A hedge against a vendor changing its terms — Kimi K3, and this is the one argument GPT-6.1 Sol cannot answer at any price.
• Choosing on rate card — neither K3 nor Sol is the cheap option in this comparison, and neither is the cheap option in the GPT-6 line. If the invoice is the deciding input, the model you want is further down the price list, not across this table.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
