
GPT-6 vs Kimi K3: Two Million-Token Models That Specialise Differently
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 65 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
GPT-6 Sol and Kimi K3 are the two models most likely to be shortlisted for the same job: a million-token context window, image input, and a price that makes long documents affordable. They got there from opposite directions. Kimi K3 is MoonshotAI's long-context specialist, released 15 July 2026, and its card is built around recall — it scores 88.7% on Artificial Analysis's long-context recall test, ahead of every model from the vendor we can compare it against. GPT-6 Sol is the vendor's mid-tier generalist, shipped 22 September 2026 at $2.00 per million input and $10.00 per million output, and its case is breadth: a slightly larger window, a higher general reasoning composite, and a cheaper output token.
So the honest framing is not better versus worse. It is recall versus reasoning, and it is decided by what you are doing with the million tokens once you have loaded them.
The windows are the same size on paper
Both cards say roughly a million tokens, and the difference is cosmetic. Kimi K3's window is 1,048,576 tokens — exactly 2^20, the binary million every long-context model quotes. GPT-6 Sol's is 1,050,000, a round decimal number about 1,400 tokens larger. For any workload you can actually fit, those are the same window. Do not choose on this.
What does differ is the rest of the envelope, and it is a short list.
• Context: Kimi K3 1,048,576 tokens, GPT-6 Sol 1,050,000 — effectively level
• Input modality: Kimi K3 accepts text and images; GPT-6 Sol accepts text, images and files
• Output ceiling: GPT-6 Sol 128,000 tokens in one response; Kimi K3's card does not publish a maximum output figure
• Standard price: Kimi K3 $3.00 in / $15.00 out per million, GPT-6 Sol $2.00 in / $10.00 out
• Cached input: Kimi K3 $0.30 per million, GPT-6 Sol $0.20 per million
• Long-context pricing: Kimi K3 lists a single rate with no documented tier; GPT-6 Sol reprices the whole request above 272,000 input tokens to $4.00 in / $15.00 out
• Reasoning controls: GPT-6 Sol exposes a configurable reasoning.effort; Kimi K3 exposes reasoning and reasoning-effort parameters through the same OpenAI-compatible surface
The missing output ceiling on Kimi K3 is worth flagging rather than smoothing over: MoonshotAI publishes a context figure and no maximum-output figure, so a system that depends on a hard output cap for budgeting cannot read one off the card. The long-context tier is the other asymmetry, and it cuts in an unintuitive direction — GPT-6 Sol's headline rate is lower, but its whole-request repricing above 272K means a 300,000-token call bills at $4.00 / $15.00 from the first token, while Kimi K3's single $3.00 / $15.00 rate is not documented as stepping at all. For a very long document, Kimi K3 can end up the cheaper of the two on input despite the higher headline.
Recall is Kimi K3's actual product
The reason to reach for Kimi K3 is not that it has a big window — several models do. It is that it retains what is in it. On Artificial Analysis's long-context recall evaluation, carried in the model catalogue, Kimi K3 scores 88.7% against GPT-6 Sol's 83.7%. That is a five-point gap on the exact test that measures whether a model can find and use a detail buried in a long input, and it is the kind of gap that shows up as real failures in production — the summariser that drops the clause on page 400, the contract reviewer that misses the indemnity section.
Kimi K3 also leads on coding, and by a wider margin than the composite index suggests. On the coding index it sits at 76.2 with a rank inside the top ten of the models evaluated, and it posts 85.0% on terminal-bench v2.1 — the agentic command-line benchmark that a coding agent's usefulness depends on. Its GPQA Diamond score, 93.5%, is in the top band of the board.
Where GPT-6 Sol answers back is on the composites and on scientific code. Its general intelligence index, 47.6, is four points clear of Kimi K3's 43.6 and near the top decile; the two are level on SciCode at 57.6% and 59.5% respectively, inside the noise of a single run; and on Humanity's Last Exam GPT-6 Sol's 47.9% edges Kimi K3's 46.9%. None of those single-benchmark differences is large enough to choose on by itself.

Cost on a workload, not on a rate card
Rate cards mislead because the ratio of input to output tokens is different for every job, and these two models disagree about which side of the trade is cheaper. Kimi K3 is cheaper on the assumption that your work is mostly input — long documents in, short answers out. GPT-6 Sol is cheaper on the assumption that your work is balanced or output-heavy, because its output rate is a third lower.
• Read-one-book-summarise-it shape (900K input, 4K output): Kimi K3 ≈ $2.76, GPT-6 Sol ≈ $3.64 before the 272K repricing applies; with repricing, GPT-6 Sol's whole call moves to the higher tier and the gap widens
• Balanced agentic shape (60K input, 20K output): Kimi K3 ≈ $0.48, GPT-6 Sol ≈ $0.32
• Code-generation shape (30K input, 60K output): Kimi K3 ≈ $0.99, GPT-6 Sol ≈ $0.66
Those are raw token arithmetic on each vendor's published rates and exclude cached input, which helps whichever model you cache against most. The shape of the answer is what matters: for pure long-input recall work Kimi K3's flat rate wins, and for anything that writes a lot back GPT-6 Sol's lower output rate wins. Neither is universally cheaper, and a comparison that names one winner on price without stating the token ratio is not measuring anything.
Provenance is part of the decision
Kimi K3 is built by MoonshotAI, a Chinese lab, and GPT-6 Sol by OpenAI. That difference is not a benchmark and it does not appear on a rate card, but for regulated workloads it decides the shortlist before price is discussed. MoonshotAI's card in the catalogue does not publish a licence field, so nothing here should be read as a claim about weights availability either way — if open weights matter to your deployment, verify it at the source rather than assuming from the family name.
For teams that need both — a Chinese-lab model for one part of a pipeline and an OpenAI model for another, with the routing, billing and residency story handled in one place — that is the reason a single gateway is worth having rather than two sets of credentials. On OrcaRouter both kimi/kimi-k3 and GPT-6 Sol are live routes in the same catalogue, at the provider's list price with 0% markup, so a MoonshotAI or OpenAI rate change is live the same day. Automatic failover means a degraded route on either provider does not take the pipeline down, and the routing DSL lets you send recall-heavy work to Kimi K3 and generation-heavy work to GPT-6 Sol by rule rather than by guess.

Which one to put in the pipeline
Choose Kimi K3 when the job is retrieval and retention over very long inputs — legal and medical document review, full-repository code understanding, transcript analysis across hundreds of documents — and when your prompts are input-heavy enough that the flat $3.00 rate beats GPT-6 Sol's repricing. Its recall lead and its top-ten coding rank are the two numbers that justify the choice.
Choose GPT-6 Sol when the job is a general assistant backed by a million-token window: mixed reasoning, serialisation to JSON, agentic loops that write a lot back, and any workflow where a 128,000-token output ceiling is a feature rather than a constraint. It is cheaper on output, it exposes explicit reasoning-effort control, and its composite index is the stronger general signal. Where a pipeline genuinely does both, the honest answer is to route rather than to pick — which is a statement about the shape of your traffic, not about either model.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
