
GPT-6 Luna vs Kimi K3: Four Times the Throughput, One Thirtieth the Price, One Real Exception
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Two models, both released within the last three months, both positioned as the cheap tier of a frontier family, and both wrong about what "cheap tier" means in the same way — which is not at all. Kimi K3 is Moonshot AI's open-weight flagship, public under a Modified MIT licence since July 27, 2026, priced at $3.00 per million input tokens and $15.00 per million output with cache reads at $0.30. GPT-6 Luna is the vendor's volume tier, generally available since September 22, 2026 at $0.10 and $0.50 with cache reads at a cent. Thirty times on input, thirty times on output, and a licence on one side that the other cannot offer at any price.
The rate card is not what decides this pairing, though, because it is not close enough to need deciding. What decides it is a number neither vendor puts in a headline: measured output speed. Kimi K3 generates 38.7 tokens per second. GPT-6 Luna generates 153.9. That is a fourfold gap, and for any workload where the model is a component inside a loop rather than an endpoint a person talks to, a fourfold throughput difference changes the architecture rather than the bill.
What these two actually are
Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model with 104 billion parameters active per token, a 1,048,576-token context window, a 1,048,576-token output ceiling, and weights published on Hugging Face under a Modified MIT licence. It is the highest-scoring open-weight model on the independent board — first among 112 open-weight models — and it holds the top position on Code Arena's WebDev leaderboard at 1,679 Elo. Moonshot reports 88.3% on Terminal-Bench 2.0, which is a vendor figure and not independently reproduced.
GPT-6 Luna is the small tier of OpenAI's GPT-6 generation, sitting below GPT-6 Sol at $2.00/$10.00 and well below the flagship GPT-6 Astra at $10.00/$50.00. Text and image in, text out, a 1,050,000-token window with 922,000 maximum input and a 128,000-token output ceiling, and reasoning effort selectable from none through low, medium, high, xhigh and max with medium as the default. It is closed, single-identifier, and available to anyone with an API key.
One is a download. One is an endpoint. Both are priced for volume. That is the shape of the comparison.
The cards, line by line
One line per dimension, both sides on each:
• Input per 1M — GPT-6 Luna $0.10 vs Kimi K3 $3.00
• Output per 1M — GPT-6 Luna $0.50 vs Kimi K3 $15.00
• Cached input per 1M — GPT-6 Luna $0.01 vs Kimi K3 $0.30, a tenth of its own input rate on both cards
• Long-context clause — GPT-6 Luna doubles input and cache and lifts output 1.5× above 272,000 input tokens, applied to the whole request vs Kimi K3 no equivalent clause published on its card
• Context window — GPT-6 Luna 1,050,000 tokens with 922,000 maximum input vs Kimi K3 1,048,576 tokens
• Maximum output — GPT-6 Luna 128,000 tokens vs Kimi K3 1,048,576 tokens, eight times the ceiling
• Intelligence Index, independent — GPT-6 Luna 37 at max effort vs Kimi K3 44 at max effort, same index revision
• Default-effort score — GPT-6 Luna 29 at its default medium vs Kimi K3 measured at its max setting
• Cost per Index task, independent — GPT-6 Luna about $0.07 vs Kimi K3 $2.00
• Output tokens on the index run — GPT-6 Luna 150M vs Kimi K3 160M, against a board median of 88M; both are far more verbose than the board's median model
• Output speed, independent — GPT-6 Luna 153.9 tokens/sec vs Kimi K3 38.7 tokens/sec
• Weights — GPT-6 Luna closed, API only vs Kimi K3 public on Hugging Face since July 27, 2026
• Licence — GPT-6 Luna not applicable vs Kimi K3 Modified MIT, which is not the same instrument as MIT
• Total parameters — GPT-6 Luna undisclosed vs Kimi K3 2,800 billion, of which 104 billion are active per token
• Standout evidence — GPT-6 Luna tool calling and agentic loops at a tenth of a cent per thousand cached input tokens vs Kimi K3 Code Arena WebDev #1 at 1,679 Elo, Terminal-Bench 2.0 88.3% vendor-reported, top open-weight score on the independent board
• On OrcaRouter — GPT-6 Luna not in our catalogue vs Kimi K3 routable at Moonshot's list price

Throughput is the spec nobody headlines
A 38.7-token-per-second model and a 153.9-token-per-second model are not the same product with different price tags. They are different products.
Take a task where the model has to produce a 1,500-token answer — a summary, a structured extraction, a code patch with its explanation. At 153.9 tokens per second that answer takes about ten seconds. At 38.7 it takes about thirty-nine. Neither figure includes reasoning time or network latency, so the real-world gap is wider than four times in proportion, not narrower.
Now put it in a loop. An agent that makes thirty model calls to complete one task — plan, select a tool, read the result, decide the next step, repeat — produces perhaps 15,000 output tokens across those calls. GPT-6 Luna spends roughly 97 seconds generating. Kimi K3 spends roughly 388. That is the difference between a task a user waits for and a task a user schedules, and no amount of licence permissiveness changes it.
The verbosity compounds the speed problem rather than offsetting it. Kimi K3 generated 160 million output tokens completing the independent index against GPT-6 Luna's 150 million — so on that evaluation the slower model was also producing slightly more tokens, not fewer. The two models are similarly verbose; they are simply not similarly fast, and on an output-priced card being verbose is expensive twice.
It is worth saying plainly that this is not a defect in Kimi K3. A 2.8-trillion-parameter model with 104 billion active is doing more work per token than a small tier model, and the throughput figure is a measurement of that work. The relevant question is whether your workload needs the work. If it does, 38.7 tokens per second is the price of the capability. If it does not, you are paying for deliberation you did not ask for, thirty times over.
Where K3's seven points live
Kimi K3 scores 44 on the independent Intelligence Index at maximum effort. GPT-6 Luna scores 37 at the same setting. Seven points, on a composite where one point is inside the noise of a re-run — and the two are separated by roughly twenty-eight times on cost per completed task, $2.00 against about $0.07.
Those seven points are not evenly distributed. Kimi K3's reputation is concentrated in coding and in agentic software work, and it is the reason the model exists: Code Arena's WebDev leaderboard has it first at 1,679 Elo, and the vendor's Terminal-Bench 2.0 figure of 88.3% is a coding-agent measurement. GPT-6 Luna's own coding position is a step backwards rather than forwards — its Coding Agent Index sits at 41, two points below the GPT-5.6 Luna it replaced, with the declines concentrated in SWE-Atlas-QnA and DeepSWE v1.1. If your work is an agent writing software against a real repository, that is a seven-point index gap and a two-point generational regression pointing the same direction, and the case for the cheap tier gets considerably weaker.
Where the seven points are not is the class of work GPT-6 Luna is priced for. Classification, extraction, routing, bulk summarisation, the inner loop of an agent that needs a fast yes-or-no — on those tasks the two models will produce the same usable output the overwhelming majority of the time, and the difference will not survive contact with your own eval set. The index measures a composite that weights long-horizon reasoning and coding heavily, and neither of those is what a routing decision requires.
One caveat worth attaching to the index figures on both sides: this board is a rolling evaluation and its revision matters. Earlier snapshots of the same leaderboard placed Kimi K3 materially higher than the 44 our capture shows today, because the composite, the harness and the model set all move. Any comparison that leans on the precise gap between two models on a rolling index is leaning on something that will not hold still. The honest reading is that these two are in adjacent capability bands at their ceilings, and that the index cannot tell you which is better for your task.
The exception: what Modified MIT actually gets you
Kimi K3's licence is the one dimension on which GPT-6 Luna has no answer at any price, and it deserves more precision than the phrase "open weights" usually gets.
The weights are public on Hugging Face and the licence is Modified MIT, not MIT. That distinction matters: a Modified MIT licence typically attaches conditions — commonly a requirement to display an attribution when the model is used in a product above some scale — and anyone planning to build a commercial product on the weights should read the actual instrument rather than the family name it resembles. This is not a reason to avoid it. It is a reason not to describe it as MIT in a procurement document.
What the download buys is real and bounded. It buys a floor on cost, in the sense that a fixed GPU footprint can beat a metered bill at sufficient volume. It buys independence from a vendor's roadmap — a model you host cannot be deprecated out from under you. It buys the ability to fine-tune on your own distribution. And it buys the option to keep prompts inside your own boundary, which for some teams ends the comparison before price is mentioned.
What it costs is the hardware. A 2.8-trillion-parameter mixture-of-experts model with 104 billion active parameters is not a model that runs on a workstation. Serving it at production throughput is a cluster-scale commitment with a capital cost, an operational cost and a latency profile that you own rather than rent. That is not an argument against self-hosting Kimi K3 — it is an argument that the correct comparison is not $15.00 per million output tokens against $0.00, but $15.00 per million output tokens against a fully-loaded cost per token that includes the silicon and the people. For a small number of organisations that number is lower. For most it is not, and the crossover point is further away than the licence makes it feel.
Running one of them, or both
Kimi K3 is on OrcaRouter at Moonshot's list price, under the pass-through pricing that means a vendor rate change is live on our side the same day. Automatic failover is the part that earns its place for this model in particular: a self-hosted or third-party Kimi K3 endpoint degrading is exactly the scenario where you want the route to move without a deployment, and a 38.7-token-per-second model has less headroom to absorb a slow upstream than a fast one does.
GPT-6 Luna is not in our catalogue as of this writing and is reached through OpenAI's own API, so a team that wants both runs one key for Kimi K3 alongside whatever it already uses for OpenAI. This is also the pairing where the routing DSL does something a hard-coded split cannot: a rule that sends coding and long-horizon work to Kimi K3 and everything program-consumed and short to GPT-6 Luna expresses a boundary that moves every time either vendor ships, and it moves as configuration rather than as a code path. Model fusion is the other configuration worth naming here — a panel of one slow careful model and one fast cheap model answering together, with the cheap member doing the volume and the expensive one adjudicating the cases that matter, is a shape this pair of models fits unusually well, because the cost asymmetry between them is large enough that spending a Kimi K3 call on a small fraction of traffic is affordable.

Which one, and when
Take GPT-6 Luna if your workload is high-volume, program-consumed and latency-sensitive. Classification, extraction, routing, bulk summarisation, the inner loop of an agent. At $0.10/$0.50 with a one-cent cached read and four times Kimi K3's throughput, the arithmetic is not close and the latency difference is not marginal. Set the effort parameter explicitly, because the default is medium and the headline 37 is a maximum-effort number, and keep long calls on the right side of the 272,000-token line.
Take Kimi K3 if you need the weights, if your work is agentic coding against a real repository where its WebDev position and vendor-reported 88.3% on Terminal-Bench 2.0 are the evidence that matters, if you need an output ceiling above 128,000 tokens, or if your organisation cannot send prompts to a closed endpoint. Accept 38.7 tokens per second as part of the deal, and read the Modified MIT terms rather than assuming them.
The mistake this comparison invites is treating the thirtyfold rate as the answer to a question about capability. It is the answer to a question about tokens. The capability question is settled by whether you need the weights or the speed, and those two answers point in opposite directions.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
