
Grok 4.7 vs Kimi K3: Half the Price, Half the Context, None of the Weights
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 219 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 117 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 40 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Take a 300-million-token month of agentic coding work and run it through both models at their list rates, and the choice stops looking like a benchmark argument. Grok 4.7, which SpaceXAI released on September 21, 2026, charges $2 per million input tokens and $6 per million output. Kimi K3, which Moonshot AI released on July 16, 2026, charges $3 and $15. On that hypothetical month — say 200M input and 100M output — Grok 4.7 comes to roughly $1,000 and Kimi K3 to roughly $2,100. The gap is not a rounding difference; it is a second model's worth of budget. Against that, Kimi K3 gives you a 1,000,000-token context window instead of 500,000, native video input as well as images, and downloadable weights. Grok 4.7 gives you a faster, cheaper, closed model that was explicitly trained for the long-horizon terminal work Kimi K3 also targets. Both are frontier-class. They are not the same purchase.
The two models, briefly
Grok 4.7 is SpaceXAI's frontier coding and knowledge-work model, built on a larger base model than Grok 4.6 and given a longer reinforcement-learning run aimed at tasks that can take hours. It is proprietary — no weights, no download — serves a 500,000-token context window, takes text and images in and returns text, and supports reasoning-effort levels from low through xhigh. Its headline claim is speed and price: SpaceXAI positions it as roughly twice as fast as comparable models at roughly half the price.
Kimi K3 is Moonshot AI's flagship open mixture-of-experts model, the first open model in the three-trillion-parameter class: 2.8 trillion parameters total with 104 billion active per token, built on Kimi Delta Attention and quantisation-aware MXFP4 weights. It takes text, images and video in, serves a 1,000,000-token context, runs reasoning always-on with configurable effort, and its weights are downloadable under the Kimi K3 License — commercial use permitted, with restrictions. It is, by design, the model you can take home with you.
That is the entire structural difference between them. One is a cheap, fast, closed endpoint. The other is a more expensive, slower, open artifact with twice the context. Everything below is a consequence of that.
What the benchmarks actually compare
This is where most Grok 4.7 versus Kimi K3 write-ups go wrong, so it is worth being blunt: the two models' published numbers largely do not measure the same things, and at least one pair that looks comparable is not.
Start with the pair that is genuinely misleading. Kimi K3's launch material cites a Terminal-Bench 2.1 score of 88.3. Grok 4.7's launch material cites Terminal-Bench 4.0 at 38.0%. Those are different benchmark versions with different task sets and different scoring, and putting them side by side would suggest Kimi K3 is more than twice the agentic model. It is not a defensible reading, and no one should repeat it. Both numbers are also vendor-reported — neither has been independently reproduced.
What can be compared is the independent composite, and there the two are close. On the current Artificial Analysis Intelligence Index, version v4.3.2, Kimi K3 scores 44 — second of 113 models, the highest placement of any open-weights model on the board. Grok 4.7 scores 46, sixteenth of 655. Two points apart, with Kimi K3's placement earned at maximum reasoning effort. On the blended composite, these models are peers.
• Independent composite — Grok 4.7 46 on AA Intelligence Index v4.3.2 vs Kimi K3 44 on the same version. A tie in practice.
• Context window — Grok 4.7 500,000 tokens vs Kimi K3 1,000,000 tokens. A real, unqualified advantage for Kimi K3, and the one spec difference with no caveat attached.
• Input modalities — Grok 4.7 text and image vs Kimi K3 text, image and video. Kimi K3 is broader, if your workload has video in it.
• Weights — Grok 4.7 proprietary, no download vs Kimi K3 open weights under the Kimi K3 License. For an on-premise or air-gapped requirement this is the only line that matters.
• Price per million tokens — Grok 4.7 $2 in / $6 out vs Kimi K3 $3 in / $15 out, with Kimi's cached-input rate at $0.30 per million. Grok 4.7 is cheaper on every axis, and dramatically cheaper on output.
• Reasoning effort control — Grok 4.7 low through xhigh vs Kimi K3 always-on with configurable effort. Functionally similar; the difference is that Grok 4.7 lets you turn reasoning down.
• Vendor-reported agentic evidence — Grok 4.7 Terminal-Bench 4.0 38.0%, CursorBench 4.0 46.3%, DeepSWE v1.1 71.0% (all vendor-reported) vs Kimi K3 Terminal-Bench 2.1 88.3%, Frontend Code Arena first place at 1,679 Elo (vendor-reported at launch). Not comparable — see above.


Where the money actually goes
The rate card undersells the gap, because the two models consume tokens differently. Grok 4.7 is a verbose reasoner — Artificial Analysis measured 240 million output tokens generated while running its Intelligence Index, against a median of 92 million across models. Kimi K3 is the opposite kind of expensive: it is a large always-on reasoning model, and at $15 per million output tokens, verbosity is what you are buying.
So the honest cost model is not "$2/$6 versus $3/$15." It is output tokens per completed task, multiplied by the output rate, plus input tokens including every retry, multiplied by the input rate. A model that solves a task in one long pass at $6 per million can beat a model that needs three attempts at $15 per million, and it can also lose to it if the cheap model's passes are long enough. That is a measurement you run on your own repository, not one anybody publishes for you.
One asymmetry is worth knowing before you run it: Grok 4.6, Grok 4.7's predecessor, applies a pricing cliff at 200,000 input tokens, where a request that crosses the line reprices the entire request at double the rate. Long-context work on the Grok side needs that checked against the current rate card for the model you deploy, because a 500,000-token window and a 200,000-token cliff interact badly. Kimi K3's pricing is flat, which is the quieter advantage of the two.
Where each one breaks
Both models have a failure mode that the launch material does not lead with, and they are different failures.
Kimi K3's is reliability under sustained load. Community and independent testing at launch reported that its agentic performance degrades as runs extend — strong single-shot results, weaker sustained pass rates — and that its output speed sits around 37 to 38 tokens per second, which is slow for interactive coding. Moonshot also had to pause new consumer subscriptions shortly after launch because demand outran capacity, which tells you something about serving headroom on the hosted side.
Grok 4.7's is the opposite shape: it is fast, cheap, and closed, which means you cannot inspect it, cannot run it yourself, and cannot hedge against a deprecation. SpaceXAI has already publicly sequenced Grok 4.8, Grok 4.9 and Grok 5 behind this release, and Grok 4.6 lasted roughly five weeks before being superseded. Anything you build on Grok 4.7 should be built so that the model ID is a configuration value, not an assumption.
There is a third consideration that is not a failure of either model: neither one is the right answer for a two-week autonomous run, and both vendors have demoed exactly that. The published 16-day and 48-hour autonomous coding demos are vendor-run, on vendor-chosen tasks, without independent reproduction. They are worth knowing about and not worth planning against.

Running both, and why that is the real answer
The cost gap and the context gap point in opposite directions, which is the argument for not picking one. Kimi K3 earns its price on repository-scale work where a 1M window means you do not chunk the input, and on anything that has to be self-hosted. Grok 4.7 earns its price on volume — high-turn agent loops, tool-call-heavy pipelines, and long terminal sessions where speed and output cost dominate.
Kimi K3 is on OrcaRouter, which makes that split a routing decision rather than a procurement one. It sits in the catalog at Moonshot's list price, and because OrcaRouter passes provider list pricing through with 0% markup, a vendor price cut is live here the same day it is announced rather than at the next contract renewal. For workloads where a long-context pass and an execution pass should collaborate rather than compete — one model holding the whole repository in view, another acting on it — the routing DSL composes several models into a single call, and model fusion lets a panel of models answer together when the task is worth the extra spend. One API key covers Kimi K3 plus 200-plus other models, so the execution half of that split can come from whichever model is cheapest and fastest for the job rather than from the one that happens to hold the context.
There is one thing to be clear about: Kimi K3's open weights are downloadable, but the hosted endpoint is what OrcaRouter routes to. If your requirement is genuinely on-premise inference, that is a self-hosting project on your own hardware, not a routing decision — and it is the one scenario where this comparison has a single correct answer.
The verdict, and the one number to hold
If you are choosing on cost and speed, Grok 4.7 wins and it is not close: half the input price, under half the output price, faster serving, and a reasoning dial you can turn down. If you are choosing on context, modality breadth, or the ability to own the model, Kimi K3 wins on the same terms. If you are choosing on capability alone, the independent composite says they are two points apart, and you should decide on your own evaluation rather than on either vendor's launch chart.
The number to hold onto is the one from the top: $2/$6 against $3/$15. Everything else in this comparison is negotiable, and that ratio is not.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
