A generated hero card titled 'DeepSeek V4 Pro vs Kimi K3' with a speed gauge icon on the left labeled 'DeepSeek V4 Pro' badged '~81 tok/s · /bin/bash.66/.98' and a brain icon on the right labeled 'Kimi K3' badged 'Index 44 · .00/5.00', with a subtle racing-track divider between them and the footer 'Speed and price against a reasoning ceiling'.
Guides & Insights

DeepSeek V4 Pro vs Kimi K3: Speed and Price Against a 44-Point Reasoning Ceiling

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

DeepSeek V4 Pro and Kimi K3 are the two open-weights reasoning flagships with a million-token window, and their competition settled into a stable shape a month ago: Kimi K3 (Moonshot AI's 2.8T-parameter model, released July 16, 2026) is the intelligence leader, and DeepSeek V4 Pro (released August 13, 2026) is the speed and price leader. What is new is that the cheap half of that pairing was nearly deleted in the first half of September — Deep​Seek put V4 Pro on a September 14 retirement path on the 8th, delayed it, and then on the 11th cancelled the retirement outright, committing to keep serving the model at unchanged billing. That matters more for this matchup than for most, because "cheap, fast, 1M context, open weights" is a combination that is hard to replace when the alternative costs five times as much.

Independent figures are Artificial Analysis measurements. Deep​Seek's and Moonshot's own launch benchmarks are labeled vendor-reported, because neither has been independently reproduced.

Kimi K3 wins the index; V4 Pro wins the bill

The independent scoreboard is unambiguous about who is smarter. Kimi K3 at maximum effort scores 44 on the Artificial Analysis Intelligence Index, ranked #2 among comparable models. DeepSeek V4 Pro scores 36, ranked #7 in its class. That nine-point gap is the entire reason anyone pays Kimi K3's prices. And the prices are the reason the matchup is not settled by the index alone: Kimi K3 lists at $3.00 per million input and $15.00 per million output tokens, while V4 Pro lists at $0.66/$1.98 off-peak ($1.32/$3.96 peak). On output, that is roughly 7.5x — before the cache discount. V4 Pro's cache reads cost about $0.02 per million tokens off-peak, a 97% discount; Kimi K3's 90% cache discount starts from a much higher base. For a token-heavy workload, the effective gap widens toward ten times.

• Intelligence Index — Kimi K3 44 (#2/113) vs V4 Pro 36 (#7/113)

• Price — Kimi K3 $3.00/$15.00 per 1M vs V4 Pro $0.66/$1.98 off-peak ($1.32/$3.96 peak)

• Speed — V4 Pro ~81 tokens/sec vs Kimi K3 ~35 tokens/sec

• Context — both 1M tokens

• Parameters — Kimi K3 2.8T total / 104B active vs V4 Pro 1.6T total / 49B active

• Weights — both open (Kimi K3 open weights, V4 Pro MIT)

• Input — Kimi K3 text and image vs V4 Pro text-only

Speed is the part everyone underestimates

The gap nobody talks about is throughput. V4 Pro outputs about 81 tokens per second measured independently; Kimi K3 manages about 35. In an agent loop where the model writes several thousand tokens per step, that is the difference between a task that finishes in minutes and one that drags across an afternoon — and the cost difference compounds it, because you are paying $15 per million output tokens at roughly half the speed. For interactive coding assistants, bulk document processing, and any workload where wall-clock time is a budget item, V4 Pro's combination of 2.3x speed at a seventh of the output price is the dominant answer.

When the nine points are worth it

The honest case for Kimi K3 is the reasoning ceiling. Its evaluation cost on the Intelligence Index was about $3,658 for the run — a reflection of genuinely long and expensive reasoning chains, not a quirk of pricing. If your workload is hard reasoning — deep mathematical proof, long-horizon planning, multi-hour agentic tasks where an extra nine index points shows up as fewer failures — Kimi K3 earns its premium. The honest case against it is everything else: the throughput, the cache economics, the sheer per-token price.

The other reason the September reversal matters here: a month ago the sensible advice was "if you are going to standardize on the cheap fast open model, hedge — it may be retired." As of September 15 that hedge is obsolete. Deep​Seek has committed publicly, with notice promised if anything changes, to keep V4 Pro on the API, and the MIT license means the model remains deployable on your own hardware regardless of what Deep​Seek does with its API. Kimi K3's open weights give the same freedom on the other side, which is what makes this a genuine two-open-weights comparison rather than a fight between a service and a product.

A screenshot of the DeepSeek API docs Models & Pricing page, captured September 15, 2026, showing the English-language pricing table for deepseek-v4-pro (DeepSeek-V4-Pro-0813) at .32 input and .96 output per million tokens at peak, and the footnote announcing that DeepSeek V4 Pro API service continues after September 14, 2026 with unchanged billing.

Both models are on OrcaRouter — DeepSeek V4 Pro and Kimi K3 both route through the platform on one key, with provider list prices passed through at zero markup — and the routing DSL is the natural way to split this matchup: route the long-horizon reasoning tasks to Kimi K3 where the nine points pay rent, and route the token-heavy throughput work to DeepSeek V4 Pro where the 7.5x price gap and 2.3x speed gap do the work for you. The two are complements as often as they are competitors, and the platform makes running both the same amount of work as running one.

A screenshot of the OrcaRouter model page for deepseek/deepseek-v4-pro, captured September 15, 2026, showing the 1M token context window, 384K max output, a reasoning-mode toggle, the input price of /bin/bash.66 and output price of .98 per million tokens, and the model description reading 'DeepSeek V4 Pro flagship MoE — 1.6T total / 49B active params, 1M context'.A generated scoreboard titled 'DeepSeek V4 Pro vs Kimi K3 — the scoreboard': left column DeepSeek V4 Pro rows AA Index 36 (#7/113), Price /bin/bash.66/.98 off-peak, Speed ~81 tok/s, Context 1M, Active params 49B/1.6T, Weights MIT; right column Kimi K3 rows AA Index 44 (#2/113), Price .00/5.00, Speed ~35 tok/s, Context 1M, Active params 104B/2.8T, Weights open; footer 'AA figures independent; DeepSeek price is vendor list off-peak.'

The bottom line

Kimi K3 is the smarter open-weights model, nine index points ahead, and its premium — 7.5x on output tokens, roughly half the speed — is a real cost, not a rounding error. DeepSeek V4 Pro is the fast, cheap, durable workhorse, and after September 11 it is no longer a flight risk. The practical answer is the one both price columns point at: if your tasks are hard reasoning, buy Kimi K3; if they are token-heavy and time-sensitive, buy DeepSeek V4 Pro; and if you have both kinds, run both and route by task.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily