
DeepSeek V4 Pro Release Date: Already 14x Cheaper Than Kimi K3, Not Yet as Smart
- metaNEWMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenNEWQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekNEWDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxNEWMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens · 2110 tok/s
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1653Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1560Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Intelligence77Coding
- grokxAI: Grok 4.52026-07-0856Intelligence72Coding
- tencentTencent: Hy32026-07-0642Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3055Intelligence72Coding
- klingKling: Kling 3.0 Turbo2026-06-1757Intelligence52Coding57Math
On 31 July 2026, the company published a change-log entry announcing that its cheap model had beaten its expensive one. The entry was about DeepSeek V4 Flash, the 284B efficiency tier, and the headline claim was "significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview." V4-Pro-Preview is DeepSeek V4 Pro — the 1.6-trillion-parameter flagship, the one that costs three times as much per token. The same entry closed with a single sentence about it: "The official release of DeepSeek-V4-Pro will follow soon." Which is the context for the forecast currently going around, that when the real V4 Pro lands it will be Kimi K3 level but ten times cheaper. One half of that is already measurable, and it is wrong in the direction almost nobody expects.
A note on what this article is and isn't. There is no leaked system card here, no internal benchmark, no date. The prediction in circulation comes from @scaling01 on X — "DeepSeek-V4 Pro will be a banger probably also Kimi-K3 level but 10x cheaper" — and it is a forecast from a well-followed practitioner, not a disclosure from anyone with access. Treating it as a leak is how bad release-date coverage gets written. What we can do instead is take the two testable halves of the claim and check them against numbers that already exist, because the preview build has been callable for three months and both it and Kimi K3 have been run by an independent lab. Everything below is labelled by source: the company's own documentation, Artificial Analysis's measurements, or arithmetic we did ourselves and show.
What DeepSeek has actually said, and what it hasn't
The confirmed surface is larger than the "unreleased model" framing suggests. V4 Pro is not vapour — you can call it right now. What has not happened is the release the company itself considers official.
• Architecture — a mixture-of-experts model with 1.6T total parameters and 49B active per token, confirmed by the company and listed identically on Artificial Analysis as 1600B/49B.
• Context and output — a 1M-token input window (1,048,576) with up to 384K output tokens. Text in, text out; no image, audio or video input.
• Licence and weights — MIT, with weights published on Hugging Face. This is an open-weights flagship, which is the whole reason the comparison to Kimi K3 is interesting.
• List price — $0.435 per million input tokens on a cache miss, $0.003625 on a cache hit, and $0.87 per million output tokens, from the company's own pricing page.
• Interfaces — OpenAI-style ChatCompletions and Anthropic-format endpoints, plus reasoning-effort levels, tool calling and structured output.
• Timeline so far — the V4 family arrived as a preview on 24 April 2026; the legacy deepseek-chat and deepseek-reasoner model names were retired on 24 July 2026; Flash got its official build on 31 July 2026; Pro did not.
• 6 August 2026 — DeepSeek told developers it plans to raise the overall pricing for DeepSeek API services "in the near future, with a significant increase expected," and reiterated that the official V4 Pro release will follow. No new rates and no effective date have been published.
And the parts that remain genuinely unknown, which is most of what people are searching for:
• The GA date — still unannounced. "Will follow soon" remains the entirety of the company's formal commitment. Unconfirmed press reports point to a mid-August window — one outlet has floated 10–20 August, and a reseller's pricing page moved V4 Pro cache prices on 3 August in a pattern that historically precedes a release — but none of that is the company speaking. No date has been confirmed.
• What the GA build will change — unstated. The company has not said whether the official V4 Pro is a re-post-train of the same weights, a new pretrain, or a repricing.
• Whether the price survives it — unaddressed, and there is a specific reason to worry about this that we come to below.
One correction worth making explicitly, because the search results for this question are contaminated: you will find pages asserting that the V4 family hit general availability on 20 July 2026. The company's own change log, eleven days after that date, says the official V4 Pro release is still pending, and adds that the 31 July update "only upgrades the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and the APP/WEB models are unchanged." When a third-party summary and a vendor's primary changelog disagree, the changelog wins.

The price claim: "10x cheaper" is the conservative reading
Kimi K3 lists at $3.00 per million input tokens and $15.00 per million output, with cache hits at $0.30. V4 Pro lists at $0.435 and $0.87, with cache hits at $0.003625. Run those against each other and "10x" turns out to be the low end of a range, not the headline:
• Input tokens — $3.00 against $0.435, so V4 Pro is 6.9x cheaper. This is the one dimension where the claim overstates.
• Output tokens — $15.00 against $0.87, or 17.2x cheaper. Output is where reasoning models spend, which makes this the number that matters most.
• A 70/30 input-output blend — $6.60 per million against $0.5655, or 11.7x.
• Artificial Analysis's 7:2:1 cache-hit/input/output blend — $2.31 against $0.18, or 12.8x.
• Cached input — $0.30 against $0.003625, roughly 83x. The company's cache-read price ranks 4th of 101 models Artificial Analysis tracks, at a 99% discount to its own cache-miss rate.
Hypothetical blends are arguable, so here is a measured one. Artificial Analysis publishes what it actually spent running its full Intelligence Index on each model — the same evaluation suite, the same harness, real token counts rather than an assumed ratio. Kimi K3 cost $2,437.41 to evaluate. V4 Pro cost $176.34. That is 13.8x, on an identical workload, in dollars someone really paid.
What makes that figure lower than the 17.2x output-price ratio is worth understanding, because it is the one place where V4 Pro's cheapness is partly self-inflicted. V4 Pro is verbose: it burned 180M output tokens getting through the index, against Kimi K3's 130M, and against a 100M median for its class. So it spends about 38% more tokens saying what it has to say, which eats into a per-token advantage. The 13.8x figure has that already priced in; the 17.2x figure does not. If you are budgeting, use the measured one.
This is also the part of the story that is easy to check yourself rather than take on faith. Because OrcaRouter passes provider list prices straight through at 0% markup, the per-token numbers on our deepseek/deepseek-v4-pro and kimi/kimi-k3 model pages are the two vendors' own — K3 reads $3.00/$15.00 on our page and $3.00/$15.00 on Moonshot's — not a resold rate with a spread hidden in it — which means the arithmetic above reproduces on one bill, and if either vendor moves a price, the change lands on our side the same day rather than after a contract cycle.
The intelligence claim: 13 points, not parity
"Kimi K3 level" is the half of the forecast that does not survive contact with the current numbers. On Artificial Analysis's Intelligence Index, as of 5 August 2026, Kimi K3 scores 57 and DeepSeek V4 Pro (reasoning, max effort) scores 44 — ranked #6 of 101 models, which is a genuinely strong placement, and still 13 points short. Both are measured by the same lab on the same composite, both at maximum reasoning effort, both against a 1M-token context.
The vendor benchmarks tell a different and more flattering story, which is exactly why they need labelling. The company reports V4 Pro-Preview at 80.6% on SWE-bench Verified and 90.1% on GPQA Diamond. Moonshot reports Kimi K3 at 76.8% on SWE-bench Verified and 93.5% on GPQA Diamond. Taken at face value, V4 Pro wins the coding benchmark outright. Neither set has been independently reproduced, they were produced on different harnesses by teams with an interest in the outcome, and on the one composite where a third party ran both models itself, the ordering reverses. That is the whole case for distrusting single-benchmark comparisons between vendors.

Two things about the index number deserve caveats of their own. First, if you find older write-ups quoting V4 Pro at 52 rather than 44, they are not wrong — Artificial Analysis's own April launch article did put V4 Pro (Max) at 52 and called it the #2 open-weights reasoning model. The index gets revised, and revisions move every model's score. The practical rule is that only same-snapshot comparisons mean anything; a 52 from April and a 57 from August do not belong in the same sentence. Second, Artificial Analysis's page is dated "Released April 2026" and does not distinguish the preview build from a future official one, so its 44 is a measurement of what the endpoint served, whenever it was tested.
Where V4 Pro does clearly beat Kimi K3 is on the axes that are not intelligence. It generates 66.5 tokens per second against K3's 37.6, in a class whose median is 65.9 — so V4 Pro is around average for its class and Kimi K3 is unusually slow. Time to first token is 1.61 seconds against 2.85. For an interactive agent loop, that difference is felt on every turn, and it is the strongest argument for the preview build as it stands today.
What Flash's official build tells us about Pro's
Here is the closest thing to real evidence about the unreleased model, and it comes from the sibling that already went through the process.
The company's 31 July change log states that "DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained." Same weights count, same architecture, same price — post-training only. Artificial Analysis scored the result at 50, against 40 for the previous Flash build. Ten points on the composite, from post-training alone, at no extra cost per token. Its output-token consumption on the index also fell 12%, from ~234M to ~206M, so it got cheaper to run at the same list price.
Apply that precedent to Pro and the arithmetic is unavoidable: 44 plus 10 is 54. Kimi K3 is at 57. Even if the official V4 Pro repeats the single most successful post-training upgrade the company has publicly demonstrated, it lands about three points short of "Kimi K3 level" — while costing roughly a fourteenth as much to run. That is a remarkable outcome and it is not the outcome the forecast predicts.
The honest caveats on that extrapolation, which is arithmetic and not a prediction: a +10 gain on a 284B model says little about what is available on a 1.6T model, and the headroom could be larger or much smaller. The company has not said the Pro build will be post-training-only. And the index is a composite, so three points can be one benchmark. The reason to run the number anyway is that it is the only quantified, vendor-confirmed precedent that exists, and it is a great deal more than the forecast is resting on.
One more detail from that change log deserves to be flagged rather than repeated: Flash's agentic scores were produced using "the company's Harness minimal mode (to be released soon)" at max effort, topp=0.95, temperature=1.0. The harness that generated the numbers has not shipped publicly — it entered closed beta in early August — so those figures still cannot be reproduced by anyone outside the company even in principle, not because they are wrong but because the instrument is unavailable. Expect the Pro release to arrive with the same asterisk.
The clauses that could undo the cheap-model thesis
Buried in the company's pricing documentation is a policy that almost no coverage of the V4 line prices in. In the company's words, the API "will soon adopt" a schedule under which "during peak hours, prices will be 2x the regular prices, applicable to all billing items." Peak hours are defined as 09:00–12:00 and 14:00–18:00 Beijing time, daily — seven hours a day, every day. The effective date is "subject to the official announcement," and as of this writing the surcharge is not active.
Note the direction. The company historically ran off-peak discounts on V3 and R1; this is a peak surcharge on the standard rate. If it activates at the stated multiplier, V4 Pro's daytime price becomes $0.87 in and $1.74 out, and the comparison changes shape: the 11.7x blended advantage over Kimi K3 falls to 5.8x, and the measured 13.8x falls to 6.9x. Still cheap. No longer "10x cheaper," and no longer cheap at the hours when a European or Asian team is actually working.
On 6 August the larger shoe dropped. DeepSeek announced on its developer platform that it plans to raise the overall pricing for DeepSeek API services "in the near future, with a significant increase expected." No new rates and no effective date have been published, and the company has not said whether the increase is separate from the peak-hour schedule or lands together with it. Press covering the company read the announcement as a signal that the official V4 Pro release is imminent — an unreleased flagship is not something you price aggressively into a discount regime you are about to leave. Anyone building a cost model on V4 Pro today should model both cases, and anyone reading a "the company is 14x cheaper" headline — including this one — should know it describes a price list with an announced expiry condition attached. This is the single most likely way the price half of the forecast stops being true, and it has nothing to do with the model.
How you will know the official build has landed
There is a detail in how the company ships that matters more than the date, and it is a genuine operational hazard.
The API model ID is deepseek-v4-pro. Not deepseek-v4-pro-preview, and not a dated build string — even though the company's own change log refers to the current build as V4-Pro-Preview and to the shipped Flash build as V4-Flash-0731. When Flash went official, the instruction was: "The API calling method remains unchanged — simply set the model name to deepseek-v4-flash to use the latest version." In other words, pinning the model ID does not pin the build. The same string silently began serving a different set of weights, with materially different behaviour and ten more index points.
So the practical answer to "when does V4 Pro release" is that if you are calling deepseek-v4-pro in production, you may find out by watching your own outputs change. Concretely, the things to watch are: a new dated entry on the company's API change log, which is where Flash's release appeared first; a build suffix such as V4-Pro-0731 in that entry's prose even while the callable ID stays bare; a re-score on Artificial Analysis's V4 Pro page, which is the first independent confirmation you will get; the announced price increase resolving into actual rates, since GA is the natural moment for both that and the peak-hour schedule to land; and the company's Responses API, whose documentation says deepseek-v4-pro support is planned for early August — the model disappearing from the "planned" list and appearing on the supported list is about as close to an official release signal as the company publishes. Two signals we cannot confirm are worth knowing anyway: a reseller's pricing page moved V4 Pro cache prices on 3 August in a pattern that historically precedes a release, and the company's in-house evaluation harness is in closed beta, which means its vendor benchmarks may finally become checkable when it ships. A string resembling deepseek-v4-pro-202606 has also circulated as a leaked build identifier; we could not verify it against any company property, and it should carry no weight until someone can.
Should you build on the preview today?
For most people evaluating this model, the release date is the wrong question. The preview is callable, priced, MIT-licensed and independently measured at #6 of 101 on intelligence. The real question is what the pending GA does to work you build now, and there are two distinct answers.
If you are choosing between V4 Pro and Kimi K3 on capability, the current evidence says Kimi K3 wins the composite by 13 points and V4 Pro wins on speed, latency and cost by wide margins — and that a pending post-training upgrade may narrow the intelligence gap without closing it. If your work is at the frontier of hard reasoning, that 13 points is the whole ballgame and the price is a distraction. If your work is high-volume and merely difficult, paying 14x for it is hard to justify.
If you have already picked V4 Pro, the risk to manage is not the release date but the silent swap: the build behind that model ID will change under you, exactly as Flash's did, and your evals should be able to detect it. Keep a regression set you can re-run on demand, and keep a second model reachable without a code change. This is the mundane reason we built failover the way we did — V4 Pro, V4 Flash and Kimi K3 all sit behind one OrcaRouter key on an OpenAI-compatible endpoint, so "re-run the eval on all three, then shift traffic" is a config edit rather than a procurement exercise, and a bad GA build is a routing decision instead of an incident. Trying an unproven flagship is much easier when reversing that choice costs nothing.

Questions people are actually asking
Is DeepSeek V4 Pro released or not?
Both, depending on what you mean, which is why the search results conflict. The model has been publicly callable since the V4 preview launch on 24 April 2026, with published pricing, open weights under MIT and independent benchmark coverage. But the company's own change log of 31 July 2026 explicitly states the V4 Pro API is "unchanged" and that its official release "will follow soon," and on 6 August DeepSeek again said the official version would be released as soon as possible while announcing a large API price increase. So: available, not yet official. No date has been confirmed; unconfirmed reports point to mid-August, which would make the wait short.
Will the official build make it as good as Kimi K3?
On the only quantified precedent available, probably not quite. Flash's official build gained 10 points on Artificial Analysis's Intelligence Index from post-training alone; V4 Pro sits at 44 and Kimi K3 at 57, so an identical gain lands at 54. That extrapolation could be wrong in either direction — a 1.6T model may have different headroom than a 284B one, and the company has not said the Pro release will be post-training-only. But anyone asserting parity today is asserting something no published measurement supports.
Why does DeepSeek's cheap model currently score higher than its flagship?
Because of release sequencing, not because Flash is the better model. Flash received a post-training upgrade on 31 July that took it from 40 to 50 on the index; Pro has not received its equivalent, so it is still being measured on April's post-training. The company says as much itself, describing Flash-0731's agentic results as "far exceeding V4-Pro-Preview." The inversion is a snapshot of an unfinished rollout, and the pending Pro release is precisely what would end it.
Is the $0.435 / $0.87 price safe to build a cost model on?
Conditionally, and less safe than it was a week ago. It is the company's current published list price and it already reflects a large permanent reduction from the V4 launch rates. But on 6 August DeepSeek announced that it plans to raise API pricing overall, "with a significant increase expected," and no new rates or date have been published yet — on top of the announced-but-inactive peak-hour policy that would double all billing items for seven hours a day. Model both scenarios, and note that even the doubled case is still cheaper than Kimi K3 by roughly 6x.
The honest read
The forecast that started this was half right, and the half it got right is the half people will find hardest to believe. "10x cheaper" understates it: on the same measured evaluation workload, DeepSeek V4 Pro's preview build cost $176.34 against Kimi K3's $2,437.41, and on cached input the gap is not tenfold but roughly eightyfold. "Kimi K3 level" is the part that is not there yet — 44 against 57 on the one composite an independent lab ran on both, with an optimistic extrapolation from the company's own best precedent landing at 54.
What would change this read is specific and worth watching for. If the official V4 Pro build arrives as a post-training upgrade and Artificial Analysis re-scores it above 57, the forecast is vindicated in full and this article's central number is stale. If it arrives alongside the peak-hour surcharge or the price increase announced on 6 August, the price argument halves and the interesting question becomes whether a 6x discount buys a 13-point deficit. And if the company ships its evaluation harness — already in closed beta — the vendor benchmarks that currently favour V4 Pro on coding become checkable for the first time. Until one of those happens, the accurate summary of DeepSeek V4 Pro is that it is the best-value serious reasoning model with weights you can download, that it is not the smartest one, and that its vendor has told us it is not finished.
Compared in this article4
Detected from this article · Benchmarks: Artificial Analysis · updated daily
