
Grok 4.7 vs DeepSeek V4 Pro: One Is New, the Other Is Being Swapped Under You
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
A note published on September 10, 2026 changes what this comparison is about. "We're phasing out V4-Pro," it says, and from 04:00 UTC on September 14, 2026 every request to the deepseek-v4-pro model string routes to DeepSeek-V4.1-Flash at V4.1-Flash rates — a state of affairs the post says continues "until V4.1-Pro launches." The model string still answers. The model behind it is not the one you benchmarked. That makes a head-to-head against Grok 4.7 — released by the vendor on September 21, 2026, one day before this was written — a comparison with an expiry date on one side. On the numbers that exist, Grok 4.7 is the stronger model and DeepSeek V4 Pro is the cheaper one. On the question that actually decides an integration, only one of them is still the thing it says it is.
What each model is, before the scores
DeepSeek V4 Pro reached general availability on August 13, 2026 as the DeepSeek-V4-Pro-0813 checkpoint, after an April 24 preview. It is open weights under MIT — the licence permits commercial use, modification and redistribution with no revenue threshold or regional carve-out — with 1.7 trillion total parameters on the GA model card, a 1M-token context window and a 384K maximum output, the largest output ceiling in this matchup. It is text-only: DeepSeek's own pricing page states vision is not supported on v4-pro, which is a real limitation for anyone feeding it screenshots or PDFs. It exposes a reasoning-effort dial at low, high and max, and it is capped at 500 concurrent requests per account, with expansion available on request.
Grok 4.7 is closed-weight and a day old. It carries a 500k-token context — half of DeepSeek's — takes text and images in, and runs its own effort dial. Its list price is $2.00 per million input tokens and $6.00 per million output tokens below 200k prompt tokens, doubling to $4.00 and $12.00 at or above that threshold, with cached reads at $0.50 and $1.00. xAI also lists a faster variant at roughly twice the output speed and twice the price. Neither vendor publishes a maximum output figure for Grok 4.7 in the documentation checked here, so treat that dimension as unknown rather than equal.
On the shared index, the gap is lopsided in one direction
Both models are scored on the Artificial Analysis Intelligence Index v4.3.2, the revision published September 19, 2026. Grok 4.7's column is measured at xhigh effort; DeepSeek's is the 0813 checkpoint at max effort. Reading them together, the composite gap understates how differently these two models are shaped.
• Composite — Grok 4.7: 46 on Artificial Analysis Intelligence Index v4.3.2. DeepSeek V4 Pro 0813: 36 on the same revision.
• Terminal agents — Grok 4.7: Terminal-Bench 4.0 26. DeepSeek V4 Pro: 14. Nearly double, on the evaluation closest to autonomous software work.
• Automation — Grok 4.7: AutomationBench-AA 66%. DeepSeek V4 Pro: 57%. Both credible; Grok leads by nine.
• Knowledge and hallucination — Grok 4.7: AA-Omniscience 32. DeepSeek V4 Pro: 1. This is the outlier on the board, and it is not a rounding artefact — the index scores both what a model knows and how often it invents an answer when it does not.
• Long context — Grok 4.7: AA-LCR v1.1 77%, window 500k. DeepSeek V4 Pro: 80%, window 1M. DeepSeek's one clear win, and it is the axis its 1M window was built for.
• Hard reasoning — Grok 4.7: HLE 43, CritPt 18. DeepSeek V4 Pro: HLE 41, CritPt 18. A tie on the physics benchmark and a two-point edge to Grok on HLE.
• Verbosity — Grok 4.7: 240M output tokens to run the index, flagged as unusually verbose. DeepSeek V4 Pro: 163M. Both are chatty; Grok is chattier.

The price story is not one number, it is a schedule
DeepSeek's pricing is the most aggressive thing about it, and the least stable to quote. The list rate for deepseek-v4-pro is $0.66 per million input tokens on a cache miss and $1.98 per million output tokens — during off-peak hours. During peak hours those double to $1.32 and $3.96. Peak, per DeepSeek's own documentation, is 01:00–04:00 and 06:00–10:00 UTC Monday through Friday excluding Chinese public holidays; everything else, including full weekends, is off-peak at exactly half price. Cached input reads are the real headline: $0.022 off-peak and $0.044 peak, thirty times cheaper than a cache miss, which makes a well-cached DeepSeek workload dramatically cheaper than its rate sheet implies.
Against that, Grok 4.7's $2.00 and $6.00 looks expensive — roughly 3x DeepSeek's off-peak input and 3x its off-peak output. The comparison narrows in ways worth knowing. Grok 4.7's price is flat within its band rather than clock-dependent, so a 09:00 UTC production spike costs the same as a Sunday-night batch; DeepSeek's does not. And above 200k prompt tokens Grok 4.7 doubles, at which point it is four to six times DeepSeek's off-peak rate rather than three. The clean way to put it: DeepSeek V4 Pro is the cheaper model on every axis, and the size of the discount depends on when you run and how well you cache.
What September 14 changed
This is the part that makes the price argument harder to bank. From September 14, 2026, DeepSeek routes deepseek-v4-pro requests to DeepSeek-V4.1-Flash and bills them at Flash rates — which are lower still. DeepSeek's change log and its pricing page tell a slightly different story than the September 10 news post does: the pricing table still lists deepseek-v4-pro as a live row with its own rates and no retirement tag, and the change log entry says API service for V4 Pro continues after September 14 with billing unchanged. What is not in dispute is the direction of travel. V4.1-Pro is confirmed in development with no announced date, and V4.1-Flash — released September 10 — is what DeepSeek's own announcement says beats V4 Pro on "performance, cost, speed and total runtime," a vendor claim that has not been independently reproduced at that breadth.
So the practical question is not "Grok 4.7 or DeepSeek V4 Pro" but "Grok 4.7 or whichever DeepSeek model that string resolves to next quarter." If you benchmarked V4 Pro in August and sized your context budget around a 1M window with a 384K output ceiling, the model you call in November may not be the one you measured — and the successor's context and output limits are not the ones you designed against. Grok 4.7 has the opposite risk profile: a one-day-old model whose numbers are vendor-reported and largely unreproduced, but whose identity is not in question.

Where a router fits, and where it does not
OrcaRouter carries DeepSeek V4 Pro, and the model page lists it at the provider's own rate — $0.66 in and $1.98 out with a 1M context window and 384k maximum output — because we pass provider list pricing through at 0% markup. That matters more than usual with a vendor whose rates move on a peak/off-peak schedule and whose September 10 announcement came with a price reduction: a DeepSeek price change is live on the same key the same day, with no repricing lag and no second contract. When a vendor reroutes a model string at their end, the change arrives through the same endpoint rather than requiring a re-integration on yours, and automatic failover keeps a rate limit or a provider-side capacity event from becoming an outage — useful when one side of your stack is capped at 500 concurrent requests. Grok 4.7 is not in the OrcaRouter catalog as of this writing; calling it means going through xAI's own API and the third-party platforms that carry it.
The split this suggests is unglamorous but defensible: keep DeepSeek V4 Pro on the workloads where cache-hit pricing and the 1M window are doing the work, and pin your expectations to V4.1-Flash rather than to the August checkpoint. Put Grok 4.7 on the tasks where the model has to know what it does not know — an AA-Omniscience index of 1 is a loud signal about a model's willingness to answer confidently and wrongly, and no price advantage survives a downstream consumer trusting a fabricated answer.
The call
Grok 4.7 is the better model here and it is not especially close: a ten-point composite lead, double the Terminal-Bench score, an automation edge, and a knowledge-honesty gap wide enough to be a category difference. DeepSeek V4 Pro is cheaper by roughly 3x off-peak and carries twice the context window, which is a real and defensible reason to keep it — but you are not buying the model you benchmarked, you are buying a model string that DeepSeek has already announced it will repoint, at rates that change by the hour, on a licence and a weight set that make self-hosting a genuine third option. If the workload is long-context, high-volume and cacheable, the DeepSeek economics are hard to beat and you should design for the successor rather than the checkpoint. If it is anything where being wrong is expensive, pay the 3x and take the model that admits uncertainty.

Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
