Hero title card for the comparison 'Grok 4.6 vs DeepSeek V4 Pro', subtitled 'A frontier price war, fought 24 hours apart,' with a 'Comparison · August 2026' badge.
Guides & Insights

Grok 4.6 vs DeepSeek V4 Pro: A Frontier Price War, Fought 24 Hours Apart

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Two frontier models shipped within 24 hours of each other, and the gap between their price sheets is the widest this generation has produced. Grok 4.6 launched on August 12, 2026 at $2 per million input tokens and $6 per million output. DeepSeek V4 Pro — the formal production release of D​eepSeek's flagship, version DeepSeek-V4-Pro-0813 — followed on August 13, 2026 at ¥3 per million cache-miss input tokens and ¥6 per million output, which is roughly $0.42 and $0.85. Same category, same agentic ambitions, an order of magnitude apart on price. This is not a comparison of two specs; it is a comparison of two strategies for what an agent-shaped model should cost, and both released this week.

This piece puts the two price sheets next to each other, reads each vendor's benchmark claims with the vendor label on, and works out what the gap actually costs on a real workload — because on paper these models are closer in capability than in philosophy, and the difference in price per task is where the decision lives.

The timing is the point

That both models landed in the same 48-hour window is not an accident of calendars. Grok 4.6 is SpaceXAI's agentic refit of its 1.5T foundation — a post-training refresh heavy on reinforcement learning across coding, kernel work and CAD, priced as the fuel for the Cursor and Grok Bot products the company owns. DeepSeek V4 Pro is the formalization of a preview that quietly redefined what "agentic" meant at a fraction of the price, now enhanced for the agent economy: D​eepSeek's release notes for the 0813 build lead with significantly improved agent capabilities, add support for the Responses API and Codex integration, and keep both thinking (default) and non-thinking modes, JSON output, tool calls, and the Anthropic API format. Two very different companies concluded simultaneously that the market that matters in late 2026 is agents — and priced their entries for very different wallets.

The two price sheets, side by side

Grok 4.6 — $2.00 / $6.00 / $0.50 (cached input) per million tokens, 500K context, a $4 / $12 / $1 step above roughly 200K input tokens, and a faster variant at double the base rate.

DeepSeek V4 Pro — ¥3.00 (≈$0.42) cache-miss input, ¥0.025 (≈$0.0035) cached input, ¥6.00 (≈$0.85) output per million tokens, with a 1-million-token context window and up to 384K output tokens. Its cheaper sibling, V4 Flash, runs ¥1 / ¥2.

Even against Grok 4.6 — already the cheapest frontier API of its generation — DeepSeek V4 Pro is roughly five times cheaper on cache-miss input and seven times cheaper on output, and it brings a 2x larger context window and a far larger max output to the same price tier. The two are not competing on price; D​eepSeek has unilaterally set a new floor and Grok is now the premium option in the same sentence. That framing matters, because the previous framing — "Grok 4.6 is the cheap frontier model" — was true for exactly one day.

Two-column scoreboard: Grok 4.6 AA Index 61 (independent), $2/$6, 500K context, DeepSWE v1.1 65.9% (vendor), Terminal-Bench 26% (v3.0), cache-hit input $0.50 vs DeepSeek V4 Pro AA Index none yet, ≈$0.42/$0.85, 1M context, DeepSWE 62.7% (vendor), Terminal-Bench 87.9% (v2.1), cache-hit input ≈$0.0035.

What DeepSeek V4 Pro actually improved

D​eepSeek's launch numbers are the most dramatic because they are measured against its own preview, and the preview-to-release jump is enormous. On the vendor's own evaluations:

DeepSWE — 62.7% on the release build, up from 12.8% on the preview — a software-engineering long-horizon score the vendor puts between Opus 4.8's 58.0% and Claude Fable 5's 70.0%.

Cybergym — 83.3%, up from 52.7%, edging Fable 5's 83.1% and beating Opus 4.8's 78.3%.

Terminal Bench 2.1 — 87.9%, within a point of Fable 5's 88.0% and ahead of Opus 4.8's 85.0%.

AutomationBench (public) — 31.8%, up from 12.8%, ahead of both Fable 5 (29.1%) and Opus 4.8 (27.2%).

Toolathlon-Verified, NL2Repo, DSBench-FullStack and DSBench-Hard — 74.1%, 61.5%, 71.1% and 67.2% respectively, clustered at or near the Opus 4.8 / Fable 5 band.

Every one of those is D​eepSeek-reported on D​eepSeek's own eval suite. The direction is unmistakable — the preview was not ready for production agent loops and the release build claims to be — but the honest label is that no independent lab has scored DeepSeek V4 Pro yet, and the models it is benchmarked against are not named by an external party.

What Grok 4.6 answers with

Grok 4.6's counter is different in kind: it is the only one of the two with an independent scoreboard. Artificial Analysis measured it at 61 on its Intelligence Index — tied with GPT-5.6 Sol, ahead of Kimi K3 (60) — before DeepSeek V4 Pro even shipped. Its vendor table claims a leading 65.9% on DeepSWE v1.1 and 69.9% on CursorBench v3.2, with an openly admitted weak spot at 26% on Terminal-Bench v3.0. Those are xAI-reported, none independently reproduced as of August 13.

The version mismatches matter when you try to compare the two directly. D​eepSeek's 87.9 is on Terminal Bench 2.1; Grok's 26% is on the harder v3.0 — different exams, so "D​eepSeek crushes Grok on terminals" is not an honest line, and neither is "Grok leads on DeepSWE" without noting the two vendors ran different eval versions and neither number is third-party. What is comparable is the independent dimension: Grok 4.6 has one verified scorecard, DeepSeek V4 Pro has none yet, and for a production decision that asymmetry is itself information.

Screenshot of the Artificial Analysis page for Grok 4.6 (high) scoring 61 — the independent scorecard DeepSeek V4 Pro does not yet have.

The total cost for a long agent run

Take a realistic agentic job — reading and reasoning over 500K tokens of context, writing 100K tokens of output, which is a mid-size refactor or a long document pipeline, within both models' windows:

Grok 4.6 — 0.5M input at $2 = $1.00, plus 0.1M output at $6 = $0.60. Total: $1.60.

DeepSeek V4 Pro — 0.5M cache-miss input at ¥3 = ¥1.50, plus 0.1M output at ¥6 = ¥0.60. Total: ¥2.10, about $0.29.

That is a 5.5x gap on the same nominal task, before any caching advantage — and D​eepSeek's 1M window plus 384K max output also means the job fits in a single pass where Grok 4.6's 500K window may force two. The catch that keeps this from being a one-line answer is reasoning cost and risk: D​eepSeek's thinking mode is on by default, so every request pays for a full reasoning trace even when you do not want one, and the vendor's own benchmark confidence is untested by anyone independent. Cheap per token with always-on reasoning is still cheap per task at these rates, but "cheap" and "verified" are different claims, and only one of these models carries the second.

Who should pick which

The decision is less "which is better" than "which risk profile fits the workload":

High-volume, cost-bound, and willing to carry unverified numbers — DeepSeek V4 Pro is the price-floor play. The 1M context and 384K output make it the stronger long-context and batch candidate, and the ¥0.025 cached-input rate makes retrieval-heavy pipelines nearly free. Just do your own evals, because nobody independent has.

Agentic depth with an independent scorecard, in an OpenAI-compatible ecosystem — Grok 4.6's 61 is verified, its coding and agentic benchmark direction is consistent, and the terminal weakness is known up front. The 42-second time-to-first-token is the price you pay for the verified quality.

Screenshot of the OrcaRouter model page for Grok 4.6, the pass-through endpoint where both models sit behind one key at provider list price.

Mixed traffic — this is the case for routing rather than choosing. On OrcaRouter, both models sit behind one key at provider-list pass-through pricing — so D​eepSeek's ¥3/¥6 stays ¥3/¥6 on the invoice, and Grok 4.6 stays $2/$6, with zero markup on either. A routing rule can send the long-context, cost-bound work to DeepSeek V4 Pro and the verified-agentic work to Grok 4.6, split traffic by request until your own evaluation settles the split, and rely on automatic failover if either vendor's first-week behavior contradicts its launch numbers. For a pair of one-day-old models on opposite ends of a price war, the ability to compare both on your own workload — and flip the split without changing code — is not a convenience. It is the only defensible way to adopt either.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube