
Grok 4.7 vs Grok 4.6: Same $2/$6 Price, a Different Model Underneath
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 1009 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1189 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 22 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 108 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
SpaceXAI shipped Grok 4.7 on September 21, 2026, and the first thing worth noticing about it is what did not change: the price. Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, exactly what Grok 4.6 has cost since it launched on August 12, 2026. Same rate card, same 500,000-token context window, same text-and-image input — and yet the two models are not close on the evaluations that matter for agentic work. On SpaceXAI's own published numbers, Grok 4.7 nearly doubles Grok 4.6 on Terminal-Bench 4.0, 38.0% against 20.3%. That is the whole shape of this comparison: a same-price generational swap where almost all of the gain lands in one place, and the parts of Grok 4.6 that annoyed production teams are still there.
What actually changed between the two models
SpaceXAI's own description of the release is narrow, and it is worth quoting the shape of it rather than the adjectives. Grok 4.7 is built on a larger base model than Grok 4.6, and it was given a longer reinforcement-learning run aimed at tasks that "can take hours to complete." The company says that run sharpened three things: self-verification, long-context handling, and behaviour inside the Grok Bot environment. There is no claim about a new architecture, a new modality, or a different serving stack.
That framing matters because it predicts where the benchmark gains show up, and they show up exactly there. A longer RL pass over hard multi-hour tasks moves terminal and repository work far more than it moves a knowledge quiz. If you were hoping Grok 4.7 would be a broader model than Grok 4.6, the release notes say otherwise — it is the same model shape, trained further into the hard end of agentic work.
The parameter story is the one piece of this that is not a vendor disclosure. Musk said in early September that Grok 4.7 carries roughly 2.1 trillion parameters against Grok 4.6's 1.5 trillion, a 40% increase, and that figure has been repeated everywhere since. It has never appeared on a SpaceXAI specification sheet. Treat it as a widely relayed claim about the base model, not a confirmed spec — and note that parameter counts say very little about serving cost or capability when the architecture is sparse, which the Grok line's is.
The benchmark gap, with the sourcing label attached
Every figure in this section comes from SpaceXAI's launch material. None of it has been independently reproduced at the time of writing, and the honest way to read it is as a claim set that happens to be unusually specific — which makes it checkable later, but does not make it verified now.
• Terminal-Bench 4.0 — Grok 4.7 38.0% (vendor-reported) vs Grok 4.6 20.3%. The single largest delta in the release, and the one that matches the "longer RL on harder tasks" claim.
• CursorBench 4.0 — Grok 4.7 46.3% (vendor-reported) vs Grok 4.6 40.4%. A real gain, but a modest one — under six points, on a benchmark closer to everyday editor work than to autonomous terminal sessions.
• DeepSWE v1.1 — Grok 4.7 71.0% at high reasoning effort (vendor-reported) vs Grok 4.6 65.2%. Again a solid, unspectacular improvement.
• EEBench — Grok 4.7 64.0% (vendor-reported). No Grok 4.6 figure was published alongside it, so this one tells you nothing about the upgrade.
• AA-Briefcase v1.1 — Grok 4.7 at 1,657 Elo (vendor-reported), on the professional knowledge-work evaluation.
Read the list as a set and the pattern is consistent: Grok 4.7 is meaningfully better where the task is long and agentic, and marginally better where the task is short. The Terminal-Bench jump is roughly a doubling; the editor-work gains are single-digit percentages. If your workload is autocomplete-shaped, you are paying the same $2/$6 for a small improvement. If your workload is "run for two hours and come back," the release is aimed directly at you.
What independent measurement says so far
One independent number exists, and it is worth stating carefully because it is easy to misread. Artificial Analysis scores Grok 4.7 at 46 on its Intelligence Index, version v4.3.2, ranking it 16th of 655 models on the board. Grok 4.6 sits at 44 on the same scale. A two-point gap on a composite index is not the doubling that Terminal-Bench shows, and that discrepancy is informative rather than contradictory: the index blends reasoning, knowledge, mathematics and coding, so a large gain in one agentic slice gets diluted by everything the model did not change.
Two further details from the same page are more useful than the headline score. First, Grok 4.7 generated 240 million output tokens while running the index, against a median of 92 million — it is a verbose reasoner, and on a $6-per-million output rate that verbosity is a line item. Second, the index's composite places it among the strongest models in its price tier, which is the comparison that actually matters for a buyer. Artificial Analysis has not published a populated price for Grok 4.7 on the model page, so do not take its cost-per-task figure from there; use the vendor's $2/$6.

The pricing cliff neither model fixed
Here is the part of the Grok 4.6 rate card that production teams learned to hate, and Grok 4.7 did not change it. Below 200,000 input tokens the rate is $2 in and $6 out. Above it, the entire request reprices to $4 in and $12 out. Not the excess tokens — the whole request. A 210,000-token call costs double what a 199,000-token call costs, for 11,000 extra tokens.
With a 500,000-token context window on both models, that cliff is easy to walk off. Long-context agent runs and whole-repository prompts are precisely the workloads Grok 4.7 was tuned for, which means the model's best use case is also the one most likely to trip the doubling. If you are moving work onto Grok 4.7 because it handles long agentic sessions better, budget for the repricing before you move it, not after.
There is also a speed tier to account for. SpaceXAI ships a fast variant of Grok 4.7 that doubles output speed and doubles the listed price. That is a legitimate trade for interactive coding, and a poor one for batch work — the same task at the same quality costs twice as much for a wall-clock saving nobody is watching.
Migrating from Grok 4.6 without a rewrite
The practical good news is that this is a model swap, not a platform migration. Both models take text and images in and return text, both use the same reasoning-effort levels from low through xhigh, and the request shape is the one SpaceXAI has been serving since Grok 4.5. Code that calls Grok 4.6 with an effort parameter does not need restructuring to call Grok 4.7.
Where the swap is not free is in evaluation. Grok 4.7's gains are concentrated in long agentic runs, so a test suite built from short single-turn prompts will show you almost nothing — you will see the CursorBench-sized improvement and conclude the upgrade was not worth the effort, when the case for it lives in tasks your suite never exercises. The measurement that would actually settle it is task completion over multi-step work: accepted patches, regressions introduced, tool calls per completed task, and how often a run needs human rescue. Those are the axes the vendor's own numbers point at, and they are the ones no published benchmark covers for your codebase.
Verbosity is the second thing to measure. A model that emits 240 million tokens on a standard index is going to emit more tokens on your workload than its predecessor did, and at $6 per million output that can quietly erase the value of a same-price upgrade. Track output tokens per completed task for a week before you commit.

Which one to call
The decision is genuinely simple, which is unusual for a model comparison. Grok 4.6 remains the correct choice for short-turn, cost-sensitive traffic: chat, classification, summarisation, and editor-scale code assistance, where Grok 4.7's gains are small and its verbosity is a tax. Grok 4.7 is the correct choice for the workload it was built for — long-horizon agentic coding and terminal work where a task runs for hours and self-verification is the difference between a finished job and a stuck one.
Running both is not a compromise here, it is the right answer, and it is cheap to do. Grok 4.6 is in OrcaRouter's catalog at SpaceXAI's list price with 0% markup passed straight through, so the $2/$6 rate card above is what you pay — and because we pass the provider's list price through rather than marking it up, any vendor price change is live on our side the same day it lands. One API key covers Grok 4.6 and 200-plus other models, so moving traffic between them by task is a config change rather than a second contract. If you want to trial Grok 4.7 on a production path before you trust it, the safer pattern is to keep the fallback on Grok 4.6 — a model you can route today — and let automatic failover across providers move the request back when the new model misbehaves.

What to watch next
Two things will settle this comparison, and neither has happened yet. The first is an independent reproduction of the Terminal-Bench 4.0 figure — 38% is a large enough jump that somebody will rerun it, and if it holds, the case for Grok 4.7 in agentic pipelines is strong on merit rather than on marketing. The second is the rest of the Grok roadmap: SpaceXAI has publicly sequenced Grok 4.8, Grok 4.9 and Grok 5 behind this release, and Grok 4.6's own life was about five weeks. Buying into Grok 4.7 as a long-term platform assumption would repeat the mistake teams made with Grok 4.5.
For now, the same-price upgrade is real and the direction of travel is clear. The number to hold onto is that $2/$6 bought you a model that roughly doubles on long-horizon terminal work and barely moves anywhere else — and that is a specific, useful, checkable claim rather than a generational leap.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
