A generated hero title card for an article titled 'Grok 4.7 vs Grok 4.6 — same price, different model', with the subtitle 'What the September 21 release actually changed' and stat chips reading Terminal-Bench 4.0 38.0% vs 20.3%, AA Index 46 vs 44, and $2/$6 on both.
Guides & Insights

Grok 4.7 vs Grok 4.6: Same $2/$6 Price, a Different Model Underneath

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

SpaceXAI shipped Grok 4.7 on September 21, 2026, and the first thing worth noticing about it is what did not change: the price. Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, exactly what Grok 4.6 has cost since it launched on August 12, 2026. Same rate card, same 500,000-token context window, same text-and-image input — and yet the two models are not close on the evaluations that matter for agentic work. On SpaceXAI's own published numbers, Grok 4.7 nearly doubles Grok 4.6 on Terminal-Bench 4.0, 38.0% against 20.3%. That is the whole shape of this comparison: a same-price generational swap where almost all of the gain lands in one place, and the parts of Grok 4.6 that annoyed production teams are still there.

What actually changed between the two models

SpaceXAI's own description of the release is narrow, and it is worth quoting the shape of it rather than the adjectives. Grok 4.7 is built on a larger base model than Grok 4.6, and it was given a longer reinforcement-learning run aimed at tasks that "can take hours to complete." The company says that run sharpened three things: self-verification, long-context handling, and behaviour inside the Grok Bot environment. There is no claim about a new architecture, a new modality, or a different serving stack.

That framing matters because it predicts where the benchmark gains show up, and they show up exactly there. A longer RL pass over hard multi-hour tasks moves terminal and repository work far more than it moves a knowledge quiz. If you were hoping Grok 4.7 would be a broader model than Grok 4.6, the release notes say otherwise — it is the same model shape, trained further into the hard end of agentic work.

The parameter story is the one piece of this that is not a vendor disclosure. Musk said in early September that Grok 4.7 carries roughly 2.1 trillion parameters against Grok 4.6's 1.5 trillion, a 40% increase, and that figure has been repeated everywhere since. It has never appeared on a SpaceXAI specification sheet. Treat it as a widely relayed claim about the base model, not a confirmed spec — and note that parameter counts say very little about serving cost or capability when the architecture is sparse, which the Grok line's is.

The benchmark gap, with the sourcing label attached

Every figure in this section comes from SpaceXAI's launch material. None of it has been independently reproduced at the time of writing, and the honest way to read it is as a claim set that happens to be unusually specific — which makes it checkable later, but does not make it verified now.

• Terminal-Bench 4.0 — Grok 4.7 38.0% (vendor-reported) vs Grok 4.6 20.3%. The single largest delta in the release, and the one that matches the "longer RL on harder tasks" claim.

• CursorBench 4.0 — Grok 4.7 46.3% (vendor-reported) vs Grok 4.6 40.4%. A real gain, but a modest one — under six points, on a benchmark closer to everyday editor work than to autonomous terminal sessions.

• DeepSWE v1.1 — Grok 4.7 71.0% at high reasoning effort (vendor-reported) vs Grok 4.6 65.2%. Again a solid, unspectacular improvement.

• EEBench — Grok 4.7 64.0% (vendor-reported). No Grok 4.6 figure was published alongside it, so this one tells you nothing about the upgrade.

• AA-Briefcase v1.1 — Grok 4.7 at 1,657 Elo (vendor-reported), on the professional knowledge-work evaluation.

Read the list as a set and the pattern is consistent: Grok 4.7 is meaningfully better where the task is long and agentic, and marginally better where the task is short. The Terminal-Bench jump is roughly a doubling; the editor-work gains are single-digit percentages. If your workload is autocomplete-shaped, you are paying the same $2/$6 for a small improvement. If your workload is "run for two hours and come back," the release is aimed directly at you.

What independent measurement says so far

One independent number exists, and it is worth stating carefully because it is easy to misread. Artificial Analysis scores Grok 4.7 at 46 on its Intelligence Index, version v4.3.2, ranking it 16th of 655 models on the board. Grok 4.6 sits at 44 on the same scale. A two-point gap on a composite index is not the doubling that Terminal-Bench shows, and that discrepancy is informative rather than contradictory: the index blends reasoning, knowledge, mathematics and coding, so a large gain in one agentic slice gets diluted by everything the model did not change.

Two further details from the same page are more useful than the headline score. First, Grok 4.7 generated 240 million output tokens while running the index, against a median of 92 million — it is a verbose reasoner, and on a $6-per-million output rate that verbosity is a line item. Second, the index's composite places it among the strongest models in its price tier, which is the comparison that actually matters for a buyer. Artificial Analysis has not published a populated price for Grok 4.7 on the model page, so do not take its cost-per-task figure from there; use the vendor's $2/$6.

A generated two-column comparison scoreboard for Grok 4.7 and Grok 4.6 with six rows: Terminal-Bench 4.0 38.0% vs 20.3%, CursorBench 4.0 46.3% vs 40.4%, DeepSWE v1.1 71.0% vs 65.2%, AA Intelligence Index 46 vs 44, context 500K on both, and price $2/$6 on both, with a footer noting all benchmark figures are vendor-reported and the index scores are per Artificial Analysis v4.3.2.

The pricing cliff neither model fixed

Here is the part of the Grok 4.6 rate card that production teams learned to hate, and Grok 4.7 did not change it. Below 200,000 input tokens the rate is $2 in and $6 out. Above it, the entire request reprices to $4 in and $12 out. Not the excess tokens — the whole request. A 210,000-token call costs double what a 199,000-token call costs, for 11,000 extra tokens.

With a 500,000-token context window on both models, that cliff is easy to walk off. Long-context agent runs and whole-repository prompts are precisely the workloads Grok 4.7 was tuned for, which means the model's best use case is also the one most likely to trip the doubling. If you are moving work onto Grok 4.7 because it handles long agentic sessions better, budget for the repricing before you move it, not after.

There is also a speed tier to account for. SpaceXAI ships a fast variant of Grok 4.7 that doubles output speed and doubles the listed price. That is a legitimate trade for interactive coding, and a poor one for batch work — the same task at the same quality costs twice as much for a wall-clock saving nobody is watching.

Migrating from Grok 4.6 without a rewrite

The practical good news is that this is a model swap, not a platform migration. Both models take text and images in and return text, both use the same reasoning-effort levels from low through xhigh, and the request shape is the one SpaceXAI has been serving since Grok 4.5. Code that calls Grok 4.6 with an effort parameter does not need restructuring to call Grok 4.7.

Where the swap is not free is in evaluation. Grok 4.7's gains are concentrated in long agentic runs, so a test suite built from short single-turn prompts will show you almost nothing — you will see the CursorBench-sized improvement and conclude the upgrade was not worth the effort, when the case for it lives in tasks your suite never exercises. The measurement that would actually settle it is task completion over multi-step work: accepted patches, regressions introduced, tool calls per completed task, and how often a run needs human rescue. Those are the axes the vendor's own numbers point at, and they are the ones no published benchmark covers for your codebase.

Verbosity is the second thing to measure. A model that emits 240 million tokens on a standard index is going to emit more tokens on your workload than its predecessor did, and at $6 per million output that can quietly erase the value of a same-price upgrade. Track output tokens per completed task for a week before you commit.

Screenshot of the Artificial Analysis model page for Grok 4.7 (xhigh), showing an Intelligence Index of 46 on index version v4.3.2, a 500k token context window, text and image input with text output, and a verbosity figure of 240M output tokens generated on the index.

Which one to call

The decision is genuinely simple, which is unusual for a model comparison. Grok 4.6 remains the correct choice for short-turn, cost-sensitive traffic: chat, classification, summarisation, and editor-scale code assistance, where Grok 4.7's gains are small and its verbosity is a tax. Grok 4.7 is the correct choice for the workload it was built for — long-horizon agentic coding and terminal work where a task runs for hours and self-verification is the difference between a finished job and a stuck one.

Running both is not a compromise here, it is the right answer, and it is cheap to do. Grok 4.6 is in OrcaRouter's catalog at SpaceXAI's list price with 0% markup passed straight through, so the $2/$6 rate card above is what you pay — and because we pass the provider's list price through rather than marking it up, any vendor price change is live on our side the same day it lands. One API key covers Grok 4.6 and 200-plus other models, so moving traffic between them by task is a config change rather than a second contract. If you want to trial Grok 4.7 on a production path before you trust it, the safer pattern is to keep the fallback on Grok 4.6 — a model you can route today — and let automatic failover across providers move the request back when the new model misbehaves.

Screenshot of the OrcaRouter model page for Grok 4.6 (model ID grok/grok-4.6), showing SpaceXAI's list pricing of $2.00 per million input tokens and $6.00 per million output tokens, a 500K token context window, and the reasoning and vision capability tags.

What to watch next

Two things will settle this comparison, and neither has happened yet. The first is an independent reproduction of the Terminal-Bench 4.0 figure — 38% is a large enough jump that somebody will rerun it, and if it holds, the case for Grok 4.7 in agentic pipelines is strong on merit rather than on marketing. The second is the rest of the Grok roadmap: SpaceXAI has publicly sequenced Grok 4.8, Grok 4.9 and Grok 5 behind this release, and Grok 4.6's own life was about five weeks. Buying into Grok 4.7 as a long-term platform assumption would repeat the mistake teams made with Grok 4.5.

For now, the same-price upgrade is real and the direction of travel is clear. The number to hold onto is that $2/$6 bought you a model that roughly doubles on long-horizon terminal work and barely moves anywhere else — and that is a specific, useful, checkable claim rather than a generational leap.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily