A hero title card for an article comparing GPT-6 Sol Pro with Grok 4.7, reading 'GPT-6 Sol Pro vs Grok 4.7' with the subtitle 'Twice the verbosity, a third of the price'.
Guides & Insights

GPT-6 Sol Pro vs Grok 4.7: Twice the Verbosity, a Third of the Price

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The two cheapest frontier-adjacent models of September 2026 landed a day apart, and the interesting thing about them is not which one is smarter. GPT-6 Sol — the model people reach for when they search GPT-6 Sol Pro, which is that model with its reasoning.mode set to pro rather than a separate product — shipped on September 22, 2026 at $2.00 per million input tokens and $10.00 per million output tokens. Grok 4.7 from the vendor shipped on September 21 at the same $2.00 input, and $6.00 output. On Artificial Analysis's neutral harness the two land two index points apart and roughly two dollars per completed task apart, and the reason for that gap is not capability. It is how many tokens each one burns getting there.

Sol generated 77 million output tokens running the Intelligence Index. Grok 4.7 generated 240 million. That ratio — a little over three to one — is the entire comparison, and it inverts the conclusion a per-token price table gives you.

The two rate cards, and where the cheap one is not cheap

Both vendors publish these numbers themselves; neither has been independently repriced.

• Input — GPT-6 Sol $2.00 per million tokens vs Grok 4.7 $2.00 per million tokens

• Output — GPT-6 Sol $10.00 per million tokens vs Grok 4.7 $6.00 per million tokens

• Cached input — GPT-6 Sol $0.20 per million tokens vs Grok 4.7 $0.50 per million tokens; Sol's cache read is cheaper by a factor of two and a half, which is the opposite of what the headline output rate suggests

• Context window — GPT-6 Sol 1,050,000 tokens vs Grok 4.7 500,000 tokens

• Long-context surcharge — GPT-6 Sol reprices a request above 272,000 input tokens at 2× input and cache rates and 1.5× output for the whole request vs Grok 4.7 doubling every rate at 200,000 input tokens, also for the whole request

• Knowledge cutoff — GPT-6 Sol April 20, 2026 vs Grok 4.7 a pretraining cutoff reported as June 2026 with supplemental training through August

• Reasoning effort — GPT-6 Sol none, low, medium (default), high, xhigh, max vs Grok 4.7 low, medium (default), high, xhigh

• Modality — both accept text and image input and return text

The cache-read line is the one a buyer should sit with. Grok 4.7's output rate is 40% lower, but its cached input is 150% higher, and cached input is where long agent histories and repeated system prompts live. A workload that is mostly re-reading a large stable context will find the two models much closer than $6 against $10 implies.

A six-row scoreboard card titled 'GPT-6 Sol Pro vs Grok 4.7 - the scoreboard', comparing output price, cached input, context window, index tokens, AA Intelligence Index and cost per task. GPT-6 Sol Pro is listed at $10.00 output and $0.20 cached input per million tokens, a 1,050,000-token context, 77M index tokens, an AA Index of 48 and $1.06 per task; Grok 4.7 at $6.00 output and $0.50 cached input, a 500,000-token context, 240M index tokens, an AA Index of 46 and $3.74 per task. A footer reads 'Vendor list prices; index and cost figures per Artificial Analysis.'

The independent record, and the one number that decides it

Artificial Analysis is the only harness that has measured both models, and its figures are worth laying out side by side because the divergence is instructive.

• Intelligence Index — GPT-6 Sol 48 at max effort vs Grok 4.7 46 at xhigh; both well above the comparable-model median of 25

• Cost per index task — GPT-6 Sol $1.06 vs Grok 4.7 $3.74

• Output tokens for the index run — GPT-6 Sol 77M vs Grok 4.7 240M, against a median of 88M

• Output speed — Grok 4.7 measured at 39 tokens per second, which the harness itself labels notably slow; Sol's figure on the same board is not published in the same summary line

• Coding Agent Index — GPT-6 Sol 57 at $2.99 per task vs Grok 4.7 56 with its first-party Grok Build harness, up nine points from Grok 4.6

Two points of index difference, three and a half times the cost per task. That is not a pricing anomaly — it is the arithmetic of verbosity. Grok 4.7 is a model that thinks out loud at length; the harness flags it as "very verbose" against a median of 88M tokens, and the cost follows the tokens, not the rate card.

The coding-agent comparison carries a caveat that the scoreboard hides. Grok 4.7's 56 is measured inside xAI's own Grok Build harness, which is the configuration xAI ships and benchmarks against. Sol's 57 is measured in a neutral harness. A first-party harness advantage of a point or two is well within the range that harness choice explains, which means the honest reading is that these two are level on agentic coding — not that Sol is one point ahead.

Screenshot of the Artificial Analysis model page for Grok 4.7 (xhigh), captured 23 September 2026, showing list pricing of $2.00 per million input tokens and $6.00 per million output tokens with a 75% cache discount, an Intelligence Index score of 46, a cost per Intelligence Index task of $3.74, 240 million output tokens generated during the index run, a 500,000-token context window, and a measured output speed of 39 tokens per second that the page labels notably slow.

What each vendor claims, and what neither has shown

xAI's launch material is specific and it is a long-horizon story: Grok 4.7 is built on a larger base model than Grok 4.6 and was given a longer reinforcement-learning run aimed at tasks that "can take hours to complete," sharpening self-verification, long-context handling and behaviour inside the Grok Bot environment. Vendor-reported numbers include Terminal-Bench 4.0 at 38.0% against Grok 4.6's 20.3%, CursorBench 4.0 at 46.3%, and DeepSWE v1.1 at 71.0%. Every one of those is xAI's own measurement and none has been independently reproduced; the independent Terminal-Bench 4.0 figure that circulates is materially lower than the vendor's, which is the usual shape of that gap.

OpenAI's claims for GPT-6 Sol are the ones already covered above, and they carry the same label. The vendor-reported figures are 33.2% on AutomationBench at xhigh effort against Claude Opus 5's 26.9% at max, at roughly 9% of the cost per task, and 68.8% on DeepSWE v1.1 at max effort — within 1.1 points of Claude Fable 5 at xhigh, at about 80% lower cost per task. Note the comparison target: OpenAI's own charts pit Sol against Anthropic's models, not against Grok 4.7. Neither vendor has published a head-to-head against the other.

Screenshot of OpenAI's developer model page for GPT-6 Sol, captured 23 September 2026, showing the model id gpt-6-sol as the default snapshot, a 1,050,000-token context window with 128,000 maximum output tokens, an April 20, 2026 knowledge cutoff, text and image input with text output, and Standard pricing of $2.00 input, $0.20 cached input, $2.50 cache writes and $10.00 output per million tokens. No pro model id appears on the page.

Where the routing layer earns its place

OrcaRouter does not carry either model in this matchup — Grok 4.7 and GPT-6 Sol are both absent from our catalogue, so nothing here is a claim about their price through us. What a routing layer is worth in this particular pairing is the failure mode rather than the rate card. The reason to reach for Grok 4.7 is its long-horizon agentic behaviour; the reason not to reach for it is that a 39-token-per-second output speed makes a long agent run slow enough to matter, and a slow run is a run that can time out. Automatic failover across providers means a long job can be routed to whichever model is actually available without that speed becoming a single point of failure, and the routing DSL expresses the model choice as a rule inside the call rather than as a migration — which matters here, because the correct answer for this pair depends on whether the task is token-hungry or latency-sensitive, and that is a property of the request, not of the quarter.

The decision rule

• Agentic coding at a fixed budget — GPT-6 Sol. Level capability on the neutral coding index, at roughly a third of the cost per task, because it gets there with a third of the tokens.

• Long-horizon tasks where the run length is the point — Grok 4.7. The release was trained specifically for multi-hour work, and the vendor numbers on SWE-Marathon and CursorBench move in exactly that direction.

• Latency-sensitive volume — neither, on the current record. Sol's cost per task is the better number, but Grok 4.7's measured 39 tokens per second is a real constraint on anything interactive.

• Mostly-cached long-context workloads — GPT-6 Sol, on the cache-read line rather than the output line. At $0.20 against $0.50 per million cached tokens, and with a window twice the size, the case gets stronger the more of your prompt is stable.

• If your comparison was built from the per-token rate card — rebuild it from cost per task. $6.00 against $10.00 output is the least informative pair of numbers in this article, and it is the pair every existing page leads with.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily