
GPT-6 Sol Pro vs Grok 4.7: Twice the Verbosity, a Third of the Price
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 312 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 112 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
The two cheapest frontier-adjacent models of September 2026 landed a day apart, and the interesting thing about them is not which one is smarter. GPT-6 Sol — the model people reach for when they search GPT-6 Sol Pro, which is that model with its reasoning.mode set to pro rather than a separate product — shipped on September 22, 2026 at $2.00 per million input tokens and $10.00 per million output tokens. Grok 4.7 from the vendor shipped on September 21 at the same $2.00 input, and $6.00 output. On Artificial Analysis's neutral harness the two land two index points apart and roughly two dollars per completed task apart, and the reason for that gap is not capability. It is how many tokens each one burns getting there.
Sol generated 77 million output tokens running the Intelligence Index. Grok 4.7 generated 240 million. That ratio — a little over three to one — is the entire comparison, and it inverts the conclusion a per-token price table gives you.
The two rate cards, and where the cheap one is not cheap
Both vendors publish these numbers themselves; neither has been independently repriced.
• Input — GPT-6 Sol $2.00 per million tokens vs Grok 4.7 $2.00 per million tokens
• Output — GPT-6 Sol $10.00 per million tokens vs Grok 4.7 $6.00 per million tokens
• Cached input — GPT-6 Sol $0.20 per million tokens vs Grok 4.7 $0.50 per million tokens; Sol's cache read is cheaper by a factor of two and a half, which is the opposite of what the headline output rate suggests
• Context window — GPT-6 Sol 1,050,000 tokens vs Grok 4.7 500,000 tokens
• Long-context surcharge — GPT-6 Sol reprices a request above 272,000 input tokens at 2× input and cache rates and 1.5× output for the whole request vs Grok 4.7 doubling every rate at 200,000 input tokens, also for the whole request
• Knowledge cutoff — GPT-6 Sol April 20, 2026 vs Grok 4.7 a pretraining cutoff reported as June 2026 with supplemental training through August
• Reasoning effort — GPT-6 Sol none, low, medium (default), high, xhigh, max vs Grok 4.7 low, medium (default), high, xhigh
• Modality — both accept text and image input and return text
The cache-read line is the one a buyer should sit with. Grok 4.7's output rate is 40% lower, but its cached input is 150% higher, and cached input is where long agent histories and repeated system prompts live. A workload that is mostly re-reading a large stable context will find the two models much closer than $6 against $10 implies.

The independent record, and the one number that decides it
Artificial Analysis is the only harness that has measured both models, and its figures are worth laying out side by side because the divergence is instructive.
• Intelligence Index — GPT-6 Sol 48 at max effort vs Grok 4.7 46 at xhigh; both well above the comparable-model median of 25
• Cost per index task — GPT-6 Sol $1.06 vs Grok 4.7 $3.74
• Output tokens for the index run — GPT-6 Sol 77M vs Grok 4.7 240M, against a median of 88M
• Output speed — Grok 4.7 measured at 39 tokens per second, which the harness itself labels notably slow; Sol's figure on the same board is not published in the same summary line
• Coding Agent Index — GPT-6 Sol 57 at $2.99 per task vs Grok 4.7 56 with its first-party Grok Build harness, up nine points from Grok 4.6
Two points of index difference, three and a half times the cost per task. That is not a pricing anomaly — it is the arithmetic of verbosity. Grok 4.7 is a model that thinks out loud at length; the harness flags it as "very verbose" against a median of 88M tokens, and the cost follows the tokens, not the rate card.
The coding-agent comparison carries a caveat that the scoreboard hides. Grok 4.7's 56 is measured inside xAI's own Grok Build harness, which is the configuration xAI ships and benchmarks against. Sol's 57 is measured in a neutral harness. A first-party harness advantage of a point or two is well within the range that harness choice explains, which means the honest reading is that these two are level on agentic coding — not that Sol is one point ahead.

What each vendor claims, and what neither has shown
xAI's launch material is specific and it is a long-horizon story: Grok 4.7 is built on a larger base model than Grok 4.6 and was given a longer reinforcement-learning run aimed at tasks that "can take hours to complete," sharpening self-verification, long-context handling and behaviour inside the Grok Bot environment. Vendor-reported numbers include Terminal-Bench 4.0 at 38.0% against Grok 4.6's 20.3%, CursorBench 4.0 at 46.3%, and DeepSWE v1.1 at 71.0%. Every one of those is xAI's own measurement and none has been independently reproduced; the independent Terminal-Bench 4.0 figure that circulates is materially lower than the vendor's, which is the usual shape of that gap.
OpenAI's claims for GPT-6 Sol are the ones already covered above, and they carry the same label. The vendor-reported figures are 33.2% on AutomationBench at xhigh effort against Claude Opus 5's 26.9% at max, at roughly 9% of the cost per task, and 68.8% on DeepSWE v1.1 at max effort — within 1.1 points of Claude Fable 5 at xhigh, at about 80% lower cost per task. Note the comparison target: OpenAI's own charts pit Sol against Anthropic's models, not against Grok 4.7. Neither vendor has published a head-to-head against the other.

Where the routing layer earns its place
OrcaRouter does not carry either model in this matchup — Grok 4.7 and GPT-6 Sol are both absent from our catalogue, so nothing here is a claim about their price through us. What a routing layer is worth in this particular pairing is the failure mode rather than the rate card. The reason to reach for Grok 4.7 is its long-horizon agentic behaviour; the reason not to reach for it is that a 39-token-per-second output speed makes a long agent run slow enough to matter, and a slow run is a run that can time out. Automatic failover across providers means a long job can be routed to whichever model is actually available without that speed becoming a single point of failure, and the routing DSL expresses the model choice as a rule inside the call rather than as a migration — which matters here, because the correct answer for this pair depends on whether the task is token-hungry or latency-sensitive, and that is a property of the request, not of the quarter.
The decision rule
• Agentic coding at a fixed budget — GPT-6 Sol. Level capability on the neutral coding index, at roughly a third of the cost per task, because it gets there with a third of the tokens.
• Long-horizon tasks where the run length is the point — Grok 4.7. The release was trained specifically for multi-hour work, and the vendor numbers on SWE-Marathon and CursorBench move in exactly that direction.
• Latency-sensitive volume — neither, on the current record. Sol's cost per task is the better number, but Grok 4.7's measured 39 tokens per second is a real constraint on anything interactive.
• Mostly-cached long-context workloads — GPT-6 Sol, on the cache-read line rather than the output line. At $0.20 against $0.50 per million cached tokens, and with a window twice the size, the case gets stronger the more of your prompt is stable.
• If your comparison was built from the per-token rate card — rebuild it from cost per task. $6.00 against $10.00 output is the least informative pair of numbers in this article, and it is the pair every existing page leads with.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
