A hero title card for an article comparing GPT-6 Sol with Grok 4.7, reading 'GPT-6 Sol vs Grok 4.7' with the subtitle 'The cheaper token that costs three times as much', showing the statistic '240M vs 77M output tokens' with Grok 4.7 at $6.00 output and $3.74 per task and GPT-6 Sol at $10.00 output and $1.06 per task.
Guides & Insights

GPT-6 Sol vs Grok 4.7: The Cheaper Token That Costs Three Times as Much

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Two models shipped one day apart in September 2026, and on the sticker they are not close. Grok 4.7 landed on September 21 at $2.00 per million input tokens and $6.00 per million output tokens. GPT-6 Sol landed on September 22 at the same $2.00 input and $10.00 output. Grok's output is 40% cheaper. That is the entire basis on which most comparisons will pick it, and it is the wrong number to look at.

The number that decides this pairing is not on either rate card. It is 240 million against 77 million — the output tokens each model generated while Artificial Analysis ran its Intelligence Index. Grok 4.7 burned a little over three times as many tokens to answer the same set of questions. Multiply that by the per-token rate and the model with the cheaper output ends up costing $3.74 per completed task against $1.06. Cheaper per token, more than three times the bill.

What each model actually is

Grok 4.7 is xAI's successor to Grok 4.6, positioned as its strongest coding and knowledge-work model, with a 500,000-token context window, text and image input, text output, and reasoning effort configurable across low, medium, high and xhigh. xAI has not disclosed a parameter count. It accepts a pretraining cutoff reported as June 2026 with supplemental training through August.

GPT-6 Sol is OpenAI's mid-tier GPT-6 model — below the flagship GPT-6 Astra, above the budget GPT-6 Luna — with a 1,050,000-token context window, a 922,000-token maximum input, a 128,000-token output ceiling and an April 20, 2026 knowledge cutoff. It is text and image in, text out, with reasoning effort from none through max, and OpenAI has described its pricing as permanent rather than promotional.

The context windows are not close. Sol's is roughly double. That matters for exactly one class of workload, and it is not the class most people run.

The rate card, and the line that inverts it

• Input — GPT-6 Sol $2.00 per million tokens vs Grok 4.7 $2.00 per million tokens; identical

• Output — GPT-6 Sol $10.00 per million tokens vs Grok 4.7 $6.00 per million tokens; Grok is 40% cheaper

• Cached input — GPT-6 Sol $0.20 per million tokens vs Grok 4.7 $0.50 per million tokens; Sol's cache read is cheaper by a factor of two and a half, which is the opposite of what the output rate implies

• Cache writes — GPT-6 Sol $2.50 per million tokens vs Grok 4.7 a rate not published on the same sheet

• Context window — GPT-6 Sol 1,050,000 tokens vs Grok 4.7 500,000 tokens

• Long-context surcharge — GPT-6 Sol reprices a request above 272,000 input tokens at 2× input and cache rates and 1.5× output for the whole request vs Grok 4.7 doubling every rate at 200,000 input tokens, also for the whole request

• Knowledge cutoff — GPT-6 Sol April 20, 2026 vs Grok 4.7 a pretraining cutoff reported as June 2026 with supplemental training through August

• Reasoning effort — GPT-6 Sol none, low, medium (default), high, xhigh, max vs Grok 4.7 low, medium (default), high, xhigh

• Modality — text and image in, text out, for both

The cache line is the one a buyer should sit with. Grok's output rate is lower, but its cached input is two and a half times Sol's, and cached input is where long agent histories and repeated system prompts live. A workload that is mostly re-reading a large stable context will find the two much closer than $6 against $10 suggests — and once the token-volume ratio is applied on top of that, the ordering reverses.

A six-row scoreboard card titled 'GPT-6 Sol vs Grok 4.7 - the scoreboard', comparing the two models on output price, cached input, context window, Intelligence Index score, index output tokens and cost per task. The left column lists GPT-6 Sol at $10.00 output, $0.20 cached input, a 1,050,000-token context, an Index of 48, 77M index tokens and $1.06 per task; the right column lists Grok 4.7 at $6.00 output, $0.50 cached input, a 500,000-token context, an Index of 46, 240M index tokens and $3.74 per task. A footer reads 'Vendor list prices; index figures per Artificial Analysis.'

Why 240 million tokens is the whole story

Artificial Analysis measured Grok 4.7 at xhigh generating roughly 240 million output tokens across its index run, against a median of about 88 million for the models on the same board. Its own summary calls the model "notably slow and very verbose." GPT-6 Sol, measured on the same harness, generated 77 million — described on the board as "fairly concise."

The board also reports the consequence directly. Grok 4.7 scores 46 on the Intelligence Index at $3.74 per task. At its high setting the score is also 46 and the cost is $2.73 per task. GPT-6 Sol scores 48 at $1.06 per task. Same index, same harness, and a two-to-three-and-a-half times cost difference that no rate card shows.

Two cautions before you treat those figures as a clean head-to-head. The two pages were captured on the same day but the boards carry different totals — Grok's page shows a 212-model class, and a separate Grok listing reports a rank out of 655 — so the two scores are not guaranteed to be on the same index build. And neither figure is a vendor number: Artificial Analysis runs its own harness, which is the point, but the reasoning settings differ between the two listings (xhigh for Grok, max for Sol), so treat the five-hundred-percent cost gap as the finding and the two-point index gap as approximate.

Screenshot of the Artificial Analysis model page for Grok 4.7 (xhigh), captured 23 September 2026, showing an Intelligence Index score of 46, list pricing of $2.00 per million input tokens and $6.00 per million output tokens with a 75% cache discount, a cost of $3.74 per Intelligence Index task, 240 million output tokens generated during the index run, a 500,000-token context window and 39.2 output tokens per second.

What xAI published, and what nobody has checked

xAI's own launch numbers are strong and, like all vendor launch numbers, unreproduced. CursorBench 4.0 at 46.3% at xhigh against Grok 4.6's 40.4%. Terminal-Bench 4.0 at 38.0% against 20.3% — the largest single gain in the release. DeepSWE v1.1 at 71.0% at high against 65.2%. EEBench at 64.0% against 53.0%. Those are xAI's harness, xAI's settings, xAI's reporting.

OpenAI's numbers for GPT-6 Sol follow the same pattern and deserve the same discount. The difference is that Sol has one more piece of independent evidence than Grok does: the Artificial Analysis cost-per-task figure, which is computed from measured token consumption rather than from a vendor's score table. That is the number that survives this comparison.

What nobody has published for either model is a head-to-head on the same tasks with the same harness and the same effort setting. If you see one, check which of those three variables moved.

Pricing a month of cached agent traffic

Take a workload that is mostly stable context: a coding agent with a 60,000-token system prompt and repository summary, re-read on 5,000 calls a month, each returning 4,000 output tokens. Assume caching works and the stable prefix bills at the cached rate.

On GPT-6 Sol: 300 million cached input tokens at $0.20 per million is $60.00. Twenty million output tokens at $10.00 per million is $200.00. Total, $260.00.

On Grok 4.7: the same 300 million cached input tokens at $0.50 per million is $150.00. Twenty million output tokens at $6.00 per million is $120.00. Total, $270.00.

Nearly identical — and that is with the token volume held constant, which is the generous assumption. Grok 4.7's measured verbosity on the index run was three times Sol's. Apply even a 1.5× multiplier to the output volume and Grok's bill goes to $330.00 against Sol's $320.00, and the gap widens from there. The cheaper output rate does not survive caching plus verbosity; it barely survives caching alone.

Nothing here is a claim about either model's price through OrcaRouter — neither is on our catalogue. GPT-6 Sol is reachable through OpenAI's own API and Grok 4.7 through xAI's; the rest of the Grok line, including Grok 4.6, is routable through us at xAI's list price under the pass-through model. The point of the arithmetic is that the escalation rule you write should be based on cost per completed task on your own traffic, not on the rate card — and a routing layer is what lets you test that rule without a second contract.

Screenshot of OrcaRouter's model page for Grok 4.6, captured 23 September 2026, showing the model id grok/grok-4.6 from provider SpaceXAI, a 500,000-token context window, text, image and file input, list pricing of $2.00 per million input tokens and $6.00 per million output tokens, and an OpenAI-compatible API served over both /v1/chat/completions and /v1/responses. Grok 4.7 itself does not appear on the catalogue.

Where each one belongs

• High-volume, cached-context agent work — GPT-6 Sol. The cached-input line is 2.5× in Sol's favour and the output-volume line is worse than that in Grok's, and they compound.

• Coding work where you want the strongest published Terminal-Bench improvement in the release — Grok 4.7 is the model with the number, and it is a big one. Just budget for the tokens, not the rate.

• Long documents above 500,000 tokens — GPT-6 Sol, and it is not a close call. Grok's window ends where Sol's is halfway.

• Latency-sensitive paths — neither. Artificial Analysis lists Grok 4.7 at 39.2 output tokens per second, which it describes as notably slow. Check Sol's own speed figure on the current board before you assume it is better.

• If you were shown a per-token comparison that ends with "so Grok is cheaper" — it ended one step too early. The next step is the token count.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily