
GPT-6.1 Sol vs Grok 4.7: The Cheapest Output Rate in the Room, and the Most Expensive Conversation
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 161 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 79 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 53 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 301 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Grok 4.7, xAI's successor to Grok 4.6, landed on 21 September 2026 at $2.00 per million input tokens and $6.00 per million output tokens. GPT-6.1 Sol, OpenAI's mid-tier refresh, landed eight days later on 29 September at $2.00 input and $10.00 output. The input rates are identical, the output rate is 40% cheaper on Grok 4.7, and that is where most comparisons stop. It is the wrong place to stop, because the same independent board has now measured both models end to end: Grok 4.7 burned roughly 240 million output tokens completing the Artificial Analysis Intelligence Index; GPT-6.1 Sol burned 67 million. Multiply the cheaper rate by three and a half times the volume and the model with the lower price per token costs $3.74 per completed task against $0.72.
Two models, eight days apart, and a rate card that inverts
Grok 4.7 is xAI's strongest coding and knowledge-work model: a 500,000-token context window, up to 450,000 output tokens in a single response according to the vendor's own listing, text, image and file input, and reasoning effort configurable across low, medium (default), high and xhigh. xAI has not disclosed a parameter count, and its pretraining cutoff is reported as June 2026 with supplemental training through August.
GPT-6.1 Sol is OpenAI's one-bump refresh of the GPT-6 Sol tier, positioned below GPT-6 Astra and above GPT-6 Luna: 1,050,000 tokens of context — roughly double Grok's — with a 128,000-token output ceiling that is roughly a third of Grok's, an April 30, 2026 knowledge cutoff, and the same text, image and file input. Its reasoning ladder runs low through max with a default of medium, and it no longer accepts none at all.
One line per dimension, both sides on each:
• Input — GPT-6.1 Sol $2.00 per 1M vs Grok 4.7 $2.00 per 1M; identical
• Output — GPT-6.1 Sol $10.00 per 1M vs Grok 4.7 $6.00 per 1M; Grok is 40% cheaper
• Cached input — GPT-6.1 Sol $0.10 per 1M vs Grok 4.7 $0.50 per 1M; Sol is five times cheaper, the opposite of what the output line implies
• Context window — 1,050,000 tokens vs 500,000
• Maximum output in one response — 128,000 tokens vs 450,000 claimed
• Long-context step — repriced whole-request above 272,000 input tokens at 2× input and cache and 1.5× output vs every rate doubled above 200,000 input tokens, also for the whole request
• Reasoning effort — low, medium (default), high, xhigh, max vs low, medium (default), high, xhigh
• Weights and parameters — closed, undisclosed vs closed, undisclosed; neither gives you a checkpoint
• Intelligence Index at maximum published effort — GPT-6.1 Sol 51.8 at max vs Grok 4.7 46.4 at xhigh
• Cost per index task, independent — $0.72 vs $3.74
• Output tokens on the index run — 67 million vs 240 million, against a board median near 81 million
Two lines in that list are the whole comparison. Grok 4.7's output rate is genuinely lower and its cache line is genuinely five times higher; and it generated three and a half times the tokens to be measured. The rate card, read alone, points one way. The measurement points the other, by a factor of five.

Where Grok 4.7 is actually ahead
It would be dishonest to run this article as a rout, because Grok 4.7 wins on real evaluations. Scored side by side at maximum published effort on Artificial Analysis v4.3.2 — Grok at xhigh, Sol at max, a configuration difference worth keeping in mind throughout:
• GDPval-AA v2.1 — Grok 4.7 Elo 1715.4 vs GPT-6.1 Sol Elo 1575.1. A 140-point Elo lead on economically valuable knowledge work, and the largest single independent win for the cheaper-rate-card model.
• AutomationBench-AA — Grok 4.7 0.6556 vs GPT-6.1 Sol 0.6487. Effectively a tie, nominally Grok's.
• SciCode — Grok 4.7 0.5741 vs GPT-6.1 Sol 0.5417. A 3.2-point lead on scientific coding.
• Terminal-Bench 4.0 — Grok 4.7 0.2576 vs GPT-6.1 Sol 0.5606. A 30-point deficit, and the biggest gap in the pairing.
• Humanity's Last Exam — Grok 4.7 0.4314 vs GPT-6.1 Sol 0.5292.
• AA-LCR v1.1, long-context reasoning — Grok 4.7 0.7667 vs GPT-6.1 Sol 0.83.
• AA-Omniscience — Grok 4.7 32.0 vs GPT-6.1 Sol 41.5.
• GDP.pdf — Grok 4.7 0.20 vs GPT-6.1 Sol 0.31.
• CritPt — Grok 4.7 0.1771 vs GPT-6.1 Sol 0.3171.
Three of nine evaluations go to Grok 4.7 or tie, including the one that is hardest to argue with: GDPval-AA is designed to price economically valuable professional work, and a 140-point Elo lead is not noise. If your workload is the kind GDPval was built to represent — document-heavy analysis, professional deliverables, judgment calls with a correct answer — the cheaper-rate-card model is also the better model on the neutral board, and the cost-per-task penalty is what you pay for it.
xAI's own launch numbers are large and, like all vendor launch figures, unreproduced. CursorBench 4.0 at 46.3% at xhigh against Grok 4.6's 40.4%; Terminal-Bench 4.0 at 38.0% against 20.3%, the biggest single claimed gain in the release; DeepSWE v1.1 at 71.0% at high against 65.2%; EEBench at 64.0% against 53.0%. Those are xAI's harness, xAI's settings, xAI's reporting — and note that the 38.0% Terminal-Bench 4.0 claim sits against the independent board's 0.2576 for the same model on the same benchmark, which is the kind of discrepancy you should expect between a vendor's tuned configuration and a neutral max-effort run.
OpenAI's figures for GPT-6.1 Sol deserve the same discount and have the same weakness. The one number in this comparison that neither vendor produced is the cost per completed task, because it is computed from measured token consumption rather than from a score table. That is the figure that survives.
Why 240 million tokens decides it
Artificial Analysis describes Grok 4.7's verbosity in its own summary, and the token count behind the index run is the evidence: 240 million output tokens against a board median near 81 million, roughly three times what a typical model on the same suite produces. GPT-6.1 Sol's 67 million sits well below the median. Put the two ratios together — 3.6× the volume at 0.6× the price — and the cheaper output rate is buying more tokens rather than a smaller bill, which is exactly what a $3.74-against-$0.72 cost-per-task difference looks like.
Caching changes the size of the effect but not its direction, and this is the detail that trips people up. Grok 4.7's cached input is $0.50 per million against Sol's $0.10 — five times more expensive on the line that long agent histories and repeated system prompts live on. A workload dominated by re-reading a large stable prefix therefore finds the two models closer than $6 against $10 suggests, and once Grok's verbosity is applied on top, further apart than the input rates suggest. The cache line and the token line point the same way here, which is unusual and makes this pairing cleaner than the rate cards imply.
The long-context clauses are worth pricing separately because they trigger at different points. GPT-6.1 Sol reprices an entire request above 272,000 input tokens; Grok 4.7 does the same above 200,000. On a 250,000-token call, Grok has already doubled — $4.00 and $12.00 with a $1.00 cache read — while Sol is still on its standard $2.00 and $10.00. Between 200,000 and 272,000 tokens the cheaper-rate-card model is the more expensive one by a wide margin, and that band is common in document work.
Where Grok genuinely wins the structure rather than the score is the output ceiling. A 450,000-token maximum response against Sol's 128,000 is a different class of artifact: an entire generated codebase, a long transcript, a book-length synthesis in one call. If you are hitting output caps today, that line decides the comparison and nothing in the cost arithmetic above applies — you cannot buy a capability the other model does not have at any rate.
Both are on one key, which is the useful part
Grok 4.7 sits on OrcaRouter at xAI's own list price — $2.00 and $6.00 per million tokens with a $0.50 cache read and the 200,000-token step to $4.00 and $12.00 passed through exactly as xAI lists it — and GPT-6.1 Sol landed on our routes on release day at OpenAI's price, $2.00 and $10.00 with cached input at $0.10 and the 272,000-token step unchanged. Nothing is added to either: the platform passes provider list price through rather than marking it up, which means a vendor price cut on either side is live here the same day it is live there, with no second invoice and no reseller in the middle.

For this particular pair that is worth more than convenience. The two models disagree about what they are good at — Grok 4.7 leads on professional-work Elo and on scientific coding, GPT-6.1 Sol leads on terminal execution, factual reliability and long-context retrieval — and the boundary between those workloads is visible in your own traffic before it is visible in a benchmark. Sending GDPval-shaped requests one way and long-horizon agent runs the other, behind one endpoint, with automatic failover across both and a routing rule you change in configuration rather than in a deployment, is a cheaper way to find that boundary than an evaluation project. Model fusion, where a panel of models answers together, is the other option the platform offers if you would rather not choose per request at all.

The call
• Long-output agentic and terminal work, cached-prefix loops — GPT-6.1 Sol. Thirty points on Terminal-Bench 4.0, a cache line five times cheaper and a fifth of the cost per completed task.
• Professional knowledge work, document analysis, judgment tasks — Grok 4.7 on the strength of a 140-point GDPval-AA Elo lead, with the token bill budgeted rather than assumed.
• Scientific coding — Grok 4.7, narrowly, on the board's SciCode split. Check it against your own suite; 3.2 points is inside the range a different prompt set can move.
• Single responses above 128,000 output tokens — Grok 4.7, and there is no arithmetic that substitutes for it.
• Requests between 200,000 and 272,000 input tokens — GPT-6.1 Sol, because Grok 4.7 has already doubled its whole request and Sol has not.
• Cheapest possible conversation — neither of these two. Both are frontier-priced models with frontier token appetites; the comparison is about which one finishes your task, not which one is cheap in the abstract.
• If a comparison ends with "Grok is cheaper per token" — it ended one step too early. The next step is the token count, and it is 240 million against 67.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
