
Gemini 4 Argon vs GPT-6 Sol: A Fifth of the Price, Four Times the Bill
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 118 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Run the Artificial Analysis index suite on Gemini 4 Argon and it bills $2,407. Run the same suite on GPT-6 Sol and it bills $1,546. Same index, same tasks, same week. The two models have identical list prices — $2.00 per million input tokens and $10.00 per million output tokens, Argon as a stated introductory rate from Google and Sol as OpenAI's standing rate — and yet the finished cost of doing identical work differs by 56%, in favour of the model that scores five points lower. That gap is the whole reason this comparison is worth writing, and it has nothing to do with the rate card.
Google announced Gemini 4 Argon on 2026-09-30, calling it the next era of frontier intelligence, priced at $2 in and $10 out with cached input at 95% off, rolling out to trusted cyber defenders through the Fairwind Program. GPT-6 Sol has been generally available since 2026-09-22, when OpenAI shipped the middle tier of the GPT-6 line at $2 in and $10 out with a 90% cache discount. One is buyable today on an OpenAI-compatible endpoint and the other is not buyable by anyone outside a named cohort, which is the practical fork in this article. But the cost inversion above survives even if you set availability aside entirely, and understanding why is the useful part.
The inversion, stated plainly
Artificial Analysis publishes both an Intelligence Index and an accounting of what it cost to run that index. Read together they say this:
• Intelligence Index — Gemini 4 Argon 52.6 against GPT-6 Sol 47.5. A five-point gap, measured with Argon charted at its high effort setting and Sol at max, so the comparison is generous to Sol rather than to Argon.
• List price per million tokens — identical. $2.00 in / $10.00 out on both. Cache reads are the one difference: $0.10 on Argon, $0.20 on Sol.
• Cost per index task — Gemini 4 Argon $1.99 against GPT-6 Sol $1.05. Argon costs roughly twice as much per task despite costing the same per token.
• Cost to run the full suite — Gemini 4 Argon $2,407 against GPT-6 Sol $1,546.
• Measured output speed — GPT-6 Sol 77.0 tokens per second against Gemini 4 Argon 48.9, a 1.6× difference that shows up as latency as well as spend.
There is one line in that list that explains all the others, and it is not the price. It is the token count. On the index's own tasks, Gemini 4 Argon generated roughly 561,000 output tokens per task; GPT-6 Sol generated roughly 282,000. Argon thinks for twice as long to answer the same question, and at an identical per-token rate that is exactly twice the bill.

Why the tokens are the price
This is the correction worth internalising before anyone budgets a migration, because the intuition runs the wrong way. When a model is priced the same as its competitor and consumes twice the output tokens to finish, the per-token price is not the price. The price is per-token multiplied by however many tokens the model decides to spend, and that second factor is a property of the model, not the rate card.
The margin by task varies enormously, which makes it worse — you cannot apply a flat multiplier and move on:
• Terminal-Bench 4.0 — Argon 165,330 output tokens per task against Sol's 97,372. Argon costs $6.78 a task here against Sol's $4.00, even though Argon scores 57.1% to Sol's 43.9%.
• CritPt — Argon 109,533 tokens against Sol's 31,848. A 3.4× token ratio on a task where Sol actually scores higher, 30.9% to 27.1%.
• GDPval — Argon 74,571 tokens against Sol's 41,567, for $1.82 a task against $1.32.
• Omniscience — a 938-versus-2,662 token ratio in Argon's favour, at $0.010 a task against Sol's $0.027. The one task family where the direction flips.
• Long-context reasoning — Argon 2,197 tokens against Sol's 954, and the only task in the suite where Sol clearly wins on score, 83.7% to 79.7%.
CritPt is the instructive one. A model that burns three and a half times the tokens to produce a slightly worse answer on the same question is not a capability story; it is a spending-habit story. Argon's reasoning runs long, and on some task shapes that length converts into score and on others it simply converts into dollars. The index's 52.6 and Sol's 47.5 are the aggregate, and aggregates hide exactly this.
Where the five points actually come from
Since the index gap and the cost gap point in opposite directions, it is worth knowing which tasks drive each.
• Terminal-Bench 4.0 — Gemini 4 Argon 57.1% versus GPT-6 Sol 43.9%. Argon's largest single win, and the task family closest to long-horizon agentic coding.
• Omniscience — Gemini 4 Argon 42.4 versus GPT-6 Sol 27.1. A fifteen-point spread on factual coverage, the largest proportional gap in the suite.
• AutomationBench-AA — Gemini 4 Argon 77.5% versus GPT-6 Sol 61.6%.
• SciCode — Gemini 4 Argon 61.8% versus GPT-6 Sol 57.6%.
• Humanity's Last Exam — Gemini 4 Argon 57.1% versus GPT-6 Sol 47.9%.
• GDPval — Gemini 4 Argon 1,611 versus GPT-6 Sol 1,487.
• ApexAgents / AA-Briefcase — GPT-6 Sol 1,482.78 Elo against Argon's 1,493.84, an 11-point gap that is inside the published confidence intervals and should be called a tie.
• Long-context reasoning — GPT-6 Sol 83.7% versus Gemini 4 Argon 79.7%. Sol's one clear win.
• CritPt — GPT-6 Sol 30.9% versus Gemini 4 Argon 27.1%. Sol wins again, narrowly.
So the shape is not "Argon is better". It is that Argon is better at the two things its launch materials led with — end-to-end business task execution and long-horizon software engineering — and Sol is level or ahead on the academic-flavoured reasoning ladder and on holding coherence across a long context. For a reader whose workload is a retrieval pipeline with a 400K-token prompt, Sol is simultaneously the better model and the cheaper one, which is an unusual and comfortable position to be in.
The specification differences that outlast the benchmark discussion
• Maximum output — Gemini 4 Argon 1,000,000 tokens in a single trajectory, which Google describes as the largest in the industry and an increase from the previous generation's 64K cap. GPT-6 Sol 128,000 tokens.
• Context window — GPT-6 Sol's vendor card states roughly 1,050,000 tokens; Argon's is 1,000,000. Artificial Analysis measured Sol at 872,000 tokens through its harness, which is a harness measurement and not a card figure — treat the card as the specification and the harness number as what the evaluator actually exercised.
• Input modalities — both take text, images and files. Argon's launch material also leans on long-video understanding, where Google claims 91.7% on LVBench and describes it as state of the art; that is a vendor-reported figure and no independent board reproduces it yet.
• Video input — Soft. Artificial Analysis charts Argon with video input disabled and names video explicitly at launch, while Sol's card does not list video at all. If video is load-bearing in your pipeline, neither card settles it and you should test before you architect.
• Reasoning effort — Sol exposes configurable effort (none through max, medium default) at the API and was charted at max. Argon was charted at high; nobody has published the same suite at its top setting.
The 1M-token output ceiling is the one that could genuinely change an architecture rather than a budget. Anything you currently split into a dozen chained calls to stay under a 128K cap — a multi-file patch set, a full audit report, a long document in one pass — becomes one call with fewer places for an agent loop to lose state. Whether that is worth twice the per-task cost depends entirely on whether you have that workload. If you do, it is probably worth it. If your calls are classify-and-return, it is money spent on headroom nobody will occupy.
The effort setting nobody matched, and why it matters
One caveat deserves its own paragraph because it cuts against reading the five-point index gap as a pure capability difference. Artificial Analysis charts Gemini 4 Argon at its high reasoning-effort setting and GPT-6 Sol at max. Those are not the same rung on the two ladders, and the direction of the mismatch is interesting: Sol was run at its most expensive setting and Argon at a middling one.
Two readings follow and the data cannot separate them. Either Argon at max would score higher than 52.6 — but also burn more tokens than 561,000 per task and cost more than $1.99 — or it would score about the same, which would mean the effort dial is not doing much on this suite. There is no published run that settles it. Anyone quoting 52.6 as Argon's ceiling is quoting a configuration, not a model, and the configuration is the one its vendor chose to publish first.
The same logic applies to Sol's 47.5, except in the opposite direction: max is the top of its ladder, so the figure is closer to a ceiling than a floor. A fair fight at matched effort could narrow the gap, and could also flip the cost comparison, since the tokens that max burns are the tokens that cost $10 per million.
The availability fork, which decides the actual answer
Gemini 4 Argon is not generally available. Google's announcement says it is rolling out to a set of trusted cyber defenders through the Fairwind Program, that the company is engaged in the US government's voluntary pre-release model access process, and that general access to developers, enterprises and consumers comes "as soon as possible", starting with paid API customers and Google AI Ultra subscribers. There is no model identifier, no endpoint, no quota and no date. It is not on any public API, ours included — check the catalogue and it 404s, because it does not exist there yet.
GPT-6 Sol is a callable product. On OrcaRouter it is openai/gpt-6-sol, on the same OpenAI-compatible endpoint as the other 200-plus models in the catalogue, at OpenAI's list price passed through with zero markup, so the $2/$10 above and the 272,000-token long-context step to $4/$15 are what you actually pay rather than what a rate card claims. Automatic failover across providers covers the case where the route degrades mid-run, and the routing DSL lets a single application send its cheap classification traffic to a smaller model and its hard reasoning traffic to Sol without a second SDK.

That last arrangement is the honest answer to this comparison for most teams. The cost argument for Argon is real but conditional on a workload where long-horizon agent work dominates; the cost argument against it is real on every other workload; and you cannot buy it this month anyway. The cheap move is to build the evaluation harness now — your own prompts, your own scoring, run against Sol at a few effort settings — so that when Argon reaches paid API customers you have a measured answer to a question everyone else will be answering from a leaderboard.
What would settle it
Three things, in order of how much they would move the verdict. General availability of Argon at a real, non-introductory price — the word "introductory" means the $2/$10 can move, and Google's own sequence puts paid API customers first, so the first published rate after launch is the one to watch. A published Artificial Analysis run of Argon at its top effort setting, which would say whether 52.6 is a ceiling or a hair off it. And one independent reproduction of the DeepSWE v1.1 figure — Google claims 77.9% and calls it a new state of the art; it is vendor-reported and unreproduced outside Google as of today.
Until then the split is clean and slightly ironic. GPT-6 Sol is the better-verified model, the faster one, the cheaper one per finished task, and the one you can call this afternoon. Gemini 4 Argon is the higher-scoring model on the board its vendor chose to publish, roughly twice as expensive to actually use, and not yet for sale. Choosing between them is not a judgement about capability. It is a judgement about which of those two sentences describes your month.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
