
GPT-6.1 Sol vs Gemini 4 Argon: Half a Point Apart, and Only One of Them Is For Sale
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 161 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 79 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 300 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Everything you can compare between GPT-6.1 Sol and Gemini 4 Argon is measurable, and none of it is actionable, because only one of the two models can be bought. Gemini 4 Argon was announced on September 30, 2026 in a DeepMind post credited to Koray Kavukcuoglu, priced at a stated introductory $2.00 per million input and $10.00 per million output tokens, and rolling out first to a named cohort of cyber defenders through the Fairwind Program. It has no model identifier, no API endpoint, no published per-token rate you can be billed against, and no page you can reach as a customer. GPT-6.1 Sol shipped September 29, 2026 at $2.00 and $10.00 and is available now. On the Artificial Analysis Intelligence Index the two are 0.8 points apart — Argon 52.6, Sol 51.8 — which makes this the closest pairing in this batch and, on the index alone, a dead heat. On the meter that decides real deployments, one of these models is not a candidate at all.
That asymmetry is not a technicality, and it is not a scoop either. It is the fact that should govern how the comparison is read: Argon's numbers describe a model in a controlled rollout to one cohort, and GPT-6.1 Sol's describe a product on a price list. Comparing them is still worth doing — Argon is what the vendor currently leads with, and its measurements say something real about where the top of the Gemini line sits — but every conclusion has to carry the phrase "if it were purchasable", and none of them carry the phrase "should you switch".
The index allows exactly one comparison, and it is a near-tie

Both models appear on Artificial Analysis, and both carry a score and an accounting of what producing that score cost.
• Intelligence Index, v4.3.2 — Gemini 4 Argon 52.6 vs GPT-6.1 Sol 51.8, a 0.8-point gap
• Cost per index task — Gemini 4 Argon $1.99 vs GPT-6.1 Sol $0.72, a 2.8× spread for those 0.8 points
• Cost to run the full suite — Gemini 4 Argon $2,407 vs GPT-6.1 Sol $1,082
• Output tokens emitted on that suite — Gemini 4 Argon 110M vs GPT-6.1 Sol 67M, a 1.6× verbosity difference
• Stated introductory price — $2.00 in and $10.00 out per million tokens on both, with cached input at 95% off on Argon
• Context window — GPT-6.1 Sol 1,050,000 tokens; Argon's not published
The fourth line explains the second. A 2.8× per-task cost difference on near-identical headline rates comes from writing 1.6× as many tokens at a comparable output price, plus whatever the reasoning volume behind them costs. Argon is not an expensive model by its card. It is a talkative one on the evaluation, and the account is settled in output.
The configurations differ, and the direction of the difference flatters the cheaper model. Artificial Analysis charts Gemini 4 Argon at its high effort setting and GPT-6.1 Sol at max — so Sol is measured at its most expensive setting and Argon at something below its ceiling. The 0.8-point gap is therefore the generous version of the comparison for Sol: give Argon its top setting and the gap likely widens, give Sol a lower setting and its cost falls further. Neither move changes the availability line, which is the only line that ends a deployment decision.
Reading a 0.8-point gap at all
A sub-point difference on a ten-evaluation composite is not a ranking. It is two models in the same band.
What the index does not publish for Argon is the component detail that would make the band legible — which evaluations it wins, where it is weak, how it behaves across the sub-scores that make up the composite. GPT-6.1 Sol's page carries those rows; Argon's carries a headline and a cost. A comparison built on the headline alone can say the two are level and nothing else, and it should say exactly that rather than manufacture a winner from a 0.8-point margin.
There is one external data point worth adding, and it points the other way. In a third-party coding evaluation circulated after the announcement, Argon placed third behind GPT-6 Astra and Claude Opus 5.5 — a strong result for a model in a controlled rollout, and one that sits consistently with a 52.6 on the composite. It is also a single observation from a single evaluation, published in the first days of an announcement, which makes it evidence rather than a track record. Label it and move on.
What the two price cards are actually for
The stated rate for Argon deserves more care than the index row, because there are two numbers in it.
The introductory rate is $2.00 in and $10.00 out with cached input at 95% off, which on a card would make Argon a like-for-like price match for GPT-6.1 Sol. The vendor's own footnote states that after an introductory period it does not define, the rate becomes $4.00 per million input and $20.00 per million output tokens. That second pair is not a hypothetical — it is the published number the model is expected to settle at — and it lands exactly on the current frontier list price rather than below it. A comparison that quotes only the introductory pair is quoting a promotional number for a model not yet generally available, which is two degrees of provisional on one line.
GPT-6.1 Sol's card is the opposite kind of object. $2.00 and $10.00 is the standing rate, not a promotion; the long-context step to $4.00 / $15.00 above 272,000 prompt tokens is published and applies to a defined boundary; cached input at $0.10 is live today and billable. There is nothing in it that requires a footnote about a period the vendor declines to define.
That is the substance behind the availability line. Not that Argon is a worse model — the index says the two are level — but that one of these two cards is a commitment and the other is an intention.
The rollout itself is the comparison

The Fairwind Program is the part of Argon's announcement that a technical comparison tends to skip, and it is the part that matters most here.
The vendor is putting the model in the hands of a named cohort of cyber defenders first, ahead of everyone else, with no announced date for a general opening, no developer waitlist, and no published identifier that a customer could pin. That is a legitimate way to release a frontier model and it is not unusual for capability of this class. It also means every claim about Argon's cost, latency, throughput and reliability on real traffic is currently unmeasured: there is no production traffic to measure. The vendor's own benchmark run exists, and Artificial Analysis's composite exists, and between them they are one lab's evaluation and one evaluator's suite. That is the whole record.
GPT-6.1 Sol's record is eight days old and thin in a different way: one independent configuration, no practitioner corpus, and a shipping product whose failure modes are not yet written down anywhere. Neither model is well-evidenced. One of them, however, has an invoice.
So what is this comparison for
Three uses, and none of them is a purchase decision.
First, calibration. If you are tracking where Argon sits against GPT-6.1 Sol, 52.6 against 51.8 at a 2.8× per-task cost spread is the current answer, and it says the Gemini 4 line is competitive on capability while still settling its commercial terms. Watch for the introductory rate to be replaced by the $4.00 / $20.00 pair, which would turn a price match into a price premium at unchanged capability.
Second, a holding position. A model that is announced, measured and unroutable is the clearest case there is for keeping an abstraction between your application and your model choice. Both models in this piece are reachable the same way from our side of the fence: GPT-6.1 Sol is on OrcaRouter at its own $2.00 / $10.00 with cached input at $0.10 and nothing added, so a workload can be built against it today at provider list price. If and when Argon gets an identifier and a rate card, evaluating it should be a config change rather than a rewrite — and until then, the correct amount of engineering to spend on Argon is the amount that keeps that true.
Third, a falsifiable watch list. Any of the following would turn this from a what-we-know-so-far piece into a shoppable comparison: a model identifier in the developer documentation, an entry on its model-card index, a general-availability post with published per-token rates, or Artificial Analysis publishing a second configuration at a different effort level. Until one of those lands, a piece claiming Argon is cheaper, faster or better than GPT-6.1 Sol in production is describing a model nobody outside one cohort has run.

The honest close is a single sentence, and it is about the shape of the question rather than the score. Gemini 4 Argon and GPT-6.1 Sol are close enough on the only independent board both appear on that the index cannot choose between them; what chooses between them is that one is available and one is a promise, and a promise is not something you route production traffic to. Track Argon, adopt Sol if the price and the modalities fit your work, and keep the abstraction thin enough that the day Argon becomes real, it is a line of configuration rather than a project.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
