
Gemini 4 Argon vs GPT-6 Astra: A Dead Heat on the Index, and a Price Step Two-Thirds of the Way Down
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 143 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 125 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 933 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 50 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 106 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 217 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Two frontier models, 0.11 points apart on Artificial Analysis's Intelligence Index. Gemini 4 Argon scores 52.56. GPT-6 Astra scores 52.67. If you had to choose between them on measured general capability, you could flip a coin, and the difference between the two scores is smaller than the confidence interval on either. That is a genuinely rare situation in a market where the top ten models span about fifteen points in total, and it makes this the easiest of Gemini 4 Argon's matchups to reason about, because the capability argument is over before it starts and everything that is left is economics and access.
The two are not twins, though. Argon was announced on 2026-09-30 by Google DeepMind and is rolling out in phases, starting with trusted cyber defenders in the Fairwind Program. GPT-6 Astra shipped in early September — our catalogue card dates it 2026-09-04, Artificial Analysis lists 2026-09-03 — and is a live route you can call this afternoon at $10 in and $50 out per million tokens, with a 1,050,000-token context window and a 128,000-token output ceiling. One of these models has a price you can be billed against today.
Where a 0.11-point gap is real and where it is noise
An index is a composite, so two models that tie on the composite can differ meaningfully on its parts. Here the parts mostly agree, with three exceptions worth naming.
• Terminal-Bench 4.0 — GPT-6 Astra 59.1% versus Gemini 4 Argon 57.1%. Astra leads, narrowly.
• Long-context reasoning — GPT-6 Astra 80.7% versus Gemini 4 Argon 79.7%.
• GPQA Diamond — GPT-6 Astra 96.1%. Argon publishes no figure on this board.
• CritPt — GPT-6 Astra 31.7% versus Gemini 4 Argon 27.1%.
• SciCode — Gemini 4 Argon 61.8% versus GPT-6 Astra 56.5%.
• Humanity's Last Exam — Gemini 4 Argon 57.1% versus GPT-6 Astra 54.7%.
• AutomationBench-AA — Gemini 4 Argon 77.5% versus GPT-6 Astra 68.5%.
• GDPval — Gemini 4 Argon 1,611 versus GPT-6 Astra 1,542.
• Omniscience — GPT-6 Astra 43.4 versus Gemini 4 Argon 42.4.
• ITBench-SRE — GPT-6 Astra 48.6% against no published Argon figure.
• Output speed — GPT-6 Astra 51.2 tokens per second versus Gemini 4 Argon 48.9. Effectively level.
• First-token latency on the index harness — Gemini 4 Argon 2.8 seconds versus GPT-6 Astra 320.2 seconds.
Astra holds advantages on scientific multiple-choice reasoning, on critical-point physics problems and on the SRE benchmark. Argon holds advantages on scientific coding, on the harder end-to-end task set and on GDP-valued work. Neither lead is large enough to be the basis of a migration on its own.

The latency line is a different matter. 320 seconds to first token is five and a half minutes of silence before anything is emitted, against under three seconds for Argon. Both are measuring the same thing — time to the first chunk on Artificial Analysis's index runs — so this is comparable. Astra's reasoning is always on and configurable from low through max effort; a max-effort run on a hard task genuinely does think for minutes before it speaks. Whether that is a defect depends entirely on your interface. For a batch job it costs nothing. For anything with a user watching a spinner, it is the most user-visible difference between these two models, larger than any benchmark gap on the page.
The price step, which is the real subject
This is where the tie breaks, and it breaks in a way that is easy to mis-model because Astra's rate card has three tiers that look alike.
• Standard, up to 272,000 input tokens — GPT-6 Astra $10.00 in / $50.00 out, cached reads at $1.00, cache writes at $12.50.
• Long context, above 272,000 input tokens — GPT-6 Astra $20.00 in / $75.00 out, cached reads at $2.00. The whole request reprices at the higher tier, not just the overflow: 2× input and 1.5× output.
• Fast mode — 2× the applicable standard rate, so $20.00 in / $100.00 out. This is the $100 figure that gets quoted as if it were the long-context rate; it is not.
• Gemini 4 Argon — $2.00 in / $10.00 out flat on the introductory basis, cached input at $0.10, with no published step and no published fast tier.
So Argon is five times cheaper on input and five times cheaper on output than Astra's standard tier, and ten times cheaper on input above 272K. Two independent cost measurements make the same point from a different direction: Artificial Analysis puts Argon at $1.99 per index task against Astra's $3.26, and $2,407 to run the whole index against Astra's $5,324. The per-task ratio is smaller than the per-token ratio for the same reason it was against DeepSeek — Argon runs fewer tokens to finish — but it is still 1.6×.
Vals's GDP-weighted board tells a third version of the same story: Argon first at 68.90% and $15.68 per run, Astra sixth at 63.13% and $18.46. Astra is more expensive and scores lower there. Google's material is explicit that the pricing is an introductory rate, which is a word that means it can move, and the honest reading is that Argon's price advantage is real and its permanence is not guaranteed.
Two architecture facts that decide more than benchmarks

Output ceiling first. Gemini 4 Argon generates up to 1,000,000 tokens in a single trajectory — Google's announcement calls this out as an expansion from 64K in the previous generation. GPT-6 Astra caps at 128,000. Both hold a context window around a million tokens, so this is purely about how much the model can write in one go. For long-form generation, full-patch code changes or one-pass document drafting, that is an eight-fold difference in what a single call can produce. For chat, extraction and short agent steps, both ceilings are effectively infinite and the difference costs you nothing.
Modality second, and here the advantage runs the other way in one respect. GPT-6 Astra accepts text, images and files. Gemini 4 Argon accepts text, images, video and files, and Google's launch material leans on long-video understanding specifically — an LVBench figure of 91.7% that is vendor-reported. If your pipeline ingests video, Astra is not a substitution. Audio is not named in Argon's published modality list either, so do not assume the upgrade carries it.
What to actually do
Astra is available and Argon is not, and that settles most decisions for the next month. The concrete path that wastes no work: build on GPT-6 Astra if your workload needs an always-on reasoning model with strong scientific accuracy and a proven agentic track record, keep your model name in configuration, and set your client timeouts from Astra's measured behaviour rather than from optimism — a five-minute first-token run will trip a default 60-second timeout and a stream that dies at 300 seconds looks like an outage. If you are cost-sensitive above 272K input tokens, model the repricing explicitly before you commit, because the step doubles input and raises output by half in one boundary.

The interesting middle path is that you do not have to pick a single default. Astra and Argon — when Argon ships — are the same shape of model aimed at the same class of work, which makes them natural failover partners rather than competitors for the same slot. On OrcaRouter, openai/gpt-6-astra is live at provider list price with zero markup, on the same OpenAI-compatible endpoint as 200-plus other models, with automatic failover across providers and a routing DSL for expressing a primary-and-backup arrangement on one key. When Argon joins the catalogue at whatever id Google publishes, the switch is a string — no second contract, no integration work, and no rebuild of whatever you built around Astra in the meantime.
What would break the tie permanently
A published Argon price after the introductory period, and an independent reproduction of Google's agentic claims — the 77.9% DeepSWE v1.1 figure and the 68% CWE-bench v1 result are both vendor-reported and neither has an outside run behind it. Either could move Argon up or down materially, because a 0.11-point index lead is not a buffer against a 1.6× cost disadvantage if the cost advantage turns out to be temporary, and it is not a deficit if Google holds the introductory rate and Astra's long-context tier bites harder than expected in production.
For now: on measured general capability these two are the same model in two wrappers. Astra is the one with a live route, a known price step and a five-minute thinking pause. Argon is the one with a cheaper card and no endpoint. The tie is real, and one side of it is buyable.
