
Gemini 4 Argon: What Google Announced, and What the Name Is Attached To
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 64 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 360 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Gemini 4 Argon is the only Gemini 4 there is. That sentence sounds trivial until you price it out, and pricing it out is what this page is for. Google announced Gemini 4 Argon on 2026-09-30, in a DeepMind post credited to Koray Kavukcuoglu; it is a real model name on a real first-party page, it has been measured by Artificial Analysis at an Intelligence Index of 52.56 on the v4.3.2 revision, and it sits inside half a point of GPT-6 Astra at 52.67 and Claude Fable 5.1 at 53.35 — while sitting five points under Claude Opus 5.5 at 57.62 and Claude Sonnet 5.5 at 56.00, and within a point of GPT-6.1 Sol at 51.83. What it does not have is the thing every one of those rivals has: a model identifier you can put in a request.
So the honest one-line answer to “what is Gemini 4 Argon” is that it is a model Google has announced and benchmarked but has not yet made callable, that the name is attached to a gated rollout rather than a product, and that “Gemini 4” on its own is attached to nothing at all. Both halves of that sentence are checkable from a terminal, which is why this page checks them rather than asserting them.
The name is real. The product is not.
Start with what an entity check actually returns. As of October 8, 2026, Google’s family page at deepmind.google/models/gemini/ carries Gemini 4 Argon as its lead card, under the tagline “Our next era of frontier intelligence”, with capability copy naming software engineering, enterprise knowledge work in legal and finance, and defensive cybersecurity. The announcement post exists at the blog.google/innovation-and-ai/models-and-research/gemini-models/ path. There is a model-family page. There is a nineteen-row vendor benchmark table. There is a five-page evals methodology PDF. That is a substantial paper trail, and it is more than a leak gets.
Now list what is missing from it, because the missing list is the article.
• A model ID — the Gemini API’s model list contains zero occurrences of “argon” and zero occurrences of “gemini-4”. The newest identifiers Google’s documentation recognises are the 3.8 family.
• An endpoint or price you can be billed against — the $2 input / $10 output per million tokens in the announcement is an introductory rate for a rollout that has not reached paying customers.
• A model card or system card — the DeepMind model-cards index has no Argon row, and the family’s model-card path returns 404.
• A context window — Google published an output limit of 1 million tokens, up from a previous 64K ceiling, and did not publish the input side of that pair. Artificial Analysis records 1M for the model it measured.
• A general-availability date — the announcement describes a phased rollout with no dates on any stage past the first.
• A parameter count, which Google has not published for any Gemini.
The rollout Google describes runs in three unnamed-duration stages. First, a named cohort of cyber defenders through the Fairwind Program. Second, paid API customers and Google AI Ultra subscribers. Third, developers, enterprises and consumers. Only the first stage is live in any sense a reader can verify, and Google attaches no date to stages two and three. That is not a footnote. It is the difference between a model and a countdown.

What Google’s own table claims, and who actually measured it
Google’s family page carries a four-column benchmark table with Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. It is worth reading, and it is worth reading with the provenance in view, because Google discloses that provenance in the methodology PDF and the disclosure changes how much the table can carry.
The claims, labelled as Google’s: Argon leads on Vals Index at 68.9 against Opus 5.5’s 67.0 and Fable 5.1’s 65.8; on AutomationBench at 51.3 against 42.5, 31.4 and 41.4; on Vals Finance Agent v2 at 65.4; on Harvey’s Legal Agent Benchmark at 19.6 against 3.8, 6.7 and 5.4 — a four-to-five-fold gap; on DeepSWE v1.1 at 77.9 against 74.2, 67.4 and 74.1; on Vibe Code Bench at 91.9, a 1.6-point margin; on LABBench 2 at 88.8 against 73.1, 68.6 and 85.4; on RiemannBench at 76.0; on both GraphWalks rows, 99.7 at ≤128k and 84.2 at 256k–1M; on Agent’s Last Exam at 39.5; on Terminal-Bench Science 0.1 at 57.6; on Chartography at 71.6; and on LVBench at 91.7 as state of the art.
And it loses, on Google’s own page: FrontierSWE v2 at 55.0 against Astra’s 65.5 and Opus 5.5’s 62.3; Terminal-bench 4.0 at 57.4 against Opus 5.5’s 66.4; PostTrainBench at 45.3 against Opus 5.5’s 49.3; and OSWorld-2.0 at 69.2 against Astra’s 72.6. CWE-bench v1 is a tie at 68.0 with Astra.
The provenance line matters more than any single row. Google’s methodology states that results are pass@1, single-attempt, at the highest thinking settings, and — critically — that the non-Gemini figures in that table are the providers’ own self-reported numbers, with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 taken at maximum thinking where available. A vendor table whose competitor columns are the competitors’ press releases is a useful document, but it is not a controlled comparison. Every one of those rows is Google-reported, and some of them are Google reporting what someone else reported. Treat the margins, not just the values, accordingly.
One model, three index numbers, and why they are all correct
Argon’s independent standing is more interesting than a single figure, because the figure moves depending on where you read it, and the movement is not an error.
Artificial Analysis measures Gemini 4 Argon (High) at 52.56 on the v4.3.2 revision of its Intelligence Index. The same index puts Claude Opus 5.5 at 57.62, Claude Sonnet 5.5 at 56.00, Claude Fable 5.1 at 53.35, GPT-6 Astra at 52.67, GPT-6.1 Sol at 51.83 and Gemini 3.8 Flash at 40.93. On the leaderboard’s displayed ordering, the top ten or so models compress into roughly six index points, and Argon lands in that scrum rather than under it.
Then a reader sees a second number — 53 — in the leaderboard’s rounded display, and a third if they compare against a different configuration of a sibling model, and concludes one of the sources is wrong. None of them is. The rounded 53 is the same measurement as 52.56. What differs between configurations is thinking effort: Artificial Analysis measures each model at several effort levels, and the same model at a higher effort level can be several points apart. That is why the leaderboard lists Opus 5.5 four separate times, at 58, 56, 54 and 51. Argon appears once, at high effort, because that is the only configuration Google has made available to measure. If Google ships an Argon tier at maximum thinking, that row will move, and the honest thing to say about it today is that it has not moved yet.
Three more of Argon’s independent figures are worth carrying, all from Artificial Analysis and all measured on the (High) configuration. Output speed is 60.69 tokens per second median. Time to first answer token is 2.92 seconds median. Cost per Intelligence Index task — the index score normalised against the token spend it takes to earn it — is $1.99. That last number is the one to hold on to, because it is where Argon’s position becomes uncomfortable: Opus 5.5 scores five points higher and costs $5.98 per Index task, so Argon is three times cheaper per unit of measured capability. But GPT-6.1 Sol scores half a point lower and costs $0.72, which is roughly a third of Argon’s figure. Argon is not the value pick in its own index band. It is the middle of it.
Two more independent readings exist that a careful reader should not confuse with each other. Artificial Analysis records Argon’s AutomationBench figure as a partial score of 0.7751 on a 0–1 scale, while Google’s table reports 51.3 out of 100 on the AutomationBench board. Those are not the same metric, and stacking them side by side is exactly the kind of comparison that produces an invented number. The same caution applies to the vendor table’s Knowledge-work groupings: “leading on knowledge work” is Google’s summary of Google’s own table.
Which “Gemini 4” page you are actually looking for
This is the confusion the name creates, and it is worth settling in one place, because the blog has now written about five different “Gemini 4” things and they are not variants of one article.
• If you want the release date and the rollout timeline — when Argon was announced, what the staged rollout is, and what the arena boards did with it — that is the release-date write-up, which is the piece that owns the timeline and the Text Arena and Code Arena readings.
• If you want “is Gemini 4 Argon a tier or a model”, and whether a Gemini 4 Ultra exists — the tier-positioning piece, which reads the ladder off Google’s own paths and explains why the Ultra path resolves to a subscription page.
• If you want the naming question itself, the argument that of “Gemini 4” and “Gemini 4 Argon” exactly one is a model and one is a generation prefix with no model ID, no endpoint and no price — that is the head-to-head of the two Google names.
• If you want the third-place FrontierSWE v2 result and the self-criticality observation that came with it, that is the FrontierSWE evaluation piece.
• If you want Argon against a specific rival on the numbers, there are thirteen comparison pages on the site covering the frontier band, and they are the right place for a two-model read.
What none of them is, and what this page is, is the pillar: the single place that says what the entity is, what each figure attached to it does and does not mean, and what you can actually do with it today. The related page “Gemini 4 vs Gemini 4 Argon” already makes the naming argument in full and this page will not repeat it. Everything below is new.

The pressure that has not told us anything yet
Two of the strongest signals in the Argon story point at an imminent developer release, and neither of them is evidence that a developer release has happened.
The first is that Google’s own API surfaces have moved on without Argon. The Gemini API changelog’s newest entry is October 6, 2026 — Gemini Nano Banana 2.1 generally available, under the identifier gemini-nano-banana-2.1. In other words, in the eight days since Argon was announced, Google has picked a naming convention for a Gemini model in production and shipped a GA identifier for a different model, and the gemini-4 namespace is still empty. A model whose rollout is real leaves a mark on the changelog. This one has not left one.
The second is the national-security framing. Google states that it is “actively engaged in the U.S. government’s voluntary process for pre-release model access”, and the first rollout stage is a named cyber-defender cohort. That is consistent with a staggered release in which the frontier-capability audience gets the model before the public API does, and it is equally consistent with the model simply not being ready for general traffic. Nothing about the framing tells you which. Anyone who says it does is reading tea leaves, and Google has not put a date on the API stage to settle it.
What you can call today, and what you cannot
This is the part of the page that gets dated the fastest, so it carries its date. As of October 8, 2026, the OrcaRouter model API returns HTTP 404 for google/gemini-4-argon. It is not on our routes. Neither is google/gemini-4, nor google/gemini-4-argon-high, nor google/gemini-4-pro — none of those identifiers resolve, because none of them exist as callable models anywhere.
What is on our routes today, from the same vendor, and what each one costs on the provider’s list price:
• google/gemini-3.8-flash — released 2026-09-02, a 1,048,576-token context window and a 65,536-token maximum output, at $0.75 input and $3.75 output per million tokens. This is the model Argon will eventually be compared against by people who own a credit card, and it is the closest thing to an Argon-adjacent workload you can run this afternoon.
• google/gemini-3.6-flash — released 2026-07-21, at $0.75 and $3.75.
• google/gemini-3.5-flash — released 2026-05-23, at $1.50 and $9.00.
• google/gemini-3.5-flash-lite — released 2026-07-21, at $0.30 and $2.50.
• google/gemini-3.1-flash-lite — at $0.25 and $1.50.
• google/gemini-3.1-pro-preview — released 2026-02-19, at $2.00 and $12.00. This is the only Pro-slot Gemini you can call, and it is three and a half generations behind the Flash line’s newest entry.
All of those run on the same key as the Claude and OpenAI frontier models in the same index band — Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1, GPT-6 Astra and GPT-6.1 Sol are all on the same routes — which means the interesting thing about Argon’s position is testable without Argon. If you want to know what 52-ish index points of knowledge work feels like at a fifth of the price, you can measure it today against google/gemini-3.8-flash at 40.93, and if you want the top of the band you can measure Claude Opus 5.5 at 57.62 on the same day, in the same account, with no second contract and no code change. That is the practical shape of the whole article: one API for 200+ models, 0% markup on provider list price passed through, so a vendor price cut reaches you the same day it lands, and automatic failover across providers for the day a model you depend on stops answering. We are describing that here because it is true of the Gemini 3.8 Flash you can call, not because it says anything about Argon. Argon is not one of them.
If you want Argon itself, your options today are the Fairwind cyber-defender cohort and waiting. Google’s own API is not serving it, and neither are we, and no third-party route can serve a model that has no identifier to address.
The open questions, and exactly what would close them
Four things are unresolved, and each one has a specific, checkable trigger. Keeping that list short is the point — a page like this is only useful while it is falsifiable.
• Does Argon get a Gemini API identifier? Watch the API model list for any gemini-4-prefixed ID. The moment one appears, the pricing question becomes real and the intro rate of $2 / $10 becomes a number you can be billed against rather than a figure in a blog post. Also watch whether the intro rate survives into the API stage at all, and whether it steps to $4 / $20 or somewhere else.
• Does the measured configuration change? Argon is measured at one thinking level and scores 52.56. If Google exposes a higher-effort configuration, that row moves; the peers it is chasing are all listed at two to four effort levels, so Argon being listed once is a fact about availability, not about the model’s ceiling.
• Do the legal and finance claims survive an independent run? Harvey’s Legal Agent Benchmark at 19.6 against a competitor’s 3.8, and Vals Finance Agent v2 at 65.4 leading the field, are the two largest relative margins in Google’s table. They are also the two rows most likely to be re-run by someone who does not work for Google. A four-to-five-fold margin on a domain benchmark is either the most important sentence in the announcement or a harness difference, and the way to find out is to watch for a second run.
• Is there a sibling? Argon’s announcement post does not name a family, does not preview a Pro and does not preview an Ultra. A second gemini-4-* path, a second column, or a model-card entry would each flip that reading inside a day. The tier-positioning piece lists the falsifiers in full.
The bottom line
Gemini 4 Argon is a genuine frontier model with a genuine benchmark table and a genuine index score of 52.56 from Artificial Analysis — and it is not a product, in the specific sense that no model identifier, endpoint, model card, context window or general-availability date exists for it, and both Google’s own documentation and the third-party measurement services agree that exactly one provider hosts it. The “Gemini 4” in its name is a generation label; the “Argon” is the part with the model attached. If you are choosing something to build on this month, the decision does not involve Argon: it involves the Gemini 3.8 Flash you can call today at $0.75 / $3.75 and the frontier models in the same index band that sit on the same key. Re-read this page when a gemini-4 identifier appears in the API list. Until then, everything about Argon is a claim Google made about a model nobody outside the first cohort can address.

Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
