Hero title card reading "Gemini 4 Argon — which rung of the Gemini 4 ladder?", with a vertical ladder of five ranked rungs listing Claude Opus 5.5 at 57.6, Gemini 4 Argon at 52.6, GPT-6 Astra at 52.7, Gemini 3.8 Flash at 40.9 and Gemini 3.1 Pro Preview at 29.7, a badge reading "No Gemini 4 Ultra — the Ultra page redirects to a subscription plan", and a footer line reading "Index figures per Artificial Analysis, v4.3.2 revision; Argon not callable on any public API."
Guides & Insights

Gemini 4 Argon's Rung in the Gemini 4 Ladder — and Why There Is No G4 Ultra

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The short answer to the question that started this piece is that Gemini 4 Argon is Goog​le’s top model and not its top tier, because the tier above it does not exist yet — and possibly not in the form anyone is expecting. Goog​le has published exactly one Gemini 4 model, Argon, and the only first-party page on its own domain under the path /models/gemini/ultra/ is not a model page at all: it redirects to the Goog​le AI subscription plans. That redirect is the single most useful fact in this article, and it is checkable in one line from a terminal. Everything else here is about where Argon actually sits once you stop reading its price as a tier label and start reading it as a number.

Two things are true at once and they are easy to confuse. Argon is the most capable G​emini that exists, and Argon is not a frontier-tier model. It leads every G​oogle model that has shipped, by a margin wide enough that the comparison is not really a comparison, and it lands on Artificial Analysis’s Intelligence Index within a fraction of a point of two competing flagships that are themselves a full five points behind the current leader. G​oogle priced it as though it sits one rung below the frontier, and the independent measurement agrees. So the price is not a marketing decision that undersells the model. It is an accurate statement about the model.

This is a what-we-know-so-far piece. Gemini 4 Argon was announced on 2026-09-30 in a Google DeepMind post credited to Koray Kavukcuoglu, and it is rolling out first to a named cohort of cyber defenders through Google’s Fairwind Program — not to developers, not to enterprises, not to consumers, and not to anyone with a credit card. It has no model ID in the Gemini API, no endpoint, and no published per-token price you can be billed against. Model identifiers, throughput and reliability measured on real traffic do not exist for it, which means the measurement below is one lab’s benchmark run and nothing more. Where a number comes from Google it is labelled as Google’s; where it comes from Artificial Analysis it is labelled as theirs. Nothing here should be read as a claim that Argon is a purchasable product.

The ladder Google actually shipped, read off its own pages

The fastest way to answer “where does Argon sit in the Gemini 4 lineup” is to stop asking about the lineup and start listing the models Google publishes pages for. A lineup is a marketing concept; a page is a fact. Here is what the first-party paths resolve to as of October 1, 2026.

• deepmind.google/models/gemini/pro/ returns 200 and its title is “Gemini 3.1 Pro — Google DeepMind”. The Pro slot on Google’s own site is still a 3.1-generation model.

• deepmind.google/models/gemini/flash/ returns 200 and its title is “Gemini 3.8 Flash — Google DeepMind”. The Flash slot has moved three times — 3.5, 3.6, 3.7, 3.8 — while the Pro slot has not moved at all.

• deepmind.google/models/gemini/argon/ returns 404. Argon does not get a per-model path. Its home is the family page at deepmind.google/models/gemini/, which is where Google puts the model it is currently leading with.

• deepmind.google/models/gemini/ultra/ returns 302 and lands on one.google.com/about/google-ai-plans/, the consumer subscription page. The same is true of the older path deepmind.google/technologies/gemini/ultra/. An “Ultra” on Google’s models path is a billing tier, not a checkpoint.

• deepmind.google/models/gemini/nano/ returns 301 to the Gemini family page. There is no Nano slot as a standalone model page either.

• deepmind.google/models/gemini/model-cards/ — the index Google maintains of every model card it has published — carries no entry for Argon and no entry with “Gemini 4” in the name. Its newest Gemini rows are 3.8 Flash, 3.8 Audio, 3.7 Flash, 3.6 Flash, 3.5 Flash, 3.5 Flash-Lite and 3.1 Pro. The card index is the slowest-moving document in this set, so its silence is not evidence that Argon will never have a card; it is evidence that the paperwork has not caught up with the announcement.

• deepmind.google/models/, the model index, lists “Gemini 4 Argon” as the flagship card with the tagline “Our next era of frontier intelligence”, directly above “Gemini 3.8 Flash — Best for tackling complex agentic tasks at scale”. Two Gemini cards, one generation apart, in that order.

Put those together and the shape of the ladder is not ambiguous. Google’s public model surface currently has one entry at the Gemini 4 rung, one entry at the Gemini 3.8 rung, a stale Pro rung three and a half generations back, and nothing above any of it. The word “Ultra” on Google’s domain resolves to a subscription. There is no Gemini 4 Pro page, no Gemini 4 Flash page, no Gemini 4 Ultra page, and no Gemini 4 model card. Google announced a model, not a family.

Is there a Gemini 4 Ultra? The evidence, and what would falsify it

This deserves its own section because it is the part of the question most likely to be answered differently in a week, and because the honest answer has a shelf life.

The negative evidence is specific rather than vague. A 302 on the Ultra path to the subscription plans page is not a soft signal; it is a redirect that Google’s own web team configured. Multiple blog.google URL patterns for a Gemini 4 Ultra post return 404 — the products path, the innovation-and-ai models-and-research path, and the plain announcement path. The Ultra naming has form here, and the form cuts against a Gemini 4 Ultra. Google’s original Gemini announcement introduced Gemini 1.0 “optimized for three different sizes: Ultra, Pro and Nano”, with Gemini Ultra described as “our largest and most capable model”. A launch post that names three sizes is what a family announcement looks like. Argon’s post names one model and no siblings at all. Google later retired the Ultra model name in favour of numeric tiers and moved the word to the subscription side of the business, where it now names the top consumer plan. A model page that redirects to a plans page is what a retired model name looks like when the marketing site has been cleaned up around it.

There is also nothing in the announcement that promises a bigger sibling. The Argon launch post describes Argon as “our new frontier model” and describes the phased rollout, the safeguards work and the internal deployments, and stops. It does not say “first in the Gemini 4 family”, it does not preview a Pro or an Ultra, and it does not name a successor. Google’s own framing treats Argon as the thing, not as the small one.

What would falsify the conclusion is easy to state and worth watching for, because any one of these would flip it inside a day.

• A 200 on deepmind.google/models/gemini/ultra/ that renders a model page rather than the plans page, or a new models/gemini/ child path appearing alongside pro and flash.

• A Gemini 4 Ultra row appearing on the model-cards index, which is how Google historically confirms a model exists as a documented artefact.

• A second model ID in the Gemini API in the gemini-4-* namespace. Argon currently has none, so a Pro or Ultra ID would be the first evidence that the family is real.

• An entry in Artificial Analysis’s model list, or a benchmark table on the family page with a fourth column. Both of those track Google’s releases closely enough to be early.

• A Google I/O, Cloud Next or DeepMind announcement post using the word “family” about Gemini 4. Argon’s post does not.

Until at least one of those happens, a piece claiming a Gemini 4 Ultra is coming is a prediction wearing the clothes of a report. Treat it accordingly, including from us.

A two-column comparison scoreboard titled "Gemini 4 Argon vs Gemini 3.8 Flash — the scoreboard", with the left column showing Gemini 4 Argon at AA Intelligence Index 52.6, 1M context window, $2 input per million tokens, $10 output per million tokens, $1.99 cost per Index task and 142,374 output tokens per Index task, and the right column showing Gemini 3.8 Flash at AA Intelligence Index 40.9, 1M context window, $0.75 input, $3.75 output, $1.24 cost per Index task and 154,230 output tokens per Index task, with a footer reading "Both columns' Index, price, context, cost-per-task and token-use figures per Artificial Analysis, v4.3.2 revision."

Reading the price ladder as a tier statement

Pricing is the one part of this that Google published on purpose, with a footnote, and it is the clearest statement of tier intent in the whole launch. Argon’s introductory price is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off, and Google’s own footnote one says that after an introductory period it could not define, “the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.”

Read those two pairs against the field and the tier claim writes itself.

• $4 / $20 is Claude Opus 5.5’s list price. Google’s post-introductory rate for Argon lands exactly on the current frontier leader’s rate card, for a model that measures five points behind it.

• $2 / $10 is the promotional band that GPT-6 Sol and Claude Sonnet 5.5 both occupy. Argon’s introductory rate is the same $2 / $10, which puts Argon’s launch pricing in the same band as a mid-tier Sol rather than in the frontier band.

• $10 / $50 is where Claude Fable 5.1 and GPT-6 Astra both sit. Argon is a fifth of Fable 5.1’s input price and a fifth of its output price, and it scores within 0.8 of it.

• $0.75 / $3.75 is Gemini 3.8 Flash. Argon is 2.7× that on input and 2.7× on output — a clean step up, not a cliff.

So Google has priced Argon as a model that belongs below $10/$50 and above $0.75/$3.75, with an eventual resting place at $4/$20. Nothing about that ladder says “frontier”. It says “top of the mid-band, with a headline number that will settle at the same rate the leader charges today, for less capability.” That is a defensible commercial choice. It is not the pricing of a model positioned two tiers ahead of anything.

One structural point worth carrying: Argon’s price step is a genuine step, defined in a footnote, and undated. The Sonnet 5.5 and GPT-6 families do something similar with long-context tiers, and Sol steps to $4/$15 past 272K. Argon’s step is on time, not on context length — the same $2/$10 applies across the whole 1M window and only the calendar changes the rate. That is an unusual shape and it makes any cost comparison against Argon a two-column exercise: which period, and which usage.

Where Argon sits against the rivals, on the numbers that exist

Artificial Analysis had a measurement of Gemini 4 Argon up quickly enough that it is the only independent read available. On the revision current on October 1 — v4.3.2 — Argon scores 52.56 on the Intelligence Index, configured at high reasoning effort. The models around it are not a random sample of the field.

• Claude Opus 5.5 sits at 57.62, measured at max effort with fallback. That is five points and change above Argon, and it is the current top of the board.

• Claude Fable 5.1 sits at 53.35, measured at max effort with fallback. Argon is 0.79 behind it.

• GPT-6 Astra sits at 52.67, measured at max. Argon is 0.11 behind it.

• Claude Sonnet 5.5 sits at 55.98, also max with fallback. That is a mid-tier model outscoring Argon by 3.4 points, and it is the number that should worry anyone reaching for “Gemini 4 Argon is the second-best model in the world”.

• GPT-6 Sol sits at 47.63, measured at max, on an 872K context window. Argon is 4.9 ahead of it.

• Gemini 3.8 Flash sits at 40.93 at high effort. Argon is 11.6 ahead of it.

• Gemini 3.1 Pro Preview sits at 29.72. Argon is 22.8 ahead of it, which is the single clearest statement of how far Google’s shipping Pro tier has fallen behind its own research.

Three things fall out of that list. First, Argon is level with the two challengers, Fable 5.1 and Astra, and clearly behind two Anthropic models — Opus 5.5 and, more awkwardly, Sonnet 5.5. Second, the gap between Argon and Google’s own best shipped model is the largest generational jump in the current lineup, by a distance. Third, the signal’s framing that “the frontier is at least two tiers ahead” is right in substance but slightly off in geometry: there is one model tier clearly ahead of Argon (Opus 5.5 at 57.6) and a second model in a cheaper tier that is also ahead of it (Sonnet 5.5 at 56.0), which is a sharper problem than being two rungs down a ladder Argon is actually on.

The compare-the-price framing in the original question — “prices suggest this is their Sol/Sonnet equivalent” — turns out to be roughly right about the launch price and wrong about the eventual one. Argon opens at Sol’s and Sonnet 5.5’s rate and settles at Opus 5.5’s. It is priced as a mid-tier model that intends to become a frontier-priced one, while measuring at neither.

The cost-per-task picture is where Argon is strongest and where the tier story gets a second dimension. On the Index task, Argon costs $1.99 and emits 142,374 output tokens. Opus 5.5 costs $5.98 and emits 338,099. Sonnet 5.5 costs $7.62 for 599,849. Fable 5.1 costs $7.63 for 187,455. Astra costs $3.26 for 68,691. Gemini 3.8 Flash costs $1.24 for 154,230. Argon is the cheapest way to buy a frontier-adjacent score, and it is cheaper than the Flash model above it in score by less than a dollar a task.

Effort settings are the asterisk on all of the above and the field does not report them uniformly. Argon is measured at high; Opus 5.5 and Fable 5.1 and Sonnet 5.5 at max with fallback; Astra and Sol at max. A high-effort run and a max-effort run are not the same measurement, and Anthropic’s own effort ladder moves a model several points between settings. If Argon measured at max effort scores meaningfully above 52.6, the tier argument changes. Nobody has published that run yet.

What Google’s own table claims, and how it differs from the independent one

The Gemini family page carries a Google-authored performance table with four columns — Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5 — and nineteen benchmark rows. It is the most detailed claim set Google has published for this model, and it is vendor-reported throughout. Reading it against the independent Index is instructive precisely because the two do not agree about the shape of the field.

• On Google’s table Argon wins thirteen of the nineteen rows outright or ties for first, including Vals Index Knowledge work at 68.9, AutomationBench at 51.3, Harvey’s Legal Agent Benchmark at 19.6, DeepSWE v1.1 at 77.9, LABBench 2 at 88.8, GraphWalks at 99.7 for the ≤128K band and 84.2 for the 256K–1M band, Agent’s Last Exam at 39.5, LVBench at 91.7 and Chartography at 71.6.

• Claude Opus 5.5 wins the rows that read like sustained engineering: FrontierSWE v2 at 62.3 against Argon’s 55.0, Terminal-bench 4.0 at 66.4 against 57.4, Terminal-Bench Science 0.1 at 63.3 against 57.6, and PostTrainBench at 49.3 against 45.3.

• GPT-6 Astra takes Terminal-Bench Science 0.1 at 68.1 and the OSWorld-2.0 offline subset at 72.6, and ties Argon on CWE-bench v1 at 68.0. Claude Fable 5.1 loses every row except a shared tie on Vibe Code Bench.

• On the independent Index, the ordering is Opus 5.5 first, then Sonnet 5.5, then Fable 5.1, then Argon and Astra effectively level. Google’s table has no Sonnet 5.5 column at all, which is the most consequential omission in it: the table’s four columns are chosen, and the model that outranks Argon on the Index at a lower price is not among them.

None of that makes Google’s table dishonest. Vendor benchmark tables are selected, not sampled, and every lab’s is. It makes it a claim rather than a measurement, and it means the tier question cannot be settled from it. The one place the table is genuinely load-bearing for this article is the negative: nineteen rows, four columns, and no fifth column for an Ultra.

A screenshot of Google DeepMind's Gemini 4 Argon page scrolled to its Performance section, showing the Performance heading and the line “Frontier performance in complex workflows” above the vendor-reported benchmark table, whose four model columns are Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5, with rows running from Vals Index and AutomationBench down through FrontierSWE v2 and Terminal-bench 4.0 to GraphWalks and Agent's Last Exam.

The one thing about Argon that is not a tier question at all

Every tier argument above assumes Argon is a model you can eventually call. It is worth separating that from the positioning question, because the two have different answers and only one of them is about capability.

Argon’s output token limit is 1 million tokens, up from 64K, and Google states that directly. That is an output ceiling, not a context window. The distinction matters because the two get conflated constantly in coverage of this launch, and because a 1M output limit is a different engineering claim from a 1M context window — the first is about how long a single trajectory can run, the second about how much you can hand the model at once. Google has been explicit about the output limit and silent about the context window. Artificial Analysis records 1,000,000 tokens for Argon’s context as well, but that is the tracker’s field and not a Google statement, so treat the two figures as coming from different rooms even though they read the same.

What is unambiguously not available is access. Argon is not generally available, it is not in the Gemini API under any identifier, and it is not on any third-party catalogue. It is available to vetted cyber defenders through the Fairwind Program and to Google’s own engineers, and the next stage Google names — “starting with paid API customers and Google AI Ultra subscribers” — has no date attached to it. You cannot benchmark it on your own workload. Every number in this article is therefore somebody else’s measurement of a model you cannot use.

What would move Argon up a rung

A tier judgement is a snapshot, and there are about five things that would change this one, in rough order of how much they would matter.

• An independently measured max-effort run. Argon is currently a high-effort measurement at 52.56. Anthropic’s own effort ladders move a model several points across settings, and if Google’s model behaves similarly at max, Argon’s real position against Fable 5.1 and Astra is better than this article says. This is the single most likely change.

• A second measurement source. One tracker is one measurement. A second independent lab scoring Argon on its own harness would do more for the tier claim than any additional benchmark row from Google.

• A Gemini 4 Pro. If Google opens the Pro rung of the 4 generation below Argon, the lineup becomes a family with a shape, Argon’s position stops being “the only one” and starts being “the top of three”, and every argument in this piece about a missing family needs rewriting. Worth watching for, and closer than the Ultra question.

• A price that does not step. If the introductory period simply never expires — Google has not defined when it ends — Argon’s effective resting price is $2/$10 rather than $4/$20, and its cost-per-task advantage over Opus 5.5 widens from roughly 3× to roughly 6×. That is a commercial event, not a capability one, but it changes what a procurement decision looks like.

• An Argon model card, system card or eval methodology page. The performance table links to a methodology page on Google’s domain; a full model card would put the context window, the deployment configuration and the safety envelope on the record, and would be the first artefact that lets someone outside the Fairwind cohort check Google’s numbers rather than read them.

Conversely, the thing that would most likely move Argon down is not a benchmark at all. It is the September 30 Bloomberg reporting, which quoted Google employees with access to the model saying it does less well on real work than its scores suggest. That claim is unverified by anyone outside Google and it is not a measurement, but it is the kind of claim that a second independent benchmark would either confirm or kill. Until then it sits in the record as an allegation with a named outlet and unnamed sources, and it belongs in the tier discussion because a model whose benchmark-to-field gap is disputed cannot be promoted on benchmarks alone.

What this means if you are choosing a model this month

Stripped of the positioning argument, the practical picture is straightforward enough that a paragraph covers it. Gemini 4 Argon is the best Gemini Google has built and it is not available to you, so the Google models you can actually select between today are the ones that have shipped: Gemini 3.8 Flash, which measures 40.93 on the Index at $0.75/$3.75, and Gemini 3.1 Pro Preview, which measures 29.72 at $2/$12. If you need a Gemini right now, that is the real choice, and Argon is not in it.

If you are picking across vendors rather than within Google, nothing about Argon’s announcement changes today’s decision. The top of the Index is Claude Opus 5.5 at 57.62, with Claude Sonnet 5.5 at 55.98 close behind it at a fifth of the rate on input and half on output. Gemini 4 Argon at 52.56 has no route to your application, so it is not a candidate, however good the score eventually turns out to be.

A screenshot of Google's own Google AI plans page, showing the three consumer Google AI plans side by side — Google AI Plus at $4.99/mo with 400 GB storage, Google AI Pro at $19.99/mo with 5 TB storage, and Google AI Ultra from $99.99/mo starting at 20 TB of storage — each with a “View plan benefits” link, illustrating that “Ultra” on Google's domain names a subscription plan rather than a Gemini model.

OrcaRouter is a single API in front of 200-plus models with provider list prices passed through at zero markup and automatic failover across providers, which means the model-switching decision does not require re-plumbing anything: the id in the request body is the whole change. The Google models on our catalogue today are the ones Google has actually shipped — google/gemini-3.8-flash, google/gemini-3.6-flash, google/gemini-3.5-flash and google/gemini-3.1-pro-preview. Gemini 4 Argon is not one of them and will not be while it has no public model identifier and no published rate card, so nobody should read a routing claim into anything above. When Google puts Argon on an API with an ID, it becomes a routing question. Until then the honest version of the OrcaRouter answer is that it does not apply yet, and the useful thing the catalogue does for this particular decision is let you price the models that are actually callable side by side without opening four accounts to find out.

The bottom line

Gemini 4 Argon is Google’s flagship and not its frontier. It measures 52.56 on Artificial Analysis’s v4.3.2 Index at high effort — level with Claude Fable 5.1 and GPT-6 Astra, five points behind Claude Opus 5.5, and behind Claude Sonnet 5.5, a cheaper model from a rival lab. Google priced it at $2/$10 introductory, stepping to $4/$20, which lands its eventual rate exactly on Claude Opus 5.5’s for measurably less capability. Its 11.6-point lead over Gemini 3.8 Flash is the widest generational gap in Google’s own lineup and it is the strongest evidence that the Gemini 4 generation is a real step up rather than a rebadge.

On the question that prompted all of this: there is no Gemini 4 Ultra. Google’s Ultra path on its own domain redirects to a subscription plans page, the model-cards index carries no Gemini 4 entry at all, every announcement URL pattern for a Gemini 4 Ultra post 404s, and the launch post announces one model without promising a family. That is a real finding and it is first-party evidence rather than inference, but it is also a fact with a short half-life — a single 200 on one path reverses it, and the Pro rung of the Gemini 4 generation is the one more likely to appear first.

Two caveats to carry out of this piece. Every Argon number here is somebody else’s measurement of a model nobody outside a vetted cyber-defender cohort can call, and Google’s own nineteen-row table is a claim rather than a measurement — useful for what it is, and not a substitute for the independent Index. And the fair verdict on the tier question is provisional: Argon is measured at high effort, not max, and a max-effort run is the cheapest possible way for this article to be wrong.

Compared in this article4

Detected from this article · Benchmarks: Artificial Analysis · updated daily