
Gemini 4 vs Gemini 4 Argon: One of These Two Models Does Not Exist
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 221 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 110 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 103 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 214 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Google announced a frontier model on September 30, 2026. It is not the one most people are picturing. Gemini 4 Argon is a real name attached to a real announcement, a real rate card and a published benchmark set. Gemini 4 is the family prefix that name sits under — and on the record available today it is not a product. There is no model ID for it, no endpoint, no model card and no price per million tokens. So this is not a comparison of two systems. It is a comparison of one announced model against a name that a lot of people, this blog included, have been treating as a model for months.
That distinction has money attached to it. Gemini 4 Argon carries an introductory rate of $2 per million input tokens and $10 per million output tokens, stepping to $4 and $20 after an introductory period Google has not defined. Those are vendor-stated figures from Google's own announcement post. If you are budgeting a 2026 pro-tier migration, $2/$10 is the number you can plan against, and "whatever Gemini 4 costs" is not a number at all — nothing has been published for it, because there is nothing yet to publish.
What "Gemini 4" actually is on the record
Search for a Gemini 4 product page and you will find one document: Google's own post titled Gemini 4 Argon: our next era of frontier intelligence, published September 30, 2026, at 20:00 UTC and credited to Koray Kavukcuoglu. Read it closely and "Gemini 4" never appears as a standalone model. Every capability claim, every price and every benchmark in the post is attached to the full name. The bare form appears only where the name is used as a prefix — "Gemini 4 Argon's capabilities" — which is the same grammatical slot Gemini 3 occupied before Gemini 3 Pro, Gemini 3.1 Pro and four Flash variants each became their own thing.
The generational history supports reading it that way. Alphabet's July 23 earnings call put the framing on the record: Sundar Pichai said the next generation of frontier models "requires much larger base models," and named coding and agentic coding as the areas needing improvement. On September 23, at an event hosted by The Information, Kavukcuoglu said pre-training had finished and the model was in the early stages of post-training, that the team had seen the results, and that he hoped to ship an early post-training version well before the end of the year. Business Insider described the September 30 announcement as the first model of the 4 series — a series, which is the tell. All three of those are vendor statements about the vendor's own work, relayed by third parties.
What that means practically: "Gemini 4" is the generation, and the generation has one named member. The tier structure that Google has shipped for two years — a Pro at the top, a Flash below it, a Cyber variant gated behind the same programme — has not been populated for generation 4 yet. A "Gemini 4 Pro" or a "Gemini 4 Ultra" is an inference from Google's own naming pattern, not a leak and not a roadmap item. Treat any write-up that compares Gemini 4 Argon to "Gemini 4" as comparing one model to its own surname.
What Gemini 4 Argon actually is
Gemini 4 Argon was announced on September 30, 2026 on Google's blog, and Demis Hassabis announced the same model in his own words on X that evening. The rollout is phased, and only the first phase has a date: Argon goes first to vetted cyber defenders through the Fairwind Program, the application-only tier Google built for governments and national cyber authorities, critical-infrastructure operators and core technology platforms. Hassabis's own post said the company was rolling it out "starting with government and trusted cyber defenders through our Fairwind Program today." After that, Google's wording puts paid API customers and Google AI Ultra subscribers next, then developers, enterprises and consumers. None of those later stages carries a date.
What is not in that list is the part that matters for anyone planning work: a general-availability date, a model ID, a context window, a parameter count, a model card or a system card. Google published an output limit, not a context window. It is 1 million tokens, up from the 64K ceiling the previous Gemini family used, and Google calls it industry-leading. An order-of-magnitude increase in how much a single call can return is a real change for long agent trajectories and single-shot rendering, and it is the one technical figure in the announcement with no ambiguity.

The benchmark set is broad and entirely unreproduced. Google reports DeepSWE v1.1 at 77.9%, first place on AutomationBench at 51.3%, LVBench at 91.7%, a tie for first on CWE-bench v1 at 68%, leading results on the Vals Index, its Vals Finance Agent v2, Harvey's Legal Agent Benchmark and the strongest showing on Gray Swan's Indirect Prompt Injection benchmark. Every one of those is the vendor's own number, published by the vendor, and none has been independently reproduced. A separate set of figures circulating in September, presented as predicted Gemini 4 benchmarks with an AutomationBench score of 48.7%, turned out to be recycled from an unrelated launch and is not a Google publication. Google's actual reported AutomationBench figure is 51.3%, which is a different number from a different source.
The scoreboard the two names produce
Put the two side by side on the six dimensions this pairing actually turns on, and the shape of the problem is obvious:
• Exists as a named product — Gemini 4: no. Gemini 4 Argon: yes, announced 2026-09-30.
• Callable today — Gemini 4: no model ID or endpoint exists. Gemini 4 Argon: no model ID or endpoint exists either; access is Fairwind-gated and application-only.
• Price — Gemini 4: none published. Gemini 4 Argon: $2 / $10 per million tokens introductory, then $4 / $20. Vendor-stated.
• Maximum output — Gemini 4: none published. Gemini 4 Argon: 1 million tokens, against 64K for the previous generation. Vendor-stated.
• Intelligence index — Gemini 4: not evaluated. Gemini 4 Argon: 53 on Artificial Analysis's Intelligence Index at high reasoning effort, with a cost of $1.99 per index task. Per Artificial Analysis.
• Independent leaderboard rank — Gemini 4: not ranked. Gemini 4 Argon: first on the Vals Index at 68.90% with a $15.68 cost per test, across 13.40B input and 142.01M output tokens. Per Vals.

Read the last two rows together and you get the only genuinely interesting disagreement in this comparison. Artificial Analysis places Gemini 4 Argon at 53, which is level with Claude Fable 5.1 and GPT-6 Astra, and five points behind Claude Opus 5.5 at 58. Vals places it first outright. Both are independent, both are measuring something real, and they are measuring different things — AA's index weights a fixed set of ten evaluations across coding, reasoning and knowledge work, while the Vals Index weights four sectors by their share of US economic value added. A model can lead one and not the other without either board being wrong.
There is a second thing to notice in the Vals row, and it is the one an enterprise buyer should care about. Vals records Argon's rate card as $4/$20 — the post-introductory price — where Google's own announcement says $2/$10. If the $15.68 cost per test was computed at $4/$20, then a buyer accepting the introductory rate would see a materially lower number for the same work, and a buyer who misses the introductory window would not. Google has not said how long that window runs. That single undefined variable moves the most decision-relevant figure in the whole launch.
What you can call today
Nothing under the Gemini 4 name, and that includes us. We checked the catalogue: google/gemini-4, google/gemini-4-argon, google/gemini-4-pro and google/gemini-3.8-flash-cyber all return "model not found," and the same is true of every other routing platform, because none of them has a model ID to route. The Gemini API's own model list — Gemini 3.8 Flash, Gemini 3.8 Live, Gemini 3.1 Pro, Gemini 3 Flash — does not carry a Gemini 4 entry of any kind. Until Argon clears the Fairwind queue and gets an identifier, there is no billed price and no endpoint anywhere.
The live Google pro tier is a different story. Gemini 3.1 Pro Preview has been in our catalogue since February 19, 2026 at $2.00 input and $12.00 output per million tokens, with a 1,048,576-token context window, a 65,536-token maximum output and multimodal input across text, image, audio, video and file. Eight months on, it is still labelled Preview on the Gemini API's own model list, which is itself the argument Argon exists. Push it past 200,000 input tokens and the rate doubles to $4 and $18.

That is the one place this comparison has an operational answer. Gemini 3.1 Pro Preview is live behind the same OpenAI-compatible endpoint as the other 200-plus models we route, at provider list price with no markup — so when Google eventually publishes an identifier and a billed rate for Gemini 4 Argon, and then a Pro-tier sibling if the naming pattern holds, the pass-through means those rates appear on our side the day they appear on Google's. In the meantime, if you want to measure the pro tier rather than read about it, that is the model to measure. If you want to try a model whose behaviour is still being characterised, automatic failover across providers is how you do it without betting a production path on a preview that Google itself describes as phasing in.
One caveat that belongs in the same breath as the price
Bloomberg reported on September 30 that some Google employees with direct access to Argon say it does less well on real work than the benchmark scores imply, with front-end design named specifically and two employees using the term "benchmaxxing." Google disputes the characterisation. It is worth holding next to the benchmark list precisely because the benchmark list is the entire public evidence base for this model. When there is one vendor's numbers and no way to reproduce them, a report that the numbers do not tell the whole story is not a footnote. It is half the evidence.
Who should pick which
If you are choosing between these two names, you are not actually choosing anything yet, and the honest advice is to plan around that rather than around a comparison.
• If you need a Google pro-tier model in production this quarter, the answer is Gemini 3.1 Pro Preview. It is callable, it is priced, and it is the tier Argon is being positioned to replace. Waiting for Argon is waiting on a queue you cannot join.
• If you are the audience Argon is gated for — a government cyber authority, a critical-infrastructure operator, a software maintainer — the Fairwind Program is the only route, and the specificity of that gate is a fact about the model's capabilities, not marketing. Google trained it for vulnerability discovery and patching and did not ship it generally.
• If you are budgeting a 2026 migration, use $2/$10 for the introductory window and $4/$20 for everything after it. Do not use a single blended figure, and do not use the $15.68 cost-per-test from Vals without checking which rate card it was computed at.
• If you are waiting for "Gemini 4," wait for a model ID. The name will get one — Google's naming pattern across three generations says the Pro tier and the Flash tier get their own identifiers — but the pattern is not an announcement, and nothing has been published about what a bare Gemini 4 would be, cost or do.
The wider point is one this blog has had to make before and will presumably have to make again: a generation name is not a model. Gemini 4 Argon is the only thing Google has put a name, a price, a benchmark set and an access plan behind. The rest of the family is a pattern, and a pattern is a thing you can forecast but not a thing you can call.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
