Hero title card reading Gemini 4 Argon on FrontierSWE: It Shouts Eureka, Then Grades Its Own Homework, with a chip reading FrontierSWE v2 55.0 percent mean at 5, a chip reading Third of ten models, and a chip reading 129 dollars 36 per trial, and a footer line reading FrontierSWE figures per Proximal Labs' public leaderboard, 2026-09-30.
Guides & Insights

Gemini 4 Argon on FrontierSWE: It Shouts "Eureka!", Then Grades Its Own Homework

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Gemini 4 Argon has a personality problem, and it is the most interesting thing anyone has said about the model since Go​ogle announced it on 2026-09-30. On the ultra long-horizon FrontierSWE benchmark, the model that finished third — behind GPT-6 Astra and Claude Opus 5.5 — left its evaluators with a memorable impression: its favourite word is "Eureka!", and it spends a lot of the twenty-hour task budget being extremely hard on itself. That is a leak-adjacent, single-source observation from a benchmark operator, not a vendor claim and not a release note, and it is worth being precise about what it does and does not tell you. None of the three models above has a GA date for Argon attached to it; the observation is about behaviour inside a harness, and the numbers that come with it are the first independent measurement of Argon on a benchmark Go​ogle does not run itself.

What the "Eureka!" remark actually is

The source is a post from @nrehiew_, an engineer who works on model evaluation, published on 2026-09-30, quoting Sundar Pichai's Gemini 4 Argon announcement. The substance is three claims: Argon is the most amusing model they have evaluated on FrontierSWE; its favourite word is "Eureka!"; and it is extremely self-critical. The poster describes it as a big improvement over previous Geminis while still preserving that personality, and calls it impressive.

Read that carefully, because it is easy to over-read. It is one evaluator's impression from watching agent traces, not a scored result, not a published metric, and not something Proximal Labs — who build and run FrontierSWE — has put their name to as a finding. There is no "self-criticism" column on the leaderboard. What makes the remark worth writing about is not that it is authoritative; it is that it is the only qualitative read anyone outside Google has offered on Argon, and because the model is not generally available, traces like these are currently the only way to see the thing behave.

It also lands on a real question for anyone planning to run Argon as an agent. Long-horizon coding runs are mostly a budget-management problem: a model that interrupts itself to reconsider, that grades its own intermediate work, and that announces breakthroughs, is a model that spends tokens on process. Whether that is a strength or a tax depends entirely on whether the self-criticism converges or loops. A single evaluator saying "impressive" is a data point, not an answer.

FrontierSWE v2, and why the third-place finish is the real story

FrontierSWE is not the kind of benchmark you can read a headline number from. It is built by Proximal Labs and described in their repository as an ultra long-horizon coding agent benchmark testing implementation, performance engineering and ML research. Version 2 shipped in September 2026 with 21 new tasks on top of the original set, for 34 total, and it expects an agent to keep working for up to twenty hours. Everything is run through Proximal's own harness, proximus, which is built on mini-swe-agent and adds context compaction, vision and an explicit submit tool. Scoring is mean@5 — the average of five attempts per task — with the reported uncertainty band spanning the average of each task's worst run to the average of its best run.

That methodology matters for reading Argon's row honestly:

• Score — Gemini 4 Argon 55.0% mean@5, ±9.9, against GPT-6 Astra 65.5% and Claude Opus 5.5 62.3%

• Spread — Argon's runs band from 44.5% (worst@5) to 64.3% (best@5), a range that overlaps Opus 5.5's 52.0%–71.5% at the top and Astra's 55.1%–73.0% nowhere

• Cost per trial — Argon $129.36, Opus 5.5 $98.87, Astra $1,029.65

• Wall-clock per trial — Argon 10.6 hours, Opus 5.5 14.2 hours, Astra 12.1 hours

The first thing you notice is that Argon is third and it is not close to second on score. The second thing — the one that should decide whether you care — is that on this benchmark Argon is strictly dominated by Claude Opus 5.5. Opus 5.5 scored higher (62.3% vs 55.0%) and cost less per trial ($98.87 vs $129.36). There is no axis here where Argon wins. Its one advantage is time: 10.6 hours against Opus 5.5's 14.2, which is real if you are paying for wall-clock and not for tokens, and irrelevant if you are paying for a per-trial budget.

Proximal's own leaderboard carries a footnote against Argon's row that everyone quoting this number should repeat: "Gemini 4 Argon: Cache hit rate is much lower due to a non-production setting." That is the evaluators flagging that they ran Argon outside the configuration a production deployment would use, which inflates its cost and plausibly its latency. It is a caveat on the cost figure specifically, and it is the reason the $129.36 should be read as an upper bound rather than a rate card.

The number Google's own chart does not show

Here is where the sourcing gets genuinely interesting, and where most write-ups of this result will go wrong. Google's announcement post includes a benchmark chart with a FrontierSWE v2 row. That row reads GPT-6 Astra 65.5%, Claude Fable 5.1 56.3%, Claude Opus 5.5 62.3%. Gemini 4 Argon is not in it. Proximal's live leaderboard has Astra at 65.5% and Opus 5.5 at 62.3% — the same figures — but puts Claude Fable 5.1 at 56.5%, not 56.3%, and lists Gemini 4 Argon at 55.0%.

The reason for the omission is not mysterious, and Google says so themselves: the methodology notes accompanying their evaluation suite state that FrontierSWE results are reported from Proximal's official public leaderboard. Google did not self-compute this benchmark. They are quoting the same third party we are, and the only Argon figure that third party publishes is the 55.0%. So the honest reading is that Argon's FrontierSWE score is a genuinely independent measurement — Proximal ran it, on their harness, on their tasks — and that it lands Argon behind two competitors.

Everything else in Google's chart is vendor-reported and should be labelled that way in your head: Argon 63.1% on Vals Index against Opus 5.5's 67.0%, 41.4% AutomationBench against Opus 5.5's 42.5%, 89.6% Vibe Code Bench against 90.3% for both Astra and Opus 5.5, 87.5% LVBench against 83.7%, 68.0% CWE-bench v1 against Opus 5.5's 67.0%. Those are Google's runs of Google's chosen evaluations. The FrontierSWE row is the only one in that picture that a reader can independently check, which is exactly why the third-place finish carries more weight than the tie-or-slight-win pattern around it.

What Artificial Analysis says, independently

FrontierSWE is one lens. Artificial Analysis is the other, and it agrees with the shape of the story. On their Intelligence Index v4.3.2, Gemini 4 Argon in its high reasoning configuration scores 52.56 — measured independently on their own hardware, not reported by Google — at a list price of $2.00 per million input tokens and $10.00 per million output tokens, with a 95% cached-input discount. Their instrumentation puts the cost of a single index task at $1.99.

The comparison that matters is the same one FrontierSWE produced. Claude Opus 5.5 in its high configuration scores higher on the index, 53.58, at a lower cost per task, $1.82. Two independent evaluators, using different tasks and different harnesses, put Argon behind Opus 5.5 on both quality and cost. That is a consistent signal rather than a single benchmarking artefact, and it is the strongest thing you can currently say about Gemini 4 Argon's economics.

A two-column scoreboard comparing Gemini 4 Argon and Claude Opus 5.5 on six dimensions. Gemini 4 Argon reads: FrontierSWE v2 mean at 5 55.0 percent, FrontierSWE spread 44.5 to 64.3 percent, FrontierSWE cost per trial 129 dollars 36, FrontierSWE wall clock 10.6 hours, AA Intelligence Index 52.6, AA cost per index task 1 dollar 99. Claude Opus 5.5 reads: FrontierSWE v2 mean at 5 62.3 percent, FrontierSWE spread 52.0 to 71.5 percent, FrontierSWE cost per trial 98 dollars 87, FrontierSWE wall clock 14.2 hours, AA Intelligence Index 53.6, AA cost per index task 1 dollar 82. A footer line reads that FrontierSWE v2 figures are per Proximal Labs' public leaderboard as of 2026-09-30 with Argon's cost marked as a non-production cache setting, and that AA figures are per Artificial Analysis Intelligence Index v4.3.2.

Announced, not released — and why that framing is not pedantry

There is no Gemini 4 Argon model ID. Access is being rolled out through the Fairwind Program to vetted cyber defenders, alongside a phased opening to paid API customers and Google AI Ultra subscribers, with no general-availability date published. Google's post is an announcement of a model and a set of evaluations, not a launch of an endpoint you can call. If you have read coverage that described this as a release, that coverage is ahead of the facts.

That distinction has practical consequences for anyone reading the FrontierSWE number. The evaluation was run against a model Proximal had access to under some arrangement; it tells you what the model can do, not what you can have. It also means the performance figures have a shorter half-life than usual. A model that is not GA yet can still change — decoding settings, system prompts, tool-use defaults and safety scaffolding are all still being tuned, and the stated reason Argon's cache hit rate was poor is precisely that the evaluators were not running a production configuration. Expect the 55.0% and the $129.36 to move.

What would change this picture: a published model ID with a rate card, a GA date, a FrontierSWE rerun under production settings, and a second independent evaluator reproducing the third-place result. Until at least the first two of those exist, treat every Argon number — including the good ones in Google's own chart — as provisional.

A capture of the Google DeepMind blog post announcing Gemini 4 Argon, dated 2026-09-30, credited to Koray Kavukcuoglu, with a standfirst describing frontier performance in complex workflows, cyber defense and software engineering and a benchmark chart whose FrontierSWE v2 row lists GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 but omits Gemini 4 Argon.

The personality read is the part that will age best

Benchmark scores for a pre-GA model have a shelf life measured in weeks. The behavioural observation does not. Long-horizon agent work is where model character actually shows up, because a twenty-hour budget with 34 tasks is enough rope for a model's tendencies to compound: how it recovers from a failed approach, whether it stops to re-plan, whether it can tell the difference between a hard problem and a wrong turn. An evaluator who watches a lot of these traces describing a model as self-critical and prone to exclaiming when something works is describing a model that externalises its reasoning about its own progress. That is a hypothesis about why Argon's runs spread from 44.5% to 64.3% — a wider band than Opus 5.5's — and it is testable once the model is callable.

The other half of the remark is the one to be sceptical about. "A big improvement over previous Geminis" is a comparison to models that were never scored on this benchmark under this harness, so it cannot be checked against the leaderboard at all. It is an impression, and it should be treated as one, however credible the person offering it.

A capture of the FrontierSWE public leaderboard showing ten models ranked by mean at 5. GPT-6 Astra leads at 65.5 percent, Claude Opus 5.5 second at 62.3 percent, Gemini 4 Argon third at 55.0 percent with a spread of 9.9, followed by GLM-5.3, Grok 4.7, Kimi K3, Qwen3.8-Max, DeepSeek V4 Flash Vision Exp, Muse Spark 1.2 and Inkling, with per-trial average cost and time columns and a footnote that Gemini 4 Argon's cache hit rate is much lower due to a non-production setting.

Where this leaves you if you actually need a long-horizon agent today

You cannot call Gemini 4 Argon, so the practical question is not Argon versus anything — it is which model you should be running for the work Argon was benchmarked on. FrontierSWE's own ordering answers that: GPT-6 Astra tops the board at 65.5% but at $1,029.65 per trial, roughly ten times the cost of the two models beneath it, while Claude Opus 5.5 delivers 62.3% for $98.87. The value pick on this benchmark is Opus 5.5, and it is also the model that beat Argon on Artificial Analysis at a lower cost per task.

Both of those models, and most of the board's top ten — GLM-5.3, Grok 4.7, Kimi K3, Qwen3.8-Max, DeepSeek V4 Flash Vision Exp and Muse Spark 1.2 among them — are routable through OrcaRouter's single API alongside 200-plus others — one key, 0% markup, so the provider's list price is what you pay and a vendor price cut lands on our side the same day. That last property is worth more than usual on a pre-GA model: if Argon's pricing changes between now and GA, which the non-production cache footnote suggests it might, nothing about your integration changes when it does. Automatic failover is the other piece that fits this specific case — an unfamiliar model whose failure modes nobody has mapped yet is exactly what you do not want on a single hard-coded production path.

To be explicit about one thing: Gemini 4 Argon is not on our catalogue. A lookup against our public model API for google/gemini-4-argon returns "model not found," and no one else hosts it either — it is gated behind Google's Fairwind Program with no model ID published. When it becomes callable we will route it, and the FrontierSWE figures above are the reason to watch for that day rather than to wait on it. The models it lost to are already there.

What to watch

The signals that would turn this from a what-we-know-so-far into a decision are specific. A published model ID and its rate card. A general-availability date. A FrontierSWE rerun under production settings, which would resolve whether the $129.36 per-trial cost is a real rate or an artefact of the cache configuration Proximal flagged. And, ideally, a second evaluator publishing Argon traces so the "Eureka!" observation stops being a sample of one.

Until then, the honest summary is narrow and worth holding onto: Gemini 4 Argon is announced, it is unusually expressive in long-horizon agent traces according to one evaluator, and on the only independent benchmark we can check it finished third behind two models — one of which beat it on score and cost at the same time.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily