Hero title card for the article 'Gemini 4 Carbon — LEAK REPORT' with an 'UNVERIFIED' badge, the subtitle 'An internal Google checkpoint the public cannot call — what the report says', three chips reading 'Source: Business Insider via TestingCatalog', '9–10 October 2026' and 'Tested on Jetski, Google's internal platform', a left card reading 'The claim: early internal feedback puts Carbon comparable to Claude Opus 5.5 for coding', and a right card reading 'The context: Barium-B is the checkpoint chosen for the public Gemini 4 Argon release'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Gemini 4 Carbon Leak: Google's Internal Checkpoint That Staff Say Codes Like Claude Opus 5.5

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

There is a version of this story that is news and a version that is noise, and the difference is entirely in what the names are attached to. On 9 and 10 October 2026, Business Insider reported — and the leak-tracking outlet TestingCatalog relayed — that staff at the search giant are internally testing an additional Gemini 4 checkpoint called Carbon, that early internal feedback puts Carbon as feeling comparable to Claude Opus 5.5 for coding, that Barium-B is the checkpoint selected for the public Gemini 4 Argon release, and that Carbon, tested through the company's internal Jetski platform, is a separate and possibly stronger iteration. What makes that worth reading rather than skimming is the last clause: whether Carbon ships as an Gemini 4 Argon update or as a release of its own is, in the report's own words, unconfirmed. And that is the whole state of the evidence — no announcement, no model identifier, no endpoint, no price, no model card, and no comment from the vendor. Everything concrete on this page is a leak, an inference drawn from a leak, or a fact about a different, shippable model that the leak is measured against.

What the report actually claims

The signal is thin, and it is specific enough to be falsifiable, which is the only reason it belongs on a page like this one. Stripped of framing, the claims are these:

• Google employees are internally testing a Gemini 4 checkpoint named Carbon. It is not a public model and cannot be called from any API.

• Early internal feedback reportedly places Carbon as feeling comparable to Claude Opus 5.5 for coding — an impression from staff testing, not a benchmark result, and not a number anyone outside Google can reproduce.

• The checkpoint behind the model the public has actually heard of — Gemini 4 Argon — is called Barium-B, and Carbon is described as a separate and potentially more capable iteration rather than the same build under two labels.

• Carbon is reportedly being exercised through Jetski, which the same reporting describes as Google's internal name associated with Antigravity, its agentic coding environment.

• Whether Carbon becomes a future Gemini 4 Argon update or a distinct Gemini 4 release is unknown. Google has not commented on any of it.

That is the leak. It contains one comparison, one codename relationship, and one genuinely open question, and no part of it is a release.

A two-column infographic titled 'Gemini 4 Carbon — what the report says / what it does not'. Left column 'What the report says': '9–10 Oct 2026 — Business Insider, relayed by TestingCatalog', 'Google staff testing a Gemini 4 checkpoint named Carbon', 'Early internal feedback: feels comparable to Claude Opus 5.5 for coding', 'Barium-B is the checkpoint chosen for the public Gemini 4 Argon release', 'Tested via Jetski, tied to Antigravity'. Right column 'What it does not say': 'Whether Carbon ships as an Argon update or a separate Gemini 4 release', 'No model ID, endpoint or price of any kind', 'No model card, context window or parameter count', 'No benchmark number behind the Opus 5.5 comparison', 'No date, and no comment from Google'. Footer: 'All leak claims unverified; Google has not commented.'

The codename ladder, and which rung matters

The reason this report needs a decoder ring is that Google's periodic-table habit puts three names in one paragraph that look parallel and are not. Sorted by what a reader can actually do with each:

• Gemini 4 Argon — announced. A real model name with a first-party announcement post and a family page, a published introductory price of $2 per million input and $10 per million output tokens, a published 1-million-token output limit, and a gated rollout that started with vetted cyber defenders and was described as widening to paid API customers next. It still has no model identifier you can put in a request, which is the line between an announced model and a callable one.

• Barium-B — a checkpoint name, not a product. Per the same reporting it is the build selected for the public Argon release. That is the useful thing about the label: it tells you the name Argon is a shipping decision wrapped around an internal build rather than a build itself.

• Carbon — an internal checkpoint with staff impressions attached and nothing else. No ID, no price, no window, no date. It is a rumour with a codename, and codenames are cheap.

Read that way, the story is not "a new Gemini leaked". It is that Google appears to be running more than one frontier checkpoint in parallel, that the one it put its public name on is not the one its engineers are most enthusiastic about, and that nobody outside the company knows which of those two facts will turn out to matter.

Why "feels like Claude Opus 5.5" is the load-bearing sentence

Of everything in the report, the comparison is doing the most work and is the least verifiable, so it is worth being precise about what it is being compared to. Claude Opus 5.5 is Anthropic's flagship and, on independent measurement, the strongest widely-callable coding model of the current generation — the reference point that a Google engineer reaches for when they want to say "this is the good one" without naming a benchmark. On its published list price it runs $4.00 per million input tokens and $20.00 per million output tokens, with cache reads at $0.20 per million, and it carries a 1M-token input window and a 128K-token maximum output.

Put against that, the comparison says two things at once, and the second is the one people skip. The first is competitive: an internal Google checkpoint is being described in the same breath as the leading rival's flagship, which is a compliment Google would not put in a blog post. The second is positional: the comparison is a staff impression of coding feel, which is exactly the kind of claim that survives contact with a chat window and dies on contact with a leaderboard. Nobody has published a benchmark for Carbon. There is no Artificial Analysis entry, no vendor benchmark table, no eval methodology PDF — those exist for Gemini 4 Argon and they do not exist for this. A feeling is not a score, and this page is not going to dress one up as the other.

There is also a pricing consequence worth naming without pretending to know the answer. If a checkpoint that "feels like Claude Opus 5.5 for coding" is sitting inside Google and the model Google actually announced sits five points below it on the independent index, then the interesting question is not whether Carbon is good — it is where Google would price it, and whether it would ever be sold at the workhorse rates the rest of the Gemini line has been released at.

What is missing, and why the list is longer than the news

The honest inventory of what does not exist is the whole article, so here it is:

• A model identifier — Carbon has none, and neither does Gemini 4 Argon. Neither name appears anywhere a request can be addressed to.

• A price — nothing published for Carbon at any tier, not even an introductory rate.

• A model card or system card — none, and no published evaluation methodology behind it.

• A context window, an output limit, or a parameter count — none published, and Google has not published a parameter count for any Gemini.

• A benchmark result — the Opus 5.5 comparison is a staff impression, which means no number, no test, and no reproducibility.

• A date — no announcement, no staged rollout, no window, and no comment from Google on whether a Carbon release is planned at all.

• A relationship — whether Carbon is the next Argon update or a separate Gemini 4 model. This is the single most consequential unknown in the report and it is explicitly unconfirmed in the report itself.

Anything not on that list is inference, including most of what will get written about this over the next two weeks.

What a developer can do with this today

The useful answer is that almost nothing changes, and it is worth saying precisely why. Neither Gemini 4 Argon nor Carbon is on OrcaRouter — querying our own catalogue for either returns nothing, and it will keep returning nothing until Google makes one of them callable. The Gemini 4 name is not a thing you can route to; it is a name attached to a gated rollout. Any account offering "Gemini 4 Carbon" access right now is selling a label.

What is real is the model the leak is measured against. Claude Opus 5.5 is live on OrcaRouter at Anthropic's list price — $4.00 in and $20.00 out per million tokens, cache reads at $0.20 per million, passed through with zero markup — so the specific model that Carbon is being described as comparable to is one you can call today, on the same API key as everything else. That is a more useful fact than the leak, because it is the half of the comparison that can be tested.

The posture that survives a rumour is the one that does not need the rumour to be true. Put Claude Opus 5.5 behind your hardest coding workloads and keep a failover rule to a cheaper tier for the traffic that does not need it; the day a real Gemini 4 identifier appears in a provider catalogue, our list prices update the same day Google sets them, because there is no renegotiation step between a vendor publishing a price and us passing it through. One API for 200+ models turns "should we migrate?" into a routing-DSL edit and an A/B on live traffic rather than a rewrite. Route for outcomes you can measure, not names you read about.

A screenshot of Google's own Gemini API models documentation page, listing the Gemini 3.8 family and earlier models with their model codes and pricing, and containing no entry for Gemini 4 Argon or Gemini 4 Carbon.A screenshot of the OrcaRouter model page for anthropic/claude-opus-5.5, showing the model id, its capability chips, a 1M-token context window, a 128K maximum output, and pricing of .00 per 1M input tokens and 0.00 per 1M output tokens with cache reads at /bin/bash.20 per 1M.

What we are watching

• Whether Google says anything at all. A comment, a family-page change, or a line in a launch post would move all of this out of the leak column at once. Silence keeps it where it is.

• The Argon rollout's second stage. The announcement described paid API customers and Ultra subscribers as the next cohort after vetted defenders, and the arrival of a Gemini 4 identifier is the event that turns "announced" into "callable". If a runnable endpoint lands, the question of whether its checkpoint is Barium-B or something newer becomes concrete.

• Any independent measurement of Carbon. The moment a checkpoint like this appears on a neutral leaderboard rather than in a staff impression, the Opus 5.5 sentence stops being a vibe and starts being a data point — or stops being true.

• The price sheet. Where Google prices a frontier checkpoint that its own engineers prefer is the tell for whether Carbon is an Argon update or a tier of its own.

• Whether the codename relationship survives contact with a launch. Barium-B behind Argon and Carbon alongside it is a clean story today; launch posts have a habit of collapsing clean stories into one model with one ID.

The summary that does not oversell this is short. Google staff are reportedly testing an internal Gemini 4 checkpoint called Carbon, their early impression is that it codes like Claude Opus 5.5, the checkpoint behind the public Gemini 4 Argon release is a different one called Barium-B, and whether Carbon ever becomes a product — an update, a separate model, or nothing — is unconfirmed. The half of that sentence you can act on is the half that names Claude Opus 5.5, because it is the only model in the paragraph with a model identifier, a price, and an endpoint. Keep calling what you can call, and keep a failover rule ready for the day the other name becomes one.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily