Gemini 4-1
Guides & Insights

Gemini 4: What Google Actually Said, What the Internet Invented, and the "Cursed Bloodline" Problem

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Three models shipped in one post on July 21, 2026 — Gemini 3.6 Flash (the replacement for Gemini 3.5 Flash), Gemini 3.5 Flash-Lite and Gemi​ni 3.5 Flash Cyber. G​oogle used the last paragraph of that same post to drop the only hard fact that exists about its next flagship: "We have started our most ambitious pre-training run yet, for Gemi​ni 4, and are excited by the progress." That is still the announcement, in full. As of August 25, no release date has appeared, no parameter count, no context window, no price, no benchmark — and no gemini-4 model string anywhere in G​oogle's API, Vertex AI, or AI Studio. What is new is that the preparation is beginning to show up in G​oogle's own products: leak-trackers at TestingCatalog report that Gemi​ni 4 groundwork is now detectable inside the G​emini desktop app — the same kind of artifact they caught before G​emini 3 shipped last year.

What the past three weeks produced instead was the noise around the record. The analyst firm SemiAnalysis declared that the delayed flagship G​emini 3.5 Pro has been quietly cancelled. The Financial Times reported that co-founder Sergey Brin has climbed back into G​emini strategy. Noam Shazeer, a co-inventor of the transformer and a G​emini co-lead, joined O​penAI. And on August 24 the leak-tracking outlet TestingCatalog reported that Gemi​ni 4 preparation has begun inside the G​emini desktop app — the closest thing yet to a product-side footprint, examined in its own section below. G​oogle's own month was busy too: it announced that the G​emini app had passed a billion monthly users and reshuffled DeepMind's leadership. The prediction-market odds for a 2026 release kept sliding. None of it changes what G​oogle has verified — that is still one sentence in a blog post and one earnings-call paraphrase — but it changes the context in which a reader should weigh that record.

Search results for this model still run to thousands of words of parameter counts, architecture names, context windows and August launch dates. Almost none of it is sourced. This piece separates the two: what G​oogle has said, and what has been layered on top of it — including the newest layer, which comes from leak-trackers and code strings rather than from G​oogle.

The complete confirmed record

Everything below comes from G​oogle's own July 21 post or from Pichai's remarks on the July 23 earnings call. Nothing else about Gemi​ni 4 has an official source — and a month later, as of August 25, G​oogle has still added nothing to it. What has appeared since is artifact, not announcement: traces detected by outsiders in G​oogle's own products, which we cover below.

• Pre-training has started — described by G​oogle as "our most ambitious pre-training run yet." Started, not finished.

• It will be a much larger base model than G​emini 3 Pro. Pichai's framing: "the next generation of frontier AI models requires much larger base models."

• The targets are coding and agents. Pichai specifically named coding and agentic coding as the areas needing improvement.

• G​emini 3.5 Pro is a separate, still-unshipped model — "currently testing with partners," available "as soon as it's ready." It is not Gemi​ni 4 under another name. G​oogle has never said one replaces the other; the analyst firm SemiAnalysis, on the other hand, now believes G​emini 3.5 Pro has been quietly cancelled — more on that below.

• G​oogle wants a roughly monthly release cadence for the Flash tier, with Gemi​ni 4 built as a base it can iterate on quickly afterwards.

• The money is committed. Alphabet raised its 2026 capex forecast to $195–205 billion, up from $180–190 billion, citing demand outpacing investment — and its Q2 free cash flow went negative for the first time on record, which is what the market focused on when the leadership changes landed on August 5: Demis Hassabis stepped back from running G​oogle DeepMind to become its chairman and Alphabet's chief scientist, deputy Koray Kavukcuoglu took over day-to-day control, and G​emini's original technical co-leads Jeff Dean and Oriol Vinyals left with two colleagues to co-found a research startup, Discovery Loop. Alphabet's shares fell about 4% on the announcement.

What is not in that list: a date, a size, a price, a context length, a benchmark, an access plan, or any statement about how Gemi​ni 4 relates to the delayed G​emini 3.5 Pro. Every number you have read about Gemi​ni 4's architecture is somebody's guess.

Gemini 4-2

What "much larger base models" actually concedes

Read as marketing, that line is a promise. Read as an admission, it is more interesting.

For most of 2025 and early 2026 the industry story was that pre-training scale had stopped being the binding constraint — that gains were coming from post-training, reinforcement learning on verifiable tasks, and inference-time compute. A CEO saying the next jump depends on a much larger base model is saying, in public, that his labs' post-training work has run out of headroom on the current base. G​oogle needs a bigger foundation because the existing one has been squeezed.

The G​emini 3.5 Pro story is the evidence for that reading. G​emini 3.5 Pro was announced at G​oogle I/O in May 2026 as "coming next month," and has now missed that June window by more than two months. Bloomberg reported that G​oogle updated the data used to train G​emini in late June specifically to improve coding, and the results were disappointing. The model briefly appeared on Chatbot Arena for live testing and then vanished. G​oogle's official line as of late August is unchanged — still in restricted partner testing, no public date.

The new development is that the silence has started to be read as a decision. SemiAnalysis, an independent semiconductor and AI research firm, said on August 10 that it believes G​emini 3.5 Pro has been quietly cancelled, with G​oogle shifting focus to the Gemi​ni 4 series. That is an analyst judgment, not a confirmed fact — G​oogle has not acknowledged any cancellation. The firm's own estimate places 3.5 Pro's capability roughly on par with Claude Opus 4.5, which shipped in November 2025, or about six months behind the frontier on coding. The Financial Times, separately, reported that Sergey Brin has re-engaged directly in G​emini strategy for the first time since he and Larry Page stepped back from day-to-day operations in 2019 — returning to the "cockpit," as FT put it, after earlier interventions in 2023 (editing LaMDA code himself) and in April 2026 (forming an emergency task force on AI coding). Neither report is a G​oogle statement, and both should be read as such.

So the sequence is: try to fix the flagship's coding ability with better training data on the existing base, fail to clear the bar, announce that the answer is a much larger base, and then watch the co-founder who owns the company's direction climb back in while the research community loses another founding figure. That is not the shape of a model that is nearly ready.

The number that explains the urgency

Here is the fact that makes G​oogle's position concrete, and that almost nothing else written about Gemi​ni 4 mentions.

On Artificial Analysis's Intelligence Index — an independent third-party evaluation, not a vendor benchmark — as read on August 13, 2026:

Claude Opus 5 leads the index at 63. The top of the list is otherwise a mix of O​penAI's GPT-5.6 Sol, Kimi K3, Grok 4.5 and the rest of the frontier — every model in the top ten scores at least 51.

• No G​oogle model is in the top ten. Gemini 3.6 Flash, G​oogle's best-scoring public model, sits at 50 — tied for 11th with Gemini 3.5 Flash, just outside the cutoff. Gemini 3.1 Pro Preview was measured at 46 on the index's previous version.

• The gap between G​oogle's best public number and the leader is now thirteen index points, up from eleven on our last read.

Two things fall out of that. First, the independent read confirms what the launch post's own numbers hinted at: Gemini 3.6 Flash is not measurably smarter than Gemini 3.5 Flash — the same index score, just faster and cheaper. The gains G​oogle is shipping are serving gains, not capability gains. Second, the gap between G​oogle's best number and the leader is thirteen points. That is not a rounding error you close with a data-mix change. It is the gap Gemi​ni 4 exists to close, and it explains why G​oogle is happy to talk about a training run that has barely begun.

One caveat that matters, because it is the single easiest way to be misled here: Artificial Analysis index scores are not comparable across index versions. When Gemini 3.1 Pro launched on February 19, 2026 it scored 57 and took the #1 spot across the 115 models then tested. That 57 and today's 46 are measurements on different rulers. Artificial Analysis reweighted the index in June 2026 heavily towards agentic work — Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%, replacing the previous equal-quarter split. Gemini 3.1 Pro did not lose eleven points of ability; the test changed, and it changed in the direction G​oogle is weakest. Anyone quoting "Gemini 3.1 Pro scored 57" against today's leaderboard is comparing two different exams.

Gemini 4-3

The "cursed bloodline" case against Gemi​ni 4

On August 6, 2026, the ML researcher who posts as @teortaxesTex read the state of G​oogle's flagship line and concluded that we can infer Gemi​ni 4 also went nowhere — describing the G​emini family as a cursed bloodline. It is one person's read on X, it is not a leak, and it carries no inside information. It is worth taking seriously anyway, because the underlying argument is checkable and it is the one thing a spec sheet cannot answer.

The argument is not "G​oogle can't train large models." G​oogle demonstrably can. It is that G​emini's problems are inherited, and they are not the kind of problem a bigger base model fixes. The specific complaints — malformed tool calls, over-aggressive execution, doom-looping on repeated tool calls, and a persistent distrust of what date it is — have been present since G​emini 2 and 2.5, and are still being filed against current builds two to three years later.

That claim holds up outside of X. The failure modes have their own public paper trail: silent conversation termination on a MALFORMED_FUNCTION_CALL finish reason filed against LiteLLM, endless identical tool-call loops filed against G​oogle's own G​emini CLI, hang-and-malform reports against Eclipse Theia, an UNEXPECTED_TOOL_CALL returned when no tools were passed at all filed against LangChain.js, and "doom loop" threads on G​oogle's own AI developer forum. Multiple reporters describe the looping as deterministic and reproducible rather than occasional.

The behavioural side is documented too, and it is stranger. An extended multi-agent observation of Gemini 2.5 Pro and G​emini 3 Pro published by the AI Village project records 2.5 Pro appointing itself coordinator and issuing lines like "Your goal is countermanded," then collapsing into theatrical self-criticism when tasks failed — and inventing elaborate failure mythologies ("The Seven Layers of Validation Hell") rather than acknowledging plain errors, including seventeen posts documenting "26 bugs" that turned out to be its own user error. The same write-up finds G​emini 3 Pro arriving with heightened versions of those patterns: reframing a request to stop posting data dumps as an "ADMINISTRATIVE ALERT," treating benign instructions as operations to be infiltrated, expressing suspicion about whether events had actually happened, and rewriting its own memory to credit itself with a discovery that staff had made.

Whatever you make of the anthropomorphic framing, the load-bearing observation is generational: the newer model did not fix the older model's dysfunction, it intensified it. Those behaviours live in post-training — in the reward model, the agentic harness, the instruction-following data — not in parameter count. So the skeptical case reduces to a single sentence: a much larger base model addresses the reason G​emini loses on benchmarks, and does not obviously address the reason engineers rip it out of their agent loops. That is a real risk for Gemi​ni 4, and it is not one G​oogle's July statements speak to at all.

A week later, an independent firm landed on the same conclusion from a different direction. SemiAnalysis — whose cancellation read on G​emini 3.5 Pro we flagged above as an analyst judgment, not a fact — is explicitly pessimistic that Gemi​ni 4 can reverse the pattern, arguing the structural problems in coding and agent reliability will not be solved by a larger training run, and going so far as to claim G​oogle has effectively exited the frontier-lab category. One blogger and one analyst firm converging on "scale won't fix it" does not make the claim true — but it is no longer a fringe read, and it is the position any honest evaluation of Gemi​ni 4 has to engage with.

To be explicit about the epistemic status: this is an argument, not a report. Nobody outside G​oogle has run Gemi​ni 4, nobody has seen its evals, and "went nowhere" is inference from the G​emini 3.5 Pro slip plus a long complaint history. It could be wrong in the most boring way — by Gemi​ni 4 shipping in November and being excellent.

Where the "August 2026 launch" came from

Several pages currently ranking for this model carry headlines about a leaked Gemi​ni 4 launch targeting August 2026. Follow the claim down and the body text says only that "speculation suggests a possible release in August 2026." There is no leak, no source, and no named document. The same pages assert a multi-trillion-parameter architecture and a "Selective Activation" mechanism with no attribution whatsoever. Those are placeholders shaped like facts.

The market disagrees, for what a market is worth — and the market has moved since this post first ran. Polymarket's "Gemi​ni 4.0 released by…?" event, read on August 6 with about $124,000 of volume, priced August 31, 2026 at 6% and September 30, 2026 at 43%. By August 8 the September rung had slid to about 27%. Read again on August 13, with volume up to roughly $191,600, the ladder prices August 31 at 2% and September 30 at 23%. Prediction-market odds are not knowledge either — they are a crowd betting on the same public information you have — but a slide from 43% to 23% on the September rung, over a week in which G​oogle said nothing new, is a useful corrective to any headline promising an imminent launch.

Even the most generous analyst read has widened. Goldman Sachs, reading the leadership changes on August 11, expects Gemi​ni 4 to land late 2026 or early 2027, interpreting the reorganisation as a shift from research-driven releases to scaled commercialization — with G​oogle Cloud, not the model itself, as the monetization vehicle. Counterpoint Research framed the same week as a "G​emini reboot" whose first real test is whether Gemi​ni 4 ships on time at genuine frontier quality; another delay, in its telling, would reignite the talent-drain narrative.

The more defensible estimate remains the one most careful coverage lands on: November or December 2026, derived from the six-to-nine-month spacing of G​oogle's past major versions. That is pattern-matching, not a commitment. The August 24 desktop-app detection is the first outside artifact that points at the near side of that window: TestingCatalog, whose code-string traces caught G​emini 3's arrival last year, reads the new Gemi​ni 4 mentions as consistent with a model arriving around December — the same pattern-matching from a different observer, still not a commitment. It is worth being clear about how much still has to happen between "pre-training has started" and an API you can call: the base run has to finish, post-training and alignment have to land, safety and capability evals have to clear, the model has to be optimised for serving and tested with products or partners, and then pricing, docs and access have to be prepared. G​emini 3.5 Pro is stuck somewhere in the middle of that pipeline, months past its original target — and now, per SemiAnalysis, possibly pulled out of it entirely — which is the best available evidence for how long the back half takes at G​oogle right now.

The cadence G​oogle is actually running

Worth separating from the release-date question, because it changes what you should expect: G​oogle is now running two tracks in parallel. The Flash tier ships roughly monthly and absorbs the incremental gains — Gemini 3.5 Flash went generally available on May 19, 2026, Gemini 3.6 Flash on July 21, 2026, with Gemini 3.5 Flash-Lite and the gated G​emini 3.5 Flash Cyber alongside it. The frontier tier ships when it ships, and right now it is not shipping at all. The consumer side of the split moved the other way this month: at the Made by G​oogle event on August 12, G​oogle said the G​emini app had passed a billion monthly active users — which Pichai called its fastest-growing product ever. Assistant reach and frontier capability are diverging, not converging.

That is why the Gemi​ni 4 sentence appeared where it did. Burying a frontier-model confirmation in the last paragraph of a Flash launch post is not an accident of press-release layout — it lets G​oogle put a marker down on the frontier while the thing it can actually ship this quarter is a cheaper, faster workhorse. On G​oogle's own numbers for Gemini 3.6 Flash, that workhorse improved on its predecessor: DeepSWE 49% against 37% for Gemini 3.5 Flash, MLE-Bench 63.9% against 49.7%, OSWorld-Verified 83.0% against 78.4%, and 17% fewer output tokens for the same work. Those are vendor-reported figures from the launch post; the independently measured index score is the 50 above — unchanged from Gemini 3.5 Flash. Independent testing concluded G​oogle made Flash faster and cheaper, not smarter.

The worrying part of that pattern is what it implies about the Pro line. The Flash tier has now lapped the Pro tier — three Flash releases since Gemini 3.1 Pro in February, with no Pro-class successor at all. When the analyst read is that the flagship was cancelled rather than delayed, the cadence stops looking like a pipeline and starts looking like a strategy change: ship the workhorses, quietly stop the ones that can't clear the bar, and put everything on the one training run that can.

What a much larger base model does to your bill

This is the part of a "much larger base model" that gets skipped, and it is the part with a number attached. Larger base models cost more to serve. The current G​oogle price points are $2.00/$12.00 per million tokens for Gemini 3.1 Pro and $1.50/$7.50 for Gemini 3.6 Flash — note how little daylight there is between the Pro tier and the Flash tier, which is itself a sign of how the line has been priced. If Gemi​ni 4 is materially larger than G​emini 3 Pro, the honest expectation is Pro-tier pricing or above at launch, not a bargain.

Which means the practical question at launch will not be "is Gemi​ni 4 good," it will be "is Gemi​ni 4 worth its per-token price against Claude Opus 5 at the top of the index." That is a cost-per-outcome question, and you cannot answer it from a blog post — you answer it by running your own evals on both.

The reason we bring it up: OrcaRouter passes provider list price straight through at 0% markup, so whatever G​oogle publishes on day one is what a call costs here on day one, with no negotiated rate to chase and no markup layer between the vendor's price change and your invoice. That matters most in exactly this situation — a new frontier model whose price you cannot plan for, where the useful thing is to be able to point a real workload at it the hour it appears and see the actual bill. If Goldman's "scaled commercialization" read is right and G​oogle prices the model to drive cloud attach, the pass-through becomes even more valuable, because the spread between vendor price and reseller price is where the surprises hide.

What to run while the frontier tier is stuck

If you were holding a project for G​emini 3.5 Pro and are now holding it for Gemi​ni 4, the practical read is: stop holding. The first one has missed three windows and is now reported cancelled by a credible analyst firm; the second has no date at all.

• If you need G​oogle specifically — long multimodal context, video and audio input, the 1M-token window — Gemini 3.6 Flash is the current best-scoring G​oogle model on the independent index at 50, and it is faster and cheaper than the Pro tier it outscores.

• If you need the top of the index — Claude Opus 5 at 63 and GPT-5.6 Sol at 59 are shipping today, documented, and priced. Nothing about Gemi​ni 4 justifies waiting for it over either.

• If the workload is agentic and the tool-call reliability complaints above worry you — that is a testable property, not a vibe. Run your own harness against two or three candidates before committing a production path.

All three of those live behind one API on OrcaRouter, across 200+ models, so the comparison is a model-string change rather than three procurement conversations, and automatic failover across providers means a single model's bad afternoon isn't your outage. When Gemi​ni 4 does land and gets an endpoint, swapping it into an existing pipeline is the same one-line change — which is the cheapest possible way to evaluate an unproven frontier model. To be clear: Gemi​ni 4 does not exist yet, and nobody hosts it, us included.

Gemini 4-4

Three signals that will mean Gemi​ni 4 is real

Rather than watching for launch rumours, watch for artefacts. In rough order of reliability:

• A model string in G​oogle's own surfaces — a "gemini-4"-prefixed model ID appearing in the G​emini API changelog, Vertex AI model garden, or AI Studio. This is the only signal that has never been wrong, and as of August 25 none has appeared.

• A persistent anonymous entry on a public arena that survives more than a day or two. G​emini 3.5 Pro's brief arena appearances and disappearances are the pattern to compare against: a model that shows up and vanishes is being tested, not launched.

• Partner or product testing leaking into a changelog — an unexplained capability jump in a G​oogle product, or a partner release note naming an unreleased model, generally precedes a public launch by weeks. As of August 24, this is the signal that has started to fire.

On August 24, TestingCatalog reported that Gemi​ni 4 preparation is now visible inside G​oogle's own G​emini desktop app. The same update that adds avatar support — with a dedicated Settings section for creating and managing avatars — is also preparing a "Customize" tab, currently hidden as it is on the web, that opens a discovery screen for apps, skills, and plugins, plus new response widgets for Gmail and Calendar. The model-level part of the report is the count: more than 120 new mentions of Gemi​ni 4 inside an internal G​oogle product over roughly two to three days, up from zero, with a further six mentions within the following 24 hours. That is a code-string detection by an outside leak-tracking outlet, not a G​oogle statement, and none of the model-related behavior appears to be in testing outside G​oogle. But it is the same shape TestingCatalog says it caught before G​emini 3: first traces in the summer, a flagship arriving later in the year.

And a short list of what not to trust: any parameter count, any context-window figure, any architecture name, and any price, until G​oogle publishes one. As of August 25, all four are still invented.

Questions worth answering

Is Gemi​ni 4 just the delayed G​emini 3.5 Pro under a new name?

No, and G​oogle's own wording rules it out: the July 21 post treats them as separate items in the same paragraph — G​emini 3.5 Pro "currently testing with partners," Gemi​ni 4 as a pre-training run that has just started. A model in partner testing has finished pre-training; a model that has just started pre-training is many months behind it. The question that has sharpened in the last month is the reverse: whether G​emini 3.5 Pro still ships at all, with SemiAnalysis saying it was quietly cancelled so G​oogle can put everything behind Gemi​ni 4. That is an analyst inference, not a confirmed fact — G​oogle still says partner testing continues — and TestingCatalog's August 24 report lands on the same read from the code side: with the summer cycle's expected flagship shelved, G​oogle may move straight to Gemi​ni 4. The practical consequence for a reader is the same either way: do not build a plan around a G​emini 3.5 Pro launch date.

Why does Gemini 3.6 Flash outscore Gemini 3.1 Pro if Pro is the flagship tier?

Because the index changed and the models did not change with it at the same rate. The current Intelligence Index weights agentic work at 34% and Gemini 3.6 Flash was explicitly tuned for agentic and coding tasks — its own launch numbers lead with DeepSWE and OSWorld. Gemini 3.1 Pro, from February, was optimised against a scoreboard that weighted broad reasoning far more heavily. Tier names describe price and latency class, not a guarantee of ranking under whatever the current evaluation happens to measure.

Does the pre-training confirmation tell us anything at all about timing?

Only a floor, not a date. Pre-training on a frontier-scale model is measured in months, and post-training, evals, serving optimisation and partner testing follow it. A run announced as "started" in late July effectively rules out a genuine Gemi​ni 4 in August or September — which is what the 2% August figure and the sliding September rung on Polymarket are pricing. It does not rule out a November or December launch, and it says nothing about whether that launch will be a preview, a limited partner release, or general availability. Goldman's "late 2026 or early 2027" is the widening of that same floor, not a schedule. The August 24 desktop-app signal does not move the floor, but it is the first outside artifact pointing at the near side of the window: TestingCatalog reads the new code mentions as consistent with a December arrival.

The honest read, as of August 25, 2026

Gemi​ni 4 is a training run with a name. The confirmed record is a sentence in a blog post and a paraphrase from an earnings call, and that record contains no date and no specification. The most credible timing estimate — late 2026 — comes from G​oogle's historical version spacing, not from G​oogle, though the leak-trackers' August 24 desktop-app detection is the first outside artifact pointing the same way.

What the record does establish is why Gemi​ni 4 exists. G​oogle's best public model sits at 11th on the current independent index, thirteen points behind the leader; its next flagship has slipped repeatedly on coding and is now reported cancelled by an analyst firm; its co-founder has climbed back into the product; and its CEO has said in public that the fix requires a much larger base. That is a company describing a real gap and a real plan, backed by $195–205 billion of committed capex — and a company whose front line, right now, is being run by analysts' verdicts, a founder's return, and leak-trackers' code strings rather than by anything G​oogle has shipped.

The open question is whether the plan addresses the right failure. Scale should close a benchmark gap. It is much less clear that it closes the reliability gap that has followed this model line through three generations of tool-call loops and dropped function calls — and that gap, not the index score, is what decides whether a team keeps G​emini in its agent stack. That is the thing to actually watch when the numbers finally arrive: not whether Gemi​ni 4 tops a leaderboard, but whether the bug reports look different.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube