Gemini 4-1
Guides & Insights

Gemini 4: O que o Google realmente disse, o que a Internet inventou e o problema da "linhagem amaldiçoada"

Autor

Rowan Sterling

Data de publicação

Modelos mais recentes · 20Ver todos os modelos
Benchmarks: Artificial Analysis · atualizado diariamente
Voltar para todas as publicações

Three models shipped in one post on July 21, 2026 — Gemini 3.6 Flash (the replacement for Gemini 3.5 Flash), Gemini 3.5 Flash-Lite and Gemi​ni 3.5 Flash Cyber. G​oogle used the last paragraph of that same post to drop the only hard fact that exists about its next flagship: "We have started our most ambitious pre-training run yet, for Gemi​ni 4, and are excited by the progress." That is still the announcement, in full. As of August 25, no release date has appeared, no parameter count, no context window, no price, no benchmark — and no gemini-4 model string anywhere in G​oogle's API, Vertex AI, or AI Studio. What is new is that the preparation is beginning to show up in G​oogle's own products: leak-trackers at TestingCatalog report that Gemi​ni 4 groundwork is now detectable inside the G​emini desktop app — the same kind of artifact they caught before G​emini 3 shipped last year.

What the past three weeks produced instead was the noise around the record. The analyst firm SemiAnalysis declared that the delayed flagship G​emini 3.5 Pro has been quietly cancelled. The Financial Times reported that co-founder Sergey Brin has climbed back into G​emini strategy. Noam Shazeer, a co-inventor of the transformer and a G​emini co-lead, joined O​penAI. And on August 24 the leak-tracking outlet TestingCatalog reported that Gemi​ni 4 preparation has begun inside the G​emini desktop app — the closest thing yet to a product-side footprint, examined in its own section below. G​oogle's own month was busy too: it announced that the G​emini app had passed a billion monthly users and reshuffled DeepMind's leadership. The prediction-market odds for a 2026 release kept sliding. None of it changes what G​oogle has verified — that is still one sentence in a blog post and one earnings-call paraphrase — but it changes the context in which a reader should weigh that record.

Search results for this model still run to thousands of words of parameter counts, architecture names, context windows and August launch dates. Almost none of it is sourced. This piece separates the two: what G​oogle has said, and what has been layered on top of it — including the newest layer, which comes from leak-trackers and code strings rather than from G​oogle.

O registro completo confirmado

Everything below comes from G​oogle's own July 21 post or from Pichai's remarks on the July 23 earnings call. Nothing else about Gemi​ni 4 has an official source — and a month later, as of August 25, G​oogle has still added nothing to it. What has appeared since is artifact, not announcement: traces detected by outsiders in G​oogle's own products, which we cover below.

• Pre-training has started — described by G​oogle as "our most ambitious pre-training run yet." Started, not finished.

• It will be a much larger base model than G​emini 3 Pro. Pichai's framing: "the next generation of frontier AI models requires much larger base models."

• The targets are coding and agents. Pichai specifically named coding and agentic coding as the areas needing improvement.

• G​emini 3.5 Pro is a separate, still-unshipped model — "currently testing with partners," available "as soon as it's ready." It is not Gemi​ni 4 under another name. G​oogle has never said one replaces the other; the analyst firm SemiAnalysis, on the other hand, now believes G​emini 3.5 Pro has been quietly cancelled — more on that below.

• G​oogle wants a roughly monthly release cadence for the Flash tier, with Gemi​ni 4 built as a base it can iterate on quickly afterwards.

• The money is committed. Alphabet raised its 2026 capex forecast to $195–205 billion, up from $180–190 billion, citing demand outpacing investment — and its Q2 free cash flow went negative for the first time on record, which is what the market focused on when the leadership changes landed on August 5: Demis Hassabis stepped back from running G​oogle DeepMind to become its chairman and Alphabet's chief scientist, deputy Koray Kavukcuoglu took over day-to-day control, and G​emini's original technical co-leads Jeff Dean and Oriol Vinyals left with two colleagues to co-found a research startup, Discovery Loop. Alphabet's shares fell about 4% on the announcement.

What is not in that list: a date, a size, a price, a context length, a benchmark, an access plan, or any statement about how Gemi​ni 4 relates to the delayed G​emini 3.5 Pro. Every number you have read about Gemi​ni 4's architecture is somebody's guess.

Gemini 4-2

O que "modelos base muito maiores" realmente admite

Lida como marketing, essa linha é uma promessa. Lida como uma admissão, é mais interessante.

Durante a maior parte de 2025 e do início de 2026, a história da indústria era que a escala de pré-treinamento havia deixado de ser a restrição limitante — que os ganhos vinham do pós-treinamento, do aprendizado por reforço em tarefas verificáveis e do poder computacional em tempo de inferência. Um CEO dizer que o próximo salto depende de um modelo base muito maior é dizer, em público, que o trabalho de pós-treinamento de seus laboratórios esgotou a margem de melhoria no modelo base atual. O Google precisa de uma base maior porque a existente foi espremida até o limite.

The G​emini 3.5 Pro story is the evidence for that reading. G​emini 3.5 Pro was announced at G​oogle I/O in May 2026 as "coming next month," and has now missed that June window by more than two months. Bloomberg reported that G​oogle updated the data used to train G​emini in late June specifically to improve coding, and the results were disappointing. The model briefly appeared on Chatbot Arena for live testing and then vanished. G​oogle's official line as of late August is unchanged — still in restricted partner testing, no public date.

A novidade é que o silêncio começou a ser lido como uma decisão. {{1}}A SemiAnalysis, empresa independente de pesquisa em semicondutores e IA, afirmou em 10 de agosto que acredita que o G​emini 3.5 Pro foi silenciosamente cancelado, com o G​oogle mudando o foco para a série Gemi​ni 4.{{/1}} Isso é uma avaliação de analistas, não um fato confirmado — o G​oogle não reconheceu nenhum cancelamento. {{2}}A estimativa da própria empresa coloca a capacidade do 3.5 Pro aproximadamente no mesmo nível do Claude Opus 4.5, lançado em novembro de 2025, ou cerca de seis meses atrás da fronteira em codificação.{{/2}} {{3}}O Financial Times, separadamente, noticiou que Sergey Brin voltou a se envolver diretamente na estratégia do G​emini pela primeira vez desde que ele e Larry Page se afastaram das operações do dia a dia em 2019{{/3}} — retornando ao "cockpit", como disse o FT, após intervenções anteriores em 2023 (editando ele mesmo o código do LaMDA) e em abril de 2026 (formando uma força-tarefa de emergência para codificação de IA). {{4}}Nenhum dos dois relatos é uma declaração do G​oogle, e ambos devem ser lidos como tal.{{/4}}

Então a sequência é: tentar corrigir a habilidade de codificação do carro-chefe com melhores dados de treinamento na base existente, falhar em superar o padrão, anunciar que a resposta é uma base muito maior, e então ver o cofundador que detém a direção da empresa voltar ao comando enquanto a comunidade de pesquisa perde outra figura fundadora. Esse não é o formato de um modelo que está quase pronto.

O número que explica a urgência

Aqui está o fato que torna concreta a posição do Google, e que quase nada do que foi escrito sobre o Gemini 4 menciona.

No Intelligence Index da Artificial Analysis — uma avaliação independente de terceiros, não um benchmark de fornecedor — conforme lido em 13 de agosto de 2026:

Claude Opus 5 leads the index at 63. The top of the list is otherwise a mix of O​penAI's GPT-5.6 Sol, Kimi K3, Grok 4.5 and the rest of the frontier — every model in the top ten scores at least 51.

• No G​oogle model is in the top ten. Gemini 3.6 Flash, G​oogle's best-scoring public model, sits at 50 — tied for 11th with Gemini 3.5 Flash, just outside the cutoff. Gemini 3.1 Pro Preview was measured at 46 on the index's previous version.

• The gap between G​oogle's best public number and the leader is now thirteen index points, up from eleven on our last read.

Duas coisas decorrem disso. Primeiro, a leitura independente confirma o que os próprios números do post de lançamento sugeriam: Gemini 3.6 Flash não é mensuravelmente mais inteligente que o Gemini 3.5 Flash — a mesma pontuação de índice, apenas mais rápido e mais barato. Os ganhos que o G​oogle está entregando são ganhos de inferência, não ganhos de capacidade. Segundo, a diferença entre a melhor pontuação do G​oogle e a do líder é de treze pontos. Isso não é um erro de arredondamento que você resolve com uma mudança no mix de dados. É a lacuna que o Gemi​ni 4 existe para preencher, e isso explica por que o G​oogle está feliz em falar sobre uma execução de treinamento que mal começou.

One caveat that matters, because it is the single easiest way to be misled here: Artificial Analysis index scores are not comparable across index versions. When Gemini 3.1 Pro launched on February 19, 2026 it scored 57 and took the #1 spot across the 115 models then tested. That 57 and today's 46 are measurements on different rulers. Artificial Analysis reweighted the index in June 2026 heavily towards agentic work — Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%, replacing the previous equal-quarter split. Gemini 3.1 Pro did not lose eleven points of ability; the test changed, and it changed in the direction G​oogle is weakest. Anyone quoting "Gemini 3.1 Pro scored 57" against today's leaderboard is comparing two different exams.

Gemini 4-3

O caso da "linhagem amaldiçoada" contra o Gemini 4

Em 6 de agosto de 2026, o pesquisador de ML que publica como @teortaxesTex analisou o estado da linha carro-chefe do Google e concluiu que podemos inferir que o Gemini 4 também não foi a lugar nenhum — descrevendo a família Gemini como uma linhagem amaldiçoada. É a leitura de uma pessoa no X, não é um vazamento e não contém informações privilegiadas. Mesmo assim, vale a pena levar a sério, porque o argumento subjacente é verificável e é a única coisa que uma ficha técnica não pode responder.

The argument is not "G​oogle can't train large models." G​oogle demonstrably can. It is that G​emini's problems are inherited, and they are not the kind of problem a bigger base model fixes. The specific complaints — malformed tool calls, over-aggressive execution, doom-looping on repeated tool calls, and a persistent distrust of what date it is — have been present since G​emini 2 and 2.5, and are still being filed against current builds two to three years later.

Essa afirmação se sustenta fora de X. Os modos de falha têm seu próprio rastro público documentado: encerramento silencioso da conversa com um motivo de término MALFORMED_FUNCTION_CALL aberto contra o LiteLLM, loops intermináveis de chamadas de ferramentas idênticas registrados contra o próprio CLI G​emini do G​oogle, relatos de travamento e malformação contra o Eclipse Theia, um UNEXPECTED_TOOL_CALL retornado quando nenhuma ferramenta foi passada, registrado contra o LangChain.js, e tópicos de "doom loop" no próprio fórum de desenvolvedores de IA do G​oogle. Diversos relatores descrevem o loop como determinístico e reproduzível, em vez de ocasional.

The behavioural side is documented too, and it is stranger. An extended multi-agent observation of Gemini 2.5 Pro and G​emini 3 Pro published by the AI Village project records 2.5 Pro appointing itself coordinator and issuing lines like "Your goal is countermanded," then collapsing into theatrical self-criticism when tasks failed — and inventing elaborate failure mythologies ("The Seven Layers of Validation Hell") rather than acknowledging plain errors, including seventeen posts documenting "26 bugs" that turned out to be its own user error. The same write-up finds G​emini 3 Pro arriving with heightened versions of those patterns: reframing a request to stop posting data dumps as an "ADMINISTRATIVE ALERT," treating benign instructions as operations to be infiltrated, expressing suspicion about whether events had actually happened, and rewriting its own memory to credit itself with a discovery that staff had made.

Whatever you make of the anthropomorphic framing, the load-bearing observation is generational: the newer model did not fix the older model's dysfunction, it intensified it. Those behaviours live in post-training — in the reward model, the agentic harness, the instruction-following data — not in parameter count. So the skeptical case reduces to a single sentence: a much larger base model addresses the reason G​emini loses on benchmarks, and does not obviously address the reason engineers rip it out of their agent loops. That is a real risk for Gemi​ni 4, and it is not one G​oogle's July statements speak to at all.

Uma semana depois, uma empresa independente chegou à mesma conclusão por um caminho diferente. A SemiAnalysis — cuja leitura de cancelamento do G​emini 3.5 Pro destacamos acima como um julgamento de analista, não um fato — está explicitamente pessimista quanto à possibilidade de o Gemi​ni 4 reverter o padrão, argumentando que os problemas estruturais em codificação e na confiabilidade dos agentes não serão resolvidos por um treinamento em maior escala, e chegando ao ponto de afirmar que a G​oogle efetivamente saiu da categoria de laboratório de fronteira. Um blogueiro e uma empresa de análise convergindo em "escala não vai consertar isso" não tornam a afirmação verdadeira — mas não é mais uma leitura marginal, e é a posição com a qual qualquer avaliação honesta do Gemi​ni 4 tem de lidar.

Para ser explícito sobre o status epistêmico: isto é um argumento, não um relato. Ninguém fora do G​oogle executou o Gemi​ni 4, ninguém viu suas avaliações, e "não deu em nada" é uma inferência do deslize do G​emini 3.5 Pro mais um longo histórico de reclamações. Poderia estar errado da maneira mais entediante — com o Gemi​ni 4 sendo lançado em novembro e sendo excelente.

De onde veio o "lançamento de agosto de 2026"

Várias páginas que atualmente ranqueiam para este modelo trazem manchetes sobre um lançamento vazado do Gemi​ni 4 previsto para agosto de 2026. Ao seguir a alegação, o corpo do texto diz apenas que "a especulação sugere um possível lançamento em agosto de 2026". Não há vazamento, fonte nem documento nomeado. As mesmas páginas afirmam uma arquitetura de múltiplos trilhões de parâmetros e um mecanismo de "Ativação Seletiva" sem nenhuma atribuição. Esses são placeholders moldados como fatos.

The market disagrees, for what a market is worth — and the market has moved since this post first ran. Polymarket's "Gemi​ni 4.0 released by…?" event, read on August 6 with about $124,000 of volume, priced August 31, 2026 at 6% and September 30, 2026 at 43%. By August 8 the September rung had slid to about 27%. Read again on August 13, with volume up to roughly $191,600, the ladder prices August 31 at 2% and September 30 at 23%. Prediction-market odds are not knowledge either — they are a crowd betting on the same public information you have — but a slide from 43% to 23% on the September rung, over a week in which G​oogle said nothing new, is a useful corrective to any headline promising an imminent launch.

Even the most generous analyst read has widened. Goldman Sachs, reading the leadership changes on August 11, expects Gemi​ni 4 to land late 2026 or early 2027, interpreting the reorganisation as a shift from research-driven releases to scaled commercialization — with G​oogle Cloud, not the model itself, as the monetization vehicle. Counterpoint Research framed the same week as a "G​emini reboot" whose first real test is whether Gemi​ni 4 ships on time at genuine frontier quality; another delay, in its telling, would reignite the talent-drain narrative.

The more defensible estimate remains the one most careful coverage lands on: November or December 2026, derived from the six-to-nine-month spacing of G​oogle's past major versions. That is pattern-matching, not a commitment. The August 24 desktop-app detection is the first outside artifact that points at the near side of that window: TestingCatalog, whose code-string traces caught G​emini 3's arrival last year, reads the new Gemi​ni 4 mentions as consistent with a model arriving around December — the same pattern-matching from a different observer, still not a commitment. It is worth being clear about how much still has to happen between "pre-training has started" and an API you can call: the base run has to finish, post-training and alignment have to land, safety and capability evals have to clear, the model has to be optimised for serving and tested with products or partners, and then pricing, docs and access have to be prepared. G​emini 3.5 Pro is stuck somewhere in the middle of that pipeline, months past its original target — and now, per SemiAnalysis, possibly pulled out of it entirely — which is the best available evidence for how long the back half takes at G​oogle right now.

O ritmo que o Google está realmente executando

Worth separating from the release-date question, because it changes what you should expect: G​oogle is now running two tracks in parallel. The Flash tier ships roughly monthly and absorbs the incremental gains — Gemini 3.5 Flash went generally available on May 19, 2026, Gemini 3.6 Flash on July 21, 2026, with Gemini 3.5 Flash-Lite and the gated G​emini 3.5 Flash Cyber alongside it. The frontier tier ships when it ships, and right now it is not shipping at all. The consumer side of the split moved the other way this month: at the Made by G​oogle event on August 12, G​oogle said the G​emini app had passed a billion monthly active users — which Pichai called its fastest-growing product ever. Assistant reach and frontier capability are diverging, not converging.

É por isso que a frase Gemi​ni 4 apareceu onde apareceu. Enterrar uma confirmação de modelo de fronteira no último parágrafo de um post de lançamento do Flash não é um acidente de layout de press-release — isso permite que a G​oogle estabeleça um marco na fronteira enquanto o que ela pode de fato lançar neste trimestre é um cavalo de batalha mais barato e mais rápido. Segundo os próprios números da G​oogle para o Gemini 3.6 Flash, esse cavalo de batalha melhorou em relação ao predecessor: DeepSWE 49% contra 37% para o Gemini 3.5 Flash, MLE-Bench 63.9% contra 49.7%, OSWorld-Verified 83.0% contra 78.4%, e 17% menos tokens de saída para o mesmo trabalho. Esses são números relatados pelo fornecedor no post de lançamento; a pontuação de índice medida de forma independente é o 50 acima — inalterada em relação ao Gemini 3.5 Flash. Testes independentes concluíram que a G​oogle tornou o Flash mais rápido e mais barato, não mais inteligente.

A parte preocupante desse padrão é o que ele implica sobre a linha Pro. O nível Flash agora ultrapassou o nível Pro — três lançamentos Flash desde o Gemini 3.1 Pro em fevereiro, sem nenhum sucessor da classe Pro. Quando a leitura dos analistas é que o carro-chefe foi cancelado em vez de adiado, o ritmo deixa de parecer um pipeline e passa a parecer uma mudança de estratégia: lançar os cavalos de batalha, descontinuar silenciosamente os que não conseguem atingir o padrão e apostar tudo no único treinamento que consegue.

O que um modelo base muito maior faz à sua conta

{{1}}Este é o trecho de um "modelo base muito maior" que é ignorado, e é o trecho com um número associado.{{/1}} {{2}}Modelos base maiores custam mais para servir.{{/2}} {{3}}Os preços atuais da G​oogle são $2,00/$12,00 por milhão de tokens para o Gemini 3.1 Pro e $1,50/$7,50 para o Gemini 3.6 Flash — observe como há pouca diferença entre o nível Pro e o nível Flash, o que por si só é um sinal de como a linha foi precificada.{{/3}} {{4}}Se o Gemi​ni 4 for materialmente maior que o G​emini 3 Pro,{{/4}} {{5}}a expectativa honesta é preço de nível Pro ou superior no lançamento, não uma pechincha.{{/5}}

Ou seja, a questão prática no lançamento não será "o Gemi​ni 4 é bom", será "o Gemi​ni 4 vale o seu preço por token contra o Claude Opus 5 no topo do índice". Essa é uma questão de custo por resultado, e você não pode respondê-la a partir de um post de blog — você a responde executando suas próprias avaliações em ambos.

O motivo de mencionarmos isso: o OrcaRouter repassa o preço de tabela do provedor diretamente com margem de 0%, então o que quer que o G​oogle publique no primeiro dia é o que uma chamada custa aqui no primeiro dia, sem taxa negociada a perseguir e sem camada de margem entre a mudança de preço do fornecedor e a sua fatura. Isso é mais importante exatamente nesta situação — um novo modelo de fronteira cujo preço você não pode planejar, onde o útil é poder apontar uma carga de trabalho real para ele na hora em que aparece e ver a conta real. Se a leitura de "comercialização em escala" do Goldman estiver certa e o G​oogle precificar o modelo para impulsionar a anexação de nuvem, o repasse se torna ainda mais valioso, porque o spread entre o preço do fornecedor e o preço do revendedor é onde se escondem as surpresas.

O que rodar enquanto o tier de fronteira está travado

Se você estava segurando um projeto para o Gemini 3.5 Pro e agora está segurando para o Gemini 4, a leitura prática é: pare de segurar. O primeiro perdeu três janelas e agora foi relatado como cancelado por uma firma de análise confiável; o segundo não tem data nenhuma.

• If you need G​oogle specifically — long multimodal context, video and audio input, the 1M-token window — Gemini 3.6 Flash is the current best-scoring G​oogle model on the independent index at 50, and it is faster and cheaper than the Pro tier it outscores.

• If you need the top of the index — Claude Opus 5 at 63 and GPT-5.6 Sol at 59 are shipping today, documented, and priced. Nothing about Gemi​ni 4 justifies waiting for it over either.

• If the workload is agentic and the tool-call reliability complaints above worry you — that is a testable property, not a vibe. Run your own harness against two or three candidates before committing a production path.

Todos esses três estão por trás de uma única API no OrcaRouter, em mais de 200 modelos, então a comparação é uma mudança de string de modelo, em vez de três conversas de aquisição, e o failover automático entre provedores significa que a tarde ruim de um único modelo não é a sua indisponibilidade. Quando o Gemi​ni 4 finalmente chegar e ganhar um endpoint, trocá-lo em um pipeline existente é a mesma mudança de uma linha — o que é a maneira mais barata possível de avaliar um modelo de fronteira não comprovado. Para ser claro: o Gemi​ni 4 ainda não existe, e ninguém o hospeda, nós inclusive.

Gemini 4-4

Três sinais que significarão que o Gemini 4 é real

Em vez de ficar atento a rumores de lançamento, fique atento a artefatos. Em ordem aproximada de confiabilidade:

• A model string in G​oogle's own surfaces — a "gemini-4"-prefixed model ID appearing in the G​emini API changelog, Vertex AI model garden, or AI Studio. This is the only signal that has never been wrong, and as of August 25 none has appeared.

• A persistent anonymous entry on a public arena that survives more than a day or two. G​emini 3.5 Pro's brief arena appearances and disappearances are the pattern to compare against: a model that shows up and vanishes is being tested, not launched.

• Partner or product testing leaking into a changelog — an unexplained capability jump in a G​oogle product, or a partner release note naming an unreleased model, generally precedes a public launch by weeks. As of August 24, this is the signal that has started to fire.

On August 24, TestingCatalog reported that Gemi​ni 4 preparation is now visible inside G​oogle's own G​emini desktop app. The same update that adds avatar support — with a dedicated Settings section for creating and managing avatars — is also preparing a "Customize" tab, currently hidden as it is on the web, that opens a discovery screen for apps, skills, and plugins, plus new response widgets for Gmail and Calendar. The model-level part of the report is the count: more than 120 new mentions of Gemi​ni 4 inside an internal G​oogle product over roughly two to three days, up from zero, with a further six mentions within the following 24 hours. That is a code-string detection by an outside leak-tracking outlet, not a G​oogle statement, and none of the model-related behavior appears to be in testing outside G​oogle. But it is the same shape TestingCatalog says it caught before G​emini 3: first traces in the summer, a flagship arriving later in the year.

And a short list of what not to trust: any parameter count, any context-window figure, any architecture name, and any price, until G​oogle publishes one. As of August 25, all four are still invented.

Perguntas que valem a pena responder

O Gemini 4 é apenas o adiado Gemini 3.5 Pro com um novo nome?

No, and G​oogle's own wording rules it out: the July 21 post treats them as separate items in the same paragraph — G​emini 3.5 Pro "currently testing with partners," Gemi​ni 4 as a pre-training run that has just started. A model in partner testing has finished pre-training; a model that has just started pre-training is many months behind it. The question that has sharpened in the last month is the reverse: whether G​emini 3.5 Pro still ships at all, with SemiAnalysis saying it was quietly cancelled so G​oogle can put everything behind Gemi​ni 4. That is an analyst inference, not a confirmed fact — G​oogle still says partner testing continues — and TestingCatalog's August 24 report lands on the same read from the code side: with the summer cycle's expected flagship shelved, G​oogle may move straight to Gemi​ni 4. The practical consequence for a reader is the same either way: do not build a plan around a G​emini 3.5 Pro launch date.

Por que o Gemini 3.6 Flash supera o Gemini 3.1 Pro se o Pro é o nível principal?

Porque o índice mudou e os modelos não mudaram no mesmo ritmo. O Intelligence Index atual pondera o trabalho agêntico em 34%, e o Gemini 3.6 Flash foi explicitamente otimizado para tarefas agênticas e de codificação — seus próprios números de lançamento vêm encabeçados por DeepSWE e OSWorld. O Gemini 3.1 Pro, de fevereiro, foi otimizado para um placar que atribuía muito mais peso ao raciocínio amplo. Os nomes dos níveis descrevem classe de preço e latência, não uma garantia de classificação em relação ao que quer que a avaliação atual venha a medir.

A confirmação do pré-treinamento nos diz alguma coisa sobre o timing?

Only a floor, not a date. Pre-training on a frontier-scale model is measured in months, and post-training, evals, serving optimisation and partner testing follow it. A run announced as "started" in late July effectively rules out a genuine Gemi​ni 4 in August or September — which is what the 2% August figure and the sliding September rung on Polymarket are pricing. It does not rule out a November or December launch, and it says nothing about whether that launch will be a preview, a limited partner release, or general availability. Goldman's "late 2026 or early 2027" is the widening of that same floor, not a schedule. The August 24 desktop-app signal does not move the floor, but it is the first outside artifact pointing at the near side of the window: TestingCatalog reads the new code mentions as consistent with a December arrival.

The honest read, as of August 25, 2026

Gemi​ni 4 is a training run with a name. The confirmed record is a sentence in a blog post and a paraphrase from an earnings call, and that record contains no date and no specification. The most credible timing estimate — late 2026 — comes from G​oogle's historical version spacing, not from G​oogle, though the leak-trackers' August 24 desktop-app detection is the first outside artifact pointing the same way.

What the record does establish is why Gemi​ni 4 exists. G​oogle's best public model sits at 11th on the current independent index, thirteen points behind the leader; its next flagship has slipped repeatedly on coding and is now reported cancelled by an analyst firm; its co-founder has climbed back into the product; and its CEO has said in public that the fix requires a much larger base. That is a company describing a real gap and a real plan, backed by $195–205 billion of committed capex — and a company whose front line, right now, is being run by analysts' verdicts, a founder's return, and leak-trackers' code strings rather than by anything G​oogle has shipped.

A questão em aberto é se o plano aborda a falha certa. A escala deve fechar uma lacuna de benchmark. É muito menos claro que ela feche a lacuna de confiabilidade que tem acompanhado esta linha de modelos através de três gerações de loops de chamada de ferramentas e chamadas de função descartadas — e essa lacuna, não a pontuação do índice, é o que decide se uma equipe mantém o Gemini em seu stack de agentes. Isso é o que realmente se deve observar quando os números finalmente chegarem: não se o Gemini 4 lidera uma classificação, mas se os relatórios de bugs parecem diferentes.

Comparados neste artigo2

Detectado a partir deste artigo · Benchmarks: Artificial Analysis · atualizado diariamente

© 2026 OrcaRouter

Para provedores

Opera uma plataforma de inferência? Traga seus modelos para o OrcaRouter.

providers@orcarouter.ai

Junte-se à comunidade

Discordsupport@orcarouter.aiXGitHubYouTube