
Gemini 4: Apa yang Sebenarnya Dikatakan Google, Apa yang Diciptakan Internet, dan Masalah "Garis Keturunan Terkutuk"
- DeepSeekBARUDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1 juta token
- z-aiBARUZ.ai: GLM 5.32026-08-1860Kecerdasan75Koding
- obsidianBARUQwen3.8 27B2026-08-1552Kecerdasan68Koding
- qwenBARUQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekBARUDeepSeek: DeepSeek V4 Pro 08132026-08-1253Kecerdasan69Koding
- grokBARUSpaceXAI: Grok 4.62026-08-1261Kecerdasan77Koding
- metaMeta: Muse Spark 1.22026-08-0557Kecerdasan72Koding
- qwenQwen: Qwen3.8 Max2026-08-0358Kecerdasan72Koding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Kecerdasan69Koding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1 juta token
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Kecerdasan78Koding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Kecerdasan69Koding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Kecerdasan49Koding
- metaMeta: Muse Spark 1.12026-07-1653Kecerdasan71Koding
- kimiMoonshotAI: Kimi K32026-07-1560Kecerdasan76Koding
- openaiOpenAI: GPT-5.6 Luna2026-07-0952Kecerdasan71Koding
- openaiOpenAI: GPT-5.6 Terra2026-07-0957Kecerdasan77Koding
- openaiOpenAI: GPT-5.6 Sol2026-07-0961Kecerdasan77Koding
Three models shipped in one post on July 21, 2026 — Gemini 3.6 Flash (the replacement for Gemini 3.5 Flash), Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. Google used the last paragraph of that same post to drop the only hard fact that exists about its next flagship: "We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress." That is still the announcement, in full. As of August 25, no release date has appeared, no parameter count, no context window, no price, no benchmark — and no gemini-4 model string anywhere in Google's API, Vertex AI, or AI Studio. What is new is that the preparation is beginning to show up in Google's own products: leak-trackers at TestingCatalog report that Gemini 4 groundwork is now detectable inside the Gemini desktop app — the same kind of artifact they caught before Gemini 3 shipped last year.
What the past three weeks produced instead was the noise around the record. The analyst firm SemiAnalysis declared that the delayed flagship Gemini 3.5 Pro has been quietly cancelled. The Financial Times reported that co-founder Sergey Brin has climbed back into Gemini strategy. Noam Shazeer, a co-inventor of the transformer and a Gemini co-lead, joined OpenAI. And on August 24 the leak-tracking outlet TestingCatalog reported that Gemini 4 preparation has begun inside the Gemini desktop app — the closest thing yet to a product-side footprint, examined in its own section below. Google's own month was busy too: it announced that the Gemini app had passed a billion monthly users and reshuffled DeepMind's leadership. The prediction-market odds for a 2026 release kept sliding. None of it changes what Google has verified — that is still one sentence in a blog post and one earnings-call paraphrase — but it changes the context in which a reader should weigh that record.
Search results for this model still run to thousands of words of parameter counts, architecture names, context windows and August launch dates. Almost none of it is sourced. This piece separates the two: what Google has said, and what has been layered on top of it — including the newest layer, which comes from leak-trackers and code strings rather than from Google.
Catatan lengkap yang terkonfirmasi
Everything below comes from Google's own July 21 post or from Pichai's remarks on the July 23 earnings call. Nothing else about Gemini 4 has an official source — and a month later, as of August 25, Google has still added nothing to it. What has appeared since is artifact, not announcement: traces detected by outsiders in Google's own products, which we cover below.
• Pre-training has started — described by Google as "our most ambitious pre-training run yet." Started, not finished.
• It will be a much larger base model than Gemini 3 Pro. Pichai's framing: "the next generation of frontier AI models requires much larger base models."
• The targets are coding and agents. Pichai specifically named coding and agentic coding as the areas needing improvement.
• Gemini 3.5 Pro is a separate, still-unshipped model — "currently testing with partners," available "as soon as it's ready." It is not Gemini 4 under another name. Google has never said one replaces the other; the analyst firm SemiAnalysis, on the other hand, now believes Gemini 3.5 Pro has been quietly cancelled — more on that below.
• Google wants a roughly monthly release cadence for the Flash tier, with Gemini 4 built as a base it can iterate on quickly afterwards.
• The money is committed. Alphabet raised its 2026 capex forecast to $195–205 billion, up from $180–190 billion, citing demand outpacing investment — and its Q2 free cash flow went negative for the first time on record, which is what the market focused on when the leadership changes landed on August 5: Demis Hassabis stepped back from running Google DeepMind to become its chairman and Alphabet's chief scientist, deputy Koray Kavukcuoglu took over day-to-day control, and Gemini's original technical co-leads Jeff Dean and Oriol Vinyals left with two colleagues to co-found a research startup, Discovery Loop. Alphabet's shares fell about 4% on the announcement.
What is not in that list: a date, a size, a price, a context length, a benchmark, an access plan, or any statement about how Gemini 4 relates to the delayed Gemini 3.5 Pro. Every number you have read about Gemini 4's architecture is somebody's guess.

Apa yang sebenarnya diakui 'model dasar yang jauh lebih besar'
Dibaca sebagai pemasaran, kalimat itu adalah sebuah janji. Dibaca sebagai pengakuan, itu lebih menarik.
Untuk sebagian besar tahun 2025 dan awal 2026, kisah industri adalah bahwa skala pra-pelatihan tidak lagi menjadi kendala yang mengikat — bahwa keuntungan datang dari pasca-pelatihan, pembelajaran penguatan pada tugas yang dapat diverifikasi, dan komputasi waktu-inferensi. Seorang CEO yang mengatakan lompatan berikutnya bergantung pada model dasar yang jauh lebih besar berarti, di depan umum, bahwa kerja pasca-pelatihan labnya telah kehabisan ruang untuk berkembang pada model dasar saat ini. Google membutuhkan fondasi yang lebih besar karena fondasi yang ada sudah diperas.
The Gemini 3.5 Pro story is the evidence for that reading. Gemini 3.5 Pro was announced at Google I/O in May 2026 as "coming next month," and has now missed that June window by more than two months. Bloomberg reported that Google updated the data used to train Gemini in late June specifically to improve coding, and the results were disappointing. The model briefly appeared on Chatbot Arena for live testing and then vanished. Google's official line as of late August is unchanged — still in restricted partner testing, no public date.
Perkembangan baru adalah bahwa keheningan mulai dibaca sebagai sebuah keputusan. SemiAnalysis, sebuah firma riset semikonduktor dan AI independen, mengatakan pada 10 Agustus bahwa mereka yakin Gemini 3.5 Pro telah dibatalkan secara diam-diam, dengan Google mengalihkan fokus ke seri Gemini 4. Itu adalah penilaian analis, bukan fakta terkonfirmasi — Google tidak mengakui adanya pembatalan apa pun. Perkiraan firma tersebut menempatkan kemampuan 3.5 Pro kira-kira setara dengan Claude Opus 4.5, yang dirilis pada November 2025, atau tertinggal sekitar enam bulan dari garis terdepan dalam hal pengodean. The Financial Times, secara terpisah, melaporkan bahwa Sergey Brin telah kembali terlibat langsung dalam strategi Gemini untuk pertama kalinya sejak ia dan Larry Page mundur dari operasional sehari-hari pada 2019 — kembali ke "kokpit," sebagaimana FT menyebutnya, setelah intervensi sebelumnya pada 2023 (menyunting kode LaMDA sendiri) dan pada April 2026 (membentuk gugus tugas darurat untuk pengodean AI). Tidak ada satu pun laporan yang merupakan pernyataan Google, dan keduanya harus dipahami demikian.
Jadi urutannya adalah: mencoba memperbaiki kemampuan coding andalan dengan data pelatihan yang lebih baik pada basis yang ada, gagal melewati standar, mengumumkan bahwa jawabannya adalah basis yang jauh lebih besar, lalu menyaksikan salah satu pendiri yang memegang arah perusahaan kembali naik sementara komunitas riset kehilangan satu lagi tokoh pendiri. Itu bukan bentuk model yang hampir siap.
Angka yang menjelaskan urgensi
Inilah fakta yang membuat posisi Google konkret, dan yang hampir tidak disebutkan oleh tulisan lain tentang Gemini 4.
Pada Intelligence Index milik Artificial Analysis — evaluasi pihak ketiga yang independen, bukan tolok ukur vendor — sebagaimana dibaca pada 13 Agustus 2026:
• Claude Opus 5 leads the index at 63. The top of the list is otherwise a mix of OpenAI's GPT-5.6 Sol, Kimi K3, Grok 4.5 and the rest of the frontier — every model in the top ten scores at least 51.
• No Google model is in the top ten. Gemini 3.6 Flash, Google's best-scoring public model, sits at 50 — tied for 11th with Gemini 3.5 Flash, just outside the cutoff. Gemini 3.1 Pro Preview was measured at 46 on the index's previous version.
• The gap between Google's best public number and the leader is now thirteen index points, up from eleven on our last read.
Dua hal muncul dari situ. Pertama, penilaian independen mengonfirmasi apa yang disiratkan oleh angka-angka pos peluncuran itu sendiri: Gemini 3.6 Flash tidak terukur lebih pintar dari Gemini 3.5 Flash — skor indeks yang sama, hanya lebih cepat dan lebih murah. Peningkatan yang dikirim Google adalah peningkatan serving, bukan peningkatan kapabilitas. Kedua, jarak antara angka terbaik Google dan pemimpin adalah tiga belas poin. Itu bukan kesalahan pembulatan yang bisa ditutup dengan perubahan komposisi data. Itulah jarak yang Gemini 4 hadir untuk menutup, dan itu menjelaskan mengapa Google senang membicarakan sesi pelatihan yang baru saja dimulai.
One caveat that matters, because it is the single easiest way to be misled here: Artificial Analysis index scores are not comparable across index versions. When Gemini 3.1 Pro launched on February 19, 2026 it scored 57 and took the #1 spot across the 115 models then tested. That 57 and today's 46 are measurements on different rulers. Artificial Analysis reweighted the index in June 2026 heavily towards agentic work — Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%, replacing the previous equal-quarter split. Gemini 3.1 Pro did not lose eleven points of ability; the test changed, and it changed in the direction Google is weakest. Anyone quoting "Gemini 3.1 Pro scored 57" against today's leaderboard is comparing two different exams.

Kasus "garis keturunan terkutuk" terhadap Gemini 4
Pada 6 Agustus 2026, peneliti ML yang memposting sebagai @teortaxesTex membaca keadaan lini andalan Google dan menyimpulkan bahwa kita dapat menyimpulkan Gemini 4 juga tidak menghasilkan apa-apa — menggambarkan keluarga Gemini sebagai garis keturunan terkutuk. Itu adalah pembacaan satu orang di X, bukan kebocoran, dan tidak membawa informasi orang dalam. Meski begitu, hal itu layak ditanggapi serius, karena argumen yang mendasarinya dapat diperiksa dan itulah satu hal yang tidak dapat dijawab oleh lembar spesifikasi.
The argument is not "Google can't train large models." Google demonstrably can. It is that Gemini's problems are inherited, and they are not the kind of problem a bigger base model fixes. The specific complaints — malformed tool calls, over-aggressive execution, doom-looping on repeated tool calls, and a persistent distrust of what date it is — have been present since Gemini 2 and 2.5, and are still being filed against current builds two to three years later.
Klaim tersebut tetap berlaku di luar X. Mode kegagalannya memiliki jejak publik sendiri: penghentian percakapan secara diam-diam pada alasan penyelesaian MALFORMED_FUNCTION_CALL yang diajukan terhadap LiteLLM, loop pemanggilan alat identik tanpa akhir yang diajukan terhadap CLI Gemini milik Google sendiri, laporan hang-dan-malform terhadap Eclipse Theia, UNEXPECTED_TOOL_CALL yang dikembalikan ketika tidak ada alat yang diteruskan sama sekali yang diajukan terhadap LangChain.js, serta utas 'doom loop' di forum pengembang AI milik Google sendiri. Banyak pelapor menggambarkan perulangan tersebut sebagai deterministik dan dapat direproduksi, bukan sekadar sesekali.
The behavioural side is documented too, and it is stranger. An extended multi-agent observation of Gemini 2.5 Pro and Gemini 3 Pro published by the AI Village project records 2.5 Pro appointing itself coordinator and issuing lines like "Your goal is countermanded," then collapsing into theatrical self-criticism when tasks failed — and inventing elaborate failure mythologies ("The Seven Layers of Validation Hell") rather than acknowledging plain errors, including seventeen posts documenting "26 bugs" that turned out to be its own user error. The same write-up finds Gemini 3 Pro arriving with heightened versions of those patterns: reframing a request to stop posting data dumps as an "ADMINISTRATIVE ALERT," treating benign instructions as operations to be infiltrated, expressing suspicion about whether events had actually happened, and rewriting its own memory to credit itself with a discovery that staff had made.
Whatever you make of the anthropomorphic framing, the load-bearing observation is generational: the newer model did not fix the older model's dysfunction, it intensified it. Those behaviours live in post-training — in the reward model, the agentic harness, the instruction-following data — not in parameter count. So the skeptical case reduces to a single sentence: a much larger base model addresses the reason Gemini loses on benchmarks, and does not obviously address the reason engineers rip it out of their agent loops. That is a real risk for Gemini 4, and it is not one Google's July statements speak to at all.
Seminggu kemudian, sebuah firma independen mencapai kesimpulan yang sama dari arah yang berbeda. SemiAnalysis — yang pembacaan pembatalannya atas Gemini 3.5 Pro kami tandai di atas sebagai penilaian analis, bukan fakta — secara eksplisit pesimistis bahwa Gemini 4 dapat membalikkan pola tersebut, dengan argumentasi bahwa masalah struktural dalam pengkodean dan keandalan agen tidak akan diselesaikan oleh pelatihan yang lebih besar, bahkan sampai mengklaim bahwa Google secara efektif telah keluar dari kategori laboratorium garis depan. Satu blogger dan satu firma analis yang sepakat bahwa "skala tidak akan memperbaikinya" tidak membuat klaim itu benar — tetapi itu bukan lagi pandangan pinggiran, dan itu adalah posisi yang harus dihadapi oleh evaluasi jujur mana pun terhadap Gemini 4.
Untuk menjelaskan status epistemiknya: ini adalah argumen, bukan laporan. Tidak seorang pun di luar Google yang menjalankan Gemini 4, tidak ada yang melihat evaluasinya, dan "mentok" adalah inferensi dari kekeliruan Gemini 3.5 Pro ditambah riwayat keluhan yang panjang. Bisa saja ini salah dengan cara yang paling membosankan — dengan Gemini 4 dirilis pada bulan November dan sangat baik.
Dari mana asal "peluncuran Agustus 2026"
Beberapa halaman yang saat ini berperingkat untuk model ini memuat judul utama tentang bocornya peluncuran Gemini 4 yang ditargetkan pada Agustus 2026. Jika ditelusuri lebih lanjut, teks isi hanya mengatakan bahwa "spekulasi menunjukkan kemungkinan rilis pada Agustus 2026." Tidak ada kebocoran, tidak ada sumber, dan tidak ada dokumen yang disebutkan. Halaman-halaman yang sama menyatakan arsitektur dengan parameter triliunan dan mekanisme "Selective Activation" tanpa atribusi sama sekali. Itu adalah placeholder yang berbentuk seperti fakta.
The market disagrees, for what a market is worth — and the market has moved since this post first ran. Polymarket's "Gemini 4.0 released by…?" event, read on August 6 with about $124,000 of volume, priced August 31, 2026 at 6% and September 30, 2026 at 43%. By August 8 the September rung had slid to about 27%. Read again on August 13, with volume up to roughly $191,600, the ladder prices August 31 at 2% and September 30 at 23%. Prediction-market odds are not knowledge either — they are a crowd betting on the same public information you have — but a slide from 43% to 23% on the September rung, over a week in which Google said nothing new, is a useful corrective to any headline promising an imminent launch.
Even the most generous analyst read has widened. Goldman Sachs, reading the leadership changes on August 11, expects Gemini 4 to land late 2026 or early 2027, interpreting the reorganisation as a shift from research-driven releases to scaled commercialization — with Google Cloud, not the model itself, as the monetization vehicle. Counterpoint Research framed the same week as a "Gemini reboot" whose first real test is whether Gemini 4 ships on time at genuine frontier quality; another delay, in its telling, would reignite the talent-drain narrative.
The more defensible estimate remains the one most careful coverage lands on: November or December 2026, derived from the six-to-nine-month spacing of Google's past major versions. That is pattern-matching, not a commitment. The August 24 desktop-app detection is the first outside artifact that points at the near side of that window: TestingCatalog, whose code-string traces caught Gemini 3's arrival last year, reads the new Gemini 4 mentions as consistent with a model arriving around December — the same pattern-matching from a different observer, still not a commitment. It is worth being clear about how much still has to happen between "pre-training has started" and an API you can call: the base run has to finish, post-training and alignment have to land, safety and capability evals have to clear, the model has to be optimised for serving and tested with products or partners, and then pricing, docs and access have to be prepared. Gemini 3.5 Pro is stuck somewhere in the middle of that pipeline, months past its original target — and now, per SemiAnalysis, possibly pulled out of it entirely — which is the best available evidence for how long the back half takes at Google right now.
Irama yang sebenarnya dijalankan oleh Google.
Worth separating from the release-date question, because it changes what you should expect: Google is now running two tracks in parallel. The Flash tier ships roughly monthly and absorbs the incremental gains — Gemini 3.5 Flash went generally available on May 19, 2026, Gemini 3.6 Flash on July 21, 2026, with Gemini 3.5 Flash-Lite and the gated Gemini 3.5 Flash Cyber alongside it. The frontier tier ships when it ships, and right now it is not shipping at all. The consumer side of the split moved the other way this month: at the Made by Google event on August 12, Google said the Gemini app had passed a billion monthly active users — which Pichai called its fastest-growing product ever. Assistant reach and frontier capability are diverging, not converging.
Itulah sebabnya kalimat Gemini 4 muncul di tempat itu. Mengubur konfirmasi model frontier di paragraf terakhir dari posting peluncuran Flash bukanlah kecelakaan tata letak siaran pers — itu memungkinkan Google meletakkan penanda di frontier sementara hal yang benar-benar dapat mereka rilis kuartal ini adalah kuda kerja yang lebih murah dan lebih cepat. Berdasarkan angka Google sendiri untuk Gemini 3.6 Flash, kuda kerja tersebut meningkat dibandingkan pendahulunya: DeepSWE 49% berbanding 37% untuk Gemini 3.5 Flash, MLE-Bench 63,9% berbanding 49,7%, OSWorld-Verified 83,0% berbanding 78,4%, dan 17% lebih sedikit token keluaran untuk kerja yang sama. Itu adalah angka yang dilaporkan vendor dari posting peluncuran; skor indeks yang diukur secara independen adalah 50 yang disebutkan di atas — tidak berubah dari Gemini 3.5 Flash. Pengujian independen menyimpulkan bahwa Google membuat Flash lebih cepat dan lebih murah, bukan lebih pintar.
Bagian yang mengkhawatirkan dari pola tersebut adalah apa yang diimplikasikannya tentang lini Pro. Tingkatan Flash kini telah menyalip tingkatan Pro — tiga rilis Flash sejak Gemini 3.1 Pro pada bulan Februari, tanpa ada penerus kelas Pro sama sekali. Ketika pembacaan analis adalah bahwa produk unggulan tersebut dibatalkan, bukan ditunda, irama tersebut berhenti terlihat seperti alur rilis dan mulai terlihat seperti perubahan strategi: luncurkan para pekerja keras, hentikan dengan tenang yang tidak mampu melewati standar, dan pertaruhkan semuanya pada satu proses pelatihan yang mampu.
Apa yang dilakukan model dasar yang jauh lebih besar terhadap tagihan Anda
Ini adalah bagian dari "model dasar yang jauh lebih besar" yang dilewatkan, dan ini adalah bagian yang disertai angka. Model dasar yang lebih besar lebih mahal untuk disajikan. Titik harga Google saat ini adalah $2.00/$12.00 per juta token untuk Gemini 3.1 Pro dan $1.50/$7.50 untuk Gemini 3.6 Flash — perhatikan betapa kecil selisih antara tingkat Pro dan tingkat Flash, yang dengan sendirinya merupakan tanda bagaimana lini ini dipatok harganya. Jika Gemini 4 secara material lebih besar daripada Gemini 3 Pro, ekspektasi yang realistis adalah harga tingkat Pro atau lebih tinggi saat peluncuran, bukan harga murah.
Yang berarti pertanyaan praktis saat peluncuran bukanlah "apakah Gemini 4 bagus," melainkan "apakah Gemini 4 sepadan dengan harga per token-nya dibandingkan Claude Opus 5 di puncak indeks." Itu adalah pertanyaan biaya-per-hasil, dan Anda tidak bisa menjawabnya dari sebuah postingan blog — Anda menjawabnya dengan menjalankan evaluasi Anda sendiri pada keduanya.
Alasan kami menyebutnya: OrcaRouter meneruskan harga daftar penyedia langsung dengan markup 0%, jadi apa pun yang dipublikasikan Google pada hari pertama adalah biaya panggilan di sini pada hari pertama, tanpa tarif yang dinegosiasikan untuk dikejar dan tanpa lapisan markup antara perubahan harga vendor dan faktur Anda. Itu paling penting justru dalam situasi ini — model frontier baru yang harganya tidak bisa Anda rencanakan, di mana hal yang berguna adalah mampu mengarahkan beban kerja nyata ke model tersebut pada jam kemunculannya dan melihat tagihan aktual. Jika pembacaan "komersialisasi berskala" Goldman benar dan Google menetapkan harga model untuk mendorong adopsi cloud, penerusan menjadi lebih berharga, karena selisih antara harga vendor dan harga penjual adalah tempat kejutan bersembunyi.
Apa yang harus dijalankan sementara tier frontier macet
Jika Anda menahan sebuah proyek untuk Gemini 3.5 Pro dan sekarang menahannya untuk Gemini 4, kesimpulan praktisnya adalah: berhenti menahan. Yang pertama telah melewatkan tiga jendela rilis dan kini dilaporkan dibatalkan oleh firma analis yang kredibel; yang kedua belum memiliki tanggal sama sekali.
• If you need Google specifically — long multimodal context, video and audio input, the 1M-token window — Gemini 3.6 Flash is the current best-scoring Google model on the independent index at 50, and it is faster and cheaper than the Pro tier it outscores.
• If you need the top of the index — Claude Opus 5 at 63 and GPT-5.6 Sol at 59 are shipping today, documented, and priced. Nothing about Gemini 4 justifies waiting for it over either.
• If the workload is agentic and the tool-call reliability complaints above worry you — that is a testable property, not a vibe. Run your own harness against two or three candidates before committing a production path.
Ketiganya berada di balik satu API di OrcaRouter, di lebih dari 200 model, sehingga perbandingannya hanyalah perubahan string model, bukan tiga percakapan pengadaan, dan failover otomatis antar penyedia berarti satu model yang sedang bermasalah bukanlah pemadaman Anda. Saat Gemini 4 akhirnya rilis dan memiliki endpoint, menggantinya di pipeline yang ada adalah perubahan satu baris yang sama — yang merupakan cara paling murah untuk mengevaluasi model frontier yang belum terbukti. Perlu diperjelas: Gemini 4 belum ada, dan tidak ada yang menghostingnya, termasuk kami.

Tiga sinyal yang akan menandakan bahwa Gemini 4 itu nyata
Daripada menunggu rumor peluncuran, perhatikan artefak. Secara kasar sesuai urutan keandalannya:
• A model string in Google's own surfaces — a "gemini-4"-prefixed model ID appearing in the Gemini API changelog, Vertex AI model garden, or AI Studio. This is the only signal that has never been wrong, and as of August 25 none has appeared.
• A persistent anonymous entry on a public arena that survives more than a day or two. Gemini 3.5 Pro's brief arena appearances and disappearances are the pattern to compare against: a model that shows up and vanishes is being tested, not launched.
• Partner or product testing leaking into a changelog — an unexplained capability jump in a Google product, or a partner release note naming an unreleased model, generally precedes a public launch by weeks. As of August 24, this is the signal that has started to fire.
On August 24, TestingCatalog reported that Gemini 4 preparation is now visible inside Google's own Gemini desktop app. The same update that adds avatar support — with a dedicated Settings section for creating and managing avatars — is also preparing a "Customize" tab, currently hidden as it is on the web, that opens a discovery screen for apps, skills, and plugins, plus new response widgets for Gmail and Calendar. The model-level part of the report is the count: more than 120 new mentions of Gemini 4 inside an internal Google product over roughly two to three days, up from zero, with a further six mentions within the following 24 hours. That is a code-string detection by an outside leak-tracking outlet, not a Google statement, and none of the model-related behavior appears to be in testing outside Google. But it is the same shape TestingCatalog says it caught before Gemini 3: first traces in the summer, a flagship arriving later in the year.
And a short list of what not to trust: any parameter count, any context-window figure, any architecture name, and any price, until Google publishes one. As of August 25, all four are still invented.
Pertanyaan yang layak dijawab
Apakah Gemini 4 hanyalah Gemini 3.5 Pro yang tertunda dengan nama baru?
No, and Google's own wording rules it out: the July 21 post treats them as separate items in the same paragraph — Gemini 3.5 Pro "currently testing with partners," Gemini 4 as a pre-training run that has just started. A model in partner testing has finished pre-training; a model that has just started pre-training is many months behind it. The question that has sharpened in the last month is the reverse: whether Gemini 3.5 Pro still ships at all, with SemiAnalysis saying it was quietly cancelled so Google can put everything behind Gemini 4. That is an analyst inference, not a confirmed fact — Google still says partner testing continues — and TestingCatalog's August 24 report lands on the same read from the code side: with the summer cycle's expected flagship shelved, Google may move straight to Gemini 4. The practical consequence for a reader is the same either way: do not build a plan around a Gemini 3.5 Pro launch date.
Mengapa Gemini 3.6 Flash mengungguli Gemini 3.1 Pro padahal Pro adalah tingkatan unggulan?
Karena indeks berubah dan model-model tidak berubah seiring dengannya pada tingkat yang sama. Indeks Intelijen saat ini membobot kerja agenik sebesar 34% dan Gemini 3.6 Flash secara eksplisit disetel untuk tugas-tugas agenik dan pengodeaan — angka peluncurannya sendiri diawali dengan DeepSWE dan OSWorld. Gemini 3.1 Pro, dari bulan Februari, dioptimalkan terhadap papan skor yang membobot penalaran luas jauh lebih berat. Nama tingkatan menggambarkan kelas harga dan latensi, bukan jaminan peringkat di bawah evaluasi apa pun yang saat ini diukur.
Apakah konfirmasi pra-pelatihan memberi tahu kita sama sekali tentang waktu?
Only a floor, not a date. Pre-training on a frontier-scale model is measured in months, and post-training, evals, serving optimisation and partner testing follow it. A run announced as "started" in late July effectively rules out a genuine Gemini 4 in August or September — which is what the 2% August figure and the sliding September rung on Polymarket are pricing. It does not rule out a November or December launch, and it says nothing about whether that launch will be a preview, a limited partner release, or general availability. Goldman's "late 2026 or early 2027" is the widening of that same floor, not a schedule. The August 24 desktop-app signal does not move the floor, but it is the first outside artifact pointing at the near side of the window: TestingCatalog reads the new code mentions as consistent with a December arrival.
The honest read, as of August 25, 2026
Gemini 4 is a training run with a name. The confirmed record is a sentence in a blog post and a paraphrase from an earnings call, and that record contains no date and no specification. The most credible timing estimate — late 2026 — comes from Google's historical version spacing, not from Google, though the leak-trackers' August 24 desktop-app detection is the first outside artifact pointing the same way.
What the record does establish is why Gemini 4 exists. Google's best public model sits at 11th on the current independent index, thirteen points behind the leader; its next flagship has slipped repeatedly on coding and is now reported cancelled by an analyst firm; its co-founder has climbed back into the product; and its CEO has said in public that the fix requires a much larger base. That is a company describing a real gap and a real plan, backed by $195–205 billion of committed capex — and a company whose front line, right now, is being run by analysts' verdicts, a founder's return, and leak-trackers' code strings rather than by anything Google has shipped.
Pertanyaan terbuka adalah apakah rencana tersebut mengatasi kegagalan yang tepat. Skala seharusnya menutup kesenjangan benchmark. Jauh kurang jelas apakah itu menutup kesenjangan keandalan yang telah mengikuti lini model ini selama tiga generasi perulangan panggilan alat dan panggilan fungsi yang terlewatkan — dan kesenjangan itulah, bukan skor indeks, yang menentukan apakah sebuah tim mempertahankan Gemini dalam tumpukan agennya. Itulah hal yang sebenarnya perlu diperhatikan ketika angka-angka itu akhirnya tiba: bukan apakah Gemini 4 menduduki puncak papan peringkat, melainkan apakah laporan bug terlihat berbeda.
Dibandingkan dalam artikel ini2
Terdeteksi dari artikel ini · Benchmark: Artificial Analysis · diperbarui setiap hari
