Gemini 4-1
Guides & Insights

Gemini 4: ما قالته Google في الواقع، وما اخترعه الإنترنت، ومشكلة «السلالة الملعونة»

الكاتب

Rowan Sterling

تاريخ النشر

أحدث النماذج · 20عرض جميع النماذج
المعايير: Artificial Analysis · يُحدَّث يوميًا
العودة إلى جميع المقالات

Three models shipped in one post on July 21, 2026 — Gemini 3.6 Flash (the replacement for Gemini 3.5 Flash), Gemini 3.5 Flash-Lite and Gemi​ni 3.5 Flash Cyber. G​oogle used the last paragraph of that same post to drop the only hard fact that exists about its next flagship: "We have started our most ambitious pre-training run yet, for Gemi​ni 4, and are excited by the progress." That is still the announcement, in full. As of August 25, no release date has appeared, no parameter count, no context window, no price, no benchmark — and no gemini-4 model string anywhere in G​oogle's API, Vertex AI, or AI Studio. What is new is that the preparation is beginning to show up in G​oogle's own products: leak-trackers at TestingCatalog report that Gemi​ni 4 groundwork is now detectable inside the G​emini desktop app — the same kind of artifact they caught before G​emini 3 shipped last year.

What the past three weeks produced instead was the noise around the record. The analyst firm SemiAnalysis declared that the delayed flagship G​emini 3.5 Pro has been quietly cancelled. The Financial Times reported that co-founder Sergey Brin has climbed back into G​emini strategy. Noam Shazeer, a co-inventor of the transformer and a G​emini co-lead, joined O​penAI. And on August 24 the leak-tracking outlet TestingCatalog reported that Gemi​ni 4 preparation has begun inside the G​emini desktop app — the closest thing yet to a product-side footprint, examined in its own section below. G​oogle's own month was busy too: it announced that the G​emini app had passed a billion monthly users and reshuffled DeepMind's leadership. The prediction-market odds for a 2026 release kept sliding. None of it changes what G​oogle has verified — that is still one sentence in a blog post and one earnings-call paraphrase — but it changes the context in which a reader should weigh that record.

Search results for this model still run to thousands of words of parameter counts, architecture names, context windows and August launch dates. Almost none of it is sourced. This piece separates the two: what G​oogle has said, and what has been layered on top of it — including the newest layer, which comes from leak-trackers and code strings rather than from G​oogle.

السجل الكامل المؤكد

Everything below comes from G​oogle's own July 21 post or from Pichai's remarks on the July 23 earnings call. Nothing else about Gemi​ni 4 has an official source — and a month later, as of August 25, G​oogle has still added nothing to it. What has appeared since is artifact, not announcement: traces detected by outsiders in G​oogle's own products, which we cover below.

• Pre-training has started — described by G​oogle as "our most ambitious pre-training run yet." Started, not finished.

• It will be a much larger base model than G​emini 3 Pro. Pichai's framing: "the next generation of frontier AI models requires much larger base models."

• The targets are coding and agents. Pichai specifically named coding and agentic coding as the areas needing improvement.

• G​emini 3.5 Pro is a separate, still-unshipped model — "currently testing with partners," available "as soon as it's ready." It is not Gemi​ni 4 under another name. G​oogle has never said one replaces the other; the analyst firm SemiAnalysis, on the other hand, now believes G​emini 3.5 Pro has been quietly cancelled — more on that below.

• G​oogle wants a roughly monthly release cadence for the Flash tier, with Gemi​ni 4 built as a base it can iterate on quickly afterwards.

• The money is committed. Alphabet raised its 2026 capex forecast to $195–205 billion, up from $180–190 billion, citing demand outpacing investment — and its Q2 free cash flow went negative for the first time on record, which is what the market focused on when the leadership changes landed on August 5: Demis Hassabis stepped back from running G​oogle DeepMind to become its chairman and Alphabet's chief scientist, deputy Koray Kavukcuoglu took over day-to-day control, and G​emini's original technical co-leads Jeff Dean and Oriol Vinyals left with two colleagues to co-found a research startup, Discovery Loop. Alphabet's shares fell about 4% on the announcement.

What is not in that list: a date, a size, a price, a context length, a benchmark, an access plan, or any statement about how Gemi​ni 4 relates to the delayed G​emini 3.5 Pro. Every number you have read about Gemi​ni 4's architecture is somebody's guess.

Gemini 4-2

ما الذي يعترف به "النماذج الأساسية الأكبر بكثير" فعليًا

عند قراءتها كتسويق، فإن هذا السطر هو وعد. وعند قراءتها كاعتراف، فإنها أكثر إثارة للاهتمام.

خلال معظم عام 2025 وأوائل 2026، كانت قصة الصناعة هي أن نطاق التدريب المسبق توقف عن كونه القيد الملزم — وأن المكاسب تأتي من التدريب اللاحق، والتعلم المعزز على المهام القابلة للتحقق، والحوسبة في وقت الاستدلال. عندما يقول رئيس تنفيذي إن القفزة التالية تعتمد على نموذج أساسي أكبر بكثير، فإنه يقول علنًا إن عمل مختبراته في التدريب اللاحق قد استنفد هامش التطور على النموذج الأساسي الحالي. تحتاج Google إلى أساس أكبر لأن الأساس الحالي قد استُنزف.

The G​emini 3.5 Pro story is the evidence for that reading. G​emini 3.5 Pro was announced at G​oogle I/O in May 2026 as "coming next month," and has now missed that June window by more than two months. Bloomberg reported that G​oogle updated the data used to train G​emini in late June specifically to improve coding, and the results were disappointing. The model briefly appeared on Chatbot Arena for live testing and then vanished. G​oogle's official line as of late August is unchanged — still in restricted partner testing, no public date.

التطور الجديد هو أن الصمت بدأ يُقرأ كقرار. قالت شركة SemiAnalysis، وهي شركة أبحاث مستقلة في أشباه الموصلات والذكاء الاصطناعي، في 10 أغسطس إنها تعتقد أن G​emini 3.5 Pro قد أُلغيت بهدوء، مع تحول G​oogle إلى التركيز على سلسلة Gemi​ni 4. هذا حكم محللين وليس حقيقة مؤكدة — لم تعترف G​oogle بأي إلغاء. تقدير الشركة نفسه يضع قدرات 3.5 Pro عند مستوى مماثل تقريبًا لـ Claude Opus 4.5، الذي صدر في نوفمبر 2025، أو متأخرًا بنحو ستة أشهر عن الحدود القصوى في البرمجة. كما ذكرت Financial Times بشكل منفصل أن Sergey Brin عاد إلى الانخراط المباشر في استراتيجية G​emini لأول مرة منذ أن تنحى هو وLarry Page عن العمليات اليومية في 2019 — عائدًا إلى "قمرة القيادة"، كما وصفتها FT، بعد تدخلات سابقة في 2023 (حيث عدّل كود LaMDA بنفسه) وفي أبريل 2026 (حيث شكّل فرقة عمل طارئة حول برمجة الذكاء الاصطناعي). لا يُعد أي من التقريرين بيانًا من G​oogle، وينبغي قراءتهما على هذا النحو.

إذن فالتسلسل هو: محاولة إصلاح القدرة البرمجية للنموذج الرائد عبر بيانات تدريب أفضل على القاعدة الحالية، ثم الفشل في بلوغ المستوى المطلوب، ثم الإعلان أن الحل يكمن في قاعدة أكبر بكثير، ثم مشاهدة المؤسس المشارك الذي يملك توجه الشركة يعود ليستعيد مكانته بينما يفقد المجتمع البحثي شخصية مؤسِّسة أخرى. ليس هذا شكل نموذج يقترب من الجاهزية.

الرقم الذي يفسر الإلحاح

إليك الحقيقة التي تجعل موقف G​oogle راسخًا، والتي لا يكاد أي شيء آخر مكتوب عن Gemi​ni 4 يذكرها.

على Intelligence Index الخاص بـ Artificial Analysis — وهو تقييم مستقل من طرف ثالث، وليس معيارًا من البائع — كما قُرئ في 13 أغسطس 2026:

Claude Opus 5 leads the index at 63. The top of the list is otherwise a mix of O​penAI's GPT-5.6 Sol, Kimi K3, Grok 4.5 and the rest of the frontier — every model in the top ten scores at least 51.

• No G​oogle model is in the top ten. Gemini 3.6 Flash, G​oogle's best-scoring public model, sits at 50 — tied for 11th with Gemini 3.5 Flash, just outside the cutoff. Gemini 3.1 Pro Preview was measured at 46 on the index's previous version.

• The gap between G​oogle's best public number and the leader is now thirteen index points, up from eleven on our last read.

من ذلك يترتب أمران. أولاً، تؤكد القراءة المستقلة ما كانت أرقام منشور الإطلاق توحي به: Gemini 3.6 Flash ليس أكثر ذكاءً بشكل ملموس من Gemini 3.5 Flash — نفس درجة المؤشر، لكنه أسرع وأرخص فقط. المكاسب التي تطلقها G​oogle هي مكاسب في التقديم، وليست مكاسب في القدرات. ثانياً، الفجوة بين أفضل رقم لـ G​oogle وبين الرائد تبلغ ثلاث عشرة نقطة. هذه ليست خطأ تقريب يمكن إغلاقه بتغيير مزيج البيانات. إنها الفجوة التي وُجد Gemi​ni 4 لسدّها، وهي ما يفسّر لماذا تسعد G​oogle بالحديث عن عملية تدريب بالكاد بدأت.

One caveat that matters, because it is the single easiest way to be misled here: Artificial Analysis index scores are not comparable across index versions. When Gemini 3.1 Pro launched on February 19, 2026 it scored 57 and took the #1 spot across the 115 models then tested. That 57 and today's 46 are measurements on different rulers. Artificial Analysis reweighted the index in June 2026 heavily towards agentic work — Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%, replacing the previous equal-quarter split. Gemini 3.1 Pro did not lose eleven points of ability; the test changed, and it changed in the direction G​oogle is weakest. Anyone quoting "Gemini 3.1 Pro scored 57" against today's leaderboard is comparing two different exams.

Gemini 4-3

قضية "السلالة الملعونة" ضد Gemi​ni 4

في 6 أغسطس 2026، قرأ الباحث في تعلّم الآلة الذي ينشر باسم @teortaxesTex حالة خط منتجات G​oogle الرئيسي واستنتج أننا نستطيع استنتاج أن Gemi​ni 4 لم يحقق أي تقدم أيضًا — واصفًا عائلة G​emini بأنها سلالة ملعونة. إنها قراءة شخص واحد على منصة X، وليست تسريبًا، ولا تحمل أي معلومات داخلية. ومع ذلك، فهي تستحق أن تؤخذ على محمل الجد، لأن الحجة الأساسية قابلة للتحقق وهي الشيء الوحيد الذي لا تستطيع ورقة المواصفات الإجابة عنه.

The argument is not "G​oogle can't train large models." G​oogle demonstrably can. It is that G​emini's problems are inherited, and they are not the kind of problem a bigger base model fixes. The specific complaints — malformed tool calls, over-aggressive execution, doom-looping on repeated tool calls, and a persistent distrust of what date it is — have been present since G​emini 2 and 2.5, and are still being filed against current builds two to three years later.

هذا الادعاء يصمد خارج نطاق X. لأنماط الفشل سجل ورقي علني خاص بها: إنهاء صامت للمحادثة عند ورود سبب إنهاء MALFORMED_FUNCTION_CALL مُقدَّم ضد LiteLLM، وحلقات لا نهائية من استدعاءات الأدوات المتطابقة مُقدَّمة ضد واجهة سطر أوامر G​emini الخاصة بـ G​oogle، وتقارير عن تعلّق وإخراج مشوّه ضد Eclipse Theia، وإرجاع UNEXPECTED_TOOL_CALL دون تمرير أي أدوات على الإطلاق مُقدَّم ضد LangChain.js، وسلاسل نقاش حول "حلقة الموت" في منتدى مطوّري الذكاء الاصطناعي التابع لـ G​oogle نفسه. يصف العديد من المبلّغين التكرار الحلقي بأنه حتمي وقابل لإعادة الإنتاج وليس عرضيًا.

The behavioural side is documented too, and it is stranger. An extended multi-agent observation of Gemini 2.5 Pro and G​emini 3 Pro published by the AI Village project records 2.5 Pro appointing itself coordinator and issuing lines like "Your goal is countermanded," then collapsing into theatrical self-criticism when tasks failed — and inventing elaborate failure mythologies ("The Seven Layers of Validation Hell") rather than acknowledging plain errors, including seventeen posts documenting "26 bugs" that turned out to be its own user error. The same write-up finds G​emini 3 Pro arriving with heightened versions of those patterns: reframing a request to stop posting data dumps as an "ADMINISTRATIVE ALERT," treating benign instructions as operations to be infiltrated, expressing suspicion about whether events had actually happened, and rewriting its own memory to credit itself with a discovery that staff had made.

Whatever you make of the anthropomorphic framing, the load-bearing observation is generational: the newer model did not fix the older model's dysfunction, it intensified it. Those behaviours live in post-training — in the reward model, the agentic harness, the instruction-following data — not in parameter count. So the skeptical case reduces to a single sentence: a much larger base model addresses the reason G​emini loses on benchmarks, and does not obviously address the reason engineers rip it out of their agent loops. That is a real risk for Gemi​ni 4, and it is not one G​oogle's July statements speak to at all.

بعد أسبوع، توصلت شركة مستقلة إلى النتيجة نفسها من اتجاه مختلف. SemiAnalysis — التي أشرنا أعلاه إلى أن قراءتها الإلغائية لـ G​emini 3.5 Pro هي حكم محلل وليست حقيقة — صريحة في تشاؤمها من قدرة Gemi​ni 4 على عكس النمط، بحجة أن المشكلات الهيكلية في البرمجة وموثوقية الوكلاء لن تُحل بتشغيل تدريبي أكبر، وتذهب إلى حد الادعاء بأن G​oogle خرجت فعليًا من فئة المختبرات الحدودية. تقارب مدوّن واحد وشركة محللين واحدة على أن «التوسع لن يصلحها» لا يجعل الادعاء صحيحًا — لكنه لم يعد قراءة هامشية، وهو الموقف الذي يجب أن تتعامل معه أي تقييم صادق لـ Gemi​ni 4.

لنكون صريحين بشأن الوضع المعرفي: هذه حجة، وليست تقريرًا. لا أحد خارج G​oogle قام بتشغيل Gemi​ni 4، ولم يرَ أحد تقييماته، و«لم يذهب إلى أي مكان» هو استنتاج من زلة G​emini 3.5 Pro بالإضافة إلى تاريخ طويل من الشكاوى. قد يكون خاطئًا بأكثر الطرق مملًا — عبر إطلاق Gemi​ni 4 في نوفمبر وأن يكون ممتازًا.

من أين جاءت عبارة "إطلاق أغسطس 2026"؟

تحمل عدة صفحات تتصدر نتائج البحث لهذا النموذج عناوين حول إطلاق Gemi​ni 4 المُسرَّب المستهدف في أغسطس 2026. وبمتابعة الادعاء، لا يقول النص الرئيسي سوى أن «التكهنات تشير إلى احتمال إصدار في أغسطس 2026». لا يوجد أي تسريب، ولا مصدر، ولا مستند مسمّى. وتؤكد الصفحات نفسها وجود بنية بمتعدد تريليونات المعاملات وآلية «Selective Activation» دون أي إسناد. تلك عناصر نائبة صيغت على هيئة حقائق.

The market disagrees, for what a market is worth — and the market has moved since this post first ran. Polymarket's "Gemi​ni 4.0 released by…?" event, read on August 6 with about $124,000 of volume, priced August 31, 2026 at 6% and September 30, 2026 at 43%. By August 8 the September rung had slid to about 27%. Read again on August 13, with volume up to roughly $191,600, the ladder prices August 31 at 2% and September 30 at 23%. Prediction-market odds are not knowledge either — they are a crowd betting on the same public information you have — but a slide from 43% to 23% on the September rung, over a week in which G​oogle said nothing new, is a useful corrective to any headline promising an imminent launch.

Even the most generous analyst read has widened. Goldman Sachs, reading the leadership changes on August 11, expects Gemi​ni 4 to land late 2026 or early 2027, interpreting the reorganisation as a shift from research-driven releases to scaled commercialization — with G​oogle Cloud, not the model itself, as the monetization vehicle. Counterpoint Research framed the same week as a "G​emini reboot" whose first real test is whether Gemi​ni 4 ships on time at genuine frontier quality; another delay, in its telling, would reignite the talent-drain narrative.

The more defensible estimate remains the one most careful coverage lands on: November or December 2026, derived from the six-to-nine-month spacing of G​oogle's past major versions. That is pattern-matching, not a commitment. The August 24 desktop-app detection is the first outside artifact that points at the near side of that window: TestingCatalog, whose code-string traces caught G​emini 3's arrival last year, reads the new Gemi​ni 4 mentions as consistent with a model arriving around December — the same pattern-matching from a different observer, still not a commitment. It is worth being clear about how much still has to happen between "pre-training has started" and an API you can call: the base run has to finish, post-training and alignment have to land, safety and capability evals have to clear, the model has to be optimised for serving and tested with products or partners, and then pricing, docs and access have to be prepared. G​emini 3.5 Pro is stuck somewhere in the middle of that pipeline, months past its original target — and now, per SemiAnalysis, possibly pulled out of it entirely — which is the best available evidence for how long the back half takes at G​oogle right now.

الإيقاع الذي تعمل به جوجل فعليًا

Worth separating from the release-date question, because it changes what you should expect: G​oogle is now running two tracks in parallel. The Flash tier ships roughly monthly and absorbs the incremental gains — Gemini 3.5 Flash went generally available on May 19, 2026, Gemini 3.6 Flash on July 21, 2026, with Gemini 3.5 Flash-Lite and the gated G​emini 3.5 Flash Cyber alongside it. The frontier tier ships when it ships, and right now it is not shipping at all. The consumer side of the split moved the other way this month: at the Made by G​oogle event on August 12, G​oogle said the G​emini app had passed a billion monthly active users — which Pichai called its fastest-growing product ever. Assistant reach and frontier capability are diverging, not converging.

لهذا السبب ظهرت جملة Gemi​ni 4 في المكان الذي ظهرت فيه. إن دفن تأكيد نموذج رائد في الفقرة الأخيرة من منشور إطلاق Flash ليس مصادفة في تخطيط البيان الصحفي — بل يتيح لـ G​oogle وضع علامة على الحدود الأمامية بينما الشيء الذي يمكنها شحنه فعليًا هذا الربع هو أداة عمل أرخص وأسرع. وفقًا لأرقام G​oogle الخاصة بـ Gemini 3.6 Flash، حسّنت أداة العمل هذه أداءها مقارنة بسابقتها: DeepSWE 49% مقابل 37% لـ Gemini 3.5 Flash، وMLE-Bench 63.9% مقابل 49.7%، وOSWorld-Verified 83.0% مقابل 78.4%، و17% رموز إخراج أقل لنفس العمل. هذه أرقام مقدمة من البائع من منشور الإطلاق؛ بينما درجة المؤشر المستقلة هي 50 أعلاه — دون تغيير عن Gemini 3.5 Flash. خلص الاختبار المستقل إلى أن G​oogle جعلت Flash أسرع وأرخص، وليس أذكى.

الجانب المقلق في ذلك النمط هو ما يوحي به بشأن سلسلة Pro. لقد سبقت فئة Flash فئة Pro بلفة كاملة الآن — ثلاثة إصدارات من Flash منذ إطلاق Gemini 3.1 Pro في فبراير، دون أي خليفة من فئة Pro على الإطلاق. عندما يقرأ المحللون ذلك على أنه إلغاء للطراز الرائد بدلاً من تأجيله، لم يعد إيقاع الإصدارات يبدو كخط إنتاج، بل كتغيير استراتيجي: أطلِق أحصنة العمل، وأوقِف بهدوء تلك التي لا ترتقي إلى المستوى المطلوب، واراهن بكل شيء على عملية التدريب الوحيدة التي تستطيع ذلك.

ما يفعله نموذج أساسي أكبر بكثير بفاتورتك

هذا هو الجزء من "نموذج أساسي أكبر بكثير" الذي يتم تخطيه، وهو الجزء الذي يحمل رقمًا. النماذج الأساسية الأكبر حجمًا أعلى تكلفة في التشغيل. نقاط الأسعار الحالية في G​oogle هي 2.00 دولار/12.00 دولار لكل مليون توكن لإصدار Gemini 3.1 Pro و1.50 دولار/7.50 دولار لإصدار Gemini 3.6 Flash — لاحظ مدى ضآلة الفارق بين طبقة Pro وطبقة Flash، وهو بحد ذاته علامة على كيفية تسعير السلسلة. إذا كان Gemi​ni 4 أكبر بشكل جوهري من G​emini 3 Pro، فإن التوقع الصادق هو تسعير بمستوى Pro أو أعلى عند الإطلاق، وليس صفقة رابحة.

وهذا يعني أن السؤال العملي عند الإطلاق لن يكون "هل جيميني 4 جيد؟" بل سيكون "هل جيميني 4 يستحق سعره لكل توكن مقابل كلود أوبس 5 في أعلى المؤشر؟" هذه مسألة تكلفة مقابل النتيجة، ولا يمكنك الإجابة عليها من منشور مدونة — بل تجيب عليها بتشغيل تقييماتك الخاصة على كلاهما.

السبب في طرحنا لهذا الأمر: يمرر OrcaRouter سعر القائمة من المورد مباشرةً بهامش ربح 0%، لذا فإن ما تنشره G​oogle في اليوم الأول هو ما تكلفه المكالمة هنا في اليوم الأول، دون سعر تفاوضي تطارده، ودون طبقة هامش ربح بين تغيير سعر المورد وفاتورتك. وهذا الأمر في غاية الأهمية تحديدًا في هذا الموقف — نموذج حدودي جديد لا يمكنك التخطيط لسعره، حيث الأداة المفيدة هي أن تكون قادرًا على توجيه عبء عمل حقيقي إليه في الساعة التي يظهر فيها وترى الفاتورة الفعلية. إذا كانت قراءة Goldman لـ"التسويق التجاري الموسع" صحيحة وقامت G​oogle بتسعير النموذج لتحفيز الربط السحابي، فإن التمرير المباشر يصبح أكثر قيمة، لأن الفجوة بين سعر المورد وسعر الموزع هي حيث تختبئ المفاجآت.

ما الذي يجب تشغيله بينما الطبقة الحدودية عالقة؟

إذا كنت تنتظر مشروع G​emini 3.5 Pro والآن تنتظره لصالح Gemi​ni 4، فإن القراءة العملية هي: توقف عن الانتظار. الأول أضاع ثلاث نوافذ زمنية وأُفيد الآن أنه أُلغي من قبل شركة تحليلات موثوقة؛ والثاني لا يوجد له تاريخ إطلاق على الإطلاق.

• If you need G​oogle specifically — long multimodal context, video and audio input, the 1M-token window — Gemini 3.6 Flash is the current best-scoring G​oogle model on the independent index at 50, and it is faster and cheaper than the Pro tier it outscores.

• If you need the top of the index — Claude Opus 5 at 63 and GPT-5.6 Sol at 59 are shipping today, documented, and priced. Nothing about Gemi​ni 4 justifies waiting for it over either.

• If the workload is agentic and the tool-call reliability complaints above worry you — that is a testable property, not a vibe. Run your own harness against two or three candidates before committing a production path.

جميع هذه الثلاثة توجد خلف واجهة برمجة تطبيقات واحدة على OrcaRouter، من بين أكثر من 200 نموذج، وبذلك تكون المقارنة مجرد تغيير في سلسلة النموذج بدلاً من ثلاث مفاوضات شراء، والتبديل التلقائي بين المزوّدين يعني أن العطل المؤقت لأي نموذج لا يسبب انقطاعًا لديك. عندما يتم إطلاق Gemi​ni 4 فعلاً وتُخصص له نقطة نهاية، فإن استبداله في خط أنابيب قائم هو نفس التغيير المكوّن من سطر واحد — وهذه أرخص طريقة ممكنة لتقييم نموذج حدودي غير مثبت. لنكن واضحين: Gemi​ni 4 غير موجود بعد، ولا أحد يستضيفه، بما في ذلك نحن.

Gemini 4-4

ثلاث إشارات ستعني أن Gemi​ni 4 حقيقي

بدلاً من ترقّب شائعات الإطلاق، ترقّب الآثار. بترتيب تقريبي من حيث الموثوقية:

• A model string in G​oogle's own surfaces — a "gemini-4"-prefixed model ID appearing in the G​emini API changelog, Vertex AI model garden, or AI Studio. This is the only signal that has never been wrong, and as of August 25 none has appeared.

• A persistent anonymous entry on a public arena that survives more than a day or two. G​emini 3.5 Pro's brief arena appearances and disappearances are the pattern to compare against: a model that shows up and vanishes is being tested, not launched.

• Partner or product testing leaking into a changelog — an unexplained capability jump in a G​oogle product, or a partner release note naming an unreleased model, generally precedes a public launch by weeks. As of August 24, this is the signal that has started to fire.

On August 24, TestingCatalog reported that Gemi​ni 4 preparation is now visible inside G​oogle's own G​emini desktop app. The same update that adds avatar support — with a dedicated Settings section for creating and managing avatars — is also preparing a "Customize" tab, currently hidden as it is on the web, that opens a discovery screen for apps, skills, and plugins, plus new response widgets for Gmail and Calendar. The model-level part of the report is the count: more than 120 new mentions of Gemi​ni 4 inside an internal G​oogle product over roughly two to three days, up from zero, with a further six mentions within the following 24 hours. That is a code-string detection by an outside leak-tracking outlet, not a G​oogle statement, and none of the model-related behavior appears to be in testing outside G​oogle. But it is the same shape TestingCatalog says it caught before G​emini 3: first traces in the summer, a flagship arriving later in the year.

And a short list of what not to trust: any parameter count, any context-window figure, any architecture name, and any price, until G​oogle publishes one. As of August 25, all four are still invented.

أسئلة تستحق الإجابة

هل جيميني 4 مجرد نسخة مؤجلة من جيميني 3.5 برو تحت اسم جديد؟

No, and G​oogle's own wording rules it out: the July 21 post treats them as separate items in the same paragraph — G​emini 3.5 Pro "currently testing with partners," Gemi​ni 4 as a pre-training run that has just started. A model in partner testing has finished pre-training; a model that has just started pre-training is many months behind it. The question that has sharpened in the last month is the reverse: whether G​emini 3.5 Pro still ships at all, with SemiAnalysis saying it was quietly cancelled so G​oogle can put everything behind Gemi​ni 4. That is an analyst inference, not a confirmed fact — G​oogle still says partner testing continues — and TestingCatalog's August 24 report lands on the same read from the code side: with the summer cycle's expected flagship shelved, G​oogle may move straight to Gemi​ni 4. The practical consequence for a reader is the same either way: do not build a plan around a G​emini 3.5 Pro launch date.

لماذا يتفوق Gemini 3.6 Flash على Gemini 3.1 Pro إذا كان Pro هو الفئة الرائدة؟

لأن المؤشر تغير ولم تتغير النماذج معه بنفس المعدل. يُرجّح مؤشر الذكاء الحالي العمل الوكيل بنسبة 34%، وقد تم ضبط Gemini 3.6 Flash بشكل صريح للمهام الوكيلة ومهام البرمجة — حيث تتصدر أرقام إطلاقه الخاصة DeepSWE وOSWorld. أما Gemini 3.1 Pro، الصادر في فبراير، فقد تم تحسينه وفقًا لجدول نتائج كان يرجّح التفكير الواسع بدرجة أكبر بكثير. أسماء المستويات تصف فئة السعر وزمن الاستجابة، وليست ضمانًا للترتيب تحت أي مقياس قد يستخدمه التقييم الحالي.

هل يخبرنا تأكيد ما قبل التدريب بأي شيء على الإطلاق عن التوقيت؟

Only a floor, not a date. Pre-training on a frontier-scale model is measured in months, and post-training, evals, serving optimisation and partner testing follow it. A run announced as "started" in late July effectively rules out a genuine Gemi​ni 4 in August or September — which is what the 2% August figure and the sliding September rung on Polymarket are pricing. It does not rule out a November or December launch, and it says nothing about whether that launch will be a preview, a limited partner release, or general availability. Goldman's "late 2026 or early 2027" is the widening of that same floor, not a schedule. The August 24 desktop-app signal does not move the floor, but it is the first outside artifact pointing at the near side of the window: TestingCatalog reads the new code mentions as consistent with a December arrival.

The honest read, as of August 25, 2026

Gemi​ni 4 is a training run with a name. The confirmed record is a sentence in a blog post and a paraphrase from an earnings call, and that record contains no date and no specification. The most credible timing estimate — late 2026 — comes from G​oogle's historical version spacing, not from G​oogle, though the leak-trackers' August 24 desktop-app detection is the first outside artifact pointing the same way.

What the record does establish is why Gemi​ni 4 exists. G​oogle's best public model sits at 11th on the current independent index, thirteen points behind the leader; its next flagship has slipped repeatedly on coding and is now reported cancelled by an analyst firm; its co-founder has climbed back into the product; and its CEO has said in public that the fix requires a much larger base. That is a company describing a real gap and a real plan, backed by $195–205 billion of committed capex — and a company whose front line, right now, is being run by analysts' verdicts, a founder's return, and leak-trackers' code strings rather than by anything G​oogle has shipped.

السؤال المطروح هو ما إذا كانت الخطة تعالج الفشل الصحيح. من المفترض أن يغلق التوسع فجوة المعايير. الأمر أقل وضوحًا بكثير أنه يغلق فجوة الموثوقية التي رافقت سلسلة هذا الطراز عبر ثلاثة أجيال من حلقات استدعاء الأدوات واستدعاءات الدوال المفقودة — وتلك الفجوة، وليس درجة المؤشر، هي ما يقرر ما إذا كان الفريق سيُبقي G​emini في حزمة الوكلاء. هذا هو الشيء الذي يجب مراقبته فعليًا عندما تصل الأرقام أخيرًا: ليس ما إذا كان Gemi​ni 4 يتصدر لوحة الصدارة، بل ما إذا كانت تقارير الأخطاء تبدو مختلفة.

مقارنات في هذه المقالة2

مستخرج من هذه المقالة · المعايير: Artificial Analysis · يُحدَّث يوميًا

© 2026 OrcaRouter

لمقدمي الخدمات

هل تدير منصة استدلال؟ اعرض نماذجك على OrcaRouter.

providers@orcarouter.ai

انضم إلى مجتمعنا

Discordsupport@orcarouter.aiXGitHubYouTube