A generated hero card titled Gemini 3.8 Live with a bold UNVERIFIED badge, the subtitle 'Two unreleased voice slugs on a Google Cloud quota page - what we know so far', chips reading 'Source: X/@testingcatalog', 'Sept 15, 2026' and 'Not public - no model card', three cards reading 'Reported slug: Gemini 3.8 Live', 'Reported slug: Live Extended Thinking' and 'Live line today: gemini-3.1-flash-live-preview, March 2026', and the footer 'Both slugs unverified - Google has not acknowledged either.'
Guides & Insights

Gemini 3.8 Live Leak: Two New Voice Slugs, and Why "Live Extended Thinking" Matters More

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Two strings that appear in no published documentation turned up this week on a GCP quota page: Gemini 3.8 Live, and something called Live Extended Thinking. They were spotted by the release-tracking account @testingcatalog and posted on September 15, with the flat caveat that neither is public. That is the entire public record so far — no model card, no pricing line, no API changelog entry, no acknowledgement. But the two names sit at the end of a very legible line. Gemini 3.1 Flash Live is still the model the company's own docs recommend for Live API work, Gemini 3.5 Live and Gemini 3.5 Live Experimental arrived on August 26 as the Gemini Audio family, and the text side of the same family — Gemini 3.8 Flash — has been shipping since September 2. A "3.8 Live" would close that gap. The second slug is the more interesting of the two.

Labels first, because a leak piece lives or dies on them. This is not a launch. Neither slug has been announced, documented, priced or confirmed by Google, and the source is a single account that tracks releases. Everything below about Gemini 3.8 Live and Live Extended Thinking specifically is unverified. Everything below about models already shipping — Gemini 3.1 Flash Live, Gemini 3.5 Live, Gemini 3.8 Flash — is documented and checkable, and I have marked which is which throughout.

What the signal says, and what it does not

• Reported — two slugs, "Gemini 3.8 Live" and "Live Extended Thinking", appearing on a Google Cloud quota page, posted September 15 with the note that they are "not public"

• Not reported — an announcement, a model card, a price, a context window, a release date, an API id, any benchmark figure, or any confirmation from Google

• Corroboration — none as of this writing. Google's Gemini API models page, last updated September 4, lists exactly two conversational Live models, gemini-3.1-flash-live-preview and gemini-3.5-live-translate-preview, and neither carries a "Live Extended Thinking" variant

• Same brief, separate claim — the same September 15 post also forecast that Grok 4.7 "should land near Opus 5.0, not 5.1" and that its multimodal work "still needs work." That is a different vendor's model and a forecast rather than an artifact; it is noted here only for completeness of the source

• The name's own implication — Google's Cloud quota entries are per-model SKUs, which argues for a billable API model rather than a feature inside the consumer Gemini app. That is inference from the surface the string appeared on, not a confirmed fact

A generated six-row scoreboard titled 'Gemini 3.8 Live - the scoreboard', rows reading 'Status: leak - quota page only, not public', 'Reported slugs: Gemini 3.8 Live + Live Extended Thinking', 'Live line today: gemini-3.1-flash-live-preview, March 2026', 'Thinking control: thinkingLevel minimal/low/medium/high', 'Text sibling: Gemini 3.8 Flash, shipped Sept 2', 'Independent score: none - nothing to test', with the footer 'Gemini 3.8 Live and Live Extended Thinking unverified; Live-line details per Google's Gemini API docs.'

Why a quota page is where a model leaks first

This is the part worth understanding, because it explains why a two-word artifact deserves a whole article. Quota pages are not marketing surfaces. They are billing and rate-limit infrastructure: every model a Cloud customer can call needs a per-project quota entry before the first paying request is served, and those entries are provisioned by the same engineering work that readies a model for general availability. A slug showing up there means someone has built the plumbing for a SKU — not that someone drafted a press release.

That cuts both ways, and the honest read is the cautious one. A quota entry is further along than a codename in a research paper, because it implies an intended customer-facing surface. It is also exactly the kind of thing that gets created, tested and quietly renamed before launch. The text side of this family has moved fast enough to make that concrete: Gemini 3.6 Flash landed July 21, Gemini 3.7 Flash on August 13, and Gemini 3.8 Flash on September 2 — three Flash releases in six weeks. A model line iterating that quickly is a line that stages more SKUs than it ships under those names.

The Live line has been running a version behind

Here is what makes "Gemini 3.8 Live" a real story rather than a naming curiosity: Google's voice models have not tracked its text models at all.

• Recommended Live API model today — gemini-3.1-flash-live-preview, latest update March 2026, still described in Google's own docs as the model to use for all Live API use cases

• Its limits — 131,072 input tokens and 65,536 output tokens, taking text, images, audio and video in, returning text and audio out

• The Gemini Audio family, announced August 26 — Gemini 3.5 Live, Gemini 3.5 Live Experimental and Gemini 3.5 Transcribe, with the Transcribe models replacing Chirp 3 for speech-to-text

• Meanwhile the text line — Gemini 3.5 Flash, Gemini 3.6 Flash, Gemini 3.7 Flash and Gemini 3.8 Flash — moved through four public versions in roughly the same window

A screenshot of Google's Gemini API documentation page for gemini-3.1-flash-live-preview, captured September 15, 2026, showing the banner 'Gemini 3.8 Flash is now available', the heading 'Gemini 3.1 Flash Live preview' described as a low-latency audio-to-audio model, model code gemini-3.1-flash-live-preview, supported data types text/images/audio/video in and text and audio out, an input token limit of 131,072 and an output token limit of 65,536, and capability chips including Live API Supported and Thinking Supported.

So the audio line sits at the 3.5 generation while the text line is at 3.8. A "Gemini 3.8 Live" slug is the first evidence that Google intends to put voice back on the same generation as text — which, for anyone building a voice agent, means the reasoning underneath a Live session could jump by three text-model generations at once rather than one.

Why "Live Extended Thinking" is the more interesting slug

Today, thinking on a Gemini Live model is a dial, not a model. The Live API exposes a thinkingLevel parameter with four settings — minimal, low, medium and high — and the default is minimal, chosen explicitly to optimise for the lowest latency. That default tells you what the product is for: a conversation in which a pause is a bug.

A separate slug called "Live Extended Thinking" implies Google is preparing to split that into two products rather than one dial, and the reasoning is straightforward. Latency and deliberation are in direct tension in a voice model — you cannot both answer instantly and think hard before answering — so a single model with a toggle forces every caller to pick one point on that line for every request. Two SKUs let a customer route the fast path (interruption handling, barge-in, quick lookups) and the slow path (a voice agent working through a multi-step task) at differently tuned endpoints, without one degrading the other.

The precedent already exists inside Google's own August launch. Gemini 3.5 Live Experimental was described as the variant that can "reason directly while speaking" and narrate its progress step by step during a task — an explicitly thinking-flavoured voice model, shipped as its own id rather than as a setting. "Live Extended Thinking" reads like that instinct graduating from an experimental label to a named product. That is inference from a string, and it is inference — but it is the kind that this particular line has rewarded before.

What you can actually call today

Nothing here changes what you can reach right now, and it is worth being blunt about that. There is no Gemini 3.8 Live endpoint to call, no Live Extended Thinking setting to switch on, and no way to buy access to either. Any account selling "Gemini 3.8 Live access" today is selling a label, not a model. The Live models you can genuinely use are gemini-3.1-flash-live-preview and gemini-3.5-live-translate-preview, both in preview through Google's own API.

The text side is a different picture, and it is the part that is actually routable. Gemini 3.8 Flash — shipped September 2 at $0.75 per million input tokens and $3.75 per million output on Google's own list price, scheduled to double to $1.50 and $7.50 on January 1, 2027 — runs on OrcaRouter at that same list price, with 0% markup on top. Pass-through pricing matters more than usual on a model with a pre-announced hike: when the January number takes effect it changes here the same day, with no second contract to renegotiate and no code change to make.

The OrcaRouter model page for google/gemini-3.8-flash, showing the Google vendor name and the 2026-09-02 date, a 1M-token context window, 65K max output, text, image, video, file and audio input with text output, a p50 time-to-first-token of 3.46s, and pricing of $0.75 per 1M input tokens and $3.75 per 1M output tokens.

The voice line is not on OrcaRouter today — no Live or native-audio SKU is, from any vendor — so nothing here should be read as us hosting it. What a routing layer is genuinely useful for is the other half of this story. When a 3.8-generation Live model does land, and when a second, thinking-flavoured variant lands beside it, choosing between a fast voice path and a deliberating one becomes a routing decision rather than a rewrite: two endpoints behind one key, with automatic failover if the preview id you built against gets retired. On a line that has already renamed itself once this year, that is not a hypothetical.

What would confirm this — and what would kill it

A leak piece should tell you how it could be wrong. For Gemini 3.8 Live and Live Extended Thinking, the confirming evidence is specific and checkable:

• The quota entry going public — the same slug appearing in a Google Cloud or Vertex AI model list with a rate limit attached, rather than only in a screenshot

• A Gemini API changelog entry or a row on the models page, which currently stops at gemini-3.1-flash-live-preview and gemini-3.5-live-translate-preview

• A DeepMind model card — the August audio models were documented there as "Gemini 3.5 Audio" even while coverage called them Live and Transcribe, so the card is the artifact that settles a name

• A pricing line, which is the one thing that turns a staged SKU into a product somebody can buy

• Disconfirmation — the slugs being renamed, folded into the existing model's thinking settings, or simply never reappearing. All three are ordinary outcomes for a quota page, and any of them means this piece was early rather than right

What to watch, and who should act

If you are building a voice agent on the Live API, nothing about your architecture needs to change this week — the model you would build on is still gemini-3.1-flash-live-preview, and the honest advice is to keep the model id behind a config value and keep a second Live endpoint warm, because this line has renamed itself once already and may do it again. If you are choosing a text model, this leak is irrelevant to you; Gemini 3.8 Flash is shipping, priced and measurable now, and the more useful question is the January price change rather than a voice slug.

The thing to actually watch is not the launch. It is whether "Live Extended Thinking" arrives as a second model or as a fifth setting on the first one. If it is a separate SKU, Google is telling developers that a voice agent which thinks and a voice agent that talks are different products — and that is a more consequential claim than any version number in the name.

What a routing layer is genuinely useful for is the other half of this story.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily