A generated hero card for 'Gemini 3.8 Live with Live Avatar' on a white background with soft blue-and-cyan gradient accents. Three stacked cards read '15 Sept 2026 — Gemini 3.8 Live and 3.8 Live Extended Thinking ship on the Live API', '24 Sept 2026 — Live Avatar announced, Gemini Enterprise', and 'Today — avatar is an enterprise capability, not a model ID'. A footer line reads 'Dates per Google DeepMind.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Gemini 3.8 Live with Live Avatar: Google Gives Its Voice Model a Face

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking shipped on September 15, 2026 — the two native live-dialogue endpoints Goog​le opened on the vendor's API, one optimised for latency and one for deliberation. On September 24 the company announced Gemini 3.8 Live with Live Avatar, which pairs near real-time video generation with the same speech loop so the model has a face while it talks. Nothing about the two model IDs changed. What changed is what the conversation is allowed to look like, and — more consequentially — what the model is allowed to do to the conversation while it is still running.

What actually shipped on 24 September

The DeepMind post is short and specific, and it is worth separating what it claims from what it means. Live Avatar, in Goog​le's own words, "brings near real-time visual presence to Gemini's conversational AI" by pairing real-time video generation with the existing speech stack. The company describes precise lip-syncing, natural expressions and fluid turn-taking, and gives two deployment examples: customer service and interactive walkthroughs. Then it says the sentence that matters most for anyone planning a build — "Starting today, Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise."

That is the whole availability story. Live Avatar is not a Gemini API model ID. The API's audio model list still contains gemini-3.8-live and gemini-3.8-live-extended-thinking and no avatar row, and the rate card has no avatar line item. The feature lands inside Gemini Enterprise, which means the audience for the announcement is enterprise buyers and the developers who integrate on their behalf, not the individual developer calling the Live API with a key today.

Two of the post's claims are worth quoting precisely, because both are stronger than the usual launch language. The first: Live Avatar "features native multilingual speech-to-speech synchronization" and "can seamlessly transition across 97 languages without degrading video fidelity or introducing visual drift." The 97 figure is the same language count the September 15 launch claimed for the underlying live-dialogue models, so this is the existing multilingual claim restated with a new constraint attached — the lip-sync and the facial animation have to keep up when the language changes mid-sentence. The second: all output is watermarked with SynthID, "woven directly into the audio and video output." Both tracks, not just the voice.

The genuinely new capability is the one in the middle

Read past the visuals and the most interesting paragraph is the one about tool calls. Goog​le says Live Avatar is backed by "asynchronous tool calling" that can "trigger tool calls and fetch data in the background while continuing active dialogue, handling complex tasks while ensuring an uninterrupted conversational flow." The worked example is a hotel guest check-in where the model calls tools in the background and keeps talking.

That is a real architectural difference from the request-response shape most voice integrations are built around, and it is the part with a cost attached. Asynchronous execution means the model is doing work the transcript does not show. In a text pipeline you would see the tool call as a turn. Here it happens underneath a conversation that is still in progress, which raises the same questions any concurrent execution raises: what happens if the tool call fails silently mid-sentence, and how does the caller know? Goog​le's post does not say, and the model card is the place to look for it. Treat the "uninterrupted conversational flow" claim as vendor-reported and untested by any third party we can find.

Preset avatars, and the allowlist that limits the interesting half

The customisation story has two tiers and only one of them is generally available. Every enterprise deployment gets "a library of diverse, preset avatars." Custom avatars — generated "from a high-quality reference image" while "preserving reference likeness, brand styling, or character identity" — are behind a gate, and Goog​le states it plainly: "Custom avatar creation is currently available only through enterprise allowlisting."

So the capability most teams will actually want, the one that lets an avatar match a brand or a named character, is the one you cannot self-serve. If your plan depends on a synthetic presenter who looks like your mascot, you are waiting on an allowlist decision and not on a model release. That is a procurement fact, not a technical one, and it is the kind of thing that quietly slips a quarter.

A screenshot of the Google DeepMind blog post 'Introducing Gemini 3.8 Live with Live Avatar', captured in English on 25 September 2026, showing the 24 September 2026 date, the authors listed as Shuo-yiin Chang and CJ Zheng on behalf of the Gemini Audio Team, the lede 'Building on the momentum of last week's Gemini 3.8 Live launch', the sentence 'Starting today, Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise', and the section heading 'More natural and multimodal conversations'.

Where it sits against the rest of the line

Live Avatar did not arrive into an empty product family, and the position it occupies is easiest to see against the siblings that shipped in the same fortnight.

• Ships as — Gemini 3.8 Live with Live Avatar is an enterprise capability inside Gemini Enterprise vs Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are Gemini API endpoints with model IDs and published per-minute rates
• Callable by developers — Live Avatar no, no model ID and no rate-card line vs the two Live endpoints yes, gemini-3.8-live and gemini-3.8-live-extended-thinking
• Output — Live Avatar audio plus synchronised video vs the Live endpoints audio only
• Language claim — Live Avatar 97 languages with lip-sync and expression continuity vs the Live endpoints 97 languages with automatic mid-conversation switching
• Watermark — Live Avatar SynthID on both the audio and the video track vs the Live endpoints SynthID on generated audio

The two Live endpoints themselves are the well-documented pair. Independent measurement has them far apart on some axes and close on others: on the Artificial Analysis Speech to Speech Index, Gemini 3.8 Live Extended Thinking (High) sits at 82.6 against Gemini 3.8 Live at 76.0, a gap built mostly on reasoning quality rather than conversational dynamics, where the ordering flips. Goog​le's own figures have them at 68.6% and 30.1% on τ-Voice task completion and 98% and 92% on Big Bench Audio. Live Avatar inherits whichever of those you call underneath it, and inherits their pricing too, because the avatar is a presentation layer over the same conversation loop.

A screenshot of the Artificial Analysis Speech to Speech leaderboard, captured in English on 25 September 2026, showing the AA Speech to Speech Index speed and cost-per-hour-of-input-audio panel. The top index values read 82.6, 81.5, 81.3, 80.1, 76.0, 73.9 and 71.5, with a cost axis running to roughly $10.75 per hour. The bar labels are set vertically and are not legible at capture size.

What this does not change

It does not make the models better. Live Avatar is a presentation and orchestration layer: the quality numbers for Gemini 3.8 Live are what they were on September 15, and anyone who needs the stronger of the two variants still has to choose Extended Thinking and accept its latency. It does not put video in the API. And it does not change the price of talking to the model, which is still billed on the Live API's per-minute audio rates — $0.005 per minute of audio in and $0.018 per minute of audio out — with the avatar's video generation unpriced in public because it is not sold separately.

What it does change is the set of products that can be built without a separate rendering layer. A support agent that previously spoke and left a blank panel on screen can now hold a face while it works, and the work it is doing in the background is no longer visible as dead air. That is a meaningful change to the shape of an interface and a modest change to the underlying model. Both of those can be true at once.

A generated summary card for Gemini 3.8 Live with Live Avatar on a white background with soft blue-and-cyan gradient accents. Four rows read 'Announced: 24 September 2026', 'Underlying models: gemini-3.8-live and gemini-3.8-live-extended-thinking, unchanged', 'Availability: Gemini Enterprise; custom avatars by enterprise allowlist', and 'Watermark: SynthID on audio and video'. A footer line reads 'Per Google DeepMind, 24 September 2026; vendor-reported and not independently benchmarked.' The OrcaRouter logo is composited in the bottom-right corner.

What to watch, and one routing note

Three things would move this from an announcement to a build decision. Custom avatar allowlisting would have to open up, because preset avatars are a demo and custom avatars are a product. The video track would have to get a price, because nobody can model unit economics on an unpriced output. And an avatar endpoint would have to appear in the API's model list, because that is the difference between an enterprise feature and something a developer can call tonight.

On our own side, one thing is worth stating precisely, since it is easy to assume otherwise: OrcaRouter does not route the Gemini 3.8 Live endpoints or Live Avatar, and nothing here should be read as a claim that it does. What we do carry is the rest of Goog​le's 3.8 line — google/gemini-3.8-flash among them — alongside a catalogue of more than 200 models from a single key at provider list price with no markup. If Live Avatar sends your architecture toward an agent that has to reason about what the customer said, the reasoning half of that stack is the half you can consolidate today.