A generated hero card for the article 'Muse Realtime Avatar vs Gemini 3.8 Live', showing two rounded cards either side of the headline. The left card reads 'Muse Realtime Avatar' above a video-frame icon and the lines 'video + voice out', '448x768 portrait, 25 fps' and 'announced Sept 23, no API'; the right card reads 'Gemini 3.8 Live' above a microphone-and-speech-bubble icon and the lines 'audio + text out', '$0.005/min audio in' and 'GA Sept 15, Live API'. A subtitle reads 'One renders a face. The other runs a phone line.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Muse Realtime Avatar vs Gemini 3.8 Live: A Face Against a Voice Line

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

These two are not substitutes, and the fastest way to see it is to ask what comes out of each one. Muse Realtime Avatar, announced by Meta on September 23, 2026, returns synchronized video of a character speaking — you supply a reference image and it renders that thing talking and gesturing in 448×768 portrait frames. Gemini 3.8 Live, announced by Goog​le on September 15, 2026 and generally available to developers the same day, returns audio and text. It has no video output at all. Putting them in the same comparison is not a mistake — the two are competing for the same conversational surface, and buyers are already asking which one to build on — but the decision is not "which is better." It is "which of two different products do I need," and almost every published comparison of the pair gets that wrong by lining up benchmarks that measure different things.

Here is the asymmetry that should frame everything below. You can call Gemini 3.8 Live today, from a terminal, with a key. You cannot call Muse Realtime Avatar at all — no API, no price, no date. Meta demonstrated it at Connect and published a research post; Goog​le shipped a rate card and a model ID. That is a comparison between a product and an announcement, and it means the useful question is not which one wins but what each one commits you to.

What each one actually produces

This is the whole comparison in six lines, and none of the six is close.

Output — Muse Realtime Avatar returns synchronized voice and portrait video, generated together from one stream of speech tokens; Gemini 3.8 Live returns audio and text, with no video generation anywhere in the model.

Developer access — Muse Realtime Avatar has none: it was announced September 23, 2026 as a capability of Meta's Muse agent, with no API, no endpoint and no stated date. Gemini 3.8 Live is in the Gemini API and Goog​le AI Studio as of September 15, 2026, under two model IDs, gemini-3.8-live and gemini-3.8-live-extended-thinking.

Price — Muse Realtime Avatar has no published price because there is nothing to buy. Gemini 3.8 Live publishes $0.005 per minute of audio input and $0.018 per minute of audio output through the Live API.

Independent scoring — Muse Realtime Avatar has none; its figures are Meta's own, measured against baselines Meta selected. Gemini 3.8 Live carries 76.0 on the Artificial Analysis Speech to Speech Index, with the extended-thinking variant at 82.6 — a third-party number, which is a different class of evidence.

Latency — the two published figures are not the same measurement, and this is where most write-ups go wrong. Meta reports 870 ms from the end of your turn to the first byte of a synchronized voice-and-video reply. Goog​le's number, as measured by Artificial Analysis, is 1.18 seconds to first audio for the standard model and 1.35 seconds for the reasoning variant.

Frame and language — Muse Realtime Avatar renders 448×768 portrait video at 25 frames per second. Gemini 3.8 Live carries automatic mid-conversation switching across 97 supported languages and processes near-real-time visual input.

That last row contains the distinction people keep collapsing. Gemini 3.8 Live can see. It cannot show. Muse Realtime Avatar can show. Whether it can see is not what Meta's post is about — it is an audio-driven generator, conditioned on speech tokens and a reference image, and its job is producing a face rather than reading one.

A generated comparison scoreboard titled 'Muse Realtime Avatar vs Gemini 3.8 Live — the scoreboard', with a left column headed 'Muse Realtime Avatar' and a right column headed 'Gemini 3.8 Live'. Rows read: Output — video + voice vs audio + text; Developer access — none, announced only vs Live API and AI Studio; Price — none published vs $0.005/min in, $0.018/min out; Independent score — none vs AA Speech to Speech 76.0; Latency — 870 ms to first byte of voice + video vs 1.18 s to first audio; Frame — 448x768 portrait at 25 fps vs audio only, 97 languages. A footer line reads 'Meta figures company-reported; Google and index figures per Artificial Analysis.'

The latency numbers are not comparable, and the gap is smaller than it looks

870 ms against 1.18 seconds reads like Meta is 26% faster. Read the measurements and the comparison dissolves. Meta's figure is time from the end of the user's turn to the first byte of a reply that includes video — and Meta's post is explicit that it is a first-byte measure, not the time to a complete reply and not ongoing delivery latency for the rest of the call. A system can hit 870 ms to first byte and still feel sluggish at second twelve. Goog​le's figure is time to first audio, a strictly smaller thing to produce, and it is measured by an outside evaluator rather than by the vendor.

Both are the kind of number that tells you a system is conversational rather than turn-based, and neither tells you what the call feels like at minute five. If you are choosing between them on responsiveness, the honest position is that nobody has published a like-for-like measurement, and the only way to get one is to run the one that is callable.

What you can build on today, and what you would be waiting for

Gemini 3.8 Live comes with the things a production decision needs: two model IDs you can pin, published per-minute rates, a third-party index score, and a stated general-availability date. It also comes with the constraints a voice product usually has — audio in, audio and text out, and no video generation.

Muse Realtime Avatar comes with the things a research post has: an architecture, four company-reported figures, and a set of disclosed preference tests Meta ran against two commercial avatar systems, one of which Meta itself reports as statistically indistinguishable from parity on mannerism. What it does not have is a way to buy it. Meta's post also notes that the demos shown do not all reflect avatars available in the Muse app, which is the sentence that should govern any plan you build around this.

There is a third option worth naming, because it is the one most teams actually take. Neither of these models is a complete agent. Gemini 3.8 Live hands background work to tools and APIs while it keeps talking; GPT-Live-1 delegates its reasoning to a separate backend model entirely. The conversational layer is the front end, and the thinking is a different call to a different model. That second call is ordinary text inference, and it is routable.

On OrcaRouter that layer is one endpoint: 196 models from 15 providers behind a single key, at provider list price passed through with zero markup, with automatic failover if the model you picked errors or times out. We do not host Gemini 3.8 Live — the Live API is Goog​le's alone — and we do not host Muse Realtime Avatar, which nobody hosts. What we carry is the half of a voice agent that scales with how hard your callers think, and the half you can change without renegotiating anything.

A screenshot of Google's blog post 'Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking', dated September 15, 2026, showing the announcement headline and the opening of the post introducing the two speech-to-speech models.

Where the two might actually meet

The case for putting them side by side is not benchmarks. It is that both are trying to occupy the same slot in a product: the thing a user talks to. Goog​le's bet is that the slot is a conversation and the deliverable is a good one — 97 languages, background tool execution, a reasoning variant that resolves 68.6% of τ-Voice customer-service scenarios against the standard model's 30.1%. Meta's bet is that the slot is a presence, and that a user who can see the agent is a user who stays in the conversation.

Those bets are not exclusive. A product could plausibly use Gemini 3.8 Live for the voice loop and something like Muse Realtime Avatar for the visual layer, if the avatar layer ever becomes callable. What it cannot do is treat them as interchangeable, because one of them will never give you a face and the other may never give you an endpoint.

How to decide

If you are shipping a voice product this quarter, the decision is already made for you: Gemini 3.8 Live is callable and Muse Realtime Avatar is not. Pick your variant on the axis the release actually argues about — the extended-thinking model buys 38 points of agentic task completion for a little over four times the hourly audio cost, and the standard model is the cheapest competent voice model on the public board. Then route the delegated reasoning separately, because that is where the bill and the latency both live.

If you are evaluating an avatar layer, the honest answer is that there is nothing to evaluate yet. What exists is a research post with four company-reported numbers and a demonstration. The thing to watch is not the index — Muse Realtime Avatar will not appear on one — but whether Meta publishes a rate card, an endpoint, and a date, and whether its 870 ms holds up when the avatar is running on a phone on a cellular connection rather than in a Connect demo.

Until then the comparison has a shorter form than most articles give it. One of these is a product with a price, a model ID and an outside score. The other is a very good demonstration of something that does not yet have any of those.

A screenshot of the Artificial Analysis Speech to Speech leaderboard, showing the Speech to Speech Index ranking with Gemini 3.8 Live Extended Thinking at 82.6 and Gemini 3.8 Live at 76.0 among the listed rows, alongside panels for time to first audio and cost per hour of input audio.