
Muse Realtime Avatar: What Meta Built, What It Runs On, and Why You Still Cannot Call It
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 1009 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1189 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 22 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 108 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Muse Realtime Avatar is Meta's embodiment technology. It takes the speech stream that Muse Realtime Voice produces and renders the matching performance: point it at a photographic portrait and the face answers with small changes of expression, point it at a full-body illustration and the figure gestures and shifts its posture while it speaks. Meta AI Research introduced it on 23 September 2026 in a post called "Bringing Your Muse to Life." It is a research result rather than a product — Meta's own page offers no API, no price, no waitlist and no published date for it — so the honest answer to "can I use it today" is no. The Muse model you can call today is a different one, Meta Muse Spark 1.2.
That answer is worth stating in the first paragraph rather than the last, because a reader who searches this name is usually deciding whether to plan around it. Plan around the research, not around an endpoint.
One thing to say plainly before anything else, because this page breaks a rule it should admit to. This blog normally writes about models in the week they launch. Muse Realtime Avatar was announced a week before this page went up, so on a strict reading of the freshness window the item would be skipped. This is not a news article about a dated event. It is the reference page for a model that is live in the catalogue conversation today: the page a reader reaches after typing the model's own name, whose warrant is standing search demand that we measured ourselves rather than the freshness of a launch. The September announcement is named here as a fact that dates the model, not as an event to report, and nothing below is a second announcement. Our coverage of that day is a separate piece and it stays separate.
Where it sits in the Muse family
Meta is explicit that these are two layers of one system, not two products. In Meta's wording, "Muse Realtime Voice and Muse Realtime Avatar form a single streaming system connecting intelligence, voice, and embodiment." The voice layer supplies the conversational intelligence and emits a stream of speech tokens — Meta calls them VQs — that carry both what is said and how it is delivered. An audio decoder turns those tokens into sound, and the avatar layer consumes the same stream to generate the picture. Sharing one token stream is what keeps voice, lip motion and expression in step; there is no separate lip-sync pass bolted on afterwards.
So: Muse Realtime Voice is the conversational layer. Muse Realtime Avatar is the embodiment layer built on top of it. Neither is the thing you can call.
Be equally careful about a neighbouring name, because it is the one most likely to be mistaken for either. Muse Voice Transcribe is a transcript endpoint, and it is a genuinely different product. Meta's developer documentation is unambiguous: "It is speech-to-text only; it does not synthesize speech or provide a speech-to-speech conversation API." It returns turn-level timestamps and not word-level ones, and Meta lists it at $0.18 per hour on that page. It is a real, priced endpoint. It is not a voice, and it certainly is not an avatar.
The Muse model that is actually callable is Meta Muse Spark 1.2, the reasoning model Meta ships with a million-token context and text, image, video, file and audio input. That one is routable here today, at the provider's list price with zero markup, on the same key as everything else in the catalogue, and its live card carries the full specification. We do not carry Muse Realtime Avatar. We checked the live model list again while writing this and the only Muse entries on it are the two Spark checkpoints, 1.2 and its 1.1 predecessor. Saying anything softer than that — that the family is "partly available" — would imply we route something we do not.

The technology, stated in Meta's own terms
Meta's description of the architecture is compact enough to quote rather than paraphrase. Muse Realtime Avatar is "an audio-driven, Diffusion Transformer conditioned on speech-token stream, reference media, and a rolling window of recent video latents," and it "generates video in short causal chunks." The second half of that sentence is the load-bearing part: as each chunk completes, its newest generated latents become the motion context for the next one. Meta's stated reason is continuity — the avatar's appearance and mannerisms carry forward across turns while the computation stays bounded, which is what lets generation continue for as long as the conversation does rather than drifting or running out of budget.
The provenance mechanism belongs in the same paragraph because Meta presents it as part of the system rather than as an afterthought: generated video carries a durable, invisible watermark applied through Meta Video Seal, which Meta says adds no latency to the real-time path.
Meta measured this thing and published four figures for it. Those figures — and the specific caveats each one needs, including the one that reads like a speed-up and is not one — belong to the launch write-up, which carries them with their context intact. We are not going to restate the bare numbers here, because a bare number is exactly what the caveated version exists to prevent.
We covered the announcement on the day in our Muse Realtime Avatar launch write-up, and it remains the place to read the measured claims at full length.
What the name is not
Three false neighbours are worth ruling out one at a time, because each of them is a reasonable guess from the name alone.
• It is not Muse Voice Transcribe. That endpoint converts speech to text and, on Meta's own description, does nothing in the other direction. If what you need is a transcript, Meta sells one and it is a different thing with a different price.
• It is not a voice-cloning service. Nothing on Meta's pages describes synthesising a specific person's voice, selling a voice model, or accepting a voice sample from a user. The voice layer is described as producing a token stream; the avatar layer is described as consuming one. The product shape that phrase usually implies — upload a sample, get a clone — is not what is described here, and the demo material Meta does publish is framed as illustration of model capability.
• It is not a consumer avatar in Meta's own app. Meta accompanies its examples with a caveat that is easy to miss and worth keeping: the demonstrations "illustrate model capability and do not all reflect avatars available in the Muse app," and Muse is for users aged 18 and over. Read that literally. Some of what the post shows is capability, not inventory.
A fourth name is worth separating for a different reason. Muse Realtime Avatar is not Muse Video, the video-generation model that has been sitting in a public leaderboard for most of this year without shipping. Our own page on that situation — Muse Video: Release Date, API Access, and What Meta Connect Might Change — carries a constraint we will repeat here verbatim rather than improve on it: any page quoting a specific launch day or a price for Muse Video without a newer Meta announcement is speculating. The same discipline applies to this page's subject, so this page quotes neither. That piece also supplies the date that is not ours to confuse: Muse Video is previewed July 7, 2026 and is not released, which is why a search for this family turns up dates that belong to a different model.

What this page is for, and what would change it
The reason a reference page is the right shape here is measurable. In the 28 days to 25 September 2026, the exact term drew 164 impressions on our search property and no clicks, with its best position 9.3. More than fifty of our pages mention the fragment somewhere, and nearly all of that traffic lands on one announcement post. A term with real, sustained demand, ranking just off the first page and converting nobody, is not a demand problem or a quality problem — it is a page problem. There was no page whose subject was the model itself. This is that page.
The position is also why nothing here is dressed up as news. A reader arriving from a name search wants three things in order: what this is, whether it is available, and what it is built on. They are answered above, in that order, with no dates invented and no competitor named.

What would change the page is a single announcement. If Meta publishes an endpoint, a price, a waitlist or a date for Muse Realtime Avatar, the availability section stops being a "no" and becomes a specification, and this page gets rewritten rather than appended to. If Meta ships Muse Video, the paragraph above changes with it. Until either happens, the useful move for a developer is the unglamorous one: keep the interface you already have pointed at the model you can actually reach. Every model in the catalogue, Meta Muse Spark 1.2 included, sits behind one OpenAI-compatible endpoint and one key, which means trying the part of this family that exists costs you a model-string change rather than a new contract. When the embodiment layer becomes callable, the switching cost should be a string as well.
