
Tavus Griffin vs Muse Realtime Avatar: Two 2026 Faces, Two Different Problems
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 129 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 103 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 212 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Both Tavus Griffin and Muse Realtime Avatar exist to solve the same embarrassing problem — a language model that can talk but has no face — and they attack it from opposite ends. Griffin is a video-to-video Human Interaction Model that watches a participant through a camera and generates the other side of the conversation in real time, and it was announced by Tavus in October 2026 as a research preview for selected testers. Muse Realtime Avatar is Meta's embodiment layer for its Muse agent, announced at Meta Connect on 23 September 2026, and it takes audio and turns it into a synchronised talking portrait. Neither can be bought. The difference is in what each one is trying to get right, and it shows up the moment you ask what each model is actually receiving as input.
The short version
• Griffin's input is the person on the other end — video and audio of a live participant, which it must read, respond to and be interrupted by.
• Muse Realtime Avatar's input is the agent's own speech — a stream of tokens from Muse Realtime Voice that it converts into a matching face.
• Tavus measures Griffin in whether people believe it, and publishes a 48% figure from a one-minute video call.
• Meta measures Muse Realtime Avatar in whether it keeps up, and publishes 870 milliseconds from end of turn to first byte.
• Neither has a public API, a published price, or a general availability date.
Read those four rows together and the shape of the disagreement is obvious. Tavus is optimising for the other person's belief. Meta is optimising for the clock.

Griffin: the model that has to be believed
Tavus's framing is that Griffin is the first Human Interaction Model, and the architecture behind the claim is a two-part system combining Continuous Conversational Modeling with Audio-Visual Generation. The first part is behavioural: it decides what the persona should be doing from one moment to the next, using what it can see and hear. The second part is rendering, and Tavus splits that again into streaming speech generation and streaming video generation.
The numbers Tavus leads with are not throughput numbers. In a study the company ran with Queen Mary University of London, 26 of 54 participants — 48% — believed Griffin was a real person after a one-minute video call, against 1 of 41 (2.4%) for its previous stack of Phoenix-4.5 face generation, Sparrow-2 turn-taking and Raven-1 vision. Participants who were not fooled usually doubted within the first 20 seconds. On NVIDIA's VideoFDB benchmark, scored in September 2026, Tavus reports Griffin first on face-to-face AI with a generation score of 3.83 against a 2.80 strongest reported baseline and a 3.92 human reference, and a perception score of 3.73 against 3.44 for the best reported baseline and a 4.20 human reference. Those figures are vendor-reported and benchmark-specific, but the human comparison points are the interesting ones.
The engineering under those figures is a codec called Tavec and an autoregressive diffusion transformer. Tavec runs at 48 kHz with 40 values per frame at 100 frames per second, with no codebooks, a fully causal decoder and packets as small as 10 milliseconds — designed so a face can start moving before a sentence has finished. The video path pushes 720p in 320-millisecond chunks, where one latent equals eight frames at 25 fps, three diffusion steps per latent and an 8× temporal compression in the VAE. A three-stage distillation process (distribution matching, then teacher-forcing autoregressive, then self-forcing) is what lets it run those three steps in real time at all. The headline latency claim is an average of 0.43 seconds from audio to video on H100s, which Tavus says is about half the next fastest method it measured. Voice cloning takes roughly a 10-second sample.
The price of all this is that Tavus cannot ship it. Griffin-Lite, the research preview, is available only to a select group of early testers; writing on the release, the company states plainly that Griffin-Lite will not be available for customer use at this time and that it anticipates releasing Griffin very soon after certain safety concerns are addressed. Griffin is not on the Tavus platform, and will arrive there "once we've worked out how to release it safely". A model that is convincing 48% of the time in a one-minute call is, from a deployment standpoint, a model that is convincing 48% of the time in a one-minute call.

Muse Realtime Avatar: the model that has to keep up
Muse Realtime Avatar was announced on 23 September 2026 at Meta Connect and written up the same day by Meta Superintelligence Labs under the title "Bringing Your Muse to Life". Meta's framing is narrower and more mechanical than Tavus's: the Muse agent already had a voice, courtesy of Muse Realtime Voice, and the avatar is the layer that gives that voice a body.
The published figure is 870 milliseconds from end of turn to first byte — that is, from the moment the user stops speaking to the moment the avatar's first video byte is ready. The avatar renders a 448×768 portrait at 25 frames per second. Meta says the system performs roughly 60× fewer evaluations per chunk than its baseline, and that a single NVIDIA GB200 can carry 12 concurrent real-time sessions with an 8× capacity gain over a two-step BF16 implementation.
The architectural difference from Griffin is where the information comes from. Muse Realtime Avatar extends Muse Realtime Voice and shares its speech-token stream — the same VQ representation the voice model produces is what drives the face, rather than a separate visual understanding of a participant. That makes it an audio-driven DiT embodiment layer: efficient, tightly coupled to its own voice model, and not attempting to read the human on the other end of the call at all. It is a mouth and a face for an agent, not a conversational partner that watches you.
Meta has published no API, no price and no release date for Muse Realtime Avatar. The research post is the whole of the public record.
Where the two designs genuinely diverge
The clearest way to see the split is to ask what happens when the user interrupts.
Griffin is full duplex and designed to be interrupted mid-utterance — it can watch and listen while speaking, which is why Tavus's marketing leans on the idea that Griffin learns conversational behaviour rather than video. That is also why its benchmark suite is built around perception and reaction, and why the 37% lead over the next best system at reacting in the moment is a claim the company bothers to make.
Muse Realtime Avatar's 870-millisecond end-of-turn number tells you the opposite story: it is a pipeline where the turn ends, the audio tokens resolve, and then the face catches up. It is fast — 870 ms is a good figure for a portrait renderer — but the measurement itself assumes a turn boundary exists, which is precisely the assumption a duplex model is built to throw away.
That is not a criticism of either design. Meta is building a face for an agent that lives inside Meta's own assistant surfaces, where a clean turn boundary is a reasonable model of the interaction. Tavus is building something that has to survive a stranger deciding, inside twenty seconds, whether the person on the screen is real.
Neither one is on a price list — and only one part of Muse is callable
Neither Tavus Griffin nor Muse Realtime Avatar can be put into production today. Griffin is limited to selected testers; Muse Realtime Avatar has no API, no price and no date. If you are planning a build, both of those facts belong in the plan.
What is worth flagging is the piece of the Muse stack that is reachable. Meta's Muse family includes Muse Spark, and one checkpoint of it — meta/muse-spark-1.2 — is live on our catalogue at $1.25 per million input tokens and $4.25 per million output tokens with zero markup on Meta's list price. It is not the avatar: it takes text, image, video, audio and PDF input and returns text, with a 1M-token context window and configurable reasoning effort. But it is the part of the Muse line you can actually call, and if you are building the agent that will one day sit behind an avatar, the language half is the half you can start on. OrcaRouter passes provider list prices through at 0% markup, so a Meta price change is reflected on our card the same day rather than the next billing cycle, and a request that fails on one provider is retried against a healthy one instead of surfacing as a timeout.

The same logic applies to the video side of the world. We do not route Muse Realtime Avatar and we do not route Tavus Griffin, and we will not claim otherwise. What we do route is a catalogue of over two hundred models across the same modalities, so a prototype that swaps between a text agent, a voice model and a video generator stays on one key while the two models in this article remain unreleased.
Which one is further from shipping
On the public record, Muse Realtime Avatar is further away. It has a research post and no API, no price and no date. Griffin also has no API and no price, but it has an explicit release intent — "very soon after these safety concerns are addressed" — a named preview tier that exists today for selected testers, and a stack of published benchmark results from an independent test harness. That is a more advanced position, even though it comes with an admission that the model is not safe enough to hand to customers.
The trap in comparing them is treating the 870-millisecond figure as directly comparable to the 0.43-second figure. They measure different things: Meta measures from the end of a turn to the first byte of a portrait, and Tavus measures audio-to-video latency inside a continuous duplex stream where turns are not clean events. The numbers are both real and they are not the same measurement, and anyone presenting them in one column is doing the reader a disservice.
Bottom line
Muse Realtime Avatar is an embodiment layer: audio in, portrait out, 870 milliseconds from end of turn to first byte, 448×768 at 25 fps, no API and no date. Tavus Griffin is a duplex Human Interaction Model: the participant in, a reacting face out, 0.43 seconds average audio-to-video latency on H100s, and a vendor willing to publish that 48% of people thought it was human while also stating that customers cannot use it. Meta is building something that keeps up. Tavus is building something that passes. Neither is buyable, which means the only decision available today is which half of the problem to start on — and for most teams, that half is the language model underneath.
