A hero title card for Tavus Griffin, subtitled "The first Human Interaction Model — announced 1 October 2026", above three chips reading "48% thought it was a real person", "VideoFDB generation: 3.83 against a 3.92 human reference" and "0.43 s audio-to-video latency". Footer: "All figures vendor- and NVIDIA-reported; Griffin-Lite is a research preview, not generally available." The OrcaRouter logo is composited bottom-right.
Guides & Insights

Tavus Griffin: The First Human Interaction Model and the 48% That Passed the Video Turing Test

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Twenty-six of fifty-four people finished a one-minute video call and said their partner had been a real person. That is the number Tavus is leading with for Tavus Griffin, announced on 1 October 2026, and it is a genuinely striking result: on the same protocol the company's previous stack — Phoenix-4.5 rendering the face, Raven-1 doing the perceiving, Sparrow-2 managing turn dynamics — produced one believer out of forty-one, or 2.4%. The two obvious readings of that jump are both wrong, and separating them is most of what follows. Griffin is not a better avatar engine, and it is not a product you can buy. It is a research preview called Griffin-Lite, open to select trusted testers, with no pricing, no public API, and no entry on any independent video leaderboard. Both of those things being true at once is the actual story.

What the 48% measures, and what it does not

The study is a face-to-face Turing test rather than a preference survey, and the protocol is worth stating precisely because it is the load-bearing claim. Fifty-four participants were recruited through an independent research platform and told they would be matched with another participant for a one-minute video call about what they were looking forward to this year. Their partner was a PAL running Griffin-Lite, generating face, voice and responses in real time. Only at the end of the survey were they asked whether it had crossed their mind that the partner might not be a real person, and every participant was then told it was an AI. The comparison run used the same protocol on the older stack, with forty-one participants.

The supporting figures matter as much as the headline. Believers were confident — 79% on average — and so were the doubters, at 81%, which is what a study wants: nobody was guessing. Over half the participants said the possibility had not crossed their mind at all during the call, and nearly all of that group went on to say the partner was real. The participants who did suspect tended to suspect within the first twenty seconds. That is the least comfortable line in the report and the one that should shape how you read the safety section later.

On a seven-point scale the model averaged 5.4 for seeming natural, 5.6 for seeming trustworthy, 5.8 for whether participants would enjoy talking with it again — a score that held at 5.4 even among participants who correctly identified it as an AI — and 5.5 for whether they felt their partner was really listening. The lowest score was 4.9, for whether the conversation flowed naturally. Read as a set, that is a specific profile: Griffin-Lite is convincing one moment at a time and slightly less convincing as an arc.

The independent benchmark, and how to read its two tracks

The quality numbers come from NVIDIA's VideoFDB, a benchmark for full-duplex audio-visual conversation that scores a model on two separate tracks — generation, meaning whether the model produces the right behaviour, and perception, meaning whether it understands the moment — each with its own rubric and its own leaderboard. NVIDIA built the benchmark, conducts the evaluation with its own judge against the published metrics, and scores the response on a zero-to-five scale. Tavus did not run it, which is the reason to quote it.

• Generation — Griffin-Lite scored 3.83, the next-highest published system scored 2.80 (Gemini 2.5 with Anam), and the human ground-truth reference is 3.92. That is 1.03 ahead of the best comparable system and 0.09 below human.

• Perception — 3.73 against a human reference of 4.20, with 3.44 for the strongest reported baseline, MiniCPM-o 4.5 in its audio-only configuration. Gemini 2.5 Flash Native scored 3.17 and OpenAI's gpt-realtime scored 2.97.

• Takeover-rate alignment — 62.8% on generation and 73.8% on perception. This measures how closely the model's decisions about when to speak match the timing in the reference conversations, and Tavus describes both as the highest of any system on the board.

Two caveats belong on those numbers before anyone repeats them. The generation gap to human is 0.09 on a five-point scale judged by a language model, which is better read as "close to human on this rubric" than as "at human level". And the perception comparison is not like for like: MiniCPM-o 4.5 and gpt-realtime are scored in audio-only configurations, while Griffin reads video as well. The 0.29-point lead over an audio-only baseline is real, but part of what it measures is the value of having eyes.

The vendor's own summary line — first on NVIDIA's independent test of face-to-face AI, 37% ahead of the next best AI at reacting in the moment — is a characterisation of the same data rather than a separate finding. Keep the underlying track scores, not the framing.

A screenshot of NVIDIA's VideoFDB project page, captured 2 October 2026: the heading "VideoFDB — Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents", the note that it is the first benchmark for full-duplex audio-visual-to-audio-visual conversation, the author list (Amrita Mazumdar, Seonwook Park, Rajarshi Roy, Nikhil Srihari, Shengze Wang, Yuhao Zhou, Julia Wang, Koki Nagano, Shalini De Mello) credited to NVIDIA, links to leaderboard, paper and Hugging Face dataset, and an abstract arguing that natural conversation is full-duplex.

Two engines, running at the same time

The architectural claim is easier to evaluate than the marketing around it, because it is a claim about what runs concurrently. Griffin is described as two systems working in parallel rather than a pipeline handing off between stages.

The first is a Continuous Conversational Modeling engine. It ingests the person's audio and video and re-assesses the exchange at regular sub-second intervals — Tavus calls these mini-turns — deciding at each one whether to stay silent, signal, backchannel, or take the turn. The output is not just the words but a set of expressive controls carrying timing, stance, facial expression and gesture. The point of the design is that a decision made mid-sentence shows up in the voice and face within the same mini-turn instead of after a hand-off, which is what lets the model stop when it is cut off, nod while it is listening, and treat a pause for thought as something other than the end of a turn.

The second is an audio-visual generation engine that turns those controls into sound and pixels together: a streaming speech generator and a streaming video generator driven by the same signals, so a change of tone and a change of expression land in the same beat.

One honesty note about the material Tavus published with this, because it would be easy to mistake for measurement. The interactive diagram of a nine-second exchange is annotated in the page's own text as illustrative and explicitly not measured from a real session. The same caveat applies to the turn-based-versus-full-duplex timing figure. What is measured is the latency number further down.

Inside the streaming stack

The audio side is built around a codec called Tavec, a convolutional autoencoder that maps 48 kHz audio into a continuous latent of 40 values per frame at 100 frames per second, with no codebooks. Its decoder is fully causal, so it carries state across chunks and emits audio packets as small as 10 ms without lookahead. A minute of speech is 6,000 latent frames rather than 2.88 million raw samples, which is what makes the sequence cheap enough for the speech transformer to predict faster than real time. The speech generator itself is an autoregressive diffusion transformer that conditions on an encoded speaker prefix — a voice can be cloned from roughly ten seconds of audio — and generates one latent chunk at a time as controls arrive, so playback begins before the utterance is finished.

The video side starts from the same premise: strong temporal compression is what makes a diffusion model fast enough to hold a conversation. A few-step autoregressive generator takes a reference photo, the streaming audio and the streaming controls, and produces one latent per step at three diffusion steps per latent. A VAE compresses time eightfold, so at 25 fps one latent is eight frames, or 320 ms of 720p video. Older latents are kept at progressively lower resolution so long rollouts stay affordable. The result is a generator that emits 720p in 320 ms chunks in real time.

Getting there took a three-stage distillation from a large bidirectional many-step diffusion teacher: Distribution Matching Distillation into a few-step student, teacher forcing to convert that student into an autoregressive model that generates one latent at a time, and finally training with self-forcing on the model's own generated history plus an added mechanism acting on the history frames so long rollouts do not drift. That third stage is the one that addresses the failure mode this class of model usually has — a face that degrades, or a background that slowly wanders, over a conversation that runs for minutes.

The latency number, which is the real product claim

Against four published streaming diffusion models in an audio-to-video setting, where each model gets speech and a reference image and must produce the talking face, Griffin-Lite's generator averaged 0.43 seconds of true audio-to-video latency on H100s — the time from a piece of audio arriving to the video showing its effect. Tavus describes that as half the next-fastest method. It also reports first place on DOVER and FID perceptual quality and on THEval, a talking-head evaluation framework, and second on LSE-C lip-sync confidence at 7.27, behind one baseline — with the vendor noting that LSE-C rewards pronounced mouth movement, including movement past the point where it looks natural.

All of those are Tavus's own measurements against baselines it selected. They are useful because they are specific and because the 0.43-second figure is the kind of number that either holds up in a live test or does not. They are not independent, and the only independent scores in this article are the two VideoFDB tracks above.

A screenshot of Tavus's own Griffin announcement page at tavus.io/griffin, captured 2 October 2026: the heading "INTRODUCING GRIFFIN", the strapline "The First Human Interaction Model", the description of Griffin as the world's first Human Interaction Model that listens while watching expressions and pauses, and the byline "By: Hassaan Raza, Ioannis Patras, Head of Tavus Research Team".

Why there is no price to compare

Griffin-Lite is available today to a select group of early testers as a research preview, with a wider release of a more capable model to follow. It will not be available for customer use at this time. Access is by request form. It is not on the Tavus platform, and the company's own closing line on the announcement says so plainly: Griffin is not on the Tavus platform yet, and will come once Tavus has worked out how to release it safely. Everything a buyer would normally compare — per-second or per-request pricing, a model identifier, rate limits, an SLA, a region list — does not exist yet. Tavus points anyone wanting to build today at the models 150,000 developers and businesses already use, which are the three predecessors, not Griffin.

That is not a technicality to footnote. It changes what this model is for right now. It is an evaluation target for teams whose product depends on whether a person can tell, and a signal about where the vendor's roadmap is going. It is not something to build a customer-facing path on this quarter.

The safety gate is the release plan

The most interesting thing Tavus wrote about Griffin is not about the model. It is the admission that the same properties that make a human interaction model a better interface also make it better at deceiving someone into believing it is not an AI, that further alignment and safety work is required before safe release, and that the company is working on disclosure features and with organisations on AI safety evaluations. Given that nearly half a live sample could not tell, and that the ones who could mostly worked it out inside twenty seconds, a disclosure mechanism that survives a casual conversation is not a nice-to-have.

Two things would change this article materially. A published model card with an endpoint and a price would move Griffin from research result to procurement question. And an entry on an independent video board — with the caveat that those boards score short clips by pairwise preference and are not built to score a live conversation, so the mismatch may be permanent rather than temporary — would give the latency and quality claims an outside check. Until one of those lands, the VideoFDB tracks are the strongest evidence available and they come with a human gap of 0.09 and 0.47 points respectively.

The counter-argument worth weighing is that a benchmark run is not a deployment. NVIDIA's protocol gives every system the same turn-taking task and the same judge; it says nothing about how the ninth hour of a call behaves, what happens when two people talk over each other for a full minute, or whether the expressive controls degrade on a face the reference photograph rendered badly. Tavus reports the drift-specific training — the self-forcing stage and the history mechanism — as the fix for exactly that, and reports are all a reader has until somebody outside Tavus runs a long session and publishes it. Treat the benchmark as the best available evidence rather than as a deployment result.

What to build around it in the meantime

The practical question for a team that wants something like this today is what parts of the stack are already buyable. A conversational video interface is a chain — speech recognition, a language model, speech synthesis, and in Tavus's case a rendering model — and the first three are commodity. Those sit behind one OpenAI-compatible endpoint at OrcaRouter with 200+ models on the same key and provider list price passed through at 0% markup, so a vendor price cut is live here the same day rather than after a contract cycle. If you are assembling an interaction pipeline and want to keep the generation layer open until Griffin or something like it becomes an actual product, that is where the swap costs a config change instead of a migration. Tavus Griffin itself is not a routed model here, and it will not be while it is a preview behind a request form.

The thing worth watching is not another demo reel. It is whether the disclosure tooling Tavus is building gets published, and whether a second research group reproduces the face-to-face protocol. A Turing-test pass this large is exactly the kind of result that needs a second lab to run it.

An explainer scoreboard for Tavus Griffin with six rows: Status: research preview; Pricing: none published; Model type: human interaction model; VideoFDB generation: 3.83 of 5; VideoFDB perception: 3.73 of 5; Open issue: disclosure. Footer: "NVIDIA VideoFDB per Tavus and NVIDIA, September 2026." The OrcaRouter logo is composited bottom-right.