Hero title card for 'Tavus Griffin vs MiniMax H3' with the subtitle 'A Duplex Conversational Model Meets a Shipped Video Generator', above four chips: Tavus Griffin research preview, testers only; MiniMax H3 open weights, shipped; MiniMax H3 $0.08/s at 768P; Griffin: no API, no price.
Guides & Insights

Tavus Griffin vs MiniMax H3: A Duplex Conversational Model Meets a Shipped Video Generator

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Tavus Griffin and MiniMax H3 both put a moving, speaking human on screen, and past that first impression they are barely in the same product category. Griffin is a Human Interaction Model — a video-to-video system built to hold a two-way conversation in real time, announced by Tavus in October 2026 and currently limited to a select group of early testers. MiniMax H3 is an omni-modal video generator: you hand it a prompt plus optional image, video and audio references and it returns a finished clip. One is being evaluated behind closed doors; the other has been downloadable, licensable and billable for months. That difference — not the resolution or the frame rate — is what decides which of them belongs in your stack.

What each one actually is

Griffin is Tavus's attempt to build a model that behaves like a participant rather than a renderer. The company describes it as the first Human Interaction Model, and the architecture it published splits the job in two: Continuous Conversational Modeling, which decides what the persona should be doing from moment to moment, and Audio-Visual Generation, which streams the speech and the face that match. Crucially, it is full duplex — it watches and listens while it is speaking, so it can be interrupted, can react mid-utterance, and does not wait for a clean end-of-turn signal before responding. Tavus's own framing is that previous systems learned to generate faces; Griffin was built to learn conversational behaviour from people.

MiniMax H3 is the Hailuo 3 generation, released on 31 July 2026 and open-sourced on 3 August 2026 under the MiniMax H3 Community License, with the H3-Base-FL2VA and H3-Base-Ref2VA variants published in the open. It reads text, image, video and audio references as one unified context and returns a 4-to-15 second clip at 768P or 2K with native stereo audio, generated through a three-stage pipeline (H3-Context-IR, then H3-Base at 768p, then H3-Regenerate-2K). It is a generator with an unusually wide input surface and a genuinely permissive licence — everything Griffin currently is not.

The specifications, side by side

Two-card board titled 'One Is a Product'. Left card, MiniMax H3: omni-modal video generator, Hailuo 3 generation; released 31 July 2026 with open weights on 3 August 2026; 4 to 15 seconds at 768P or 2K with native stereo audio; text, image, video and audio inputs in one context; $0.08 per second at 768P and $0.13 at 2K; hosted and billable today. Right card, Tavus Griffin: duplex Human Interaction Model, video to video; announced October 2026 as a research preview; 720p at 25 fps in 320-millisecond chunks; the live participant's video and audio as input; no published price; selected testers only, not for customer use.

Why "duplex" is the interesting word

Every video generator, H3 included, is judged on what a clip looks like once it is finished. Griffin is judged on something a clip cannot show: what happens in the 200 milliseconds after you interrupt it. That is the problem the Human Interaction Model framing is pointing at, and it is why Tavus's headline numbers are behavioural rather than visual.

The most quoted figure is that 48% of people believed Griffin was a real person after a one-minute video call — 26 of 54 participants — against 2.4% for Tavus's previous stack of Phoenix-4.5, Sparrow-2 and Raven-1, which managed 1 of 41. Tavus also reports that people who were not convinced tended to suspect within the first 20 seconds, which is a more useful detail than the headline: it means the illusion either holds early or collapses early.

The second set of numbers comes from NVIDIA's VideoFDB benchmark, scored in September 2026, where Tavus says Griffin placed first on an independent test of face-to-face AI. On the generation side Griffin scored 3.83 against 2.80 for the strongest reported baseline (a Gemini 2.5 plus Anam configuration) with a human reference of 3.92, and a task-objective rate of 62.8%. On perception it scored 3.73 against 3.44 for MiniCPM-o 4.5, 3.17 for Gemini 2.5 Flash Native and 2.97 for an audio-only OpenAI gpt-realtime configuration, against a human reference of 4.20 and a task-objective rate of 73.8%, across fifteen models evaluated. Tavus also claims Griffin leads by 37% over the next best system at reacting in the moment. These are vendor-reported figures resting on a benchmark NVIDIA built, and they should be read that way — but the human comparison points are the ones worth keeping, because a 3.83 against a 3.92 human score is a much smaller gap than video generation has historically shown.

What you can buy today

Here the two models stop being comparable at all.

MiniMax H3 is a product. It is hosted, it is priced per second of generated output, and its open variants are sitting on a model hub waiting to be downloaded. If you want a fifteen-second, 2K, stereo-audio clip of something that does not exist, you can have one this afternoon, and you can run the weights yourself if you would rather not rent them.

Tavus Griffin is not a product. Tavus is explicit in the Griffin write-up that Griffin-Lite, the research preview, is available today to a select group of early testers, that a wider release of a more powerful model will follow, and that Griffin is not on the Tavus platform yet and will only arrive once the company has worked out how to release it safely. That is an unusual amount of candour from a vendor about its own flashiest result, and it matters for planning: you cannot put Griffin in a product roadmap as a dependency, and you cannot put a price on it, because there isn't one.

So the honest planning question is not "which is better" but "which one exists at the layer I need".

Screenshot of Artificial Analysis's 'AA-Video-T2V v2.0' leaderboard, the Overall board for text-to-video, captured 2 October 2026: Wan 3.0 first at Elo 1,157, Utopai X (based on MiniMax H3) second at 1,150, Dreamina Seedance 2.5 third at 1,144, MiniMax H3 (768p) fourth at 1,139 and MiniMax H3 Max fifth at 1,134, with samples, release month and per-minute API pricing columns. The board lists MiniMax H3 and MiniMax H3 Max fourth and fifth with audio, though the with-audio filter itself is not applied in this capture.

Where OrcaRouter fits

MiniMax H3 is on our catalogue. The model page sits at minimax/minimax-h3, and it is billed the way the provider bills it: $0.08 per second of generated output at 768P and $0.13 per second at 2K. OrcaRouter passes provider list prices through at 0% markup, so a MiniMax price change shows up on our card the same day it is announced rather than at the next billing cycle. Griffin is not on the catalogue and will not be until Tavus ships it to customers — we do not host a model we cannot serve, and a research preview limited to selected testers is not something a router can route to.

The practical consequence is that one of these two models is a single line of code away from your application and the other is a research paper with a phone number. If you already route several video and language models, the per-second billing on H3 also makes it the sort of route worth putting behind automatic failover: a request that hits a saturated provider retries against a healthy one instead of surfacing a timeout, and a routing DSL lets you pin 2K jobs to a specific provider while letting 768P drafts fall wherever capacity is cheapest. Model fusion is the other lever worth mentioning here — H3's per-second price is low enough that spending a text model to rewrite and re-cut a prompt before generation is often cheaper than the wasted seconds from a bad first attempt.

Screenshot of the OrcaRouter model page for MiniMax-H3, showing the model id minimax/minimax-h3, the description 'omni-modal video generation model (the Hailuo 3 generation), released July 31, 2026', a price of $0.08 per second, p50 time to first token of 375 ms and p95 of 416 ms over seven days.

Where the comparison could flip

Three things would change this comparison, and only three.

First, if Tavus puts Griffin on a commercial API with a published rate, the whole "conversation versus clip" framing becomes a genuine buying decision rather than a scheduling decision, and the 48% face-to-face figure starts competing directly with the quality of generative output. Second, if H3's real-time story improves — and MiniMax has been shipping aggressively — a per-second generator that can stream would blunt Griffin's core advantage. Third, if NVIDIA or another neutral party re-runs VideoFDB with H3 or its successor included, the claim that Griffin is first at face-to-face AI stops being unfalsifiable. Tavus says it anticipates releasing Griffin very soon after its safety concerns are addressed; that sentence is the whole timeline.

Bottom line

MiniMax H3 and Tavus Griffin are not rivals right now, because only one of them is buyable. H3 is a shipped, open-weight, per-second omni-modal generator with a published price and a 4-to-15 second output window; Griffin is a duplex conversational model with the best face-to-face numbers anyone has published, no API, no price, and an explicit statement from its own vendor that customers cannot use it yet. If you need video generation today, H3 is the answer and it is measurable in cents per second. If you need a real-time face that can be interrupted, watch Tavus — but build the rest of the product first, because the release date is a sentence, not a date.