
Tavus Griffin vs HeyGen Video 1: Thirty-Six Hours Apart, and Only One Is Buyable
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 129 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 103 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 212 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Two launches landed within thirty-six hours of each other in the same corner of the market, and the calendar is the most interesting thing about the pairing. HeyGen Video went live on 30 September 2026 at $0.01 per second through October, tuned from MiniMax H3 and shipping under the identifier heygen-video-1. Tavus Griffin was announced on 1 October 2026 with a live face-to-face study in which 26 of 54 participants believed their partner was a real person, and it is available only to select trusted testers behind a request form, with no identifier, no endpoint and no price. Both are about generating a convincing human being. The honest reason to compare Tavus Griffin and HeyGen Video 1 is not that a buyer would choose between them — today only one of them can be bought — but that the difference explains what each was built to do, and what a second of generated human is actually worth.
One sells a clip, the other sells a conversation
HeyGen Video is a general-purpose video model that happens to be unusually good at people. It takes a prompt, or a prompt plus an image, or a prompt plus up to nine reference images, three reference videos and three reference audio clips, and returns a clip between 5 and 15 seconds long at 480p or 768p, 24 fps, with AAC audio at 32 kHz stereo carrying generated dialogue, ambience and effects. Three modes cover the conditioning — text_to_video, image_to_video, and reference_to_video, which is the default — and list order becomes the label you address in the prompt, so the first entry of reference_images is <Picture 1>. Unknown fields in the request are rejected rather than ignored, and a seed is an unsigned 32-bit integer that makes a shot repeatable when you hold it and change one clause at a time. It is a production tool with a deterministic envelope.
Griffin inverts nearly every one of those properties. It generates 720p in 320 ms chunks at 25 fps from a single reference photograph and plays each chunk as it decodes, so there is no clip length to specify and no file to hand back. A Continuous Conversational Modeling engine ingests audio and video and re-assesses the exchange at sub-second intervals, choosing whether to stay silent, signal, backchannel or take the turn, and its expressive controls drive a streaming speech generator and a streaming video generator in parallel. A decision made mid-sentence shows up in the voice and the face within the same mini-turn. You do not write a prompt; you talk to it, and it responds to a pause, a glance away, or an interruption.
That is why the two are not substitutes even though both generate a human. HeyGen Video is asked what a described moment looks like. Griffin is asked what should happen next.
What each second costs, on the one where you can buy seconds
HeyGen Video's launch price is $0.01 per second through October, and HeyGen's own prose describes that as 50% off a standard $0.02. The cost chart on the same page, footnoted as list price per second at 768p with audio before launch promotions, puts it at $0.03. That is a 50% gap between the prose and the chart on a vendor's own launch material, and until HeyGen publishes an API pricing page the safe reading is that $0.01 is confirmed because it is live, and the post-October number is somewhere between $0.02 and $0.03.
Run the numbers for a thousand seconds — a hundred ten-second clips — which is the unit a bulk team actually thinks in:
• HeyGen Video in October — $10 at $0.01 per second.
• HeyGen Video after October — $20 at $0.02, or $30 at the chart's $0.03.
• MiniMax H3, the base model it was tuned from — $80 at MiniMax's list price of $0.08 per second.
• Kling 3.0 Pro — $168 at $0.168 per second.
• Veo 3.1 — $400 at $0.40 per second.
The reason that arithmetic is the headline of HeyGen's launch is the discard rate. Teams making social cutdowns, ad variants and product loops generate far more than they ship, and at a dime a ten-second attempt the economics of throwing candidates away change shape entirely. The catch is the ceiling: 768p at 24 fps is not a 2K delivery format, and MiniMax H3 reaches 2K through a separate regeneration module on MiniMax's own hosted API at $0.13 per second, with no equivalent on HeyGen Video. If your delivery target is above 768p, the cheap second is not the second you need.
Griffin's column in that list is empty, and not in a promising way. Tavus has not published a rate, because Griffin-Lite is a research preview that its own announcement says will not be available for customer use at this time and is not on the Tavus platform. Under a thousand-second model there is nothing to put down. Anyone quoting a figure for Griffin is making it up.
The quality numbers, and who is grading
Neither of these models has an independent score, which is worth stating up front because both vendors have charts.
HeyGen's is an Elo table with HeyGen Video normalised to 1000, then Dreamina Seedance 2.0 at 955, the H3 Max family spread from 952 down to 860, Kling 3.0 Pro at 857 and Veo 3.1 at 744. The footnote does most of the work: the scores come from an internal evaluation set of Artificial Analysis Arena queries with 4,800 votes, and HeyGen's own score is normalised to 1000 for simpler comparison. A vendor normalising itself to the top of its own chart is a presentation choice, not evidence that it leads. The relative spacing beneath it is the informative part.
More useful is the head-to-head, which is the only place HeyGen tests against something it did not build. HeyGen says blind testers preferred its model to the charted H3 Max post-train in 55.8% of comparisons across 650 votes, with a range of 49–62%. That range is the honest detail: at the low end it straddles 50%, so on the vendor's own numbers the preference over that rival is not clearly separated from noise. The per-axis breakdown is mixed in a way that reads like a specific failure profile — 65% on staying true to the source image and 66% on timing, against 43% on consistency and 45% on music.
Griffin's numbers come from an instrument built by somebody else, which is the difference that matters. NVIDIA's VideoFDB is a benchmark for full-duplex audio-visual conversation scored on two tracks, and NVIDIA built it and conducts the evaluation with its own judge against a published rubric. Griffin-Lite scored 3.83 out of 5 on generation against a 3.92 human reference and 2.80 for the next-highest published system, and 3.73 on perception against a 4.20 human reference. Tavus also reports 0.43 seconds of true audio-to-video latency for the generator against four published streaming diffusion baselines on H100s. The caveats still apply — a 0.09-point gap to human on a five-point language-model-judged rubric is "close", not "equal", and part of the perception lead is measured against systems running without video input — but the grader is not the vendor.
• Independent quality scores — neither has one. HeyGen Video appears on none of Artificial Analysis's three video boards as of 2 October 2026, and so does Griffin. Griffin may never: pairwise preference voting over short clips cannot score a live conversation.
The same boards do carry the base model. MiniMax H3 sits 4th on text-to-video (with audio) at Elo 1,139 and 2nd on image-to-video at 1,181, both at $4.80/min — so HeyGen's post-train is competing on a substrate that an independent arena already rates near the top, which is the strongest argument in its favour that HeyGen did not publish itself.

Both are cheap or unpriceable for the same reason
The structural fact underneath both launches is that making a convincing human got cheap, and neither company's cost advantage comes from the thing it is selling.
HeyGen Video is a post-train of MiniMax H3, which MiniMax released with weights available under its own community licence. That turned H3 into a substrate other companies can specialise and resell, and HeyGen is the second product to do it in five weeks for a different vertical — business video, product demos, training, property walkthroughs. The pattern is why HeyGen can price below the base: it is not carrying the cost of the base model, and the base is somebody else's open-weights release.
Griffin is the opposite bet. Tavus built the whole system — the codec, the streaming speech generator, the streaming video generator, the conversational model — around the interaction itself, distilling a large bidirectional diffusion teacher into a few-step autoregressive generator in three stages so that long rollouts do not drift. That is expensive research with no open base to stand on, and it is a plausible part of why the release is gated: there is no cheap substrate underneath it to amortise.
Both companies also arrive at the same uncomfortable place from different directions. HeyGen's own documentation lists what the model is worst at, which is rare and worth reading: long on-screen text, soft organic motion like petals, paper and hair, and close hand work such as assembling or operating equipment. It is best on short, contained shots — one subject, one place, one action, a locked or barely moving camera. Griffin's limitation is not a list of failure modes but an unresolved safety problem: Tavus states plainly that the properties which make a human interaction model a better interface also make it better at persuading someone they are talking to a human, and that it is holding the preview while it builds disclosure mechanisms.
What a buyer can actually do today
HeyGen Video is callable now. It runs on POST /v3/models/videos with an x-api-key header, on paid API keys only, with download links that are signed and expire — so polling refreshes them — and callbacks attempted once per job, which HeyGen itself tells you to treat as a latency optimisation rather than a guarantee. If you are testing it this month, the $0.01 promotional second is live, the output ceiling is 768p, and the prompt enhancement mode is a real cost lever: turbo by default, quality for a longer pass, disabled to send your text through untouched.
Griffin is not callable. Access is a request form and a research preview. There is no key, no identifier, no endpoint and no SLA, and the only published statement about availability is that it will come once Tavus has worked out how to release it safely.
One thing both have in common is that neither is a model a router can help you reach, and in Griffin's case that is a category problem rather than a temporary one — a real-time full-duplex session is not a request-response call. What a router does help with in either pipeline is the layer around the video. Prompt expansion, reference-image inventory, shot descriptions, caption generation and review summarisation are all text calls, and the ones from OpenAI, Google and the rest sit on the OrcaRouter catalogue behind one OpenAI-compatible endpoint at 200+ models on the same key, at provider list price with 0% markup added, so a vendor price change is live the same day. Where the video model itself is concerned there is one useful fact: MiniMax H3 — the base HeyGen Video was tuned from, and the thing whose per-second price HeyGen's $0.01 is undercutting — is routed here as minimax/minimax-h3 at $0.08 per second, with 2K at $0.13. HeyGen Video itself is not in our catalogue and neither is Griffin; both are served through their own vendors' APIs and several third-party platforms.

Which one, and for whom
• Use HeyGen Video if the deliverable is short, contained, person-centric footage with built-in sound and you can live with 768p at 24 fps. Budget the post-October rate as somewhere between $0.02 and $0.03 per second rather than the promotional dime, and probe long on-screen text and close hand work before committing a workflow to it, because HeyGen has already told you those are where it breaks.
• Do not plan around Griffin. Request access if your question is whether a person can tell a generated human from a real one on a one-minute call, or if you need to see where a real-time interaction layer is heading before you commit a roadmap to rendered clips.
• Watch the disclosure question, because it is the one thing that will move both. HeyGen Video's customers have a straightforward answer available — the model is a post-train of an open-weights base and the output is a file with a known provenance — while Tavus is holding a model back specifically because it does not yet have one. Whichever company publishes the better disclosure mechanism will have settled the more valuable problem.

