A hero card for "Tavus Griffin vs Google Veo 3.1" with three chips reading "Veo 3.1 — SynthID on every frame", "Griffin — 48% of a live study fooled" and "Veo 3.1 Standard — from $0.40 per second", above six rows: Griffin provenance: none published; Veo provenance: SynthID plus C2PA; Griffin price: none published; Veo price: $0.40 per second; Griffin status: research preview; Veo status: generally available. Footer: "Veo 3.1 rates are Google Gemini API list prices; Griffin is a research preview with no pricing."
Guides & Insights

Tavus Griffin vs Google Veo 3.1: One Watermarks Every Frame, the Other Fooled 48% of a Room

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put Tavus Griffin and Google Veo 3.1 side by side and the conventional comparison — price per second, Elo, clip length — falls apart immediately, because only one of them is priced and only one of them is scored on anything a buyer can look up. But there is a real axis underneath, and it is not quality. Google's answer to synthetic media is to make it detectable: every Veo 3.1 output carries SynthID, an invisible watermark written into the pixels frame by frame and into the audio spectrum, plus C2PA Content Credentials recording file provenance, and Google has shipped a detector into the Gemini app so a user can ask whether Google AI made a clip. Tavus Griffin's stated reason for existing in preview rather than as a product is the mirror image: in Tavus's own live study, 26 of 54 participants finished a one-minute video call believing their partner was a real person, against 1 of 41 on the company's previous stack, and Tavus says it is holding the model back until it has worked out how to disclose what it is. One of these vendors sells detection. The other is currently the reason detection matters, and knows it.

Two products built on opposite assumptions

Veo 3.1 is a clip generator with a deliberately narrow envelope. On Google's Gemini API it produces 4, 6 or 8 seconds per pass at 24 fps, natively at 720p, with 1080p and 4K arriving through an upscaling pass added in a January 2026 update. Extension adds seven seconds per pass and can repeat up to twenty times for a theoretical ceiling around 148 seconds, but the input must be a Veo-generated clip of no more than 141 seconds at 720p in 16:9 or 9:16, and the eight-second length is mandatory for extension, for reference-image input, and for anything above 720p. You cannot make a four-second 1080p clip. The output is a file you own, watermarked, with a signed provenance record attached.

Griffin does not produce a file. It runs a conversation. A Continuous Conversational Modeling engine takes in audio and video and re-assesses the exchange at sub-second intervals, choosing whether to stay silent, signal, backchannel or take the turn, and emits expressive controls that drive a streaming speech generator and a streaming video generator together. The video generator produces 720p in 320 ms chunks, one latent at a time, at 25 fps from a single reference photograph, and plays each chunk as it decodes. There is no clip length because there is no clip; there is a session that continues until someone stops talking.

That distinction decides almost every other line in this comparison. Veo 3.1's envelope is a set of constraints a production pipeline can plan around — shot lists of eight seconds, an extension ladder for longer takes, a fixed 24 fps, a known resolution ladder. Griffin's output is not storyboardable in advance at all, because what it renders depends on what the person does next.

What Veo 3.1 costs, at every tier

Google's published Gemini API rates are audio-inclusive, per second of output, and there is no free tier at any Veo 3.1 level. Audio cannot be disabled on that API — it is generated on every request — which matters because third-party sources describe a Vertex AI rate card where silent video runs at roughly half the audio-inclusive figure. Google's own documentation does not expose that toggle. If a silent rate matters to your bill, confirm it against your own Vertex billing rather than a comparison page, including this one.

• Veo 3.1 Standard — $0.40/s at 720p and 1080p, $0.60/s at 4K.

• Veo 3.1 Fast — $0.10/s at 720p, $0.12/s at 1080p, $0.30/s at 4K.

• Veo 3.1 Lite — $0.05/s at 720p, $0.08/s at 1080p, with no 4K tier.

Run the eight-second rule against those rates and the per-attempt cost is what a creative team actually budgets: 32 cents for an eight-second 1080p take on Standard, 80 cents for a four-second one at the same resolution because the eight-second floor does not apply below 720p, and $4.80 for a single eight-second 4K take. That last number is the one to sit with, because a four-attempt search for one usable 4K shot is just under twenty dollars, and the eight-second requirement means you cannot test the idea at half the length for half the price.

Griffin has no price, no free tier, and no rate card of any kind. Griffin-Lite is available to select trusted testers as a research preview by request form, is not on the Tavus platform, and the announcement is explicit that it will not be available for customer use at this time. Tavus's closing line on the page is that Griffin will arrive once the company has worked out how to release it safely. Every claim in this article that compares the two on anything resembling a purchasing decision has a hole in Griffin's half of it, and no amount of research fills it.

What the two scoreboards actually measure

The two models have been evaluated by different instruments for the same reason they are hard to price against each other.

Veo 3.1 lives on Artificial Analysis's video arena, where Elo comes from blind pairwise human preference over short clips. Read on 2 October 2026 on the text-to-video board with audio:

• Veo 3.1 — 18th–21st, Elo 968 (±11) over 2,857 votes, $24.00/min.

• Veo 3.1 Fast — 18th–21st, Elo 961 (±10) over 3,595 votes, $7.20/min.

• Veo 3.1 Lite — 21st–24th, Elo 949 (±11) over 2,762 votes, $4.80/min. • Veo 3.1 in image-to-video (with audio) — 11th of 34, Elo 1,082 over 7,049 votes, $24.00/min. Its best placement on any board, and the one with the most votes behind it.

Two things fall out of that. All three tiers sit within nineteen Elo points of each other with overlapping confidence intervals, so the standard tier's four-times-higher per-second price does not buy measurable blind preference. And Google's own newer Gemini Omni Flash line outranks every Veo 3.1 tier on the same board — 7th–8th at Elo 1,117 over 7,801 votes and $9.12/min, ahead of all three tiers while costing between a quarter and a fifth as much per minute — which is a question every Veo 3.1 buyer should be asking, and one this article is not the place to answer.

Griffin's numbers come from NVIDIA's VideoFDB, a benchmark for full-duplex audio-visual conversation that NVIDIA built and scored independently in September 2026. Griffin-Lite scored 3.83 out of 5 on the generation track against a 3.92 human reference and 2.80 for the next-highest published system, and 3.73 on the perception track against a 4.20 human reference. Tavus also reports 0.43 seconds of true audio-to-video latency for the generator against four published streaming diffusion baselines on H100s, describing it as half the next-fastest.

Neither set transfers. An Elo from short-clip preference voting says nothing about whether a model can be interrupted; a judge's rubric score on conversation says nothing about whether a clip looks right at 4K. And the caveats run in both directions: Griffin's 0.09-point gap to the human reference is small enough on a five-point language-model-judged rubric that "close to human" is the honest reading, and part of its perception lead is measured against systems running without video input at all.

A screenshot of the middle of Artificial Analysis's text-to-video-with-audio board, captured 2 October 2026, covering rows eleven to twenty-four: grok-imagine-video-1.5 at Elo 1,046, HappyHorse-1.1 at 1,045, Wan2.7-260612 at 1,032 with 4,645 votes and $9.00/min, HappyHorse-1.0 at 1,026, Kling 3.0 Omni 1080p Pro at 1,018, Kling 3.0 1080p Pro at 1,000 with 4,844 votes, PixVerse V6 at 987, Veo 3.1 at 968 with 2,857 votes and $24.00/min, Veo 3.1 Fast at 961 with 3,595 votes and $7.20/min, MAGI-2 Preview (0912) at 960, LTX-2.5 Fast at 950, Veo 3.1 Lite at 949 with 2,762 votes and $4.80/min, LTX-2.5 Pro at 944 and Vidu Q3 Pro at 941.

Provenance, which is the real product difference

For a business with any disclosure obligation, this is the only section that matters, and it is not close.

Veo 3.1 output carries SynthID frame by frame in the pixels and in the audio spectrum, engineered to survive compression, resizing, cropping and screen recording, and C2PA Content Credentials recording file provenance and edit history, interoperable with Adobe's and Microsoft's implementations. Google has shipped detection into the Gemini app, so a user can upload a video and ask whether Google AI made it, with the detection reporting which section of the audio or video track carried the mark. The limit is real and worth stating: that detector recognises Google's models only. Nothing on Alibaba's or Kuaishou's books is comparable, and nothing from Tavus is either.

A screenshot of Google DeepMind's SynthID page, captured 2 October 2026: the Google DeepMind wordmark, the "SynthID" heading, the strapline "A tool to watermark and identify content generated through AI", the section tabs Overview / How it works / Hands-on, the heading "What is SynthID?", and a "Build with Gemini" panel with a Try Gemini link.

Tavus has not published a watermark, a detector, or a content-credentials pipeline for Griffin. What it has published is the reason it needs one. The face-to-face study is unusually blunt about the failure mode: over half the participants said the possibility of an AI had not crossed their mind during the call, nearly all of that group concluded their partner was real, and the people who did suspect tended to work it out inside the first twenty seconds. Believers were on average 79% confident and doubters 81%. That is a model whose deceptiveness is not a side effect of the generation quality — it is the generation quality.

Tavus's response is in the announcement rather than a footnote: the same properties that make Human Interaction Models powerful interfaces for natural communication allow them to deceive a human into believing it is not AI, further alignment and safety procedures are required for safe release, and the company is working on safe disclosure features and with organisations tackling AI safety. Which is why the gate exists. Griffin-Lite will not be available for customer use at this time.

The enterprise terms Google brings and Tavus does not

Veo 3.1 is served through Vertex AI with the regional infrastructure, IAM and enterprise contract surface that implies, and Google constrains what the model will do by jurisdiction: through a person-generation parameter, the EU, UK, Switzerland and MENA permit only adult subjects, while other regions can allow all. That is a documented, region-aware policy surface, and for a regulated buyer it can outweigh a leaderboard position.

There is a cost to the shape of it. Veo 3.1's model IDs on the Gemini API are still preview IDs — veo-3.1-generate-preview and its siblings — while the Vertex SKU carries general-availability naming, and third-party reports put the preview quota at roughly a fifth of the production quota. Google retired the older Veo 3.0 SKUs on 30 June 2026, which tells you how generational turnover is handled and is worth pricing into anything built on a preview ID. The consumer surface is also genuinely restricted: Flow requires a Google AI Pro or Ultra subscription and is blocked in the EU, UK, India, Indonesia, mainland China and Russia, with no VPN workaround offered.

Griffin has no terms of service to read, because there is no service. Access is a request form and a research preview. Tavus has said it is working with safety organisations on evaluations and wants people to participate in them, which is a research posture rather than a commercial one.

Getting either of them

Veo 3.1 is reachable through the Gemini API, Vertex AI, Google AI Studio and Flow, and through the Gemini app at 720p only. Outside China it is by a wide margin the easier of the two to obtain a key for, which cuts against the quality argument above and is the main reason this comparison stays interesting despite the numbers.

Griffin is reachable by submitting an access request and being selected as a trusted tester. That is the entire procurement path.

Where a router genuinely helps in a Veo pipeline is adjacent to the video model rather than in it. The Google text models — Gemini 3.5 Flash, Gemini 3.1 Pro Preview, Gemini 3.6 Flash and the rest of the line — are in the OrcaRouter catalogue and callable on the same OpenAI-compatible endpoint as 200+ other models at provider list price with 0% markup added. Shot lists, prompt expansion, caption generation, continuity notes between eight-second segments, and review summarisation are all text work, and that is the layer where being able to put a cheap model on a bulk pass and a frontier model on the final one costs a config change rather than a migration. Veo 3.1 itself we do not route, and neither is Griffin; it is worth being explicit about that rather than letting a paragraph like this imply otherwise.

Which one, honestly

• Choose Veo 3.1 when the deliverable needs a provenance trail, when the buyer is regulated, when distribution has to be a file with credentials attached, or when 4K is non-negotiable. Budget for the eight-second floor in every cost model, and check the preview-ID lifecycle before building a customer-facing path on it.

• Do not choose Griffin, because you cannot. Request access if your question is whether a person can tell an AI from a human on a one-minute call, or if you need to know where the interaction layer is going before you commit a roadmap to the clip layer.

• Watch the detector, not the demo. If Tavus publishes a disclosure mechanism that survives a casual conversation, the gap between these two vendors narrows considerably. If it does not, Griffin is a result worth citing and Veo 3.1 remains the only one of the two you can put in front of a compliance team.

A two-column scoreboard titled "Tavus Griffin vs Google Veo 3.1 — the scoreboard". Left column "Tavus Griffin": Provenance: none published; Price: none published; Clip length: no clip - a session; Resolution: 720p at 25 fps; Loop: full duplex; Status: research preview. Right column "Google Veo 3.1": Provenance: SynthID plus C2PA; Price: $0.40 per second; Clip length: 4, 6 or 8 seconds; Resolution: 720p native; Loop: one generation; Status: generally available. Footer: "Veo 3.1 per Google's Gemini API pricing, 2 October 2026."