Hero title card for 'Tavus Griffin vs Runway Gen-4.5' with the subtitle 'Conversation and Rendering Are Two Different Bills', above four chips: Gen-4.5 $0.12 per second of footage; Gen-4.5 2-10 s clips, no audio; Griffin full duplex, interruptible; Griffin no published price.
Guides & Insights

Tavus Griffin vs Runway Gen-4.5: Conversation and Rendering Are Two Different Bills

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put Tavus Griffin and Runway Gen-4.5 next to each other and the first thing you notice is that they do not share a unit of measurement. Griffin, announced by Tavus in October 2026 as a research preview, is a Human Interaction Model — a video-to-video system that holds a live face-to-face conversation and generates the other participant's face and voice as the call runs. Runway Gen-4.5, announced on 1 December 2025 and rolled out to paying subscribers from 12 December, is a text-to-video and image-to-video generator that charges you by the second of footage it renders. One of them is sold by the turn and the other by the second, and the only honest way to compare them is to notice that a minute of Griffin and a minute of Gen-4.5 are not the same minute.

Two clocks running in opposite directions

Gen-4.5 has been in paying customers' hands for the better part of a year. It produces clips of two to ten seconds from a text prompt or a still image, at 720p and 24 or 25 frames per second, across six aspect ratios. There is no audio track, no multi-shot storyboarding, and no dialogue — it is a rendering engine, and a disciplined one. Its API is asynchronous, which is the right shape for a job that takes longer than a request should. Runway's own API also serves third-party video models alongside its own, and the company ships a world model, GWM-1, which tells you where Runway thinks the interesting problem is: not in holding a conversation, in modelling a scene.

Griffin points the other way. Every design decision in Tavus's published architecture is about latency rather than fidelity. The Tavec codec runs at 48 kHz with 40 values per frame at 100 frames per second, no codebooks, a fully causal decoder and packets as small as 10 milliseconds. The video path emits 720p in 320-millisecond chunks, where one latent covers eight frames at 25 fps and each latent costs three diffusion steps. Audio-to-video latency averages 0.43 seconds on H100s, which Tavus says is roughly half the next fastest method it measured. Nothing in that stack is trying to render ten beautiful seconds. Everything in it is trying to render the next 320 milliseconds in time.

So the specification comparison is real but it is not a scoreboard:

• Billing unit — Runway Gen-4.5 is billed in seconds of output; you know your cost before you press generate. Griffin has no published price at all, because it has no public product.

• Clip length — Gen-4.5 produces two to ten seconds; Griffin runs for as long as the call runs.

• Resolution and rate — Gen-4.5 renders 720p at 24 or 25 fps; Griffin streams 720p at 25 fps in 320-millisecond chunks.

• Audio — Gen-4.5 produces no audio track; Griffin generates speech and video together and clones a voice from about a 10-second sample.

• Interactivity — Gen-4.5 is a one-way asynchronous request; Griffin is full duplex and can be interrupted mid-utterance.

• Motion and grading extras — Gen-4.5 offers ProRes and PNG export at a 5-credit-per-second surcharge and HDR at 20 credits per second, rising to 40 above roughly four megapixels; Griffin has no equivalent because it is not exporting files.

• Availability — Gen-4.5 has been live for paying subscribers since 12 December 2025; Griffin is in research preview for selected testers and Tavus states it is not available for customer use.

Board titled 'Two Clocks, Two Bills' with six labelled rows comparing Runway Gen-4.5 (left) against Tavus Griffin (right): billing unit 'Runway Gen-4.5 - 12 credits per second of video' versus 'Tavus Griffin - no price, no product'; clip length 2 to 10 seconds versus as long as the call runs; resolution 720p at 24 or 25 fps versus 720p at 25 fps in 320 ms chunks; audio no audio track versus speech and face generated together; interactivity one asynchronous request versus full duplex, interruptible; availability live for subscribers since 12 December 2025 versus research preview, not sellable.

The Gen-4.5 bill, in plain numbers

Runway prices Gen-4.5 at 12 credits per second of generated video. At the standard one cent per credit that works out to $0.12 per second, or roughly $7.20 for a full minute of footage if you could generate one continuously — which you cannot, because the model caps at ten seconds, so a minute means six or more separate generations stitched together with whatever continuity errors that implies. ProRes and PNG output adds 5 credits per second, and HDR adds 20 credits per second, rising to 40 above about four megapixels. API credits are billed separately from app-subscription credits, which is a detail that catches teams out when they prototype in the web app and then move to the API.

Board titled 'What One Minute Costs'. Left card, a minute of Runway Gen-4.5: rate 12 credits per second, about $0.12; hard limit 10 seconds per generation; so a minute means six or more separate generations; list price if it could run straight through about $7.20; extras ProRes or PNG +5 credits per second; HDR +20 credits per second, up to +40 above 4 MP. Right card, a minute of Tavus Griffin: rate none published; product research preview for selected testers; so a minute means not obtainable at any price; published evidence 0.43 s average audio to video on H100s; vendor statement not available for customer use at this time; availability no API, no price, no date.

Where Gen-4.5 is reached matters too. It is available through Runway's own API and through several third-party platforms. It is not an OrcaRouter route — we do not host it, and we will not suggest you can get it here.

What Griffin costs, and why that is the wrong question

There is no price for Tavus Griffin. Tavus's release material states that Griffin-Lite, the research preview, is available today to a select group of early testers, that a wider release of a more powerful model will follow, that Griffin-Lite will not be available for customer use at this time, and that Griffin is not on the Tavus platform yet — it will come once the company has worked out how to release it safely.

That paragraph is doing a lot of work. It means the comparison in this article is not a procurement decision and cannot become one until Tavus names a rate. It also means Tavus is being unusually direct about a result most vendors would have shipped as a product page with a "contact sales" button.

What Griffin does publish is evidence about behaviour. In a study run with Queen Mary University of London, 26 of 54 participants — 48% — believed Griffin was a real person after a one-minute video call, against 1 of 41 (2.4%) for Tavus's previous Phoenix-4.5, Sparrow-2 and Raven-1 stack. On NVIDIA's VideoFDB benchmark, scored in September 2026, Tavus reports first place on face-to-face AI: generation 3.83 against a 2.80 strongest reported baseline and a 3.92 human reference, perception 3.73 against 3.44 for the best reported baseline and 4.20 for a human, with a 73.8% task-objective rate on the perception side across fifteen models evaluated. Those are vendor-reported numbers resting on a benchmark NVIDIA built, and they are about conversation, not cinematography.

Screenshot of the Tavus website's Griffin announcement, captured 2 October 2026: the headline 'The First Human Interaction Model' under 'Introducing Griffin', a paragraph describing a video-to-video system that listens while the other person talks, and three statistic tiles reading 48% of people believed Griffin was a real person after a one-minute video call, #1 on NVIDIA's independent test of face-to-face AI, and 37% ahead of the next best AI at reacting in the moment, dated October 1st, 2026.

The different jobs, stated plainly

If you want a shot of a city at dusk with a camera move that matches a reference frame, Gen-4.5 is a tool that exists and can do it today for about twelve cents a second. There is no scenario in which Griffin helps you, because Griffin does not accept a prompt and does not return a file.

If you want a face that listens, gets interrupted, and keeps a conversation coherent over several minutes, Gen-4.5 cannot help, because it does not accept input after the request is sent. A ten-second clip with a talking head in it is not a conversation; it is a video of a conversation, and the distinction is exactly the one Tavus is selling.

The overlap is narrower than the "AI video" label suggests: both output pixels of a person, and both run at 720p and 25 fps. Everything above that is different, including the unit on the invoice.

Why a router is useful in exactly this spot

Teams that end up needing both usually discover it the hard way — a product wants a generated establishing shot and a talking avatar, and the two arrive from different vendors with different auth, different credit systems and different failure modes. This is the case where a single endpoint earns its keep. OrcaRouter carries over two hundred models behind one key, passes provider list prices through at 0% markup so a vendor price change is live on the card the same day, and retries against a healthy provider when one route fails instead of handing you a timeout. A routing DSL lets you send low-stakes drafts wherever capacity is cheapest and pin anything customer-facing to a specific provider, and model fusion lets you put a cheap text model in front of an expensive generation call to rewrite a prompt before you spend the seconds.

To be exact about what that does and does not cover in this matchup: Gen-4.5 is not one of our routes, and neither is Griffin. The catalogue's video side includes minimax/minimax-h3, MiniMax's omni-modal generator, at $0.08 per second of output at 768P and $0.13 per second at 2K — a per-second price on the same axis as Runway's, from a model we can actually serve you today. That is the honest recommendation if what you need is generated seconds and you want them on one bill with everything else.

Which one to plan around

Gen-4.5 is the safer dependency by a wide margin: it has been in paying customers' hands since December 2025, its API is documented, its credits are priced, and its limits — ten seconds, no audio, no storyboarding — are published rather than discovered. A team that needs generated footage can start this week.

Griffin is the higher-ceiling bet with the lower floor. Its 48% and its 0.43-seconds average latency describe something genuinely new, and the fact that Tavus is refusing to sell it yet is a signal about the maturity of that something rather than a marketing decision. Do not put it in a roadmap as a dependency; put it on a watch list and check back when the safety note disappears.