A hero card for "Tavus Griffin vs Kling 3" with three chips reading "Kling 3 — $0.084 per request on OrcaRouter", "Griffin — no price, request access" and "Kling 3 — up to six shots", above six rows: Kling price: $0.084 per request; Griffin price: none published; Kling shots: up to six; Griffin shots: none to plan; Kling routing: routed as kling/kling-v3; Griffin routing: not routable. Footer: "Kling per-request price is OrcaRouter's live card; Griffin has no published rate."
Guides & Insights

Tavus Griffin vs Kling 3: A Real-Time Conversation Against a Ten-Second Clip

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The honest shape of this matchup is a routing asymmetry. Kling 3 is a callable model with a price, a parameter surface and an entry in our own catalogue; Tavus Griffin is a research preview for select trusted testers behind a request form, announced on 1 October 2026 with no identifier, no endpoint and no published rate. Kling 3.0 is in the OrcaRouter catalogue as kling/kling-v3, priced at $0.084 per request with provider list price passed through and nothing added. Griffin is not routable, and not because nobody has got round to it — a real-time full-duplex session in which the model generates 720p in 320-millisecond chunks while listening is not a request-response call, and it would not fit the same shape even if Tavus opened the gates tomorrow. So this is not a comparison a buyer resolves by price. It is a comparison of two answers to the question of what a generated human is for.

What Kling 3.0 actually is

Kling 3.0 launched on 5 February 2026 as a four-model series — Video 3.0, Video 3.0 Omni, Image 3.0 and Image 3.0 Omni. Its defining feature is automatic multi-shot sequencing: the model plans shot transitions itself, and a custom mode lets you set the number of shots and the duration of each, with reviewers reporting usable sequences of up to about six distinct shots from a single prompt and shot size, camera movement and narrative beat specifiable per shot. The ceiling is 15 seconds per generation with a three-second floor. On the OrcaRouter side it appears as a flagship text-to-video and image-to-video model with multi-shot, subject and motion control, 3–15 second clips and up to native 4K claimed, served on POST /v1/video/generations.

Audio is generated in the same pass, natively. Kuaishou's version supports Chinese, English, Japanese, Korean and Spanish, handles mixed-language dialogue within a clip, binds a distinct voice to each character in multi-character scenes, and offers regional accents — American, British and Indian English; Northeastern, Beijing, Taiwanese, Cantonese and Sichuanese Chinese. Lip-sync consistency is the single most-flagged weakness in independent hands-on reviews, with one test reporting roughly two of five dialogue clips needing a retake, and reported artefacts clustering around dialogue in noisy ambiences where water and wind backgrounds leave speech muffled.

What Griffin actually is, in one paragraph

Griffin is not a video generator with a conversation mode. It is two engines running concurrently. A Continuous Conversational Modeling engine ingests a person's audio and video, re-assesses the exchange at sub-second intervals, and decides at each one whether to stay silent, signal, backchannel or take the turn — emitting expressive controls that carry timing, stance, facial expression and gesture, not just words. An audio-visual generation engine turns those controls into sound and pixels together: an autoregressive diffusion transformer that generates speech one latent chunk at a time from an encoded speaker prefix, and a few-step autoregressive video generator that produces 720p in 320 ms chunks at 25 fps from a single reference photograph. There is no prompt, no clip length, and no file. To audio it reports 0.43 seconds of true audio-to-video latency on H100s against four published streaming diffusion baselines, which Tavus describes as half the next-fastest method.

Price, and the one-sided table

Kling 3.0's pricing is the one place in this article where there is a hard number on both sides of a comparison and one side is empty.

• Kling 3.0 on OrcaRouter — $0.084 per request, provider list price passed through at 0% markup, so a Kuaishou rate change or promotional cut is live here the same day rather than after a contract cycle.

• Kling 3.0 on the independent arena — the text-to-video with-audio board read on 2 October 2026 lists the 1080p Pro tier at $10.08/min and the Omni 1080p Pro tier at $8.40/min. The same board rates Kling 3.0's four tiers across four different prices, so "the price of Kling 3.0" is not a single figure.

• Griffin — no price. Griffin-Lite is available to select trusted testers as a research preview, is not on the Tavus platform, and its own announcement says it will not be available for customer use at this time. There is no per-request rate, no per-second rate and no free tier to quote, and the model will arrive only after Tavus has worked out how to release it safely.

That is the whole price section, and it is honest to leave it there rather than invent a per-second estimate from the compute footprint.

The scoreboards measure different things and do not meet

Both models have been scored by somebody other than their vendor, on instruments that cannot be compared.

Kling 3.0 sits on Artificial Analysis's video arena, where Elo comes from blind pairwise human preference over short clips. Two readings on the same day tell different stories, which is the trap worth naming: on the text-to-video with-audio board the Kling tiers place mid-table, while on the image-to-video with-audio board they place materially higher — but the Kling tiers do not all appear on both. Only the Pro and Omni Pro 1080p builds are on the text board at all; the 720p Standard tier is scored on image-to-video and nowhere else.

• Kling 3.0 in text-to-video (with audio) — 1080p Pro at 16th, Elo 1,000 over 4,844 votes, $10.08/min; Omni 1080p Pro at 16th, Elo 987 over 2,741 votes, $6.90/min. The more expensive of the two tiers scores lower, which is the same non-result the tier spread shows everywhere else on this board.

• Kling 3.0 in image-to-video (with audio) — 1080p Pro at 18th–20th, Elo 1,055 (±6) over 14,896 votes, $20.16/min; 720p Standard at 18th–20th, Elo 1,051 (±6) over 14,692 votes, $15.60/min; Omni 1080p Pro at 21st–22nd, Elo 1,044 (±8) over 2,945 votes, $16.80/min; Omni 720p Standard at 21st–24th, Elo 1,036, $13.44/min.

The reading is that Kling 3.0 is stronger when it starts from a supplied image than when it starts from text, and the image-to-video board carries three times the vote count, so its placement there is the better-resolved one. That is also the mode most commercial work uses. Note what the four tiers buy: within image-to-video they span sixteen Elo points across a price range from $13.44 to $20.16 a minute, and the cheapest of the four is not the lowest-ranked. Paying more for Kling 3.0 does not move it up this board.

A screenshot of Artificial Analysis's image-to-video-with-audio board, captured 2 October 2026, covering rows fifteen to thirty-one: PixVerse V6 at Elo 1,070, SkyReels V4 at 1,070, Veo 3.1 Fast at 1,066, Vidu Q3 Pro at 1,056, Kling 3.0 1080p Pro at Elo 1,055 with 14,896 votes and $20.16/min, Kling 3.0 720p Standard at 1,051 with 14,692 votes and $15.60/min, Kling 3.0 Omni 1080p Pro at 1,044 with 2,945 votes and $16.80/min, LTX-2.5 Fast at 1,038, Kling 3.0 Omni 720p Standard at 1,036 with 6,198 votes and $13.44/min, Agnes-Video-2.5 at 1,031, Vidu Q3 Turbo at 1,029, LTX-2.5 Pro at 1,007, Seedance 1.5 Pro at 1,000 and Kling 2.6 Pro (January) at 988.

Griffin's scores come from NVIDIA's VideoFDB, which scores full-duplex audio-visual conversation on two tracks and was built and run by NVIDIA with its own judge against a published rubric. Griffin-Lite scored 3.83 out of 5 on generation against a 3.92 human reference and 2.80 for the next-highest published system, and 3.73 on perception against a 4.20 human reference, with takeover-rate alignment of 62.8% and 73.8%. Those numbers say nothing about whether a clip looks right at 1080p, and Kling's Elo says nothing about whether a model can be interrupted mid-sentence. The only honest use of either is inside its own question.

The gap nobody is measuring: latency against length

There is a real engineering contrast here that neither leaderboard captures, and it is the most useful thing in this comparison for anyone planning a product.

Kling 3.0 optimises for length and structure inside a bounded generation: up to fifteen seconds, up to about six planned shots, per-shot control. Its cost scales with output duration, and its quality scales with how specific the prompt is. Nothing about it runs while the user is still talking, and that is not a flaw — it is the shape of the category.

Griffin optimises for latency inside an unbounded session: 320-millisecond chunks, 25 fps, a decision made every sub-second interval, no clip to plan because the conversation has no end. Its cost, whatever it turns out to be, would scale with conversation duration rather than with rendered seconds, and its quality scales with how natural the person on the other end is rather than with the brief. The three-stage distillation into an autoregressive generator trained on its own history exists precisely because the failure mode of this architecture is drift over a long rollout rather than a bad shot.

Put the two side by side and the practical consequence is that they fail in opposite directions. Kling 3.0 fails visibly and early — a borderline take you can delete and regenerate for a couple of cents, with the retake budget built into the workflow. Griffin, if it worked as reported, would fail invisibly and late, by not being what it appears to be, which is the reason Tavus is holding it.

Where a router genuinely helps, and where it does not

Kling 3.0 is in the OrcaRouter catalogue as kling/kling-v3 at $0.084 per request, on the same key and the same OpenAI-compatible endpoint as 200+ other models. That is a real difference from holding a Kuaishou account plus a separate integration for whichever text model your pipeline needs: because we add nothing to the provider price, a Kuaishou rate change is live on our side the same day, and if the upstream provider degrades or rate-limits, automatic failover moves the request before the response starts rather than returning an error to your application. For a video pipeline where one call can cost more than a thousand text calls, having the request survive a bad minute upstream is worth more than the markup you are not paying.

Griffin is not in our catalogue, is not in anyone's catalogue, and cannot be — a full-duplex session model is not a routed request. It is reachable only by submitting an access request. Tavus's own pointing of would-be builders at its older models rather than Griffin is the tell: the three predecessors power the PALs that 150,000 developers and businesses already use, and Griffin is not among them.

A screenshot of OrcaRouter's own model page for kling/kling-v3, captured 2 October 2026: the description "Kling 3.0 — flagship text-to-video and image-to-video, multi-shot + subject + motion control, 3-15s clips, up to native 4K", the /v1/video/generations endpoint, a price of $0.08, a p50 of 1.00 s, and an OpenAI-compatible code sample against api.orcarouter.ai/v1.

Which one, for what

• Choose Kling 3.0 for planned multi-shot sequences from one prompt, for image-to-video work where its no-audio arena placement is strong and heavily voted, for multilingual dialogue with per-character voice binding across five languages, for element consistency across camera moves, and for 4K if you confirm the tier in writing — Kuaishou's own documentation claims native 4K while its user guide lists only 720p and 1080p, and third-party credit rates for 4K vary by more than a factor of two between sources.

• Test Kling 3.0 on lip sync first if dialogue is the point, because that is where independent hands-on reviews say it breaks, and it is the failure that costs the most retakes.

• Do not plan around Griffin. Request access if the question you need answered is whether a person can tell a generated human from a real one on a one-minute call, or if knowing where the interaction layer is going should shape what you build on the clip layer now.

• Watch two things, not one. On the Kling side, whether the tier spread starts to track price, because the four tiers currently sit within nineteen Elo points of each other on image-to-video while differing in price by half again, and the cheapest tier is not the worst-placed. On the Griffin side, whether a disclosure mechanism and a model card arrive; if they do, the comparison stops being a design essay and starts being a procurement decision, and this article will need rewriting.

A two-column scoreboard titled "Tavus Griffin vs Kling 3 — the scoreboard". Left column "Tavus Griffin": Status: research preview; Price: none published; Loop: full duplex; Output: 720p in 320 ms chunks; Shots: none to plan; Routing: not routable. Right column "Kling 3": Status: generally available; Price: $0.084 per request; Loop: one generation; Output: 720p and 1080p; Shots: up to six; Routing: routed as kling/kling-v3. Footer: "Kling price per OrcaRouter's live card; arena Elo 1,000 to 1,055 per Artificial Analysis, 2 October 2026."