A hero card for "Tavus Griffin vs Alibaba Wan 2.7", with two chips reading "Tavus Griffin — no published price" and "Alibaba Wan 2.7 — $9.00 per minute at 1080p", above six rows: Makes: a live conversation; Against: a 2 to 15 second clip; Griffin price: none published; Wan 2.7 price: $9.00 per minute; Griffin status: research preview; Wan 2.7 status: generally available. Footer: "Wan 2.7 rates are Alibaba Cloud list prices; Griffin is a research preview with no rate card."
Guides & Insights

Tavus Griffin vs Alibaba Wan 2.7: A Conversational Model Against a Generative One

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The tempting way to compare Tavus Griffin and Alibaba Wan 2.7 is per-second of video, and that comparison cannot be made. Wan 2.7 has a published rate card — $9.00 a minute at 1080p on Alibaba Cloud's international Singapore endpoints, with a cheaper Beijing and Tokyo tier — and Tavus Griffin has no rate card at all, because Griffin-Lite shipped on 1 October 2026 as a research preview for select trusted testers behind a request form. What the two do share is a word: video. Past that they are answers to different questions. One is asked what should happen next in a conversation; the other is asked what a described scene looks like. The interesting comparison is not quality, it is what each one needs from you before it will produce anything, and what you get back.

The difference is the loop, not the resolution

Wan 2.7 generates clip by clip from a prompt. You write a description, you get a piece of footage, you look at it, you write another. Wan 2.7 is a four-model suite — text-to-video, image-to-video, reference-to-video and video-edit — and the interaction is fundamentally transactional: an input, a rendered output, and a decision about whether to keep it. Nothing about it runs while you are still deciding what to ask for.

Griffin is built the other way round. It is a full-duplex system: a Continuous Conversational Modeling engine that ingests audio and video and re-assesses the conversation at sub-second intervals, deciding whether to stay silent, signal, backchannel or take the turn, and emitting expressive controls that drive a streaming speech generator and a streaming video generator in parallel. It generates 720p in 320 ms chunks from a single reference photograph, and the claim that matters is that a decision made mid-sentence shows up in the voice and the face within the same mini-turn. The benchmark for this is not a leaderboard of clips. It is whether you can interrupt it.

Put plainly: Wan 2.7 answers the question "what does this scene look like in motion?" Griffin answers "what should this face and voice do in the next two hundred milliseconds, given that the person just paused."

Price, on the one side where there is one

Wan 2.7's published rates are per second of generated video, and the tiers are not uniform across the suite, which is the detail most comparison pages skip.

• wan2.7-t2v and wan2.7-i2v — billed on output duration only. International Singapore: $0.10/s at 720p and $0.15/s at 1080p, which is the $9.00/min figure that appears on the leaderboards. Beijing and Tokyo list $0.086/s and $0.143/s for the same work.

• wan2.7-r2v (reference-to-video) — billed on the input video duration, capped at five seconds, plus the output. Feed a five-second reference into a fifteen-second generation and you are billed for twenty seconds.

• wan2.7-videoedit — billed on output only, with the input footage free.

Two lifecycle facts belong next to those numbers. Alibaba put a limited-time 30% launch discount on its own Bailian and Qwen Cloud surfaces running through 23 September 2026, which is now past. And Alibaba Cloud's Model Studio retirement notice (ID 118434) takes several older models offline on 10 October 2026, including wan2.7-r2v and wan2.7-image. That is variant-level rather than a retirement of the generation — wan2.7-t2v and wan2.7-i2v remain listed with live rates and pages updated as recently as 11 September — but if reference-to-video is why you were looking at 2.7, that route has a date on it.

The comparison against Griffin has one honest line: there is no price to put next to Wan 2.7's, because Tavus has published none. Griffin-Lite is not on the Tavus platform, the announcement says it will not be available for customer use at this time, and the company's own closing line is that Griffin will come once it has worked out how to release it safely. There is no per-second rate, no per-request rate, no free tier, no enterprise tier. Anyone quoting a price for Griffin is inventing it.

Quality scores: two boards that do not meet

The two models have been scored on entirely different instruments, which is why no Elo comparison between them exists.

Wan 2.7 has two entries on Artificial Analysis's video arena, and they are not in the same place — which is the trap in quoting one number for it. Elo there is derived from blind pairwise human preference votes, a methodology built for short generated clips, and both readings below are from boards read on 2 October 2026.

• Wan2.7-260612 (the June build) in text-to-video (with audio) — 13th–14th, Elo 1,032 (±10) over 4,645 votes, $9.00/min.

• Wan 2.7 (April build) in image-to-video (with audio) — 12th of 34, Elo 1,077 over 4,571 votes, $9.00/min, a board that carries roughly the same vote count but places it materially higher.

• Wan 3.0, the newer sibling — 1st at Elo 1,157 on text-to-video and 6th at 1,164 on image-to-video, both at $12.00/min. That is the number worth knowing before you compare Wan 2.7 to anything: Alibaba's own next generation is a table topper and 2.7 is a mid-table build.

A screenshot of the top of Artificial Analysis's "AA-Video-T2V v2.0" leaderboard, text-to-video with audio, captured 2 October 2026: Wan 3.0 first at Elo 1,157 with 7,575 votes and $12.00/min; Utopai X (based on MiniMax H3) second at 1,150; Dreamina Seedance 2.5 third at 1,144; MiniMax H3 (768p) fourth at 1,139 and $4.80/min; MiniMax H3 Max fifth at 1,134; FLUX 3 sixth at 1,127; Gemini Omni Flash eighth at 1,117 and $9.12/min; Wan2.7-260612 thirteenth at 1,032 with 4,645 votes and $9.00/min; and the two Kling 3.0 1080p tiers sixteenth, at Elo 1,018 and Elo 1,000.

Griffin sits on NVIDIA's VideoFDB, a benchmark for full-duplex audio-visual conversation that NVIDIA built and scored independently in September 2026, using a language-model judge against a zero-to-five rubric. Griffin-Lite scored 3.83 on the generation track against a 3.92 human reference and 2.80 for the next-highest published system, and 3.73 on the perception track against a 4.20 human reference. Tavus also reports 0.43 seconds of true audio-to-video latency for the generator against four published streaming diffusion baselines on H100s.

Both sets of numbers are legitimate and neither travels. Elo on the arena is a vote count about clip preference; VideoFDB is a judge's rubric score about conversational behaviour. There is no bridge between them, and the low-sample trap runs the other way too: the arena's ±10 band on Wan 2.7 is wide enough that its placement against neighbours is not tightly resolved, and NVIDIA's 0.09-point generation gap to human is small enough on a five-point rubric that "close to the human reference" is the honest reading, not "at human level". The perception comparison is also not like for like — several of the systems Griffin is placed ahead of, including OpenAI's gpt-realtime and MiniCPM-o 4.5, are scored in audio-only configurations while Griffin perceives video as well. Part of the 0.29-point lead measures having eyes.

What each needs from you

This is where the comparison becomes useful, because the inputs are the products.

• Input — Wan 2.7 takes text, an image, a video and audio. Griffin takes a live person on camera and microphone, and does not generate a clip on request at all.

• Output — Wan 2.7 produces a video file, 2 to 15 seconds per generation, up to 1080p, with native audio. Griffin produces an ongoing conversation: 720p video at 25 fps in 320 ms chunks, generated in real time and played as it is decoded.

• What steers it — Wan 2.7 is steered by the prompt and by up to five subject images and five reference clips, within a ceiling of ten reference assets per request. Griffin is steered by what the person does: a pause, a glance away, a gesture, an interruption.

• Time horizon — Wan 2.7's unit of work is a clip of at most fifteen seconds, which you then assemble. Griffin's unit of work is a conversation, which is why the architecture had to solve drift — the three-stage distillation into an autoregressive generator, trained on its own generated history so long rollouts stay stable — rather than fighting for a longer single generation.

• Output container — Wan 2.7 hands back a file you own and can edit, grade and distribute. Griffin hands back a rendered session. There is no file to edit, because the pixels are generated per chunk from a reference photo and the model's own history, and the background, the shadows and the chair the person sits in are part of the generation.

One structural difference worth naming: Wan 2.7's costs scale with output duration and its quality scales with how specific your prompt is. Griffin's costs are unmeasured and its quality scales with how natural the person on the other end is. You can improve one by writing better briefs and the other only by using it with real people.

A screenshot of the top of Artificial Analysis's "AA-Video-I2V v1.0" leaderboard, image-to-video with audio, captured 2 October 2026: MiniMax H3 Max first at Elo 1,195; MiniMax H3 second at 1,181; Gemini Omni Flash third at 1,178; Dreamina Seedance 2.0 720p fourth at 1,176; HiDream-O1-Video-1.0 fifth at 1,175; Wan 3.0 sixth at 1,164 with 9,302 votes and $12.00/min; Veo 3.1 eleventh at 1,082; and Wan 2.7 twelfth at Elo 1,077 with 4,571 votes and $9.00/min. A note above the table says "AA-Video-I2V v2.0 is coming soon".

Reference and consistency, where the two designs collide

Wan 2.7 controls subject consistency through supplied references: five subject images, five reference clips, ten assets total. It is a photographic approach — you show it what the subject looks like and it holds that appearance across the generation.

Griffin controls the same thing the opposite way. It generates every pixel in every frame from one reference image in real time, and the announcement is explicit that it controls the whole scene rather than the face: the arms and fingers, the movement of the chair, the shadows cast and the background behind. Consistency there is achieved by making the environment a generation target rather than a composited plate.

Neither approach is strictly better and they fail differently. Reference-conditioned generation fails by drifting away from the supplied subject over a long clip. Per-chunk generation with an autoregressive history fails by wandering over a long session if the drift handling is wrong, which is exactly the failure the third distillation stage exists to prevent. Wan 2.7 is the safer choice when the deliverable is a specific person in a specific look. Griffin is not competing for that brief.

Safety, provenance and the release posture

Wan 2.7 carries no published provenance system comparable to a watermark that survives compression — the announcement names audio quality and on-screen text accuracy as areas still needing work, which is a fidelity caveat rather than a safety one.

Griffin's safety story is inverted: Tavus's stated reason for keeping it in preview is that the same properties that make it a better interface make it better at persuading someone they are talking to a human, and the numbers support the concern. In the company's own live study, 26 of 54 participants believed their partner was a real person, against 1 of 41 on the previous Phoenix-4.5 / Raven-1 / Sparrow-2 stack. Participants who did suspect tended to suspect within the first twenty seconds. Tavus says further alignment and disclosure work is required before release, that it is building disclosure features, and that it is working with organisations on safety evaluations. That is the most substantive thing anyone has published about this model, and it is also why there is no product to buy.

For anyone comparing the two as tools, the practical consequence is simple: one is generatable on demand and shippable today, and the other is an evaluation target whose release is gated on a problem the vendor has not yet solved.

Which one, for what

• Choose Wan 2.7 if the deliverable is footage. Instruction-based editing of existing material is the workload where it is unusually strong and unusually cheap, because the editing variant bills on output only and treats the input footage as free. Its reference-to-video variant is worth checking against the 10 October retirement before you commit to it.

• Do not choose Griffin for anything this quarter, because you cannot. Request access if you need to know whether a person can tell, or if your roadmap depends on the interaction layer rather than the footage layer.

• Watch the gap, not either model. Griffin's arrival makes the more interesting comparison a future one: a system that generates the whole conversation against a suite that generates the whole scene. The first product that does both will be the one worth re-reading this against.

Where a router fits, and where it does not

Neither of these models is in the OrcaRouter catalogue, and that is worth stating before anything else in this section. Wan 2.7 is reachable through Alibaba's own platforms — Model Studio on Alibaba Cloud and Qwen Cloud — and through third-party creative tools that integrated it after launch; Tavus Griffin is reachable only by request form. If either becomes routable they will appear at provider list price.

The part of a video pipeline that is a routing problem is the text around it. Shot lists, prompt expansion, caption generation, quality review summaries, and the transcript or dialogue work that feeds a reference-to-video pass are all text calls, and those run on the same OpenAI-compatible endpoint at OrcaRouter as 200+ other models at provider list price with 0% markup added, so a vendor price cut is live the same day rather than after a contract cycle. Splitting a cheap model across a bulk pass and a frontier model across the final one is a config change there rather than a second vendor relationship — which is the same argument that applies to the generation layer, once the generation layer is something you can call.

What to watch: whether the Wan 2.7 suite's reference and editing variants survive the 10 October Model Studio retirement unchanged, and whether Griffin's safety gate produces a published model card with an endpoint on it. Those two events are what would turn this comparison from a design essay into a procurement decision.

A two-column scoreboard titled "Tavus Griffin vs Alibaba Wan 2.7 — the scoreboard". Left column "Tavus Griffin": Makes: a live conversation; Price: none published; Output: 720p in 320 ms chunks; Input: a person on camera; Clip length: none - a session; Score: VideoFDB 3.83 of 5. Right column "Alibaba Wan 2.7": Makes: 2 to 15 second clips; Price: $9.00 per minute; Output: up to 1080p; Input: text, image, video; Clip length: 2 to 15 seconds; Score: arena Elo 1,032. Footer: "Wan rates per Alibaba Cloud; Elo per Artificial Analysis, 2 October 2026. Griffin figures Tavus-reported."