A generated title card for Vidu Q4 Preview vs Alibaba Wan 2.7, subtitled 'Two modes against four variants, and only one board they share', with chips reading 'Vidu Q4 Preview — #3, Elo 1,179', 'Alibaba Wan 2.7 — #13, Elo 1,077' and 'Gap on image-to-video: 102 Elo'. The OrcaRouter logo sits in the bottom-right padded strip.
Guides & Insights

Vidu Q4 Preview vs Alibaba Wan 2.7: Two Modes Against Four, and One Board Where It Counts

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Here is the whole comparison in one line: Vidu Q4 Preview beats Ali​baba Wan 2.7 by 102 Elo on the Artificial Analysis image-to-video board, and cannot be entered against it on text-to-video or video editing, because Vidu Q4 Preview does not do either.

Both models launched in 2026 and both generate video with sound. Vidu Q4 Preview arrived on 7 October 2026 from Shengshu Technology as a first public preview ahead of the full Q4 model, in two modes — Image-to-Video and Reference-to-Video. Ali​baba Wan 2.7 shipped in April 2026 as a four-variant suite: text-to-video, image-to-video, reference-to-video and video edit, billed per second of generated video. So the tie-breaker in most head-to-head pages — a score — only exists on one of the axes, and the honest way to judge these two is to start by asking which axes you actually need.

The scoreboard, with the gaps left visible

Artificial Analysis runs three separate video leaderboards. Each carries its own Elo scale and its own vote pool, so a score from one does not transfer to another. Read on 8 October 2026, again with audio kept in:

• Image-to-video (AA-Video-I2V v1.0) — Vidu Q4 Preview 3rd, 1,179 ±10, 5,543 samples, Oct 2026, $7.20/min; Ali​baba Wan 2.7 13th, 1,077 ±8, 4,571 samples, Apr 2026, $9.00/min

• Text-to-video (AA-Video-T2V v2.0) — Vidu Q4 Preview not present; Wan2.7-260612 13th-14th, 1,030 ±9, 5,571 samples, Jun 2026, $9.00/min

• Video editing (AA-Video-Editing v1.1) — Vidu Q4 Preview not present; Wan 2.7 7th, 1,074, 20,021 samples, Apr 2026, $16.90/min

Two things follow from that table. Wan 2.7 is the more versatile of the two by a wide margin — it is the only one of them scored on all three boards, and its editing entry carries the largest vote count of any Wan 2.7 score anywhere, twenty thousand samples deep. And on the single axis where they meet, Vidu Q4 Preview wins by 102 Elo, which is roughly five times the width of either model's confidence interval.

The 102-point gap is not a small result. For scale: the whole spread across the top six places on the image-to-video board is 21 Elo. Vidu Q4 Preview is not marginally ahead of Wan 2.7 on image-to-video; it is a different tier, and it costs $1.80 per minute less while also being the more recent release. If image-to-video is your workload, the decision is not close and the independent board is the reason it is not close.

A generated two-column scoreboard for Vidu Q4 Preview and Alibaba Wan 2.7. Left column rows read 'I2V rank: #3, Elo 1,179', 'Modes: image, reference', 'Max clip: 16 seconds', 'Max resolution: 4K', 'Text-to-video: none' and 'Measured rate: $7.20 per minute'; right column rows read 'I2V rank: #13, Elo 1,077', 'Modes: four variants', 'Max clip: 15 seconds', 'Max resolution: 1080p', 'Text-to-video: yes' and 'Measured rate: $9.00 per minute'. A footer reads 'Ranks, Elo and rates per Artificial Analysis, 8 October 2026; modes, clip length and resolution are vendor-stated.' The OrcaRouter logo sits in the bottom-right padded strip.

Where Wan 2.7 is the only one of the two that shows up

Wan 2.7's own documentation names two areas as still needing work: audio quality, and accuracy of on-screen text. That is the vendor's assessment, not a reviewer's, and it is worth carrying into any comparison, because both models make the same headline claim — native audio generated with the picture rather than bolted on afterwards.

The difference is in what happens by default. Vidu Q4 Preview's image-to-video endpoint carries an audio parameter that defaults to true: you get sound unless you turn it off, and turning it off is how you get a silent clip. Wan 2.7 does not document the same switch in the same place. If silent output matters to your pipeline — and for anything you plan to score separately, or d​ub into another language, it does — that asymmetry changes what your first integration pass looks like.

On the editing board, Wan 2.7's seventh place at 1,074 over 20,021 samples is the strongest single result either model holds on any board by vote count. Twenty thousand votes is a measurement with real weight behind it, and no Vidu model appears on that board to be compared against it. If your work is re-timing or restyling existing footage rather than generating new shots, Wan 2.7 is the only one of the two that applies.

Alibaba Wan 2.7's Video Editing board entry, the single strongest result either model holds by vote count.

The specification gap on the generation side

Where they overlap, the two differ on almost every dimension, and the differences are not symmetrical in Wan's favour.

• Maximum resolution — Vidu Q4 Preview 4K (2K and 4K at 10-bit colour) vs Wan 2.7 up to 1080p

• Maximum clip length — Vidu Q4 Preview 16s (Image-to-Video 3-16s, Reference-to-Video 1-16s) vs Wan 2.7 up to 15s

• Reference images — Vidu Q4 Preview up to 15 in one setup vs Wan 2.7 up to ten assets

• Reference audio — Vidu Q4 Preview up to 3 clips for voice consistency vs not published to the same detail by Ali​baba

• Task coverage — Vidu Q4 Preview two modes (image-to-video, reference-to-video) vs Wan 2.7 four variants (adding text-to-video and video edit)

• Published list rate — Vidu Q4 Preview from $0.014/s promotional; measured $7.20/min on the independent board vs Wan 2.7 $0.10-$0.15/s by variant, $9.00/min measured

The resolution line is the one to weigh carefully. 4K output at 10-bit is a genuine deliverable difference for anyone finishing for a screen or a large display, and it is not a feature Wan 2.7 offers on any variant. But 4K is also where per-second pricing usually stops looking like the floor advertised, and Vidu's own launch material carries the caveat that final pricing, supported resolutions, feature availability and usage terms vary by plan and region.

Vidu Q4 Preview's product page open on Reference-to-Video, showing the reference-image and audio upload surface.

Which one, by workload

If you are animating stills — product shots, character reference frames, storyboards — Vidu Q4 Preview is the better model on the only independent evidence available, at a lower measured rate. That is the straightforward case.

If you need one contract that covers text-to-video, image-to-video and editing, Wan 2.7 is the only one of the two that answers, and it does so with a seventh-place editing result built on the largest vote count either model has. Consolidating on one vendor is a real benefit and this comparison does not pretend otherwise.

If you need both, the cost of running two vendors is usually not the licence — it is the second integration, the second key, the second set of retry semantics, and the second schema to keep current. OrcaRouter collapses that: one OpenAI-compatible endpoint across 200-plus models, provider list prices passed through with no markup added, and automatic failover across providers. On the video side specifically we serve MiniMax H3, which is the model currently ranked first and second on the image-to-video board above — callable on the same endpoint and the same key as the text models. Neither Vidu Q4 Preview nor Ali​baba Wan 2.7 is an OrcaRouter route today, and this page should not be read as implying otherwise; what we can do is keep the rest of the workload on one integration so the video experiment is a line item rather than a project.

What would change this

Vidu Q4 Preview is a preview. The pricing is promotional, the build sits ahead of the full Q4 release by the vendor's own account, and the API documentation still carries a mistyped model alias alongside the correct one. Two outcomes matter: if the full Q4 release adds text-to-video, the axes collapse and this becomes a direct comparison on three boards instead of one; if it does not, the split stays as it is and the choice is purely about workload shape.

On the other side, Wan 2.7's samples on the image-to-video board are five months older than Vidu's and about a thousand fewer. Both intervals will narrow. Neither model's rank should be treated as fixed beyond the date printed beside it.

The narrow answer: they share exactly one competitive axis, Vidu Q4 Preview leads it clearly at a lower measured rate, and Wan 2.7 wins by default everywhere Vidu does not compete. The wide answer is that a two-mode preview model and a four-variant suite are not really alternatives to each other — they are alternatives to different jobs, and the useful part of this comparison is knowing which job you have.