A generated title card for Vidu Q4 Preview vs Google Veo 3.1, subtitled 'Three Veo tiers, one Vidu build', with chips reading 'Vidu Q4 Preview — Elo 1,179', 'Veo 3.1 — Elo 1,082' and 'Veo 3.1 Lite — Elo 1,071'. The OrcaRouter logo sits in the bottom-right padded strip.
Guides & Insights

Vidu Q4 Preview vs Google Veo 3.1: Veo Sells Three Tiers, and the Gap Between Them Is Smaller Than the Gap Up

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Goo​gle's Veo 3.1 comes in three tiers. On the Artificial Analysis image-to-video board they score 1,082, 1,071 and 1,066 — sixteen Elo across the range — while their per-minute prices span five times, from $4.80 to $24.00. Vidu Q4 Preview is one model, not three, scores 1,179 on the same board, and lists at $7.20 per minute. Read those numbers in order and the practical question stops being "which Veo tier" and becomes "why buy any of them" — which is not the comparison most pages write, and it deserves the specific answer below rather than a shrug.

Vidu Q4 Preview shipped on 7 October 2026 from Shengshu Technology as a first public preview of its next-generation flagship, ahead of the full Q4 release, in two modes: Image-to-Video and Reference-to-Video. Veo 3.1 has been on the boards since January 2026 with a Lite tier added in March, and it takes text prompts as well as images. Different shapes, different vintages, and one board where they can be measured against each other.

A generated two-column scoreboard for Vidu Q4 Preview and Google Veo 3.1. Left column rows read 'Image-to-video Elo: 1,179', 'Tiers: one build', 'Text-to-video: none', 'Measured rate: $7.20 per minute', 'Provenance: none stated' and 'Max resolution: 4K'; right column rows read 'Image-to-video Elo: 1,082 / 1,071 / 1,066', 'Tiers: three', 'Text-to-video: yes', 'Measured rate: $24.00 / $9.00 / $4.80 per minute', 'Provenance: SynthID plus C2PA' and 'Max resolution: not stated'. A footer reads 'Elo and rates per Artificial Analysis, 8 October 2026; provenance per Google DeepMind.' The OrcaRouter logo sits in the bottom-right padded strip.

Veo's tier ladder, read off the board

All figures below are Artificial Analysis, read 8 October 2026, with audio kept in. Each board has its own Elo scale and its own vote pool, so the two blocks are not comparable to each other — only within themselves.

Image-to-video (AA-Video-I2V v1.0):

• Vidu Q4 Preview — 3rd, 1,179 ±10, 5,543 samples, Oct 2026, $7.20/min

• Veo 3.1 — 12th, 1,082 ±7, 7,040 samples, Jan 2026, $24.00/min

• Veo 3.1 Lite — 15th, 1,071 ±8, 5,122 samples, Mar 2026, $4.80/min

• Veo 3.1 Fast — 18th, 1,066 ±6, 14,688 samples, Jan 2026, $9.00/min

Text-to-video (AA-Video-T2V v2.0):

• Veo 3.1 — 19th-21st, 962 ±10, 2,998 samples, Jan 2026, $24.00/min

• Veo 3.1 Fast — 20th, 961 ±10, 4,325 samples, Jan 2026, $7.20/min

• Veo 3.1 Lite — 22nd-25th, 948 ±10, 2,903 samples, Mar 2026, $4.80/min

• Vidu Q4 Preview — not present on this board

The two blocks tell the same story from different angles. On image-to-video, the three Veo tiers sit inside a sixteen-point band while the price varies by a factor of five — so paying five times more for Veo 3.1 over Veo 3.1 Lite buys you eleven Elo, well inside the ±7 and ±8 intervals those two scores carry. On text-to-video the spread is fourteen points across the same five-times price range, and there the ordering is generous to the cheap tier: Veo 3.1 Fast is one Elo behind the flagship for a third of the price, with more votes behind its number.

Goo​gle's own tiering is therefore mostly a latency and resolution decision rather than a quality ladder, which is a reasonable thing for a vendor to sell. It is just not the thing most "which Veo" articles claim it is.

The Artificial Analysis image-to-video board showing all three Veo 3.1 tiers and Vidu Q4 Preview on one scale.

The comparison that does hold

Vidu Q4 Preview is 97 Elo above Veo 3.1's best tier on the image-to-video board, at $16.80 per minute less, and 113 Elo above Veo 3.1 Fast. Those margins are five to seven times the width of the intervals involved, so they are not noise.

The obvious objection is that this compares an October 2026 model against a January 2026 one, and it does — that is the entire finding. Veo 3.1 is eight months old on this board, its tiers were last meaningfully reshuffled when Lite arrived in March, and it has been overtaken by a preview build from a smaller lab at a lower price. The board's own notice says two models were added in the last 30 days, and Vidu Q4 Preview is one of them.

The equally obvious counterweight: Vidu Q4 Preview has no text-to-video entry at all, which is where two of Veo's three tiers live. If your workload starts from a written prompt rather than a frame, none of these comparisons apply and the Veo ladder is the only one of the two that answers.

The Artificial Analysis text-to-video board, on which Vidu Q4 Preview has no entry.

What Veo gives you that Vidu does not

There is one dimension where the gap is structural rather than numeric, and it has nothing to do with Elo.

Goo​gle's Veo 3.1 output is watermarked with SynthID on every frame and carries C2PA content credentials recording that it was AI-generated, with a detector available in the Gem​ini app. That is not a feature you switch on; it is what the model produces. For anyone shipping into a market with disclosure requirements, working with a platform that screens for synthetic media, or simply wanting an audit trail of which shots were generated, this is a procurement difference that no benchmark will show you. It is also, on some workflows, a reason not to use it — a watermark you cannot disable is a constraint if your deliverable needs to pass as unmarked footage.

Vidu Q4 Preview publishes no watermarking or provenance mechanism of comparable specificity. That is an absence of a claim, not a claim of absence, and it should be recorded as such.

Two smaller asymmetries worth holding:

• Audio default — Vidu Q4 Preview's image-to-video endpoint has an audio parameter defaulting to true, so you get sound with the picture unless you turn it off; Veo 3.1 prices audio as its own tier decision, with roughly $0.40/s at Standard including it against substantially less for silent output

• Reference budget — Vidu Q4 Preview takes up to 15 reference images and up to 3 reference audio clips in Reference-to-Video; Veo 3.1's reference handling is not documented to that level of detail

Where this leaves the choice

If you are generating from a still — a product frame, a character reference, a storyboard panel — the independent board puts Vidu Q4 Preview in a different tier from every Veo 3.1 variant at a fraction of the top tier's price, and the tier ladder inside Veo does not rescue it. That is the finding, and it is grounded in 5,543 votes against 7,040, 5,122 and 14,688 respectively. The vote counts are real; the rank will move as intervals narrow.

If you need provenance built in, text-to-video, or a vendor with Goo​gle's distribution and enterprise terms, Veo 3.1 is the answer for reasons the Elo table cannot express and should not be read as overriding.

If you need both, the arithmetic is one integration against two. OrcaRouter runs one OpenAI-compatible endpoint across 200-plus models with provider list prices passed through and no markup added, so a rate change at either vendor shows up in your invoice the same day rather than at renewal, and automatic failover across providers means a single route going unavailable does not take the request with it. Neither Vidu Q4 Preview nor Veo 3.1 is an OrcaRouter route today — the video models we do serve, including MiniMax H3, are callable on that same endpoint — so the honest framing is that we can hold the orchestration layer steady while you decide which render path wins, not that we host either of these two.

The question that actually matters

Veo 3.1 is not a bad model. It is a January 2026 model that has been priced like a flagship while three of its own tiers converge on the same score, and it has now been measured against an October 2026 preview that beats its best variant by a margin wider than Veo's entire internal spread.

What would change the picture: AA-Video-I2V v2.0 is announced as coming, aligned with the T2V v2.0 taxonomy and covering 1080p throughout. A re-run on a new board with new vote pools can reshuffle any of these places, upward or downward, and both models should be re-read on that board rather than trusted to the numbers on this page.

The narrow answer for a reader who wants one: on the only board where both appear, the single Vidu Q4 Preview build outranks all three Veo 3.1 tiers, and it does so at $7.20 per minute against Veo 3.1 Standard's $24.00 and Veo 3.1 Fast's $9.00 — though not against Veo 3.1 Lite, which is the one Veo tier that undercuts it, at $4.80 per minute and eighteen places lower on the same board. Veo's own tier ladder is not a quality ladder, and that is the part of this comparison worth carrying forward. The wide answer is that they are aimed at different jobs, and the deliverable that decides it is whether your input is a photograph or a sentence.