
Vidu Q4 Preview Debuts Third on the Image-to-Video Board — and Is Absent From the Other One
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 62 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 358 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Vidu Q4 Preview entered the Artificial Analysis AA-Video-I2V v1.0 leaderboard at number three with an Elo of 1,179, and the model it replaced as Vidu's best-placed entry, Vidu Q3 Pro, sits at nineteenth with 1,056. That is a 123-point move in one release, on a board where the confidence intervals are roughly twenty points wide, and it happened at a price of $7.20 per minute of generated video rather than the $9.60 the older model carries. Shengshu Technology shipped Vidu Q4 Preview on 7 October 2026 as the first public preview of its next-generation flagship, explicitly ahead of the full Q4 model, and is asking creators to use it in real projects while the final release is still being shaped.
The debut figure is the headline. The absence is the more interesting fact, and it is the reason this is a piece rather than a chart.
What actually shipped on 7 October
Vidu Q4 Preview is a video generation model with native audio, sold through Vidu's own web product and API platform. Shengshu's launch material describes three areas of work: character performance (facial expression, emotion, body movement and voice coordinated rather than layered), camera and cut coherence in fast action, and visual effects — explosions, particles, fireworks — that sit inside a scene instead of on top of it.
The specification, as the vendor states it:
• Maximum resolution — 4K, with 2K and 4K carrying 10-bit colour depth; the full ladder is 540p, 720p, 1080p, 2K, 4K
• Maximum clip length — 16 seconds, split unevenly: Image-to-Video runs 3 to 16 seconds, Reference-to-Video runs 1 to 16
• Reference budget — up to 15 reference images (characters, wardrobe, props, products, environments in one setup) and up to 3 reference audio clips
• Launch pricing — from $0.014 per second, which is the vendor's figure and is explicitly promotional
The vendor also makes a comparative claim it cannot easily be held to: that under comparable output specifications and billing conditions, Vidu Q4 Preview delivers "up to five times as much output for the same budget." That is an efficiency claim, not a quality claim, and because the phrase "comparable output specifications and billing conditions" does the entire work in that sentence, it is worth treating as a marketing construction until someone reproduces it. Shengshu's own caveat, published alongside, is that final pricing, supported resolutions, feature availability and usage terms may vary by plan and region.

The two boards disagree about what this model is for
Artificial Analysis runs separate video leaderboards with separate Elo scales and separate vote pools, and the difference between them is the whole story here.
On AA-Video-I2V v1.0 — image-to-video, ranked from pairwise human preference votes with audio kept in — Vidu Q4 Preview is third at 1,179 ±10 over 5,543 samples, dated October 2026, priced at $7.20 per minute. Only two models are above it: MiniMax H3 Max at 1,195 ±9 over 5,894 samples ($4.80/min, August 2026) and MiniMax H3 at 1,181 ±8 over 7,284 samples ($7.80/min, July 2026). Gemini Omni Flash is fourth at 1,178 ±7 over 12,327 samples. The board's own notice says two models were added in the last 30 days: Vidu Q4 Preview and HiDream-O1-Video-1.0, which is sixth at 1,175 ±10.
On AA-Video-T2V v2.0 — text-to-video, a different board built on a different use-case taxonomy — Vidu Q4 Preview does not appear at all. The best Vidu entry there is Vidu Q3 Pro, twenty-fifth at 941 ±10 over 4,088 samples. Wan 3.0 leads that board at 1,156 ±9.
Read those two sentences together and the shape of the release becomes clear. This is an image-and-reference model. Vidu's own product page lists exactly two modes for it, Image-to-Video and Reference-to-Video, with no text-to-video tab, and the API documentation matches. A launch that looks like a general flagship on the vendor's marketing is, on the independent boards, a targeted one.
One further caveat on the rank itself. The I2V board's ten-point intervals overlap heavily from third through sixth — 1,179, 1,178, 1,176, 1,175 — so "third" is a position within noise, not a clear separation from fourth. What is not within noise is the gap to the model it replaced. 123 Elo between Vidu Q3 Pro and Vidu Q4 Preview is larger than the entire spread across the board's top six places.

What you actually call, and what it costs
The API is unglamorous and specific. Image-to-Video posts to /ent/v2/img2video with the model name viduq4-preview, takes a single start-frame image (png, jpeg, jpg or webp, up to 50MB), a prompt of up to 20,000 characters, a duration from 3 to 16 seconds defaulting to 5, and a resolution from 540p to 4K defaulting to 720p. Reference-to-Video posts to /ent/v2/reference2video, takes one to fifteen images plus zero to three mp3 audio references of three to twelve seconds each, and offers aspect ratios of 1:1, 9:16, 16:9, 3:4 and 4:3 defaulting to 16:9.
The parameter worth noticing is audio, which defaults to true. Vidu Q4 Preview generates sound with the picture by default — dialogue and sound effects included — and you switch it off deliberately. Returned creation URLs stay valid for 24 hours, which is short enough to matter if you are generating in a batch and rendering later.
On price, the honest arithmetic is that the two numbers floated — "$0.014 per second" and the $7.20 per minute that Artificial Analysis records — are not the same product. $0.014 per second is $0.84 per minute, roughly a ninth of the board's figure, which is what a promotional floor at the lowest resolution looks like next to a measured rate at a production one. Anything you budget from the launch number should be re-derived from your own resolution and duration, not from the headline.

The practical problem with a preview SKU
Vidu Q4 Preview is, by the vendor's own framing, not finished. It shipped ahead of Q4 so that feedback on performance, creative control and day-to-day production needs would inform the final release. Pricing is promotional. Feature availability varies by plan and region. The API documentation even carries a mistyped alias — viduq4-previerw appears in the model parameter description alongside the correct viduq4-preview — which is a small thing that tells you something true about where this build sits in the pipeline.
That is not an argument against using it. It is an argument against putting it on a critical path without a plan for what happens when the promotional price ends or the preview build is superseded, which for a model released in this way is a question of weeks, not quarters.
If you want to try it without betting a production pipeline on it, OrcaRouter is built for exactly this case: one OpenAI-compatible endpoint across 200-plus models, automatic failover across providers, and provider list prices passed through with no markup added. The failover half matters more than usual with a preview build, because it is what lets you route a request to a stable model on the same key when a preview route is unavailable. To be plain about the routing: Vidu Q4 Preview is not an OrcaRouter route today. The video models we do serve — including MiniMax H3, the top-ranked model on this very board — are callable on the same endpoint as the text models, and a per-second video rate is a line item, not a separate contract.
What to watch
Three things would move this story.
First, the text-to-video board. Artificial Analysis has said AA-Video-I2V v2.0 is coming, aligned with the T2V v2.0 taxonomy and covering 1080p video throughout. If Vidu Q4 Preview is an image-and-reference model, it will not appear there either — and if it does, that tells you Shengshu has shipped a wider model than the preview implies.
Second, the full Q4 release. Preview-to-final is where promotional pricing typically ends and where the feature set either holds or quietly narrows. The 15-image and 3-audio reference budget is the part of the specification most likely to survive, because it is the differentiator the whole launch is built on.
Third, sample count. 5,543 votes is a real measurement but a young one; the intervals will narrow and the rank will move, in either direction, as the board fills. Anyone quoting "third" as a settled fact should say the date they read it.
The question worth holding onto is narrower than the marketing suggests. Vidu Q4 Preview is measurably the best image-to-video model Vidu has produced, by a margin larger than the board's own top-six spread, at a lower per-minute rate than its predecessor. Whether that becomes a durable position or a preview bump is not something any of the current numbers can answer.
