
Vidu Q4 Preview vs Grok Imagine Video: The Model That Led the Arena in January Sits Ninth Now
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 62 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 358 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Three numbers describe this matchup, and none of them are a spec sheet. On the Artificial Analysis image-to-video board, Vidu Q4 Preview sits third at 1,179 with 5,543 votes behind it. Grok Imagine Video 1.5 sits ninth at 1,098. The January 2026 build of Grok Imagine Video, still listed and still callable under its own name, sits fourteenth at 1,072 — and the two Grok entries are priced at $8.40 and $4.20 per minute respectively, which means the newer one costs exactly twice as much for twenty-six more Elo.
Vidu Q4 Preview is one build, released 7 October 2026 by Shengshu Technology as a first public preview ahead of the full Q4 model, in two modes: Image-to-Video and Reference-to-Video. Grok Imagine Video is a family with at least three entries across two boards. The comparison is really about what happens to a leaderboard position over nine months, and the Vidu release is the thing that arrived to make the point.
Two builds of the same model, twenty-six Elo apart
Artificial Analysis, read 8 October 2026, audio kept in, image-to-video board (AA-Video-I2V v1.0):
• Vidu Q4 Preview — 3rd, 1,179 ±10, 5,543 samples, Oct 2026, $7.20/min
• Grok Imagine Video 1.5 — 9th, 1,098 ±8, 5,267 samples, May 2026, $8.40/min
• grok-imagine-video — 14th, 1,072 ±7, 14,371 samples, Jan 2026, $4.20/min
And on text-to-video (AA-Video-T2V v2.0), a separate board with its own scale:
• Grok Imagine Video 1.5 — 11th, 1,045 ±9, 4,056 samples, May 2026, $15.00/min
• Grok Imagine Video 1.5 Lite — 17th-18th, 993 ±10, 5,715 samples, Aug 2026, $8.40/min
• Vidu Q4 Preview — not present
Two things jump out of the first block. The January build of Grok Imagine Video carries the most votes of any model in this comparison — 14,371, nearly three times Vidu Q4 Preview's 5,543 and nearly three times the 1.5 build's. When a model launched as the arena leader in January, that is what the sample count looks like in October: the community voted on it extensively, and then stopped. The score is stable and well-measured and no longer near the top.
The second thing is the price inversion. Grok Imagine Video 1.5 is twice the per-minute rate of its January predecessor and twenty-six Elo ahead of it. Twenty-six Elo is just outside the two intervals combined — ±8 and ±7 — so the improvement is probably real, but it is the kind of margin that comes from a generational refresh rather than a new architecture, and it is being sold at double. Meanwhile Vidu Q4 Preview is eighty-one Elo above the 1.5 build and priced below it, at $7.20 against $8.40. That is a fresher model beating an older one on both axes, which is what is supposed to happen and rarely does this cleanly.
The third oddity sits between the boards. The same Grok Imagine Video 1.5 that is $8.40 per minute on image-to-video is $15.00 per minute on text-to-video, where it also scores lower by the other board's scale. If you are planning a Grok workload, the input modality you choose changes the rate by roughly eighty per cent.

What Grok Imagine Video actually gives you
It is not a bad model and this piece is not going to pretend otherwise. Two things it does that Vidu Q4 Preview does not:
It takes text prompts. Vidu Q4 Preview's product page lists exactly two modes — Image-to-Video and Reference-to-Video — and its API documents /ent/v2/img2video and /ent/v2/reference2video and nothing else. There is no text-to-video endpoint. Grok Imagine Video 1.5 and its Lite sibling both appear on the text-to-video board, so if your pipeline starts from a written prompt, one of these two models is available to you and the other simply is not.
It has a cheap tier. On the image-to-video board the January build is $4.20 per minute against Vidu Q4 Preview's $7.20 on that same board; the Lite tier's $8.40 sits on text-to-video, a separate board with its own scale. If you are generating high volume and can accept a fifth-place-tier score, the low end of the Grok family is the cheapest thing in this article — the January build at $4.20 carries 1,072 Elo, which is 107 below Vidu Q4 Preview but also $3.00 per minute cheaper.

What Vidu Q4 Preview gives you
The Vidu side of this is a stronger generation model on the axis they share, and a much larger reference budget.
• Resolution — Vidu Q4 Preview up to 4K at 10-bit colour (2K and 4K) vs Grok Imagine Video not documented at 4K on these boards
• Clip length — Vidu Q4 Preview 3-16s for Image-to-Video and 1-16s for Reference-to-Video vs Grok Imagine Video not published to the same detail
• Reference inputs — Vidu Q4 Preview up to 15 reference images and up to 3 reference audio clips vs not published to that level for Grok Imagine Video
• Audio default — Vidu Q4 Preview's image-to-video endpoint carries an audio parameter defaulting to true, so sound arrives with the picture unless you disable it
• Text-to-video — absent, and that is the one line on this list that goes the other way
Every Vidu figure above is the vendor's own and should be read as such. The one number in this article that is independently measured is the Elo, and on that Vidu Q4 Preview is eighty-one points clear of Grok Imagine Video 1.5 with a thousand fewer votes. Eighty-one Elo is roughly five times the combined interval width. It is a real gap today and it is dated October 2026, which means it is a gap about a different moment than the Grok scores sit in.

Is a preview build a fair opponent?
Shengshu's framing is unusually direct: Vidu Q4 Preview shipped ahead of the full Q4 model so that creators could use it in real projects and help shape the final release. Pricing is described as promotional. Final pricing, supported resolutions, feature availability and usage terms may vary by plan and region. The API documentation carries a mistyped alias, viduq4-previerw, sitting beside the correct viduq4-preview, which is the sort of detail a shipping product usually cleans up and a preview usually does not.
So the eighty-one Elo was earned by a build that its own vendor says is not final, at a price its own vendor says is promotional. Both of those cuts go in Vidu's favour — the model is likely to improve and the price is likely to rise — but they also mean neither number is a commitment. Anyone budgeting from this page should plan for the rate to move on the Q4 release, and should re-read the board rather than the rank when they do.
If the preview build turning over every few weeks is the problem, that is exactly the case automatic failover exists for: one OpenAI-compatible endpoint across 200-plus models, with provider list prices passed through and no markup added, so a promotional rate ending is a line item change rather than a renegotiation, and a route going unavailable falls through instead of failing the request. Being straight about coverage — Vidu Q4 Preview is not an OrcaRouter route, and neither is Grok Imagine Video. The video models we do serve, including MiniMax H3, sit on that same key and endpoint as the text models.
The narrow answer
If you want the best image-to-video output and your input is a frame, Vidu Q4 Preview is eighty-one Elo ahead of Grok Imagine Video's current build at a lower measured rate, and there is no strong argument for the Grok side beyond price floor. If you want the cheapest thing that scores respectably, Grok Imagine Video's January build at $4.20 with 1,072 Elo behind 14,371 votes is a well-measured, well-voted, deliberately cheap option and Vidu has nothing at that price point. If your input is a sentence, the Grok family is the only one of the two families in the room.
What would change it: the full Vidu Q4 release, which could add modes or narrow the price gap; and AA-Video-I2V v2.0, announced as coming with a new taxonomy and a fresh vote pool, which would reset every number on this page rather than adjust it. Until either lands, the position worth remembering is the one the sample counts tell you — Vidu's 5,543 votes are one day old, Grok's January build's 14,371 are nine months old, and a leaderboard is a photograph, not a standing.
