
Vidu Q4 Preview vs Seedance 2.5: The Reference Budget Is Not the Real Difference
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 62 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 358 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The published reference budgets read like a straight win for Seedance 2.5: up to 30 images, 10 videos and 10 audio clips against Vidu Q4 Preview's 15 images and 3 audio clips. That is twice the images and three times the audio, and on its own it would settle the question. It does not, because the number that matters is not the count — it is the second modality. Seedance 2.5 accepts video as an input reference. Vidu Q4 Preview accepts images and audio, and has no documented path to take an existing clip in.
Seedance 2.5 launched 31 July 2026 as ByteDance's long-form entry, generating a 30-second clip in a single run with multi-turn extension for longer sequences. Vidu Q4 Preview launched 7 October 2026 from Shengshu Technology as a preview of its next-generation flagship, with two modes — Image-to-Video and Reference-to-Video — and two documented API endpoints. Both are reference-heavy models. They were built to hold different things.
What each input modality actually unlocks
A model that takes images as references can hold a cast and a look. Vidu Q4 Preview's documentation describes combining 1–15 images across characters, scenes and style so that subjects stay consistent across a multi-shot take, with up to 3 mp3 audio references of 3–12 seconds cloning a voice and keeping delivery aligned. Fifteen slots across a single scene's cast, wardrobe, props and lighting is a coherent budget, and it is the reason the model holds a tie at the top of the image-to-video preference board.
A model that also takes video as a reference can be pointed at footage. Seedance 2.5 accepting up to 10 video inputs alongside 30 images and 10 audio clips means the reference set can include motion, pacing and camera behaviour drawn from an existing clip rather than described in a prompt. That is a different capability, not a larger version of the same one, and it is what puts Seedance 2.5 on the video-editing board at Artificial Analysis — second, at 1,161 Elo over 5,162 votes as read on 8 October 2026 — editing a video from a text instruction with the audio kept. Vidu Q4 Preview is not on that board, and the reason is its endpoint list, not its score.
The practical contrast:
• Image references — Vidu Q4 Preview up to 15 vs Seedance 2.5 up to 30
• Video references — Vidu Q4 Preview none documented vs Seedance 2.5 up to 10
• Audio references — Vidu Q4 Preview up to 3 mp3 clips of 3–12 seconds vs Seedance 2.5 up to 10
• Single-run length — Vidu Q4 Preview 3–16 seconds on Image-to-Video, 1–16 seconds on Reference-to-Video vs Seedance 2.5 30 seconds in one run, with multi-turn extension beyond that
• Modes — Vidu Q4 Preview Image-to-Video and Reference-to-Video vs Seedance 2.5 text-to-video, reference-to-video and video-in reference, with an editing route on the board
• Output resolution — Vidu Q4 Preview 540p through 4K as priced generation tiers, 2K and 4K at 10-bit colour vs Seedance 2.5 480p and 720p on the hosts that publish its settings, with higher tiers reaching delivery resolution through an upscale
Put those side by side and the models are not competing for the same brief. A fifteen-reference 4K sixteen-second take and a thirty-reference 720p thirty-second take with video input are answers to two different production questions.
The pricing units do not convert, and that is the whole story
This is where most comparisons of these two go wrong, and the fault is in the unit rather than the arithmetic.
Seedance 2.5 is not sold by the second by its vendor. ByteDance meters the model in video tokens on the official rate card, published through Volcano Engine Ark in China and BytePlus ModelArk internationally. What a clip costs depends on resolution, duration and actual token usage, so a per-second figure is only valid for one resolution and one duration — and failed takes bill the same as successful ones. Third-party hosts that expose Seedance 2.5 add their own margin on top and publish their own per-second or per-clip rate, which is why a dozen pages give a dozen different numbers and none of them is the vendor's.
Vidu Q4 Preview is priced the other way: a flat per-second credit rate that varies only with the resolution tier you select, published openly by Shengshu. Nine credits per second at 540p, 19 at 720p, 24 at 1080p, 38 at 2K, 78 at 4K, against a stated platform credit rate of $0.005, with limited-time bundle discounts of up to 30% currently running. At list rate the curve runs from roughly $0.045 per second at 540p to about $0.39 at 4K. The $0.014 per second figure from the launch announcement is the 540p tier with the discount applied.
Which means the two rate cards cannot be compared line by line, and any page that does it is comparing a vendor's open list to a host's markup. What can be compared is what the vendor publishes about the shape of each: Vidu's cost scales with resolution and duration and is knowable before you press generate; Seedance's is token-metered and therefore depends on how the model chooses to spend tokens on your prompt. One is predictable, the other is metered. Neither is wrong, but they suit different buyers — and the second is the more dangerous of the two to budget against, because the variance is on the vendor's side of the line.
On the independent board, the recorded per-minute figures are worth reading as an order of magnitude only. Artificial Analysis records Vidu Q4 Preview at $7.20 per minute on the image-to-video board, and Dreamina Seedance 2.5 at $34.12 per minute on text-to-video and $40.82 on editing. Those columns are the cost of one minute of 1080p output at each model's default settings — and Seedance 2.5 defaults to the longer, heavier configuration its rate card prices for, while Vidu Q4 Preview's default is a single-generation clip. They measure different default behaviours, not a four-to-five-fold efficiency gap.
Sixteen seconds against thirty is a real difference
There is one comparison that does survive the unit problem, and it is length. Thirty seconds in a single run is nearly twice Vidu Q4 Preview's sixteen-second ceiling, and Seedance 2.5's multi-turn extension means a sequence can be built past that. For narrative work — a scene with an arc, an ad with a setup and a payoff — the difference between sixteen and thirty seconds is the difference between a shot and a scene.
It cuts the other way at the fine end. Vidu Q4 Preview generates from three seconds, and its 540p tier is priced at roughly a third of its own 720p rate, which makes blocking a composition cheap before committing to a full-resolution render. Nothing in the published Seedance 2.5 rate structure plays that role: its 480p host tier costs less than 720p, but the vendor's own billing is token-based, so a short low-resolution test does not have a clean published price to plan against.
Which one, by brief
Take Seedance 2.5 when the job involves footage you already have. If the work is extending, restyling or re-cutting an existing clip, or if the reference set needs to carry motion rather than only appearance, the video input is the deciding capability and Vidu Q4 Preview cannot be substituted in. Add the thirty-second single run, and the case is straightforward for narrative and advertising work with a longer arc.
Take Vidu Q4 Preview when the job is a cast that has to hold and a delivery specification that goes past 720p. Native 2K and 4K generation at 10-bit colour is a tier Seedance 2.5 does not publish at the vendor level, and the per-second price makes a 4K take a quantity you can actually budget rather than a metered unknown. Fifteen image references and three voice references are enough for a single-scene cast, and the 540p blocking tier makes iteration cheap.
What would change this: a full Vidu Q4 release that adds a video input modality would collapse the structural difference and make these two direct competitors, and Shengshu's own material says the preview exists precisely to gather the feedback that shapes the final build. On the other side, if ByteDance publishes a token-rate example table at more settings, the metered-versus-predictable asymmetry gets much easier to plan around. Until either lands, the correct reading of "30 references versus 15" is that one of these models takes video in and the other does not.
One key for the video models you can actually call
OrcaRouter puts video and language models behind a single OpenAI-compatible endpoint on one key, at provider list prices passed through with no markup added, so a vendor rate change lands the same day rather than at renewal. Vidu Q4 Preview and Seedance 2.5 are not currently OrcaRouter routes, and this page is not suggesting otherwise. The video routes we do serve — Kling 3.0 and its Turbo and v2.6 variants among them — sit beside the language models on the same endpoint and the same request shape, with automatic failover between routes rather than a separate contract per model.
That is the structure that makes a comparison like this one finishable: run the candidates on your own brief, at your own settings, and let the acceptance rate decide rather than two rate cards in different units.



