
Kandinsky 6.0 Video vs Seedance 2.5: The Version in the Paper Is Not the Version You Buy
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 150 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 126 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 250 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The most quoted sentence in this matchup is a comparison against the wrong model. The Kandinsky 6.0 Video technical report, published on 6 October 2026 with the MIT release of the weights, benchmarks against Dreamina Seedance 2.0 — the previous generation. Seedance 2.5, ByteDance's current video-and-audio model, has been shipping since 31 July 2026 and is the one on sale. So when the report concludes that Seedance "is preferred on the visual criteria, with statistically significant margins on overall visual quality, artifacts, motion realism, and visual prompt following," it is describing a build you cannot license today, evaluated by the lab that wrote Kandinsky 6.0 Video. Both halves of that sentence matter, and neither is a reason to ignore the comparison — the version gap sets a floor on how much it can tell you, and the framing tells you who to trust for what.
What changed between the two Seedance versions
The 2.0-to-2.5 step is not a repaint, which is exactly why the substitution is not cosmetic. Seedance 2.5 generates up to thirty seconds in a single pass — six times the envelope of the model in the report, and six times Kandinsky 6.0 Video's fixed five. It adds timestamp-level targeting and region-level post-generation edits, meaning a change can be applied to a window of the clip or a portion of the frame rather than by regenerating the whole thing. It reads text alongside roughly thirty images, ten videos and ten audio clips as reference material in one request, against Kandinsky 6.0 Video's single masked tail frame. And it produces synchronized audio with dialogue, which is the only reason this pair belongs in the same article at all.
Every one of those is a capability the report's evaluation did not test. The visual-criteria result therefore tells you that Kandinsky 6.0 Video trailed Seedance's previous generation on overall visual quality, artefacts, motion realism and prompt following. It does not tell you what the current build would do, and the honest assumption is that a newer generation does not get worse on the axes its own release notes emphasise.
What the report's own evaluation does and does not establish
The method is disclosed thoroughly enough to be weighed. Side-by-side pairs generated from the same prompt by two models, both clips anonymised, left/right positions randomised per task, task order shuffled, audio first disabled for the visual questions and then enabled for the audio ones, a symmetric tie option and an explicit N/A option where a criterion cannot be judged, majority-vote aggregation across annotators, and approximately 200 or more pairwise comparisons for each evaluated criterion.
On that method, the Seedance 2.0 result splits in a way that will look familiar to anyone who has read the same report's Kling comparison. The visual criteria go to Seedance with margins that clear significance. The audio domain is much closer: audio-video synchronization lands near parity, and speech quality favours Kandinsky 6.0 Video Pro. Between those two statements sits the whole decision, because it means the newest open audio-video checkpoint is measurably behind on how a clip looks and measurably ahead on how the speech in it sounds.
The limits on that result are the ordinary ones and none of them are fatal to it. It was run by Kandinsky's authors on prompts of their choosing with annotators they recruited. No public blind arena scores Kandinsky 6.0 Video at all, so nothing external has tested whether a stranger's prompts reproduce the split. The report also discloses the ceiling it is measuring against: clips of up to five seconds, base generation at SD resolution with Full HD reached through a super-resolution stage, and a gap remaining to the strongest proprietary systems in visual quality and overall audio quality.
Three structural gaps that no benchmark would capture
Duration is the first and bluntest. Kandinsky 6.0 Video is trained and evaluated at 121 frames, runs inference at 121 frames, 5 seconds at 24 fps, and ships with no extension mode — a thirty-second piece is six generations, five joins, and continuity risk you own. Seedance 2.5 does thirty seconds in one pass. For a single dialogue exchange, a product walkthrough, or anything where the join would be visible, that is not a benchmark difference; it is a difference in whether the job is one render or an assembly step.
The second is reference conditioning, and the asymmetry is stark. Kandinsky 6.0 Video takes one reference image and applies it as a masked tail frame — a keyframe to interpolate toward. Seedance 2.5 takes a working set of roughly thirty images, ten video clips and ten audio clips as one context. Character consistency across a sequence, style transfer from a mood board, a performance carried in from an existing clip: those are workloads the reference count enables and the single-frame mode does not attempt.
The third is editability, and it is the one that changes a workflow rather than a single shot. Region-level and timestamp-level editing means a flawed three-second window is repaired in place. Kandinsky 6.0 Video has no edit surface — a bad five seconds is regenerated, and if it was joined to four others, the joins are reconsidered too. Teams pricing a re-render habit should price that difference into the model choice rather than discovering it.
• Duration — Kandinsky 6.0 Video: 5 s fixed, 121 frames at 24 fps, no extension mode. Seedance 2.5: up to 30 s in a single pass.
• Reference conditioning — Kandinsky 6.0 Video: one image, applied as a masked tail frame. Seedance 2.5: roughly thirty images, ten videos and ten audio clips per request.
• Editing — Kandinsky 6.0 Video: none; regenerate and re-join. Seedance 2.5: timestamp-level targeting and region-level post-generation edits.
• Resolution — Kandinsky 6.0 Video: SD at 512×768, Full HD 1920×1080 through a built-in super-resolution stage. Seedance 2.5: mode-dependent, with the price per clip set by the resolution you choose.
• Audio — Kandinsky 6.0 Video: 44 kHz with lip-sync, generated jointly with the picture. Seedance 2.5: synchronized audio including dialogue, generated with the video.
• Weights — Kandinsky 6.0 Video: MIT, Lite 3B and Pro 29B, with code and diffusers integration. Seedance 2.5: no.
Two meters that do not convert into each other
Seedance 2.5 is metered in tokens, which lands at roughly $0.51 for a five-second clip at 480p and roughly $1.16 at 720p. The reason to state it that way rather than as a per-second rate is that the token count scales with the footage, so a thirty-second generation is not thirty seconds of a fixed rate — it is six times the tokens of a five-second one at the same settings, and the editing operations price on what they touch. A pipeline that habitually regenerates will find the meter responds to that habit.
Kandinsky 6.0 Video has no meter because there is nothing to buy. Its cost is a machine and a clock, and the report supplies the clock: roughly 402 seconds of working time on an H100 for a Pro Full HD five-second clip, 854 on an RTX 5090, 1,247 on an RTX 4090, block offloading required below a 72.8 GiB peak allocation, and a 16 GB VRAM preset that quantises the text encoder to NF4 — the one preset change the report says alters the clip itself. The distilled checkpoint, measured at parity with the full model at a mean preference of 51% to 49% with no criterion reaching statistical significance, is the one that makes those numbers workable.
• Cost per five seconds — Kandinsky 6.0 Video: no list price; six to twenty-one minutes of one GPU depending on the card. Seedance 2.5: roughly $0.51 at 480p and $1.16 at 720p in tokens.
Nothing in those two lines is commensurable, and the usual attempt to reduce them to a single figure — dividing a GPU-hour rate by five-second clips — silently assumes full utilisation, zero idle time and a card you already own. The more useful comparison is qualitative: Seedance 2.5's cost scales with how much you generate and vanishes when you stop; Kandinsky 6.0 Video's cost is incurred whether or not you do.

Where each one is reachable
Neither model is on OrcaRouter, and neither claim is being made. Seedance 2.5 is reached through the vendor's own platforms and several third-party services; Kandinsky 6.0 Video is a checkpoint you download under MIT and run on your own hardware. Any page that tells you both are one API call away on a multi-model platform is describing different models.
The reason a routing layer is still worth mentioning in a Kandinsky or Seedance pipeline is that these systems are mostly text by call count and video by cost. Prompt expansion turns a shot note into the long structured caption that a T2AV model expects. Dialogue is drafted before a lip-sync pass attempts it, and is rewritten when the pass mangles it. Six candidate generations of the same beat get reviewed and one selected — a text judgement about a video artefact, which is the cheapest way to make the expensive call correctly. On a single OpenAI-compatible endpoint across 200-plus models at provider list prices passed through with 0% markup, those calls cost what the provider charges and no more, including the day after a vendor changes its rates. The routing DSL earns its place when the cheap reviewer and the careful reviewer should be different models at the same call site, and automatic failover earns it because a caption that dies while clip four of six is rendering should reroute rather than end the session.
What would move this matchup
The version gap is the whole story and also the shortest path to resolving it. A Seedance 2.5 run against Kandinsky 6.0 Video would settle whether the visual-criteria losses the report records against 2.0 have widened, narrowed or reversed — the report cannot answer that, and its authors had no way to. For now, the defensible position is narrower than either side's marketing wants: on the evidence published by the model's own laboratory, Seedance is ahead on how the clip looks and Kandinsky 6.0 Video is ahead on how the speech in it sounds, and the version of Seedance that claim rests on is one generation behind the one you would buy.
Pick accordingly. If the work is long, reference-heavy or edit-driven, the structural gaps decide it before quality enters the room. If the work is a five-second talking shot where the dialogue has to sound right the first time, Kandinsky 6.0 Video's measured advantage is on exactly that criterion — and the MIT licence means the answer does not depend on anyone's roadmap. The thing to watch is not a leaderboard. It is a rerun of a comparison the report was never able to run.


Cheapest way to make the expensive call correctly: a routing DSL that swaps the reviewer per stage, with automatic failover so a dead caption does not cost the clip.
