
Kandinsky 6.0 Video vs Google Veo 3.1: Three Tiers Against Two Sizes
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 150 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 126 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 250 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Open Sber's technical report for Kandinsky 6.0 Video and find the head-to-head section, and you will notice something the comparison pages miss: the opponent is not Veo 3.1. It is Veo 3.1 Fast. Sber ran its human evaluation against the middle of Veo's three tiers, and that single choice tells you more about this matchup than any quality claim in it. Veo 3.1 has been generally available since 17 November 2025 in three service tiers; Kandinsky 6.0 Video arrived in the first week of October 2026 in two downloadable sizes, Pro at 29B and Lite at 3B, under MIT. One of these is a service you buy by the second, the other is a checkpoint you run — and the tier question, not the quality question, is where the decision actually lives.
Veo 3.1 is three products wearing one name
Google's model is generally available in Standard, Fast and Lite, all three producing the same format family: 720p, 1080p and native 4K at 24 fps, native durations of 4, 6 or 8 seconds with longer outputs reported at 1080p and scene-extension chaining beyond that, up to three reference images for character and scene consistency, plus seed and negative-prompt controls. Output carries SynthID and C2PA provenance metadata.
What separates the tiers is price and speed, and the independent board says something uncomfortable about the spread. On Artificial Analysis's text-to-video leaderboard the three sit inside 14 Elo of each other — Veo 3.1 at 962 (±10) over 2,998 votes, Veo 3.1 Fast at 961 (±10) over 4,325, Veo 3.1 Lite at 948 (±10) over 2,903 — while their listed rates run $24.00, $7.20 and $4.80 per minute. Read those together and the arena is telling you it cannot reliably distinguish the tiers by preference. That does not mean the tiers are identical; it means a buyer's real question is which tier clears their bar, because the preference board will not answer it for them.
Kandinsky 6.0 Video takes the opposite approach to variation. There are two sizes — Pro at 29B and Lite at 3B — sharing one architecture and one fixed output format: 5 seconds, 121 frames at 24 fps, synchronized 44 kHz audio with lip-sync, from text or a reference image, with a separate super-resolution model raising the result to Full HD. There is no Fast variant and no extension mode. You pick a size and a GPU.
What the lab's own evaluation actually claims
This is the only head-to-head data that exists between these two systems, and it is vendor-reported on both sides of the slash — Sber ran it, Sber published it, and nobody outside Sber has reproduced it. Stating it accurately matters more than stating it favourably.
In the text-to-audio-video mode, the report has Kandinsky 6.0 Video Pro preferred on artifacts and camera motion with statistically significant margins, and slightly ahead on motion realism in a field dominated by ties. Veo 3.1 Fast leads on visual prompt following, speech quality, audio-video synchronization, overall audio quality, sound-prompt following and user-task solving, all but speech reaching significance. Overall visual quality is near parity with a slight edge to Veo.
In image-to-video the pattern flips on the visual half. Kandinsky 6.0 Video Pro is preferred on overall visual quality, reference-image alignment, artifacts and camera motion, all significant. Veo 3.1 Fast still leads on visual prompt following, audio-video synchronization, overall audio quality, sound-prompt following and user-task solving, also significant. Motion realism and speech quality come out near parity.
Two things follow. The image-to-video result is the genuine finding: an open model being preferred on reference-image correspondence and overall visual quality against a Google tier, in the lab's own study, is a real claim worth weighing. And the audio result is where Veo keeps its lead regardless of mode — which brings us to the flag nobody mentions.
The audio flag that decides whether that evaluation is fair
Veo 3.1 generates native 48 kHz stereo audio — dialogue, ambience, foley — jointly with the picture. It is one of the model's strongest features. On the API, though, the audio parameter defaults to off, and silent generation is priced below audio generation: roughly $0.20 per second silent against roughly $0.40 per second with audio on Google's published Standard tier.
Kandinsky 6.0 Video has no such switch. Audio is part of the output, and the serving pipeline merged into vLLM-Omni emits 44.1 kHz AAC by default. So the two models do not face the same default: one is audio-first, the other is silent-first with sound as an upgrade.
That asymmetry has a specific cost in comparisons. Any head-to-head that pits a silent Veo 3.1 against a competitor whose audio always ships, and then scores the pair on an audio-aware board, has handicapped Veo by an accident of a default. When you read the Sber numbers above — or anyone else's — the first question to ask is whether the Veo side had audio on. The second is what the price difference was, because at Standard tier turning sound on doubles it.
Five seconds against eight, and 1080p against native 4K
The format gap is where the two models' design briefs are visible.
• Clip length — Kandinsky 6.0 Video: 5s fixed, 121 frames at 24 fps, no extension. Veo 3.1: 4, 6 or 8s native, longer at 1080p, scene-extension chaining beyond that.
• Resolution — Kandinsky 6.0 Video: SD, HD and Full HD (1920×1080) via a super-resolution pass. Veo 3.1: 720p, 1080p and native 4K.
• Provenance — Kandinsky 6.0 Video: none stated in the release. Veo 3.1: SynthID and C2PA metadata on output.
• Reference input — Kandinsky 6.0 Video: one image, applied as a masked tail frame. Veo 3.1: up to three reference images plus seed and negative prompt.
• Format — Kandinsky 6.0 Video: 44.1 kHz AAC with the picture, always on. Veo 3.1: 48 kHz stereo, optional, priced two ways.
The provenance line is the one with procurement consequences. If your deliverable has to be traceable — brand work, broadcast, anything a legal team will look at — Veo 3.1 attaches metadata that says the footage was generated, and Kandinsky 6.0 Video does not. That is not a quality difference and no benchmark will show it, but it is the kind of requirement that decides a vendor on its own, and an open model without it is not a drop-in replacement for a team that needs it.
The 4K line matters in the same way, and for the same reason. Kandinsky 6.0 Video's Full HD is real — there is a dedicated SR model, and the repository's SR package handles x2, x4 and x2.25 upscales of five-second clips — but it tops out where Veo 3.1's native path starts at its upper end. For large-screen work, that ceiling is not negotiable.
What a clip costs on each side
Veo 3.1 is a per-second invoice. At Standard tier a five-second 1080p clip with audio is about $2.00, and the same clip silent is about $1.00; Fast and Lite sit below that, and because their arena scores land within the confidence interval of Standard's, the cheaper tiers are hard to argue against on the evidence.
Kandinsky 6.0 Video has no invoice. It has GPU time. The repository's figures for the non-distilled model, measured after warmup with weight loading and MP4 encoding excluded, put a five-second Pro Full HD clip at about 402 seconds on an H100, 854 seconds on an RTX 5090, 1,247 seconds on an RTX 4090 and 3,530 seconds on an RTX 5060 Ti. Lite Full HD is 284 seconds on an H100. Peak allocation on a Pro SD run reaches 72.8 GiB, so anything under that needs block offloading — and the 16 GB preset quantizes the text encoder to NF4, the one preset change the paper confirms alters the clip itself rather than introducing rounding-level noise.
Those two cost structures do not compare. One is a bill you forecast per finished second; the other is a machine you buy, a render measured in minutes, and the option to fine-tune — which Veo 3.1 does not offer at any tier.
The tier question is a routing question
OrcaRouter serves neither Google Veo 3.1 nor Kandinsky 6.0 Video, and neither is implied here — Veo 3.1 is reached through Google's own surfaces and Kandinsky is yours to download. But the structure of the Veo decision is worth naming, because it is the problem our routing layer exists to solve on the text side: three tiers whose independent scores are indistinguishable and whose prices differ five-fold is exactly the case where "route the cheap one and escalate only when the bar is not cleared" beats "standardise on the flagship". OrcaRouter's adaptive routing and routing DSL express that as a configured policy across 200-plus text models on one OpenAI-compatible endpoint, at provider list price with 0% markup, with automatic failover when a route goes down. The same instinct applied to video spend — Fast by default, Standard for the shots that matter — is a procurement decision you can make today without a router at all.
Which one, for whom
Choose Veo 3.1 if the deliverable needs sound taken seriously, if it needs 4K or a clip longer than five seconds, or if someone downstream requires provenance metadata. Pick the tier deliberately rather than defaulting to Standard, and budget for the audio flag instead of discovering it on the invoice.
Choose Kandinsky 6.0 Video if you want weights, if five seconds suits your unit of work, and if image-to-video fidelity against a supplied still is the job — the one dimension where the lab's own study puts it ahead of a Veo tier on significant margins. Go in knowing the audio is always on, the provenance is absent, and the quality claim rests entirely on a report its author wrote.


The asymmetry to hold onto is that Veo 3.1 has been in production for eleven months and has therefore accumulated the thing an open model released this month cannot: a mapped set of failure modes its users have already published, and a price curve a finance team can model. Against that, Kandinsky's MIT licence means the model on your disk does not depend on anyone's roadmap. Which of those is worth more is not a benchmark question.

