A generated title card reading 'Kandinsky 6.0 Video Lands in vLLM-Omni' with the subtitle 'A checkpoint becomes a service', above an icon row labelled 'Video Generation', 'Synchronized Audio' and 'Open-Weight Deployment'. The OrcaRouter logo sits in the bottom-right corner.
Engineering & Research

Kandinsky 6.0 Video Lands in vLLM-Omni: What a Serving Layer Adds to an Open Model

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Pull request #8537 merged into vLLM-Omni at 07:55 UTC this morning, and it does exactly one thing: it makes Kandinsky 6.0 Video servable. The model itself arrived from Sber's Kandinsky Lab the same day as a pair — Kandinsky 6.0 Video Pro at 29B parameters and Kandinsky 6.0 Video Lite at 3B — generating five-second clips with synchronized 44 kHz audio, lip-sync included, from a text prompt or a reference image, with code, weights and diffusers integration all under MIT. What it did not have until this morning was a way to run it inside an inference server instead of a one-shot command line.

That distinction is worth more than it sounds, and it is the reason this is a separate story from the release. An open checkpoint you can run on your own card is a research artifact; an open checkpoint with a server-side pipeline, shared entry points and a serving recipe is something you can put behind a queue. This week's release gave you the first. Today's merge is the beginning of the second.

What actually merged

The PR adds text-to-video-and-audio and image-to-video-and-audio to vLLM-Omni from the checkpoint kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers. The engineering choices are the interesting part, because they tell you what the architecture really looks like when somebody other than the authors has to serve it.

• One DiT denoises both modalities. The pipeline is registered as Kandinsky6TI2VAPipeline, and video latents and audio latents are denoised jointly by a single transformer rather than by two models run in sequence. That mirrors the paper's dual-stream CrossDiT, where a pretrained video stream and a from-scratch audio stream are connected by bidirectional cross-attention.

• Weights load folder by folder. The Hub checkpoint is read as transformer/, vae/, text_encoder/, text_encoder_2/ and audio_vae/. The published checkpoint's internal names — videoT and audioT — are mapped onto video_dec_block and audio_dec_block at load time, which is the sort of detail that eats a day if you discover it yourself.

• Audio is on by default. The pipeline emits 44.1 kHz AAC rather than treating sound as an opt-in flag, so a request with no audio parameters still comes back with synchronized speech, music or ambience.

• An image input is a masked tail frame, not a separate mode. Reference-image conditioning is applied by masking, which is how the same pipeline serves both T2AV and I2AV without branching.

The submitter's smoke test is a useful floor: a single NVIDIA H100 80GB with FlashAttention-3, CPU offload on, 864×480, 125 frames, 50 steps, CFG 5.0, seed 42. That is a 5-second clip at just under HD, not a Full HD render, and the PR does not claim otherwise.

A checkpoint is not a service

The reason to care about a merge like this is what it removes from your list. Running the reference implementation means cloning the repository, running just setup, letting it compile SageAttention on anything below Hopper, downloading the checkpoint into a cache directory and invoking a generate command whose output lands in a timestamped folder. That is fine for a first look and awkward for a product.

vLLM-Omni gives you the model inside a serving process, with the same offline-inference entry points the project uses for its other diffusion models (examples/offline_inference/text_to_video/text_to_video.py and its image-to-video sibling) and a serving recipe at recipes/Kandinsky/Kandinsky6-TI2VA.md. Two practical consequences follow. Your deployment story becomes a container that speaks an HTTP interface rather than a script you shepherd, and the model stops being a special case in your stack — it sits next to whatever else you run in that server.

Note what this is not. It is not a hosted endpoint, it is not a price, and it is not an independent quality measurement. It is one more place the model can run, contributed by someone outside the lab, three days after the weights appeared.

What a five-second clip costs in wall-clock, on real cards

The repository's own table is the honest starting point. These are the non-distilled base model's working times for a single 5-second clip after warmup, with weight loading and MP4 encoding excluded. Full HD here means the super-resolution pass is running too.

• Full HD (Pro) — 402 s on an H100, 765 s on an RTX PRO 6000, 854 s on an RTX 5090, 1,106 s on an A100 80GB, 1,247 s on an RTX 4090, and 3,530 s on an RTX 5060 Ti.

• Full HD (Lite) — 284 s on an H100, 387 s on an RTX PRO 6000, 406 s on an RTX 5090, 664 s on an A100 80GB, 578 s on an RTX 4090, and 1,774 s on an RTX 5060 Ti.

• SD (Pro) — 356 s on an H100 against 936 s on an RTX 4090, which is where the consumer-card penalty is smallest and the economics are least bad.

The shape of that table matters more than any single cell. An H100 is roughly three times faster than a 4090 on Pro Full HD, and roughly six times faster than a 5060 Ti on the same job — but the SR stage is the same cost on either model size, about 64 seconds at HD and about 96 seconds at Full HD on a 5090. Lite does not buy you a shorter super-resolution pass, because the SR model is trained separately and does not care which base model fed it.

The 16 GB path, and the one setting that changes your picture

Peak allocated memory on a Pro SD run reaches 72.8 GiB, which is past every consumer card. The repository answers with block offloading — only two transformer blocks resident on the GPU at a time, the rest streamed from host memory — plus three presets for 32 GB, 24 GB and 16 GB cards. The 32 and 24 GB presets are identical, because above 24 GB the extra headroom does not translate into speed.

The caveat is buried and worth repeating: on the 16 GB preset the text encoder is quantized to NF4, and the paper is explicit that this is the only preset change that alters the clip itself. A different textual representation produces a different video. Everything else — segment length, spatial strips, offloading — changes output at the level of rounding noise, around 12% of pixels differing by one brightness level at roughly 57 dB PSNR. If you are reproducing someone else's result on a 16 GB card and it does not match, that is the first thing to check.

Three things the release notes do not settle

• Five seconds is the ceiling, not a default. The model was trained at 121 frames and 24 fps, and both the paper and the serving recipe are built around that. There is no extension mode in the release. Anything longer is a stitching problem you own.

• The distilled checkpoint is what you will actually serve, and it was measured at parity. The paper's ablation puts the distilled model preferred or tied on the majority of criteria with no individual difference reaching significance, at a mean preference of 51% to 49%. That is the trade you take on for the speed.

• There are two word-error-rate numbers in the paper and they are not comparable. The 47% reduction in generated-speech WER that accompanies the reinforcement-learning stage is scored with Whisper-large-v3, one of the RL reward models; the WER reported alongside VABench uses Qwen3-ASR-1.7B on the speech subset. The authors say so themselves. Quoting one as the other is the easiest mistake to make with this release.

Where a router sits in a stack like this

OrcaRouter does not serve Kandinsky 6.0 Video, and nothing here should be read as a claim that it does — the model is yours to run from the Hub, or reachable through the lab's own surfaces. What a pipeline like this needs from a router is the text around it. The prompt-expansion pass that turns a shot description into the long, medium or short caption the model was trained on, the caption and transcript work that feeds an image-to-video pass, the review summaries that decide which of four generations to keep: those are ordinary text calls, and they run on one OpenAI-compatible endpoint across 200-plus models with provider list prices passed through at 0% markup, so a vendor price change is live the same day. Automatic failover matters more than usual here, because a five-minute render that dies at the caption stage should reroute rather than restart the job.

The next thing to watch is not another merge but a number. Kandinsky 6.0 Video Pro has vendor-reported results on VABench and a lab-run human evaluation against Kling 2.6, Veo 3.1 Fast, MiniMax H3 and Seedance 2.0 — and no independent arena score yet. When one appears, the open-weights claim stops being an argument about access and starts being an argument about quality, which is the comparison that decides whether this is a model teams download or a model they admire.

A generated scoreboard titled 'Kandinsky 6.0 Video — the scoreboard' with a single column of six rows: 'Sizes: Pro 29B, Lite 3B', 'Clip length: 5s fixed', 'Audio: 44 kHz lip-sync', 'Resolution: Full HD plus SR', 'Licence: MIT' and 'Serving: vLLM-Omni day 0'. A footer reads 'All figures vendor-reported; no independent Elo yet. October 2026.' The OrcaRouter logo sits in the bottom-right corner.A screenshot of the GitHub repository page for kandinskylab/kandinsky-6, showing 11 commits, 35 stars and 4 forks, an MIT license badge, and a README titled 'Kandinsky 6.0: A family of diffusion models for Video + Audio Generation' with badges for KANDINSKYLAB, REPORT, DIFFUSERS and COMFYUI; recent commits include 'README update', 'Added source code', 'Some improvements for just generate' and 'Keep ComfyUI user documentation in README'.

The repository also ships two ComfyUI nodes — kandinsky6 and kandinsky6-sr — installable through ComfyUI Manager, and two Hugging Face Spaces: one for the Pro distill checkpoint and one for the video super-resolution demo. Between the CLI, the ComfyUI nodes, the diffusers bundle and now a vLLM-Omni pipeline, the deployment surface is broader on day three than most open video models manage in a quarter. That breadth is the actual story of the week: Sber released a family, not a paper with a demo attached.

A screenshot of the Artificial Analysis AA-Video-T2V V2.0 leaderboard, ranking text-to-video models by Elo from human preference votes. Wan 3.0 is first at 1,156 over 8,300 samples and $12.00 per minute (Aug 2026); MiniMax H3 (768p) is fourth at 1,137 over 8,315 samples at $4.80 per minute (Jul 2026); grok-imagine-video-1.5 is eleventh at 1,045 over 4,056 samples at $15.00 per minute (May 2026); Wan2.7-260612 is thirteenth at 1,030 over 5,571 samples at $9.00 per minute (Jun 2026); Veo 3.1 is eighteenth at 962 over 2,998 samples at $24.00 per minute, Veo 3.1 Fast nineteenth at 961 over 4,325 samples at $7.20 per minute, and Veo 3.1 Lite twenty-first at 948 over 2,903 samples at $4.80 per minute.