A generated title card reading 'Kandinsky 6.0 Video vs Grok Imagine Video' with the subtitle 'The leaderboard turned over', above an icon row labelled 'Video Generation', 'Synchronized Audio' and 'Open-Weight Deployment'. The OrcaRouter logo sits in the bottom-right corner.
Guides & Insights

Kandinsky 6.0 Video vs Grok Imagine Video: The Leaderboard Turned Over

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

In January 2026, when Grok Imagine Video launched, it topped the Artificial Analysis video arena in both modes. Today the same board places its current build eleventh or twelfth in text-to-video. That is not a scandal and it is not a failure — it is what happens to a fast-moving category in nine months, and it is the honest reason to compare xAI's generation model against Kandinsky 6.0 Video right now. One of them is a service optimised for how quickly you can produce another clip; the other is an MIT-licensed checkpoint, 29B parameters in the Pro size and 3B in the Lite, released by Sber's Kandinsky Lab in the first week of October 2026, generating five seconds of video with synchronized 44 kHz audio and lip-sync. The interesting question is not which one is ahead on a board — it is what each one is structurally able to do next.

The board that moved under it

Grok Imagine Video is xAI's generation model, first released on 29 January 2026 and stable at version 1.5 since 16 June. On Artificial Analysis's text-to-video leaderboard its 1.5 build sits 11th–12th at Elo 1,045 (±9) over 4,056 votes, listed at $15.00 per minute. The original build is not on that board.

Image-to-video is where the family is stronger, and where the two builds separate in a way worth reading carefully:

• grok-imagine-video-1.5 — 7th–9th, Elo 1,098 (±8), 5,267 votes, $8.40/min.

• grok-imagine-video (the January build) — 12th–17th, Elo 1,072 (±7), 14,371 votes, $4.20/min.

• For scale, the top of that I2V board — MiniMax H3 Max — sits at Elo 1,195 (±9) over 5,894 votes, and Wan 3.0 is sixth at 1,164.

Two readings follow, and both are true. xAI's newer build is genuinely ahead of its own predecessor on image-to-video, by 26 Elo, and it costs twice as much per minute to be. And the launch-week story — a model that arrived at the top — has not held: the field moved, the arena accumulated tens of thousands of votes across a dozen newer systems, and a mid-table position is where nine months of competition puts a fast, general-purpose generator.

Kandinsky 6.0 Video has no Elo on any board. Its evidence is a VABench run and a lab-run human evaluation, both published by Sber, neither reproduced outside it. Placing an unaudited open model beside a scored commercial one is exactly the comparison to treat with care: the scored model's position is real and unflattering, and the unscored model's position is unknown. "Unknown" is not "better".

Two kinds of fast, and only one of them is measured in seconds

Grok Imagine Video's design goal is throughput per creator. It is wired into X, it hands back a finished clip with sound, and its feature set reads like a list of things people actually do in a feed: multiple aspect ratios, camera motion, object add, remove and swap, style transfer, reference-conditioned generation from several images for character consistency, native audio with dialogue, effects and music. Clip length runs up to fifteen seconds, resolution tops out at 1080p. The whole product is built so that a second attempt costs you a couple of minutes and a few cents.

Kandinsky 6.0 Video is fast in a different sense — not per attempt, but per unit of ownership. Nothing about it is instantaneous. The repository's own working times for the non-distilled model, after warmup and excluding weight loading and MP4 encoding, put a five-second Pro Full HD clip at roughly 402 seconds on an H100, 854 on an RTX 5090, 1,247 on an RTX 4090 and 3,530 on an RTX 5060 Ti. Lite Full HD is 284 seconds on an H100. The tool that makes that bearable is the distilled checkpoint, which the paper measured at parity with the full model — preferred or tied on the majority of criteria, mean preference 51% to 49%, no individual difference reaching significance. What you get in exchange for the slow render is a model on your own disk with no per-clip meter running at all.

That is the real shape of this matchup. Grok Imagine Video optimises the cost of the tenth attempt. Kandinsky 6.0 Video optimises the cost of the ten-thousandth, and charges you in hardware and patience for the first.

Both generate audio — and they are not the same claim

This is the axis where a lazy comparison would say "tie" and be wrong.

Grok Imagine Video produces native audio with dialogue, effects and music as one feature among many. It is a general-purpose generator that happens to include sound.

Kandinsky 6.0 Video is built around sound. The architecture is a dual-stream CrossDiT pairing a video stream initialised from Kandinsky 5.0 Video with an audio stream trained from scratch on large audio corpora, joined by bidirectional cross-attention and trained jointly. The output is 44 kHz with explicit lip-sync, and the paper's reinforcement-learning stage is reported to cut the word error rate of generated speech by 47% — measured with Whisper-large-v3, which is one of the RL reward models, a detail that matters because the paper reports a second, non-comparable WER alongside VABench using Qwen3-ASR-1.7B. Speech is the thing the model was post-trained on.

Against that, Sber's own study is candid about where it loses. In the lab's side-by-side evaluation, MiniMax H3 and Dreamina Seedance 2.0 are preferred on visual criteria with significant margins, and the model is described as competitive rather than ahead. The claim worth taking seriously is narrower and more useful: against Kling 2.6 and Veo 3.1 Fast the results trade criterion by criterion, and Kandinsky 6.0 Video stays competitive on speech quality and audio-video synchronization where a large fraction of comparisons are ties.

Fifteen seconds against five, and what each costs to reach

• Max clip — Grok Imagine Video: up to 15s. Kandinsky 6.0 Video: 5s fixed, 121 frames at 24 fps, no extension mode in the release.

• Resolution — Grok Imagine Video: up to 1080p. Kandinsky 6.0 Video: SD, HD and Full HD through a dedicated super-resolution model, x2, x4 or x2.25 upscales of five-second clips.

• Reference input — Grok Imagine Video: reference-conditioned generation from several images for character consistency, plus object add/remove/swap. Kandinsky 6.0 Video: one image, applied as a masked tail frame.

• Pricing — Grok Imagine Video: per second of output, from the low end of the range up to $0.30/s at 1080p, with an additional charge per image input; free access sits behind X subscription tiers. Kandinsky 6.0 Video: none — no service, no rate, only the GPU time you supply.

• Weights — Grok Imagine Video: no. Kandinsky 6.0 Video: MIT, Pro and Lite, with diffusers integration.

The premium for the newer Grok build is the number to circle. Going from the January build to 1.5 costs twice as much per minute on the image-to-video board and buys 26 Elo. That is a real improvement at a real price, and it is the same trade every video vendor is asking buyers to make right now.

Where a router fits, and where it does not

OrcaRouter does not serve Grok Imagine Video or Kandinsky 6.0 Video, and nothing here claims otherwise: Kandinsky is yours to download and run, Grok Imagine Video is reached through xAI's own surfaces. What a pipeline built on either model needs from a router is the text that surrounds the render — the shot list, the prompt expansion that turns an idea into the long or short caption these models were trained on, the captioning and transcription work, and the review summaries that decide which generation is kept. Those run on one OpenAI-compatible endpoint across 200-plus text models at provider list price with 0% markup, so a vendor cut is live on the same key the day it lands, with automatic failover so a failed call in the middle of a render queue reroutes instead of stalling the batch. The orchestration around a video model is a routing problem even when the video model itself is not routable, which is the case for both of these today.

Which one, and for what

Reach for Grok Imagine Video when the work is volume: social clips, fast iterations, a shot you will regenerate six times before it is right, and a deadline that does not stretch. Its price per attempt is the product, and the fact that it now sits mid-table rather than at the top is the normal cost of a category that moves this fast.

Reach for Kandinsky 6.0 Video when the work is a pipeline: a fixed five-second unit you can batch, an image-to-video pass where fidelity to a supplied still is what matters, speech you intend to be intelligible, and a deliverable you would rather not rent. Budget for the render, expect a slow one on anything consumer-grade, and treat every quality figure in this piece as Sber's own until somebody outside the lab scores it.

A generated two-column scoreboard titled 'Kandinsky 6.0 Video vs Grok Imagine Video — the scoreboard'. The Kandinsky 6.0 Video column reads 'Max clip: 5s fixed', 'T2V Elo: not scored', 'I2V Elo: not scored', 'Audio: 44 kHz, always on', 'Weights: MIT, Pro and Lite', 'Price: GPU time only'. The Grok Imagine Video column reads 'Max clip: up to 15s', 'T2V Elo: 1,045, 11th', 'I2V Elo: 1,098, 7th', 'Audio: native, optional', 'Weights: none', 'Price: $4.20 to $15.00/min'. A footer reads 'Grok figures per Artificial Analysis; Kandinsky unscored. October 2026.' The OrcaRouter logo sits in the bottom-right corner.A screenshot of the Artificial Analysis AA-Video-I2V V1.0 leaderboard, ranking image-to-video models by Elo from pairwise human preference votes with audio. MiniMax H3 Max is first at 1,195 over 5,894 samples at $4.80 per minute (Aug 2026); MiniMax H3 is second at 1,181 over 7,284 samples at $7.80 per minute (Jul 2026); Gemini omni Flash is third at 1,178 over 12,327 samples at $6.00 per minute (May 2026); Wan 3.0 is sixth at 1,164 over 9,302 samples at $12.00 per minute (Aug 2026); grok-imagine-video-1.5 is eighth at 1,098 over 5,267 samples at $8.40 per minute (May 2026); Wan 2.7 is twelfth at 1,077 over 4,571 samples at $9.00 per minute (Apr 2026); grok-imagine-video is thirteenth at 1,072 over 14,371 samples at $4.20 per minute (Jan 2026). The board carries a note that AA-Video-I2V v2.0 is coming soon.

What makes the pairing worth watching is which one can absorb a bad quarter. If Grok Imagine Video slips further down the arena, the answer is a better model and a new price — that is how the category works and xAI has already shown it will ship one. If Kandinsky 6.0 Video turns out to be mid-table once it is scored, the checkpoint still exists, still runs, and still fine-tunes; nothing about the licence changes with the number. Those are different kinds of durability, and only one of them is measured by Elo.

A screenshot of OrcaRouter's own model page for kling/kling-v3-omni, titled 'kling/kling-v3-omni' with a FLAGSHIP badge and credited to Kling, showing the route endpoint /v1/video/generations and a per-request price of $0.08, with a description reading that Kling 3.0 Omni is a unified text-to-video and image-to-video API with multi-shot, subject control and video-reference, 3-15 second clips and up to native 4K.