
Kandinsky 6.0 Video vs Grok Imagine Video: The Leaderboard Turned Over
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 150 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 126 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 250 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
In January 2026, when Grok Imagine Video launched, it topped the Artificial Analysis video arena in both modes. Today the same board places its current build eleventh or twelfth in text-to-video. That is not a scandal and it is not a failure — it is what happens to a fast-moving category in nine months, and it is the honest reason to compare xAI's generation model against Kandinsky 6.0 Video right now. One of them is a service optimised for how quickly you can produce another clip; the other is an MIT-licensed checkpoint, 29B parameters in the Pro size and 3B in the Lite, released by Sber's Kandinsky Lab in the first week of October 2026, generating five seconds of video with synchronized 44 kHz audio and lip-sync. The interesting question is not which one is ahead on a board — it is what each one is structurally able to do next.
The board that moved under it
Grok Imagine Video is xAI's generation model, first released on 29 January 2026 and stable at version 1.5 since 16 June. On Artificial Analysis's text-to-video leaderboard its 1.5 build sits 11th–12th at Elo 1,045 (±9) over 4,056 votes, listed at $15.00 per minute. The original build is not on that board.
Image-to-video is where the family is stronger, and where the two builds separate in a way worth reading carefully:
• grok-imagine-video-1.5 — 7th–9th, Elo 1,098 (±8), 5,267 votes, $8.40/min.
• grok-imagine-video (the January build) — 12th–17th, Elo 1,072 (±7), 14,371 votes, $4.20/min.
• For scale, the top of that I2V board — MiniMax H3 Max — sits at Elo 1,195 (±9) over 5,894 votes, and Wan 3.0 is sixth at 1,164.
Two readings follow, and both are true. xAI's newer build is genuinely ahead of its own predecessor on image-to-video, by 26 Elo, and it costs twice as much per minute to be. And the launch-week story — a model that arrived at the top — has not held: the field moved, the arena accumulated tens of thousands of votes across a dozen newer systems, and a mid-table position is where nine months of competition puts a fast, general-purpose generator.
Kandinsky 6.0 Video has no Elo on any board. Its evidence is a VABench run and a lab-run human evaluation, both published by Sber, neither reproduced outside it. Placing an unaudited open model beside a scored commercial one is exactly the comparison to treat with care: the scored model's position is real and unflattering, and the unscored model's position is unknown. "Unknown" is not "better".
Two kinds of fast, and only one of them is measured in seconds
Grok Imagine Video's design goal is throughput per creator. It is wired into X, it hands back a finished clip with sound, and its feature set reads like a list of things people actually do in a feed: multiple aspect ratios, camera motion, object add, remove and swap, style transfer, reference-conditioned generation from several images for character consistency, native audio with dialogue, effects and music. Clip length runs up to fifteen seconds, resolution tops out at 1080p. The whole product is built so that a second attempt costs you a couple of minutes and a few cents.
Kandinsky 6.0 Video is fast in a different sense — not per attempt, but per unit of ownership. Nothing about it is instantaneous. The repository's own working times for the non-distilled model, after warmup and excluding weight loading and MP4 encoding, put a five-second Pro Full HD clip at roughly 402 seconds on an H100, 854 on an RTX 5090, 1,247 on an RTX 4090 and 3,530 on an RTX 5060 Ti. Lite Full HD is 284 seconds on an H100. The tool that makes that bearable is the distilled checkpoint, which the paper measured at parity with the full model — preferred or tied on the majority of criteria, mean preference 51% to 49%, no individual difference reaching significance. What you get in exchange for the slow render is a model on your own disk with no per-clip meter running at all.
That is the real shape of this matchup. Grok Imagine Video optimises the cost of the tenth attempt. Kandinsky 6.0 Video optimises the cost of the ten-thousandth, and charges you in hardware and patience for the first.
Both generate audio — and they are not the same claim
This is the axis where a lazy comparison would say "tie" and be wrong.
Grok Imagine Video produces native audio with dialogue, effects and music as one feature among many. It is a general-purpose generator that happens to include sound.
Kandinsky 6.0 Video is built around sound. The architecture is a dual-stream CrossDiT pairing a video stream initialised from Kandinsky 5.0 Video with an audio stream trained from scratch on large audio corpora, joined by bidirectional cross-attention and trained jointly. The output is 44 kHz with explicit lip-sync, and the paper's reinforcement-learning stage is reported to cut the word error rate of generated speech by 47% — measured with Whisper-large-v3, which is one of the RL reward models, a detail that matters because the paper reports a second, non-comparable WER alongside VABench using Qwen3-ASR-1.7B. Speech is the thing the model was post-trained on.
Against that, Sber's own study is candid about where it loses. In the lab's side-by-side evaluation, MiniMax H3 and Dreamina Seedance 2.0 are preferred on visual criteria with significant margins, and the model is described as competitive rather than ahead. The claim worth taking seriously is narrower and more useful: against Kling 2.6 and Veo 3.1 Fast the results trade criterion by criterion, and Kandinsky 6.0 Video stays competitive on speech quality and audio-video synchronization where a large fraction of comparisons are ties.
Fifteen seconds against five, and what each costs to reach
• Max clip — Grok Imagine Video: up to 15s. Kandinsky 6.0 Video: 5s fixed, 121 frames at 24 fps, no extension mode in the release.
• Resolution — Grok Imagine Video: up to 1080p. Kandinsky 6.0 Video: SD, HD and Full HD through a dedicated super-resolution model, x2, x4 or x2.25 upscales of five-second clips.
• Reference input — Grok Imagine Video: reference-conditioned generation from several images for character consistency, plus object add/remove/swap. Kandinsky 6.0 Video: one image, applied as a masked tail frame.
• Pricing — Grok Imagine Video: per second of output, from the low end of the range up to $0.30/s at 1080p, with an additional charge per image input; free access sits behind X subscription tiers. Kandinsky 6.0 Video: none — no service, no rate, only the GPU time you supply.
• Weights — Grok Imagine Video: no. Kandinsky 6.0 Video: MIT, Pro and Lite, with diffusers integration.
The premium for the newer Grok build is the number to circle. Going from the January build to 1.5 costs twice as much per minute on the image-to-video board and buys 26 Elo. That is a real improvement at a real price, and it is the same trade every video vendor is asking buyers to make right now.
Where a router fits, and where it does not
OrcaRouter does not serve Grok Imagine Video or Kandinsky 6.0 Video, and nothing here claims otherwise: Kandinsky is yours to download and run, Grok Imagine Video is reached through xAI's own surfaces. What a pipeline built on either model needs from a router is the text that surrounds the render — the shot list, the prompt expansion that turns an idea into the long or short caption these models were trained on, the captioning and transcription work, and the review summaries that decide which generation is kept. Those run on one OpenAI-compatible endpoint across 200-plus text models at provider list price with 0% markup, so a vendor cut is live on the same key the day it lands, with automatic failover so a failed call in the middle of a render queue reroutes instead of stalling the batch. The orchestration around a video model is a routing problem even when the video model itself is not routable, which is the case for both of these today.
Which one, and for what
Reach for Grok Imagine Video when the work is volume: social clips, fast iterations, a shot you will regenerate six times before it is right, and a deadline that does not stretch. Its price per attempt is the product, and the fact that it now sits mid-table rather than at the top is the normal cost of a category that moves this fast.
Reach for Kandinsky 6.0 Video when the work is a pipeline: a fixed five-second unit you can batch, an image-to-video pass where fidelity to a supplied still is what matters, speech you intend to be intelligible, and a deliverable you would rather not rent. Budget for the render, expect a slow one on anything consumer-grade, and treat every quality figure in this piece as Sber's own until somebody outside the lab scores it.


What makes the pairing worth watching is which one can absorb a bad quarter. If Grok Imagine Video slips further down the arena, the answer is a better model and a new price — that is how the category works and xAI has already shown it will ship one. If Kandinsky 6.0 Video turns out to be mid-table once it is scored, the checkpoint still exists, still runs, and still fine-tunes; nothing about the licence changes with the number. Those are different kinds of durability, and only one of them is measured by Elo.

