
Odyssey-3 vs Kandinsky 6.0 Video: Open Weights and Audio Against a Closed Interactive Preview
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 111 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 55 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 378 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Kandinsky 6.0 Video arrived on October 6, 2026 with the thing open model releases rarely land on day one: support in Diffusers, ComfyUI, vLLM-Omni, SGLang and FastVideo, all documented in the repository's own update log, alongside an arXiv report submitted on October 4 and MIT-licensed checkpoints on Hugging Face. Odyssey-3 arrived two days later, on October 8, as a public research preview of an autoregressive diffusion transformer that builds an environment from a prompt and predicts changes in real time as you move through it. One is a 29-billion-parameter model you can download and run on your own GPU tonight. The other is an interactive system you can only reach through Odyssey's browser preview and a contact form. They are the two poles of what "video model" meant in the same week of October 2026, and the choice between them is less about benchmarks than about who holds the weights.
They also want different things from you. Kandinsky 6.0 Video is a generation model: prompt in, a five-second clip with synchronized audio out. Odyssey-3 is a prediction model: act, and watch the world respond. Neither substitutes for the other, and the fact that they shipped 48 hours apart is the most useful context either of them has.
The two architectures, in the vendors' own terms
Kandinsky 6.0 Video is a family of foundation diffusion models for synchronized audio-video generation, comprising Kandinsky 6.0 Video Lite at 3B parameters and Kandinsky 6.0 Video Pro at 29B. Both generate five-second clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video and image-to-audio-video modes, and a plug-in super-resolution model raises the output to Full HD 1920×1080. The technique is a dual-stream CrossDiT that connects a pretrained video stream to a newly trained audio stream through bidirectional cross-attention, so the two modalities align in time and semantics rather than being generated separately and muxed. The training recipe is continuous: the audio stream is trained from scratch on large audio corpora, then both streams are trained jointly on paired audio-video data while preserving unimodal fidelity, followed by supervised fine-tuning, reinforcement-learning post-training and distillation.
Odyssey-3 is an autoregressive diffusion transformer described by Odyssey as a multi-step video diffusion model extended through teacher forcing and causal masking, trained to continue from preceding observations and predict future states conditioned on action inputs, then distilled with distribution matching and adversarial distillation into a few-step variant fast enough to interact with. It runs at 832×480, with Odyssey-3 Pro at 1280×720, produces no audio, and its whole purpose is that its output depends on what you did a moment ago.
Day-one support is the difference nobody benchmarks

Read the Kandinsky 6.0 update log again, because it is the most concrete fact in this comparison. On October 6 the project recorded availability in Diffusers, the ComfyUI node registry, vLLM-Omni, SGLang and FastVideo, and a Hugging Face Space for the distilled Pro model. That is five independent inference stacks in one day. For anyone who has waited months for a checkpoint to reach a serving framework, that is the number that decides whether the model is usable.
Now hold it next to Odyssey-3's access story. The preview runs at experience.odyssey.systems with first-person navigation, third-person navigation and independent camera movement. API access is a contact form. There is no pricing page, no rate limit documentation, and no self-serve key. The site's API terms were last updated 2026-01-22 and still govern "the prototype of Odyssey-2 API" — the commercial surface has not been rewritten for the model that shipped this month.
• Where it runs — your own GPUs, with weights and code under MIT, against Odyssey's servers only.
• How you integrate it — Diffusers pipeline, ComfyUI nodes, vLLM-Omni, SGLang, FastVideo, against a browser preview and a contact form.
• What you can inspect — architecture, training recipe and checkpoint sizes in a public report, against a vendor description of the training pipeline with no weights and no paper.
• What you can modify — fine-tunes, LoRAs, quantization and distillation, all of which the community was already publishing on Hugging Face the day after release, against nothing at all.
Dimension by dimension

• Output — five-second clips with synchronized 44 kHz audio for both Kandinsky tiers, against an interactive environment with no audio for Odyssey-3.
• Resolution — Full HD 1920×1080 via the plug-in super-resolution model for Kandinsky 6.0 Video, against 832×480 for Odyssey-3 and 1280×720 for Odyssey-3 Pro.
• Modalities in — text and image for Kandinsky, in text-to-audio-video and image-to-audio-video modes; a prompt plus live control input for Odyssey-3.
• Parameters — 3B and 29B published for the Lite and Pro tiers, against undisclosed for both Odyssey-3 tiers.
• Hardware — an NVIDIA GPU and Python 3.13 or 3.14, with presets shipped for H100/H200 through RTX 5090-class cards for Kandinsky; a browser for Odyssey-3.
• Price — $0 per clip for Kandinsky 6.0 Video if you have the GPU, against no published price of any kind for Odyssey-3. The board's normalized cost estimates for Odyssey-3 and Odyssey-3 Pro are $0.139 and $0.267 per generated video, excluding prompt handling.
• Third-party scoring — neither model's own generation is on an independent head-to-head. Kandinsky's report relies on side-by-side human evaluation against its predecessor; Odyssey-3's figures come from a benchmark submission plus a vendor-run evaluation.
• License — MIT for Kandinsky 6.0 Video, against no published license because there are no weights to license.
What the benchmarks do and do not say
Kandinsky 6.0 Video does not appear on Physics-IQ Verified, the Anates Labs and DeepMind board that scores continuations of real physical experiments across 41 models. Neither does Odyssey-3's head-to-head partner here, because the board scores models that generate clips and both of these do. What the board does carry is Kandinsky's earlier world-model line, Kandinsky-WM 1.0, at rank 20 and −8.02 pp against the track mean — a different family, a different purpose, and the only Kandinsky entry on the board.
Odyssey-3's rows, for contrast: rank 8 at +2.15 pp on the benchmark's own prompts and $0.129 per video, and rank 3 at +9.99 pp on custom prompts, with Odyssey-3 Pro second overall at +11.15 pp. Those are better numbers than anything the Kandinsky line has on that particular test, and they are also measuring a different capability — Odyssey-3 is scored on what it predicts after a disturbance, which is the same thing its preview sells.
Kandinsky's own evidence is the human evaluation in the technical report: Kandinsky 6.0 Video Pro clearly outperforms its predecessor Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly on speech quality. That is a vendor-run side-by-side, and it is the kind of claim that a released checkpoint can be checked against by anyone with a 5090 and an afternoon. Odyssey-3's equivalent claim cannot be checked by anyone without an API key.

Reaching them
OrcaRouter does not host either model, and it never will imply otherwise: Odyssey-3 has no published rate for anyone to route, and Kandinsky 6.0 Video is open weights, which means the correct way to run it is the repository's own setup on your own GPU, not a hosted endpoint.
Where one OrcaRouter key fits is the layer around the experiment. Kandinsky's repository assumes you supply the GPU, the environment and the comparison; what it does not supply is a way to call the models you want to benchmark it against. More than 200 models sit behind a single OrcaRouter key at provider list price with 0% markup — including minimax/minimax-h3, the open-weight video model MiniMax released on July 31, 2026, at $0.08 per second of 768P output — so a vendor price change lands on your invoice the same day, automatic failover keeps an evaluation sweep from dying on one upstream's bad hour, and the routing DSL lets you put a hosted video model and a vision model behind one interface when the job is generating and then scoring. You still clone the repo and run the setup yourself. You just do not have to build the other half of the rig.
Which one you actually want
If you need output you can own — commercial terms you can read, weights you can fine-tune, a pipeline that runs without someone else's API being up — Kandinsky 6.0 Video is the answer, and the day-one framework support means you are not betting on a research artifact. If you need sound on the video, it is the only one of the two that generates any.
If the problem is prediction rather than production — a policy that has to react, a training environment, anything where the next state depends on the last action — Odyssey-3 is the only one of the two that does it, and it is also the one you cannot buy. That is the honest shape of this matchup in the week both landed: an open model you can download and an interactive model you can only try, with the benchmark evidence favoring the closed one and the practical evidence favoring the open one.
