Title card for Odyssey-3 vs Grok Imagine Video, with a left panel headed Odyssey-3 reading an interactive environment, no published price, 832x480 or 1280x720 Pro, no audio; and a right panel headed Grok Imagine Video reading xAI Imagine API, metered per second, a clip you author, native 1080p on version 1.5, Physics-IQ -4.04 pp rank 13.
Guides & Insights

Odyssey-3 vs Grok Imagine Video: A Metered Video API Against an Unpriced Simulator

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

G​rok Imagine Video comes with a price list and no way to steer it. Odyssey-3, which Odyssey opened as a public research preview on October 8, 2026, comes with a way to steer it and no price list at all. That asymmetry, rather than any benchmark row, is the real shape of this matchup: x​AI sells generated video by the second with a published rate for every resolution, while Odyssey-3 ships as an interactive world model you can drive in a browser and cannot yet buy. One of these you can put in a product this afternoon. The other one you can only evaluate.

They also disagree about what video generation is for, and the disagreement shows up in the vendor pages. G​rok Imagine Video's documentation leads with reference images, pinned frames, lip-synced speech and voice references — the vocabulary of production control. Odyssey-3's lead is an environment that predicts what happens next as you move through it or trigger an event. Both are diffusion-based generators that emit video. Past that sentence there is almost no overlap.

What each one is

Odyssey-3 is an autoregressive diffusion transformer. Odyssey describes it as a multi-step video diffusion model extended autoregressively with teacher forcing and causal masking, so it continues from preceding observations and predicts future states conditioned on action inputs, then distilled through distribution matching and adversarial distillation into a few-step variant capable of real-time interaction. It runs at 832×480, with Odyssey-3 Pro at 1280×720, and the public preview exposes first-person navigation, third-person navigation and independent camera movement. There is no audio.

G​rok Imagine Video is x​AI's video model in the Imagine API. The current generation is G​rok Imagine Video 1.5, which x​AI's release notes record as supporting text-to-video, image-to-video and reference-to-video with optional preset voices and native 1080p for the first two — with text-to-video implemented as text-to-image followed by image-to-video under the hood. It accepts up to fourteen reference images and three voice references, supports pinned frames and lip-synced speech, and is metered per second of generated output. The earlier grok-imagine-video, still documented and still priced, sits at $0.05 per second at 480p and $0.07 at 720p, with input video billed at $0.01 per second and input images at $0.002.

The rate card, which is the part you can actually plan around

xAI's Grok API documentation page for grok-imagine-video, read 9 October 2026, showing the model header with a “Latest” badge, an at-a-glance table listing modalities as text, image and video under the heading “Text > Video”, a pricing heading, and release notes tabs, with the navigation marked “APIs / Grok / Build”.

• Price — G​rok Imagine Video is billed per second: $0.05/sec at 480p and $0.07/sec at 720p on the original model, and G​rok Imagine Video 1.5 from $0.08/sec across 480p, 720p and 1080p. G​rok Imagine Video 1.5 Lite starts at $0.02/sec for up to 1080p and fifteen seconds. Odyssey-3 has no published price at all; its only public cost figures are the leaderboard's normalized estimates of $0.139 per generated video for Odyssey-3 and $0.267 for Odyssey-3 Pro.

• A ten-second clip, worked out — $0.50 at 480p or $0.70 at 720p on G​rok Imagine Video; $0.80 on G​rok Imagine Video 1.5; $0.20 on the Lite tier. Odyssey-3 cannot be priced this way, because you cannot buy ten seconds of it.

• Resolution — G​rok Imagine Video 1.5 does native 1080p for text-to-video and image-to-video. Odyssey-3 tops out at 1280×720 in the Pro tier.

• Duration — G​rok Imagine Video 1.5 Lite reaches fifteen seconds. Odyssey-3's unit is not a clip; it is a session, bounded by how long the preview keeps its predictions coherent.

• Audio — G​rok Imagine Video generates lip-synced speech and accepts voice references. Odyssey-3 generates none.

• Control — reference images, pinned frames, first and last frame, and voice references, all specified before generation. Against live navigation and mid-generation events during generation. G​rok's is authoring; Odyssey-3's is interaction.

• Availability — a self-serve API key today for G​rok Imagine Video. A contact form and a browser preview for Odyssey-3.

Where the physics board puts them

Physics-IQ Verified, the Anates Labs and DeepMind benchmark that scores continuations of real physical experiments, currently ranks 41 models, and both of these are on it — which makes it the only third-party surface where the two can be compared at all.

• G​rok Imagine Video — rank 13, net improvement −4.04 pp against the track mean, on the benchmark's own prompts, 24 fps, normalized cost $0.352 per video.

• Odyssey-3 — rank 8 on the same prompt set at +2.15 pp and $0.129, and rank 3 on custom prompts at +9.99 pp.

The gap is not small, and it is not a resolution artifact: the board normalizes cost to 24 fps and 1280-wide output where the inputs allow it, and it lists the prompt regime for every row. A negative net-improvement score is not a statement that G​rok Imagine Video generates bad video — it clearly does not, and it is in production at a price that makes it easy to use. It is a statement that on the specific task of continuing a physical experiment correctly, it does worse than the average model on that board, while Odyssey-3 does better than the average model on that board twice over, under two different prompt regimes.

Odyssey-3's own headline claim is stronger still: 66.1 on the video-to-video track for Odyssey-3 Pro, the highest score Odyssey reports, plus 54.7 for the same tier on image-to-video. Those come from Odyssey's writeup rather than from an independent run, and they should be read that way.

Two-column scoreboard headed “Odyssey-3 vs Grok Imagine Video — the scoreboard” with six shared dimension labels. Odyssey-3: published price none; 832x480, 1280x720 Pro; a session, not a clip; no audio; Physics-IQ board prompts +2.15 pp rank 08; custom prompts +9.99 pp rank 03. Grok Imagine Video: $0.05/sec at 480p and $0.07/sec at 720p; native 1080p on version 1.5; up to 15 seconds on the Lite tier; lip-synced speech and voice references; Physics-IQ -4.04 pp rank 13; no custom-prompt entry. Footer: “Odyssey-3 figures vendor-reported and unreproduced; Grok Imagine Video rates per xAI documentation; Physics-IQ Verified board read Oct 9 2026”.

The thing the price list does not cover

There is a category difference hiding behind the per-second rates that is worth naming, because it decides which of these you want.

G​rok Imagine Video's billing model — dollars per second of output — is a promise about throughput. Every dollar buys a fixed quantity of finished video, and the only question is whether fifteen seconds at 1080p with lip-sync is worth $1.20. For an ad unit, a social template, a marketing pipeline, that is an excellent question with a clean answer, and the Lite tier at $0.02/sec makes the answer easy more often than not.

Odyssey-3 is not billed that way and cannot be, because its output is not a fixed quantity. An interactive environment is consumed by how long someone stands inside it and how many times they change their mind, and that makes the per-second model meaningless. When Odyssey publishes a price — and there is no signal of when — it will have to be a session price, a compute price, or something else that has not been invented yet for this category. That is not an argument against Odyssey-3. It is the reason a developer comparison between these two is currently impossible to finish.

OrcaRouter's own model page for minimax/minimax-h3, showing the model header “text + image + video + audio”, output video, a p50 time-to-first-token of 406 ms, performance and vision/audio badges, and a Python code sample on the left with the model list and playground navigation above.

Reaching the metered one today

G​rok Imagine Video is x​AI's own API, and OrcaRouter does not host it. Odyssey-3 is not on any hosted platform we can verify, and there is no published rate for anyone to route even if it were.

What one OrcaRouter key does reach is the rest of the video and model stack you would build this comparison around: MiniMax-H3 at $0.08 per second of 768P output, the K​ling video line, and more than 200 models behind a single endpoint at provider list price with 0% markup — so when a lab cuts a per-second rate, the change is live on your usage the same day rather than at your next contract renewal. Automatic failover across providers keeps a batch that mixes several video models from failing on one upstream's bad hour, and the routing DSL lets you put a hosted video model and a text or vision model behind one interface, which is what a pipeline that generates and then evaluates actually needs. Neither Odyssey-3 nor G​rok Imagine Video is in that catalogue today.

Which one you want

If you need video output with sound, at 1080p, at a price you can put in a spreadsheet, and you need it this week, G​rok Imagine Video is a working answer and Odyssey-3 is not a candidate — it has no audio, a lower resolution ceiling, and no way to buy it.

If you are building anything that has to predict what happens after an action — a policy, a control loop, a training environment — the price list is irrelevant and the physics board is not. Odyssey-3 is being scored on exactly that capability and it is near the top of the board; G​rok Imagine Video, excellent at producing a finished shot, is below the track mean on the same test. Different questions, and the answers do not overlap.