
Odyssey-3 vs Google Veo 3.1: Frontier Physics Against Ten Months of Production
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
The comparison between Odyssey-3 and Veo 3.1 is really a comparison between two dates. Veo 3.1 has been generally available since November 17, 2025 — ten months of production use, published per-second pricing, native 4K, and a serving stack that has been through an enterprise review at least once. Odyssey-3 was previewed by Odyssey on September 15, 2026, has no published price, no API documentation, no independent benchmark, and a release window of "the coming weeks." Setting them side by side is still worth doing, because the reason they get mentioned together is real: Veo 3.1 renders beautiful video, and Odyssey-3 is trying to simulate a world that behaves correctly under control. Those are the two different products teams keep conflating under the word "video," and the gap between them is exactly where the purchasing decision lives.
What ten months of shipping buys you
Veo 3.1's advantage is not a feature, it is elapsed time. The model has been in paying customers' hands since October 2025 as a preview and since November 17 as a GA product, with a Lite tier added on March 31, 2026 and a free tier on April 2. That history produces things a preview cannot: a documented API surface, an established set of failure modes that practitioners have already mapped, and a cost curve a finance team can model.
The spec sheet reflects the same maturity. Veo 3.1 generates at 720p, 1080p, and native 4K, at 24 frames per second, in native durations of 4, 6, or 8 seconds, with longer outputs reported at 1080p and scene-extension chaining beyond that. It accepts up to three reference images for character and scene consistency, plus seed and negative prompt controls. Output carries SynthID and C2PA provenance metadata. For brand, broadcast, or large-screen work, the native 4K path is the single capability that most of the competing field simply does not have — and it is the reason Veo 3.1 keeps winning projects it does not win on price.

The audio default that quietly skews every comparison
Veo 3.1 generates native 48kHz stereo audio — dialogue, ambient sound, foley — jointly with the picture, which is genuinely one of its strongest features. On the API, though, the audio parameter defaults to off, and silent generation is priced below audio generation. On Google's published Standard tier that works out to roughly $0.20 per second silent against roughly $0.40 per second with audio. Two consequences follow. First, the effective price of a Veo clip depends entirely on a flag most people never change, so a "$0.20 per second" figure and a "$0.40 per second" figure can both be correct for the same model. Second, and more subtly, it means a great many head-to-head tests you have read pitted a silent Veo 3.1 against competitors whose audio is always on, and then scored them on an audio-aware leaderboard. If your output needs sound, the honest number to plan against is the audio-enabled one.
What Odyssey-3 is actually claiming
Odyssey-3 is not a better video generator and does not claim to be. It is an autoregressive diffusion transformer trained on visual observation of the world, and the capability at the centre of the announcement is an action decoder — "a learned output component attached to the world model" that turns the model's internal scene representation into motor commands for specific hardware. Odyssey reports it controlling robot arms from tens of hours of demonstrations, driving humanoid policies built with Flexion in real time, producing closed-loop driving waypoints from a frozen backbone trained on 20 hours of simulated data, flying drones indoors around obstacles, and playing Grand Theft Auto V with transfer to Red Dead Redemption 2 and Sleeping Dogs without extra policy training there. The single comparative number released is that simulation-trained driving policies traveled about 77% as far between safety-driver interventions as policies trained on real footage.
Note what kind of claims those are. None of them is about how the output looks. Every one is about what a program downstream of the model can do with it.
The dimensions that do not overlap
• Output — Google Veo 3.1: a rendered clip, up to native 4K, with optional 48kHz stereo audio. Odyssey-3: a simulated world state, plus motor actions via the action decoder.
• Consumer — Google Veo 3.1: a viewer or an editor. Odyssey-3: a robot controller, a driving policy, or a training loop.
• Clip length — Google Veo 3.1: 4/6/8s native, longer at 1080p, extension chains beyond. Odyssey-3: continuous rollout, no clip concept.
• Price — Google Veo 3.1: ~$0.20/s silent, ~$0.40/s with audio at Standard; Fast and Lite tiers below that. Odyssey-3: none published.
• Availability — Google Veo 3.1: GA since 2025-11-17. Odyssey-3: previewed 2026-09-15, public "in the coming weeks."
• Evidence — Google Veo 3.1: ten months of production use and independent leaderboard placement. Odyssey-3: vendor demonstrations, no benchmark table, no technical report.
Where the physics claim actually pays off — and where it does not
It is tempting to assume a model with better physics makes better video, and that is not what Odyssey is selling. A video team's problem is almost never that the world in the clip is physically inconsistent — it is that the shot needs to be 4K, or the character needs to look the same in three different setups, or the deliverable needs provenance metadata attached before legal will sign off. Veo 3.1 answers those questions today. A physics-accurate simulator answers a different question: can I train something in here that will work out there. If you are not training anything, the physical accuracy of Odyssey-3 is a feature with no buyer in your workflow.
Where the two genuinely converge is previsualization. A director who wants to explore a space interactively — walk a camera through a set, change a blocking choice, see the consequence — wants exactly what a world model provides and what a rendered clip cannot, because the clip is fixed the moment it is generated. That is a real use case, and it is also one where an unreleased preview is not yet an option.
Running this in production without betting on one endpoint
The reason production teams care about the architecture around the model is that a video pipeline is rarely one call. A prompt model drafts the shot list, an image model produces the start frame, a vision model checks the result against the brief, a transcription model handles the voiceover. Each of those is a dependency that can fail at 2am, and each is a separate integration if you buy them separately. Putting that layer on one API across 200-plus models at provider list price with 0% markup collapses the integrations, and automatic failover means a prompt-model outage reroutes instead of stalling the render queue. The routing DSL matters more here than in a single-model workflow, because composing several models into one call is what lets a multi-stage pipeline fail gracefully stage by stage rather than as a whole. To be clear about scope: OrcaRouter does not serve Google Veo 3.1 or Odyssey-3, and neither model is reachable through it.

Who should act now, and who should wait
If your deadline is this quarter and your deliverable is a file, Google Veo 3.1 is the defensible choice, and the price — roughly $0.40 per second at 1080p with audio, less on the Fast and Lite tiers — is the cost of a model that has been in production long enough to be boring. Budget for the audio flag deliberately rather than discovering it later.
If your problem is a policy that has to work on hardware, Odyssey-3 is aimed at you and there is still nothing to buy. Wait for three things: a published price, a release date that is a date rather than a window, and one independent number. The 77% sim-to-real figure is Odyssey's own, and until someone outside the company reruns it, the correct posture is interest rather than planning. In the meantime the surrounding stack is the part you can build, and building it so a new model can be slotted in later is the cheapest position to hold.

