
Odyssey-3: Odyssey's New World Model Takes the Physics-IQ Crown and Opens a Public Preview
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 111 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 55 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 378 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Odyssey-3 Pro scored 66.1 on Physics-IQ Verified's video-to-video track, and Odyssey published that number on October 8, 2026 next to the thing that actually matters to a developer: a public research preview of Odyssey-3, an autoregressive diffusion transformer built to generate an environment from a prompt and then keep predicting it in real time while someone walks through it, moves a camera, or triggers an event. Odyssey-3 is the model; Odyssey-3 Pro is the 720p tier of it, running at 1280×720 against the base model's 832×480. Both go up on the same day, into a preview you can drive yourself at experience.odyssey.systems, with a separate contact route for API access.
That pairing is what makes this more than a benchmark post. Interactive world models have been demo-ware for two years; the ones that generate a few frames of plausible motion on a fixed trajectory are not the same artifact as a model that keeps its predictions coherent after you jab the controls. Odyssey is claiming the second thing, and it is letting anyone test the claim.
What shipped, precisely
Two checkpoints, one architecture, no weights.
• Odyssey-3 — the 480p tier, 832×480 output, 16 fps on the benchmark harness, listed at $0.139 per generated video in the leaderboard's normalized cost column.
• Odyssey-3 Pro — the 720p tier, 1280×720, listed at $0.267 per video on the same basis.
• Access — a research preview in the browser with first-person navigation, third-person navigation, and independent camera movement. API access is a contact form, not a self-serve key.
• Openness — none. No weights, no license, no per-token price published. The API terms page on the same site was last updated 2026-01-22 and still governs "the prototype of Odyssey-2 API," which tells you how new the commercial surface is.
The architecture is described in Odyssey's own writeup as a multi-step video diffusion transformer extended autoregressively through teacher forcing and causal masking, with a post-training pipeline that combines distribution matching and adversarial distillation into a few-step distilled variant — the part that makes real-time interaction possible at all. Diffusion gives you fidelity; autoregression gives you a model that survives an action sequence; the distillation gives you the clock speed. All three are load-bearing, and all three are the vendor's account of its own system, not an independent one.
The benchmark, read the way the board actually reads it
![The Physics-IQ Verified leaderboard from Anates Labs and Google DeepMind, headed “Dynamic ranking of video models by how well they understand physical principles”, read 9 October 2026, with separate image-to-video and video-to-video columns. The visible ranking runs FLUX 3 [large] at 01, Odyssey-3 Pro at 02, Odyssey-3 at 03, then Cosmos-class entries at 04 to 06, Seedance 2.5 at 07, Odyssey-3 at 08, MiniMax H3 at 09, Cosmos3 Nano at 10 and MiniMax H3 Max in the low teens.](https://cms.orcarouter.ai/api/media/file/2-1765.png)
Physics-IQ Verified is run by Anates Labs with DeepMind, and it is a genuinely awkward benchmark to game: it shows models video of real physical experiments — fluid dynamics, optics, solid mechanics, magnetism, thermodynamics — and scores the continuation against what actually happened. It currently ranks 41 models from 16 labs. Odyssey-3 Pro's 66.1 is the highest raw score on the video-to-video track that Odyssey reports, and the same board's net-improvement column, which measures the gain over the track mean in percentage points, tells a subtly different story:
• Net improvement over track mean — FLUX 3 [large] leads at +12.27 pp; Odyssey-3 Pro is second at +11.15 pp; Odyssey-3 is third at +9.99 pp.
• Cost per generated video — Odyssey-3 Pro $0.267, Odyssey-3 $0.139, against FLUX 3 [large] at $0.868 and Seedance 2.5 at $2.838. (Costs are normalized to 24 fps and 1280-wide output where the leaderboard can, so they are comparable across models that do not share a frame rate.)
• Submetric leadership — Odyssey-3 takes first on the spatiotemporal submetric at 44.70, ahead of Odyssey-3 Pro at 42.97; FLUX 3 [large] takes spatial at 64.36 and weighted spatial at 53.25, with the two Odyssey entries second and fourth on spatial.
So the honest reading is not "Odyssey-3 is the best video model in the world." It is: a world model built for interaction has landed near the top of a physical-plausibility board that was previously the property of cinematic generators, and it did so at roughly a third of FLUX 3's normalized cost. Odyssey's separate claim — 54.7 for Odyssey-3 Pro on the image-to-video track, best-of-8 — is on the same board, and the same caveat applies: these are self-submitted configurations, and the leaderboard mixes entries submitted with custom prompts and entries run on the benchmark's own prompts. Odyssey appears in both columns, which is unusually transparent and also means its two rows are not measuring the same thing.

Why "world model" is not a synonym for "video model"
The distinction is the whole product argument, so it is worth stating plainly. A clip generator takes a prompt and returns a fixed object; the output is the same for everyone who asks. Odyssey-3 takes a prompt and returns a system you act inside — the preview's first-person and third-person camera modes and its ability to accept an event mid-generation are the mechanism by which it predicts what happens next given what you just did, not given what the prompt said.
That property is what makes the physical-system results in Odyssey's launch material look the way they do. The company reports that an action decoder attached to the frozen world model, trained on tens of hours of robot demonstrations, produced recovery behaviors that were never in the demonstrations — reorienting a gripper after a missed grasp, retrieving an object dropped in an unusual orientation. It reports a humanoid policy built by Flexion on top of Odyssey-3 with tens of hours of teleoperation data that continued executing under lighting changes which broke the VLA baselines it was compared against. It reports a driving policy trained on 20 hours of simulated data that drove in closed loop on real roads, traveling about 77% as far between safety-driver interventions as a policy trained on real footage. It reports drone navigation from simulated flight data and gameplay in GTA V that transferred movement to Red Dead Redemption 2 and motorcycle riding to Sleeping Dogs without further training.
Every one of those is a vendor-reported result with no independent replication, and two of them are explicitly framed by Odyssey as qualitative observations rather than measurements. But they are the right kind of evidence for the claim: the interesting part is not that the model can do a task it was trained for. It is that a frozen backbone plus a small adapter gets meaningful behavior out of a few tens of hours of data, and occasionally generalizes past it.
What is still unproven

Three things, and they are not small.
First, no independent evaluator has run Odyssey-3. The Physics-IQ entry is a submission, and the second scoring source in Odyssey's own writeup — WorldMark, on which Odyssey-3 ranks first in three of four splits — is a vendor evaluation using the benchmark's own captions and, in Odyssey's phrasing, "the mean of its 13 reported metric scores." That is the company scoring itself on someone else's dataset. It is more than most launches offer. It is not a third party.
Second, the one split Odyssey-3 does not win in that evaluation is the one closest to the marketing claim. In first-person real environments it places third at 80.6, behind Lyra 2.0 at 84.4 and AlayaWorld at 83.0. The model's edge in the same evaluation is concentrated in stylized and third-person environments — 77.2 in first-person stylized, 79.0 in third-person real, 76.3 in third-person stylized. If you are building for photoreal first-person embodiment, the ranking you should care about is the one where Odyssey-3 is not on top.
Third, availability. "Research preview" is doing real work in that sentence. There is no pricing page, no rate limit documentation, and no self-serve key. If you want this in a product, you are in a queue.
Comparing it without a second contract
Here is the practical problem with a launch like this: Odyssey-3 is the most interesting thing in video generation this week, and you still cannot call it. The models you would actually benchmark it against are a different story.
OrcaRouter does not host Odyssey-3, and there is no published price to pass through even if we did. What one key does reach is the rest of the comparison set — minimax/minimax-h3 at $0.08 per second of generated video at 768P, the Kling video line, and more than 200 models behind a single endpoint — at provider list price with 0% markup, so a vendor's price change shows up on your invoice the same day rather than at the next renewal. Automatic failover across providers means an evaluation sweep does not die because one upstream is having an afternoon, and the routing DSL lets you put several models behind one interface so a head-to-head is a config change rather than a second integration. If Odyssey opens a self-serve API, that is the shape you want to slot it into.
What would settle it
An independent run of Odyssey-3 on a harness it did not choose. A dozen practitioners with preview access describing where the environment holds together after a minute of aggressive camera work and where it does not. A price. Any of those converts this from a strong launch into a reference point.
What is already true is narrower and still notable: on the one public board that measures whether a generative model understands physics rather than whether it photographs well, a model whose purpose is interaction is sitting second and third, at a normalized cost below the model in first.
