
Odyssey-3 vs Wan 3.0: A World Model Meets a Video Factory
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 111 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 55 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 347 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 377 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Odyssey-3 and Wan 3.0 both get filed under "video model," and that shared label is carrying more weight than it can bear. Odyssey-3 is a learned dynamical system — a simulator that builds an environment from a prompt and predicts how that environment changes as someone moves through it or triggers an event — released by Odyssey on 8 October 2026 as a research preview. Wan 3.0 is a production video generator from Alibaba Cloud, generally available since its 24 August 2026 launch, billed by the second, and already scored on two independent leaderboards. One of these wants to be a place you put a camera. The other wants to be the thing that renders the footage you publish.
Put the two launch pages side by side and the split is immediate. Odyssey leads with physical accuracy — how well its model understands what objects do when they meet. Alibaba leads with throughput, duration and reference control: 30-second takes, native audio, up to ten images and five videos supplied per generation. Neither is a worse version of the other. They are answers to different questions, and the answers change what a team should do this week.
One framing note before the numbers. Most figures in this article come from the vendors' own launch material and have not been reproduced by anyone outside those labs. Where a number comes from an independent board rather than a press page, the text says so. That distinction matters more than usual here, because the two models publish wildly different amounts of checkable information — one sells by the second on an open API, the other has not published a price at all.
What each one actually is
Odyssey's own description of Odyssey-3 is unusually specific for a launch post, and worth quoting rather than rounding off: it is "a learned dynamical system, implemented as an autoregressive diffusion transformer, that predicts how objects move and interact through space and how situations evolve over time." The thing it returns is not a clip in the ordinary sense. It is a state that keeps evolving while it is driven. The model ships in two sizes — Odyssey-3 at 832×480 and Odyssey-3 Pro at 1280×720 — and Odyssey says its preview supports first-person and third-person navigation alongside independent camera movement. The build also includes a few-step distilled variant, produced through a combination of distribution-matching and adversarial distillation, that the vendor describes as real-time interactive.
Wan 3.0 is a generator on the established model: a prompt and one or more references go in, a finished piece of video comes out. Alibaba's headline capabilities are a 30-second single-run generation, native audio, and reference inputs spanning text, image, video and audio — plus documents in doc, xls, ppt, pdf, md, txt, key, pages and numbers formats, up to 100 MB and 50 pages. Editing is instruction-based, meaning a generation can be modified in visuals, plot or dialogue without regenerating from scratch. Alibaba also states plainly that audio quality and on-screen text accuracy still need work, which is a rarer admission in a launch post than it should be and a useful one to have on the record.
• What the output is — a simulated environment you act inside (Odyssey-3) versus a rendered clip you publish (Wan 3.0) • Advertised sizes — 832×480, or 1280×720 on Odyssey-3 Pro, versus 480p, 720p and 1080p • Single-run length — as long as the environment keeps being driven, versus 30 seconds per generation • Audio — not a headline capability on either size, versus native audio with vendor-flagged quality limits • Inputs — a prompt that defines the environment, versus text, image, video, audio and documents up to 100 MB and 50 pages • Editing model — change the state and re-simulate, versus instruction-based edits to a finished generation • Documented API — none published, versus open and metered since 24 August 2026 • Published price — none, versus 0.3 / 0.6 / 1.2 yuan per second at 480p / 720p / 1080p
What one second of output costs, and why the units do not line up
Wan 3.0 is priced in the most legible unit available to a buyer: time. Alibaba's rates are 0.3 yuan (about $0.05) per second of generated video at 480p, 0.6 yuan (about $0.10) at 720p, and 1.2 yuan (about $0.20) at 1080p. A full 30-second run therefore lands near 9 yuan ($1.50) at 480p, 18 yuan ($3.00) at 720p, or 36 yuan ($6.00) at 1080p. Those are the August launch rates; the launch promotion that cut them by 30% ran only through 23 September, so the numbers above are the ones a buyer meets today. Providers that resell Wan 3.0 typically quote per-minute figures instead, which is worth checking against the per-second arithmetic before any budget is signed off.
Odyssey-3 does not have a comparable number, because Odyssey has not published a price, a rate card, a licence or API documentation. What exists instead is a cost figure inside a third-party benchmark, computed on an assumption the vendor states openly: $1 per MI355X GPU-hour. On that basis Physics-IQ Verified lists Odyssey-3 Pro at $0.267 per generated video and Odyssey-3 at $0.139, both normalized to 24 FPS and 1280-wide output. The same board separately notes that prompt rewriting through Odyssey's API is estimated at $0.01 per submitted video. Read that carefully — it is a modelling basis, not a price. It tells you what a simulation would cost at the vendor's assumed hardware rate, not what Odyssey will invoice.
Those two facts point in opposite directions for a buyer. Wan 3.0 is a metered line item whose worst case can be computed before anything is rendered. Odyssey-3 has a lower modelled cost per video and no invoice to attach it to — which is precisely the stage where a procurement conversation stops. For a comparison to be fair, the reader has to hold both halves: a real rate with a real bill, and an estimated rate with nothing behind it yet.

The one scoreboard they share, and the name that is missing from it
Physics-IQ Verified, a dynamic ranking from Anates Labs and DeepMind of how well video models handle physical principles, is the only independent board this pairing has in common — and it is where the comparison gets genuinely interesting, because it is also where it breaks down.
Odyssey's launch post states that Odyssey-3 Pro "sets a new state of the art on Physics-IQ Verified's benchmark, and ranks 1st in 3 of WorldMark's 4 categories." The WorldMark half of that checks out on the numbers Odyssey publishes: first place in First-Person Stylized (77.2), Third-Person Real (79.0) and Third-Person Stylized (76.3), with First-Person Real second at 80.6 behind AlayaWorld's 83.0 and Lyra 2.0's 84.4. Every one of those is a vendor-reported figure from Odyssey's own evaluation run.
The state-of-the-art half is where a reader should slow down. On the version of the Physics-IQ Verified board available now, ranked by net improvement in percentage points above the track mean, Black Forest Labs' FLUX 3 [large] sits at rank 01 with +12.27 pp. Odyssey-3 Pro is rank 02 at +11.15 pp, and Odyssey-3 is rank 03 at +9.99 pp. So the vendor's "new state of the art" is, on the board the vendor itself cites, second place — either because the claim predates a board update or because it is scoped more narrowly than the sentence suggests. Either way, the honest reading is that Odyssey-3 Pro is a top-three physics model, not an outright leader. FLUX 3 [large] also carries a much higher $0.868 cost per video against Odyssey-3 Pro's $0.267, which is the trade-off the board's cost view exists to expose.
Odyssey does hold one outright first place in the submetric breakdown: Spatiotemporal, at 44.70, ahead of Odyssey-3 Pro at 42.97 and Physis-Lang (Cosmos3 Super) at 41.57. On Spatial it is fourth (59.17) behind FLUX 3 (64.36), Odyssey-3 Pro (61.78) and Physis-Lang (59.87); on Weighted Spatial and MSE it is fourth and third respectively. That pattern — outright leadership in how well motion holds together over time, second tier in single-frame spatial fidelity — is a coherent signature of a model built to simulate evolution rather than to render a beautiful still.
Now the missing name. Wan 3.0 does not appear on Physics-IQ Verified at all. The only Wan entries are Wan 2.2 14B at rank 17, −6.64 pp, and Wan 2.2 5B at rank 22, −11.13 pp — both below the track mean, both open-source, both a generation behind. The board's own scope note says the leaderboard "includes benchmarked models and models known from preprints or announcements, even if public access is unavailable or no release date is confirmed," so Wan 3.0's absence is not evidence that it scores badly. It is evidence that nobody has measured it on this axis. The practical consequence is blunt: any claim that Odyssey-3 beats Wan 3.0 on physical plausibility is unverified in both directions, and any claim that it does not is equally unverified. Wan 3.0's independent record lives elsewhere — it placed #2 overall on OpenArt Arena and #1 in Video Editing with Audio on Artificial Analysis — on boards that measure perceived quality and edit quality, not physics.
What you can call today, and what you can only ask about
Wan 3.0 has been callable since 24 August 2026, when Alibaba Cloud's launch turned the metered, application-gated beta into a fully open API. The model is reachable through the vendor's own API and several third-party platforms, all of them metered, with the per-second rates above. If a team needs generated video this afternoon, that is a real option and it comes with a bill.
Odyssey-3 does not offer that. Odyssey's page says the research preview is "available now" and invites physical-AI developers to get in touch — which describes a demo and a conversation, not an endpoint behind a key. There is no published price, no licence, no API documentation and, as the section above shows, no independent evaluation. A developer reading the launch post cannot today compute what a production build on Odyssey-3 would cost, or check whether the physics claims hold outside the vendor's own harness. That is normal for a launch-day world model, and it is also the single most important fact in this comparison.
Since neither model is on our catalogue — the OrcaRouter model list carries no Odyssey-3 and no Wan 3.0 — the useful routing point is the layer that sits above them. Odds are a video build does not call one model. It calls a prompt-rewriting model, a video model, sometimes an audio model, and it re-calls all of them every time a vendor changes a price or a rate limit. On OrcaRouter those sit behind one API covering 200-plus models at provider list price with 0% markup, so a per-second price cut from a video vendor is live here the same day rather than waiting on a config change, and automatic failover means one model's outage does not stall a render queue. The video models we do serve — MiniMax H3, Kling 3.0, Kling Video O1 and others — are on that same key, which is what makes the substitution cheap when a Wan 3.0 preview or an Odyssey-3 trial finally does open up.

Which one belongs in your stack
The choice is not close, because there is no overlap in what the two produce. The honest decision rule is about the artifact you need at the end.
Choose Wan 3.0 if the deliverable is footage a person watches: an ad, a shot insert, a social cut, a teaching segment. Thirty seconds per run with native audio, reference control across images, video and audio, instruction-based editing and documents as input is a production feature list, and the metered API means the cost of a test is knowable in advance. It is a generation, not a simulation — the model's job ends when the clip does. Budget from the vendor's per-second rates, and treat the vendor's own caveats about audio and on-screen text as design constraints rather than footnotes.
Choose Odyssey-3, when you can get it, if the deliverable is a behaviour rather than a file: an environment a policy learns in, a scene whose response to an action has to stay consistent over a long rollout, a navigation or manipulation task where the camera is part of the problem. That is where the outright first place in Spatiotemporal plausibility and the $0.139–$0.267 modelled cost per video actually matter. What is missing is everything a production team needs to commit — a price, a licence, an endpoint and an outside measurement — and none of those are things a demo can stand in for.
The thing to watch is specific and close at hand. If Wan 3.0 appears on Physics-IQ Verified, the physics question becomes answerable in one direction. If Odyssey publishes a rate card or an API, the cost question resolves in the other. Until one of those happens, this comparison has a strange shape: a production system with a published price and no physics score, next to a simulator with published physics scores and no price. Anyone who tells you which is better today is guessing about at least one of those two numbers.
![Capture of the Physics-IQ Verified leaderboard from Anates Labs and DeepMind, ranked by net improvement in percentage points above the track mean, showing FLUX 3 [large] first at +12.27 pp, Odyssey-3 Pro second at +11.15 pp and Odyssey-3 third at +9.99 pp, with Wan 2.2 14B at rank 17 and Wan 2.2 5B at rank 22 and no entry for Wan 3.0.](https://cms.orcarouter.ai/api/media/file/4-1741.png)
