Odyssey-3 与 Wan 3.0 对比标题卡:Odyssey-3 为 2026 年 10 月 8 日发布的研究预览版,未公布价格或 API,Physics-IQ Verified 净提升 +11.15 个百分点,排名第 02;Wan 3.0 为按量计费 API,自 2026 年 8 月 24 日起正式可用,720p 每秒 0.6 元。
Guides & Insights

Odyssey-3 对比 Wan 3.0:当世界模型遇上视频工厂

作者

Elias Hawthorne

发布日期

最新模型 · 20查看全部模型 →
基准测试:Artificial Analysis · 每日更新
返回全部文章

Odyssey-3 and Wan 3.0 both get filed under "video model," and that shared label is carrying more weight than it can bear. Odyssey-3 is a learned dynamical system — a simulator that builds an environment from a prompt and predicts how that environment changes as someone moves through it or triggers an event — released by Odyssey on 8 October 2026 as a research preview. Wan 3.0 is a production video generator from A​libaba Cloud, generally available since its 24 August 2026 launch, billed by the second, and already scored on two independent leaderboards. One of these wants to be a place you put a camera. The other wants to be the thing that renders the footage you publish.

Put the two launch pages side by side and the split is immediate. Odyssey leads with physical accuracy — how well its model understands what objects do when they meet. A​libaba leads with throughput, duration and reference control: 30-second takes, native audio, up to ten images and five videos supplied per generation. Neither is a worse version of the other. They are answers to different questions, and the answers change what a team should do this week.

在给出数字之前,先说明一点背景。本文中的大多数数字都来自厂商自己的发布材料,尚未被这些实验室之外的任何人复现。如果某个数字来自独立评测榜单而非新闻稿页面,文中会予以注明。这一区别在这里比通常更为重要,因为这两个模型公开的可核查信息量差异极大——一个通过开放 API 按秒计费出售,另一个则根本没有公布价格。

每一个到底是什么

Odyssey 自己对 Odyssey-3 的描述,对一篇发布文章而言具体得不同寻常,值得直接引用而非粗略概括:它是“一个学习得到的动力系统,以自回归扩散 transformer 实现,能够预测物体如何在空间中移动和相互作用,以及各种情境如何随时间演变”。它返回的东西并非通常意义上的片段。它是一种在被驱动时持续演化的状态。该模型提供两种规格——Odyssey-3 为 832×480,Odyssey-3 Pro 为 1280×720——Odyssey 表示其预览版支持第一人称和第三人称导航,以及独立的摄像机运动。该版本还包含一个少步蒸馏变体,通过分布匹配与对抗蒸馏相结合的方式生成,厂商称其具备实时交互能力。

Wan 3.0 is a generator on the established model: a prompt and one or more references go in, a finished piece of video comes out. A​libaba's headline capabilities are a 30-second single-run generation, native audio, and reference inputs spanning text, image, video and audio — plus documents in doc, xls, ppt, pdf, md, txt, key, pages and numbers formats, up to 100 MB and 50 pages. Editing is instruction-based, meaning a generation can be modified in visuals, plot or dialogue without regenerating from scratch. A​libaba also states plainly that audio quality and on-screen text accuracy still need work, which is a rarer admission in a launch post than it should be and a useful one to have on the record.

• 输出是什么——你在其中行动的模拟环境(Odyssey-3),对比你发布的渲染片段(Wan 3.0) • 宣称尺寸——832×480,或在 Odyssey-3 Pro 上为 1280×720,对比 480p、720p 和 1080p • 单次运行时长——只要环境持续被驱动就可一直运行,对比每次生成 30 秒 • 音频——在两种尺寸上都不是核心卖点,对比原生音频,但厂商标注了质量限制 • 输入——一段定义环境的提示词,对比文本、图像、视频以及最大 100 MB、最多 50 页的文档 • 编辑方式——改变状态并重新模拟,对比对已完成的生成结果进行基于指令的编辑 • 有文档的 API——未发布任何文档,对比自 2026 年 8 月 24 日起开放并按量计费 • 公布价格——无,对比在 480p / 720p / 1080p 下每秒 0.3 / 0.6 / 1.2 元

一秒的输出需要付出什么代价,以及为什么单位对不上

Wan 3.0 is priced in the most legible unit available to a buyer: time. A​libaba's rates are 0.3 yuan (about $0.05) per second of generated video at 480p, 0.6 yuan (about $0.10) at 720p, and 1.2 yuan (about $0.20) at 1080p. A full 30-second run therefore lands near 9 yuan ($1.50) at 480p, 18 yuan ($3.00) at 720p, or 36 yuan ($6.00) at 1080p. Those are the August launch rates; the launch promotion that cut them by 30% ran only through 23 September, so the numbers above are the ones a buyer meets today. Providers that resell Wan 3.0 typically quote per-minute figures instead, which is worth checking against the per-second arithmetic before any budget is signed off.

Odyssey-3 没有可比的数字,因为 Odyssey 尚未公布价格、价目表、许可证或 API 文档。取而代之的是第三方基准测试中的一个成本数字,它基于厂商公开声明的假设计算得出:每 MI355X GPU 小时 1 美元。在此基础上,Physics-IQ Verified 将 Odyssey-3 Pro 列为每个生成视频 0.267 美元,Odyssey-3 为 0.139 美元,两者均按 24 FPS 和 1280 宽输出归一化。同一榜单还另行指出,通过 Odyssey 的 API 进行提示词改写,估计为每个提交视频 0.01 美元。请仔细读这句话——这是一种建模依据,而不是价格。它告诉你的是在厂商假设的硬件费率下,一次模拟将花费多少,而不是 Odyssey 会开出多少账单。

这两个事实对买家而言指向相反的方向。Wan 3.0 是一项按量计费的明细项,其最坏情况可以在渲染任何内容之前就算出来。Odyssey-3 每段视频的建模成本更低,却没有任何账单可以与之挂钩——而恰恰是在这一步,采购层面的对话就进行不下去了。要让比较公平,读者必须同时握住这两半:一个有真实账单支撑的真实费率,以及一个背后尚无任何依据的估算费率。

Two-column scoreboard comparing Odyssey-3 and Wan 3.0 across six dimensions: what you get, advertised sizes, single-run length, published price, documented API, and Physics-IQ Verified. Odyssey-3 returns a simulated environment, ships at 832x480 or 1280x720 on Odyssey-3 Pro, runs as long as it is driven, has no published price or documented API, and scores +11.15 pp at rank 02 on Physics-IQ Verified; Wan 3.0 returns a rendered clip, ships at 480p, 720p and 1080p, renders 30 seconds per run, costs 0.3 / 0.6 / 1.2 yuan a second, has been open and metered since 24 August 2026, and is not listed on that board.

他们共有的那唯一一块记分牌,以及上面缺失的那个名字

Physics-IQ Verified 是 Anates Labs 与 DeepMind 推出的一个动态排名,用来衡量视频模型处理物理原理的表现,也是这对组合唯一共同拥有的独立榜单——而比较真正变得有趣的地方正在于此,因为也正是在这里,它开始失效。

Odyssey 的发布帖称,Odyssey-3 Pro“在 Physics-IQ Verified 的基准测试中刷新了最先进水平,并在 WorldMark 的 4 个类别中位列 3 个类别的第 1 名”。其中关于 WorldMark 的那一半与 Odyssey 公布的数据相符:在第一人称风格化(77.2)、第三人称真实(79.0)和第三人称风格化(76.3)中排名第一,而第一人称真实以 80.6 位居第二,落后于 AlayaWorld 的 83.0 和 Lyra 2.0 的 84.4。这些全都是来自 Odyssey 自家评测运行的厂商报告数据。

最先进的部分是读者应该放慢速度的地方。在目前可用的 Physics-IQ Verified 榜单版本上,按高于赛道均值的净提升百分点排名,Black Forest Labs 的 FLUX 3 [large] 以 +12.27 个百分点位列第 01 名。Odyssey-3 Pro 以 +11.15 个百分点位列第 02 名,Odyssey-3 以 +9.99 个百分点位列第 03 名。因此,厂商所谓的“新的最先进技术”,在厂商自己引用的榜单上,是第二名——要么是因为该声明早于榜单更新,要么是因为其范围比这句话所暗示的要窄。无论如何,诚实的解读是,Odyssey-3 Pro 是一个排名前三的物理模型,而不是绝对的领先者。FLUX 3 [large] 的每视频成本也高达 $0.868,而 Odyssey-3 Pro 为 $0.267,这正是榜单的成本视图旨在揭示的权衡。

Odyssey 确实在细分指标中拿下了一个绝对第一:时空性,得分 44.70,领先于 Odyssey-3 Pro 的 42.97 和 Physis-Lang(Cosmos3 Super)的 41.57。在空间性上它位列第四(59.17),落后于 FLUX 3(64.36)、Odyssey-3 Pro(61.78)和 Physis-Lang(59.87);在加权空间性和 MSE 上,它分别位列第四和第三。这种模式——在运动随时间保持连贯方面绝对领先,而在单帧空间保真度上处于第二梯队——是一个为模拟演化而非渲染精美静帧而构建的模型的一致特征。

现在来说说缺失的名字。Wan 3.0 在 Physics-IQ Verified 上完全没有出现。唯一的 Wan 条目是排名第 17 的 Wan 2.2 14B,−6.64 pp,以及排名第 22 的 Wan 2.2 5B,−11.13 pp——两者都低于赛道均值,两者都是开源的,两者都落后一代。该榜单自己的范围说明写道,排行榜“包含已基准测试的模型以及从预印本或公告中得知的模型,即使公共访问不可用或未确认发布日期”,因此 Wan 3.0 的缺席并不是它得分很差的证据。它证明的是没有人在这个轴上测量过它。实际后果很直白:任何声称 Odyssey-3 在物理合理性上击败 Wan 3.0 的说法在两个方向上都是未经证实的,任何声称它没有击败的说法同样未经证实。Wan 3.0 的独立记录在别处——它在 OpenArt Arena 上总排名第 2,在 Artificial Analysis 上的 Video Editing with Audio 中排名第 1——这些榜单衡量的是感知质量和编辑质量,而不是物理。

今天你可以怎么称呼,以及你只能询问什么

Wan 3.0 has been callable since 24 August 2026, when A​libaba Cloud's launch turned the metered, application-gated beta into a fully open API. The model is reachable through the vendor's own API and several third-party platforms, all of them metered, with the per-second rates above. If a team needs generated video this afternoon, that is a real option and it comes with a bill.

Odyssey-3 并不提供这些。Odyssey 的页面称该研究预览版“现已可用”,并邀请物理 AI 开发者与其联系——这描述的是一次演示和一次对话,而不是一个藏在密钥背后的端点。没有公布价格,没有许可证,没有 API 文档,而且正如上一节所示,也没有独立评估。开发者阅读发布博文后,今天无法计算出基于 Odyssey-3 构建生产版本的成本,也无法核实那些物理声明在厂商自家测试框架之外是否成立。对于发布当日的世界模型而言,这很正常,而这也是本次对比中最重要的一个事实。

由于这两款模型都不在我们的目录中——OrcaRouter 的模型列表里既没有 Odyssey-3,也没有 Wan 3.0——真正有用的路由节点是位于它们之上的那一层。一个 API 以供应商标价、0% 加价覆盖 200 多个模型,因此视频供应商的每秒价格下调当天就能在这里生效,而无需等待配置变更,并且自动故障转移意味着某个模型中断不会让渲染队列停摆。我们实际提供的视频模型——MiniMax H3、Kling 3.0、Kling Video O1 等——都使用同一个密钥,正因如此,当 Wan 3.0 预览版或 Odyssey-3 试用版终于开放时,替换成本才会如此之低。

Capture of Odyssey's own 'Meet Odyssey-3' launch page, dated October 8th 2026, describing Odyssey-3 as a learned dynamical system implemented as an autoregressive diffusion transformer and inviting physical-AI developers to get in touch about the research preview.

哪一个适合你的技术栈?

这个选择并不接近,因为两者所产生的结果毫无重叠。诚实的决策规则取决于你最终需要的那件成果。

如果交付物是给人看的素材——一条广告、一个插入镜头、一段社交媒体剪辑、一个教学片段——就选 Wan 3.0。每次运行三十秒并带原生音频,跨图像、视频和音频的参考控制,基于指令的编辑,以及以文档作为输入,这是一份制作功能清单;而按量计费的 API 意味着测试成本可以提前预知。它是生成,而不是模拟——片段结束之时,模型的工作也就结束了。按照供应商的每秒费率来编预算,并把供应商自己关于音频和屏幕文字的提醒当作设计约束,而不是脚注。

选择 Odyssey-3——当你能拿到它的时候——前提是交付物是一种行为而非一个文件:一个策略在其中学习的环境,一个对动作的响应必须在长时间推演中保持一致的场景,一个相机本身就是问题一部分的导航或操作任务。正是在这里,时空合理性上的绝对第一名,以及每段视频 0.139–0.267 美元的建模成本,才真正重要。缺失的,是生产团队要下定决心投入所需的一切——一个价格、一份许可证、一个端点,以及一项外部测量——而这些都不是一场演示所能替代的。

值得留意的事情很具体,也近在眼前。如果 Wan 3.0 出现在 Physics-IQ Verified 上,物理问题就能朝一个方向得到答案。如果 Odyssey 发布价目表或 API,成本问题就会在另一个方向上得到解决。在这些事情之一发生之前,这项比较呈现出一种奇怪的形态:一个已公布价格、却没有物理评分的生产系统,旁边是一个已公布物理评分、却没有价格的模拟器。今天任何告诉你哪个更好的人,至少都是在猜那两个数字中的一个。

Capture of the Physics-IQ Verified leaderboard from Anates Labs and DeepMind, ranked by net improvement in percentage points above the track mean, showing FLUX 3 [large] first at +12.27 pp, Odyssey-3 Pro second at +11.15 pp and Odyssey-3 third at +9.99 pp, with Wan 2.2 14B at rank 17 and Wan 2.2 5B at rank 22 and no entry for Wan 3.0.