
Odyssey-3 vs Seedance 2.5: Seconds of Footage vs Hours of Simulation
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
Ask what Seedance 2.5 costs and you get an answer in dollars per second. Ask what Odyssey-3 costs and you get nothing at all — no rate card, no endpoint, no licence, because the model announced on 15 September 2026 has not been released yet and its maker has only promised it publicly "in the coming weeks." That gap looks like a maturity difference and is mostly a category difference. Seedance 2.5, released by ByteDance's Seed team on 31 July 2026, is a production engine sold by the second of finished footage to people who need footage. Odyssey-3 is a world model with an action interface, and the buyer it is aimed at does not want footage at all — it wants an environment that responds correctly when something acts inside it. The two are being pointed at some of the same problems, including robot and driving data, which is precisely why the distinction is worth drawing carefully.
Seedance 2.5 is a production engine, and the launch partners show it
The headline capability is duration. Seedance 2.5 generates thirty seconds in a single pass, double its predecessor's fifteen, with multi-round extension for longer sequences and a reported ultra-long beta mode reaching further still. That alone changes what the model is for: thirty seconds is a scene rather than a shot, and ByteDance's marketing leans on the phrase "one-take" for exactly that reason.
Around the duration sit the features that make a generation usable in a real edit. The model accepts heavy multimodal referencing — reports put it at up to thirty images, ten video clips and ten audio clips per generation, with reference video and audio each capped around thirty seconds, and audio-only references are new in this version. There are new reference modes for white-model or clay renders, green-screen plates, motion and creative referencing. Control is timestamp-level, letting a prompt target narrative, camera or motion within a specific time range, and region-level editing after generation preserves continuity rather than forcing a re-roll. Audio and video are generated jointly in one pass, so lip sync follows quoted dialogue, and more than ten languages are supported natively.
The adoption list is the tell. On launch day XCMG, XPeng, Lingchu Intelligence, Weifen Zhifei and Qiongche Intelligence confirmed partnerships, and the stated use cases run from embodied-AI training data and autonomous-driving corner cases to industrial SOP training videos, education and e-commerce advertising. Some of that is creative work. Some of it is data generation for systems that will later act on the world — which is the part that overlaps with Odyssey-3.
What a second of Seedance 2.5 costs
Pricing is published, and it is metered in tokens rather than seconds, which produces a conversion step worth doing before you commit. The figures reported for the ModelArk endpoint are $10.70 per million tokens with no video input and $6.40 per million tokens when video input is included. In practice that works out to roughly $0.51 for a five-second 16:9 clip at 480p and about $1.16 for the same clip at 720p. Those are the vendor's published rates as reported at launch; treat them as current list rather than as a measured cost per deliverable, because resolution and duration both scale it.
Availability has real edges. Consumer access runs through ByteDance's own apps, while API access splits between Volcano Engine Ark in China and BytePlus ModelArk internationally. At least one distributor states the model is not available in the United States, and sources disagree on whether the supported resolutions top out at 720p or extend to 1080p — a discrepancy worth resolving against the vendor's own documentation before you design around a resolution.
This is the kind of model where a routing layer earns its place for a specific reason: rates move. A per-token video price that changed once at launch will change again, and the difference between a cost model built on today's number and one built on last quarter's is the difference between a margin and a surprise. OrcaRouter passes provider list price through with 0% markup, so a vendor cut is live on our side the same day rather than at the next contract renewal, and automatic failover keeps a generation queue moving when one provider endpoint degrades — a scheduling concern rather than a theoretical one on video workloads. Seedance 2.5 itself is not routed by OrcaRouter; it is served through ByteDance's own platforms and its ModelArk endpoints.

Odyssey-3 is not sold by the second
Odyssey-3 has no unit of sale yet, which is the point. It is described as a foundation world model that can power robots, drive cars, train AIs, pilot drones and play video games — an autoregressive diffusion transformer trained on a large collection of visual observations, which its maker says yields a learned understanding of physics, dynamics, cause-and-effect and human behaviour. Adaptation to a new task is done by training an action decoder on observation-action pairs and keeping the pretrained world model frozen, so a small policy reads the frozen representation and emits motor commands.
The demonstrations, all vendor-reported and none independently reproduced, span six embodiments: robot arms from tens of hours of demonstrations, including recovery behaviours that were not in the data; humanoid policies built with Flexion from tens of hours of teleoperation, which Odyssey says generalise better than the VLA baselines it tested against without publishing a figure; closed-loop driving in India from twenty hours of simulated data; indoor drone flight from simulated data; Grand Theft Auto V policies that transferred to Red Dead Redemption 2 and Sleeping Dogs with no extra training, roughly two hours of GTA footage producing horseback movement in RDR2; and environment generation for training other AIs through the PROWL work.
The single hard number is 77% — simulation-trained driving policies travelled about that fraction as far between safety-driver interventions as policies trained on real footage. That is a vendor figure with no technical report and no external verification behind it.

The unit of sale decides what gets optimised
Once you see the two models as meters rather than as capabilities, their differences stop looking like one being ahead of the other.
A model billed per second of output is optimised for output that satisfies a viewer. That is not a lower ambition; it is a different target with its own hard problems. Seedance 2.5 producing thirty coherent seconds with synchronised dialogue and timestamp-level control is a genuinely difficult result, and the failure modes it is judged on — a character's face drifting, a cut that does not land, a line of dialogue out of sync — are the ones a person in an edit bay would notice.
A model asked to predict what happens next when a controller acts is optimised for something a viewer will never check: whether the dynamics hold up under a sequence of decisions, including decisions nobody demonstrated. That is why Odyssey reports recovery behaviour emerging from tens of hours of data rather than reporting fidelity, and why its only quantitative claim is measured in safety-driver interventions instead of in visual quality. A renderer can be wrong in ways an audience forgives. A simulator can be wrong in ways a controller acts on, and the second kind of error compounds over a rollout.
The overlap in declared use cases — robot training data, autonomous-driving edge cases — is where these two meet, and it is a real competition for budget rather than for benchmark position. The question a robotics team actually faces is whether to buy rendered seconds of a plausible world or to buy a dynamics model with an action interface, and the answer depends on whether the downstream consumer is a human reviewer or a training loop.
Side by side
• Unit of sale — Seedance 2.5: tokens, converting to roughly $0.51 per 5s clip at 480p and $1.16 at 720p. Odyssey-3: none published.
• What you get per unit — Seedance 2.5: finished footage, up to 30s in one pass, with synchronised audio and dialogue. Odyssey-3: a predicted next world state plus motor actions for attached hardware.
• Inputs — Seedance 2.5: text plus up to ~30 images, 10 videos and 10 audio clips. Odyssey-3: an action decoder trained on observation-action pairs, over a frozen backbone.
• Editing control — Seedance 2.5: timestamp-level targeting and region-level post-generation edits. Odyssey-3: not applicable; there is no edit surface.
• Stated partners — Seedance 2.5: XCMG, XPeng and three others at launch. Odyssey-3: Flexion on humanoid policies; no commercial customers announced.
• Availability — Seedance 2.5: shipping since 31 July 2026, subject to regional limits. Odyssey-3: announced 15 September 2026, public release "in the coming weeks," no date and no price.
Which one belongs in your stack
If the deliverable is a video a person will watch, Seedance 2.5 is a shipping product with a published price and a wide feature set, and the decision is a normal build-versus-buy one — with two caveats worth checking first: whether the resolutions you need are actually available on the endpoint you can reach, and whether your region is served at all. Its longer single-pass duration and timestamp control are the features that most change workflow, because they remove stitching steps rather than improving a frame.
If the deliverable is a trained policy or a controller, Seedance 2.5 is not a substitute regardless of how good the footage looks, and Odyssey-3 is pointed at your problem with nothing to buy. The claim worth tracking is the frozen backbone: one pretrained model reaching arms, humanoids, a car, a drone and three games through a small decoder is a statement about how much physical understanding transfers between embodiments, and if it holds up it changes what a robotics team has to collect. It has not been tested by anyone outside Odyssey, and there is no technical report, no licence and no price.
What to watch is unglamorous: for Seedance 2.5, a stabilised price and a clear regional availability statement; for Odyssey-3, a repository, a licence, and a third party reproducing the 77%. Until the second list exists, the difference between these two models is not quality. It is that one of them is a product, and the other is an argument.

OrcaRouter keeps 200+ models behind a single API key, so the video model you pick is not the only one you can call.
