Title card for Odyssey-3 vs ByteDance Seedance 2, with a left panel headed Odyssey-3 reading “a world you steer live”, 832x480 or 1280x720 Pro, Physics-IQ +9.99 pp rank 03; and a right panel headed ByteDance Seedance 2 reading GA since Feb 12 2026, metered platform, a finished clip up to 15s, resolution not published, Physics-IQ not submitted.
Guides & Insights

Odyssey-3 vs ByteDance Seedance 2: An Environment You Act Inside Against a Clip You Watch

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put Odyssey-3 and B​yteDance Seedance 2 side by side and the first thing to notice is that they are not competing for the same request. Odyssey-3, released by Odyssey on October 8, 2026, generates an environment from a prompt and then keeps predicting it in real time while you move through it or trigger an event — it is a simulator with a camera you drive. B​yteDance Seedance 2 is a clip generator with a director's chair: text, image, audio and video references in, a finished video out, up to fifteen seconds of it, generated in one pass and delivered as a fixed object. One of them has a control input; the other has a timeline. The comparison is worth making anyway, because the two are converging on the same buyer, and because they take opposite bets on where the value of a generative video model lives.

There is one shared scoreboard and it is thinner than it looks. Seedance 2 itself does not appear on Physics-IQ Verified, the Anates Labs and DeepMind board that scores how well a model predicts real physical experiments. Seedance 2.5, the July 31, 2026 successor, does. So the only third-party number that puts anything from this family next to Odyssey-3 is a comparison across two model generations on one side, which is worth stating before anyone quotes it as a fair fight.

The two mechanisms

Odyssey-3 is an autoregressive diffusion transformer. Odyssey's own description: a multi-step video diffusion model extended autoregressively with teacher forcing and causal masking, so it continues from preceding observations and predicts future states conditioned on action inputs, then a post-training pass of distribution matching and adversarial distillation produces a few-step variant fast enough to interact with. The output is not a scene that plays; it is a scene that responds. The preview exposes first-person navigation, third-person navigation and independent camera movement — three different ways to push on the same prediction and watch what it does.

B​yteDance Seedance 2 is a unified multimodal audio-video joint generation model. It accepts text, image, audio and video as references and generates the clip in one pass, up to fifteen seconds, with audio generated jointly rather than dubbed afterward. B​yteDance's page describes images, audio and video references giving control over performance, lighting, shadow and camera movement, which is the language of production rather than simulation: you are specifying a shot, not steering a world. Released February 12, 2026, it went viral within days on clips of real actors and existing films — and that is also what produced the copyright storm, with cease-and-desist letters from Disney and Netflix, condemnation from the Motion Picture Association and SAG-AFTRA, and a March 2026 letter from two US senators asking B​yteDance to shut it down. B​yteDance said in February 2026 that it was strengthening safeguards. That history is part of the model's specification if you are shipping anything commercial.

Dimension by dimension

ByteDance's BytePlus ModelArk developer documentation in English, showing the Ark SDK quick-start with a Python client using model seed-2-0-lite-260228 against base_url https://ark.ap-southeast.bytepluses.com/api/v3, a sidebar covering Models, video generation, image generation and the Model list, and three endpoint cards for Dola Seed 2.0, Dreamina Seedance 2.5 (described as the mainline video generation model with 30s extended storytelling, multimodal references and improved generation quality) and Dola Seedream 5.0.

• What it outputs — an interactive environment you act inside, against a fixed clip of up to fifteen seconds. Everything downstream of this row changes.

• Control surface — live navigation and mid-generation events, against prompt plus image, audio and video references specified up front. One is steered during generation; the other is directed before it.

• Audio — Seedance 2 generates it jointly with the video, in the same model. Odyssey-3's launch material says nothing about audio at all.

• Resolution — Odyssey-3 runs at 832×480 with Odyssey-3 Pro at 1280×720. B​yteDance's page for Seedance 2 does not publish an output resolution, and the 1080p figure that circulates belongs to the earlier Seedance 1.0 line.

• Cost — Odyssey-3 is listed at $0.139 per generated video on Physics-IQ's normalized cost column and Odyssey-3 Pro at $0.267, both excluding prompt handling. Seedance 2 publishes no per-second or per-clip list price we could verify; its successor Seedance 2.5 is listed on the same board at $2.838 per video.

• Independent scoring — neither model's own generation has a third-party head-to-head. Seedance 2.0's published results are B​yteDance's internal SeedVideoBench-2.0; Odyssey-3's are a benchmark submission plus a vendor-run evaluation on someone else's dataset.

• Weights — closed on both sides. Odyssey has released no checkpoints, and Seedance has never been open.

• Commercial posture — Odyssey-3 is a research preview with a contact form for API access and no price list. Seedance 2 is sold through B​yteDance's own platforms and, outside China, through CapCut as Dreamina Seedance 2.0.

The one number that actually compares them

Physics-IQ Verified scores models on the net improvement in physical plausibility over the track mean, in percentage points, and it records which prompt set was used. On the benchmark's own prompts, Seedance 2.5 sits seventh at +3.59 pp, with a normalized cost of $2.838 per video. Odyssey-3, on the same prompt set, sits eighth at +2.15 pp, at $0.129 per video. Read carefully that is a 1.44-point gap in Seedance's favor and a cost difference of roughly twenty-two times against it.

Odyssey-3's headline rows use a different column. Submitted with custom prompts, Odyssey-3 scores +9.99 pp and Odyssey-3 Pro +11.15 pp — third and second on the board, behind FLUX 3 [large] at +12.27 pp. The same model family appears twice on the same leaderboard under two different prompt regimes with very different scores, which is the single most useful thing to understand about how to read this board: the number is partly a property of the model and partly a property of who wrote the prompt.

Odyssey also reports 66.1 on the video-to-video track for Odyssey-3 Pro, the highest score it claims, and 54.7 for Odyssey-3 Pro on image-to-video. Both are the vendor's own reading of a board it submitted to, not an independent run.

Two-column scoreboard headed “Odyssey-3 vs ByteDance Seedance 2 — the scoreboard” with six shared dimension labels. Odyssey-3: an environment you act inside; live navigation and mid-run events; no audio; 832x480, 1280x720 Pro; /bin/bash.139 a video, Pro /bin/bash.267; no independent scoring yet. ByteDance Seedance 2: a fixed clip up to 15 seconds; text, image, audio and video references upfront; audio generated jointly with the video; resolution not published; no per-second list price published; not on Physics-IQ. Footer: “Odyssey-3 figures vendor-reported and unreproduced; Seedance 2 has no third-party score; Seedance 2.5, a later model, is on Physics-IQ at +3.59 pp”.

Where each one actually wins

If your problem is "I need a shot," Seedance 2's design is answering exactly that: multimodal references mean you can hand it the actor's likeness, the audio bed and the camera move, and get a finished clip with sound. Nothing in Odyssey-3's preview does that, and nothing in its launch material suggests it will. The fifteen-second ceiling is a real constraint, but so is a five-second world-model rollout with no audio.

If your problem is "I need to know what happens next," the trade runs the other way. Seedance 2 cannot be asked a counterfactual — you got one clip, and to see another outcome you prompt again and get an unrelated one. Odyssey-3's entire premise is that the next state depends on the action you just took, which is what makes it usable for a driving policy, a robot arm or a training environment where the sequence matters more than the shot.

The physical-accuracy evidence favors Odyssey-3 on the mechanism even where the score column does not: the model is being evaluated on continuations after a disturbance, which is the same capability the interactive preview is selling. Seedance 2.5's +3.59 pp on the same board is a strong result for a clip generator and, notably, better than Odyssey-3's on the matched prompt set — which is why the honest summary is that B​yteDance's video model is more physically plausible than a video model has any right to be, not that Odyssey-3 has been beaten at its own game.

OrcaRouter's own model page for minimax/minimax-h3, showing the model header “text + image + video + audio”, output video, a p50 time-to-first-token of 406 ms, performance and vision/audio badges, and a Python code sample on the left with the model list and playground navigation above.

Getting to either one

Neither model is callable from here today. OrcaRouter does not host Odyssey-3 or B​yteDance Seedance 2 — Odyssey-3 has no published price for anyone to route, and Seedance is sold through B​yteDance's own platforms.

What a single OrcaRouter key does reach is the rest of the video stack: MiniMax-H3 at $0.08 per second of 768P output, the K​ling video line at $0.084 to $0.28 per request, and more than 200 models in total, all at provider list price with 0% markup, so a vendor's price movement is live on your invoice the same day rather than at renewal. Automatic failover keeps an evaluation sweep alive when one upstream degrades, and the routing DSL puts several video models behind one interface so testing a new generator is a config change rather than an integration. When either of these two opens a self-serve endpoint, that is the shape to drop it into.

What to watch

Two things would change this comparison more than any benchmark update. The first is audio: if Odyssey-3's next generation generates sound jointly the way Seedance does, the "environment you act inside" stops being a narrower category. The second is an independent harness — a third party running Seedance 2.0 and Odyssey-3 side by side on the same clips, rather than one appearing on a board the other has never been submitted to. Until then, the correct reading of Odyssey-3 against B​yteDance Seedance 2 is that they are not two candidates for the same job. They are two answers to different questions, and the question you are asking picks the model for you.