A generated two-card infographic titled "Simulator vs Renderer". The left card, headed Odyssey-3, reads "a world you can act in" with rows for output (world state + actions), consumer (robot and driving policies) and price (not published). The right card, headed FLUX 3 Video, reads "a file you can watch" with rows for output (1080p SDR MP4), consumer (viewers and editors) and price ($0.06-$0.29 per second).
Guides & Insights

Odyssey-3 vs FLUX 3 Video: A Simulator and a Renderer Are Not the Same Tool

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Odyssey-3 and FLUX 3 Video get compared because both produce video of a world that does not exist, and that is roughly where the similarity ends. FLUX 3 Video is Black Forest Labs' generation model, generally available since August 4, 2026, that turns a prompt or a start frame into up to twenty seconds of 1080p SDR MP4 with native audio, priced per second of output. Odyssey-3 is Odyssey's foundation world model, previewed September 15, 2026 and not yet publicly available, that simulates how a scene evolves under actions and attaches a learned decoder to turn that simulation into motor commands for a robot arm, a humanoid, a car, or a drone. One is a renderer: it produces observations for a person to watch. The other is a simulator: it produces a state that a program can act on. Putting them head to head is useful anyway — not to crown a winner, but because the two things they output are the two different things a lot of teams are currently trying to buy with one budget line.

The difference is what comes out of the model

The cleanest way to hold the distinction is the taxonomy World Labs published for world models, which sorts systems by which part of the perception-action loop they output. A renderer outputs observations — pixels, for human eyes, judged on visual fidelity. A simulator outputs state — a representation of the world that is geometrically and physically faithful enough for both humans and programs to compute on. A planner outputs actions. FLUX 3 Video is squarely a renderer, and an excellent one. Odyssey-3 is aiming at the simulator column with one foot in the planner column, because the action decoder is what lets the simulation drive hardware.

That is not a marketing distinction, and it has a practical test. Render a room with a video model, walk your viewpoint out of it, and turn around: the furniture may not be where you left it, because the model predicted frame by frame and never held a persistent world. Ask a simulator the same thing and the answer is supposed to be stable, because the point of the model is the state underneath the frames. Odyssey's own announcement leans on this — the reason to build a dynamics model rather than a video generator is that a policy trained inside it can be evaluated and improved, not just watched.

A generated two-column scoreboard for Odyssey-3 vs FLUX 3 Video using the same six dimension labels on both sides: output, consumer, max length, audio, price and availability. Odyssey-3 reads world state + actions, robot and driving policies, continuous rollout, none, not published, preview only. FLUX 3 Video reads 1080p SDR MP4, viewers and editors, 20s single pass, native, $0.06-$0.29 per second, GA 2026-08-04. The footer reads "Odyssey-3 vendor-reported and unreleased; FLUX 3 Video specs per Black Forest Labs."

FLUX 3 Video, specifically

Black Forest Labs shipped FLUX 3 Video on August 4, 2026, and its spec sheet is aimed at production output rather than simulation. It generates up to twenty seconds in a single pass at up to 1080p, in SDR, as an MP4, with native audio generated alongside the picture. Pricing is per second of generated video: a draft 720p tier at $0.06 per second, a standard 720p tier at $0.17, and 1080p at $0.29, with video-to-video work priced above those. Those are vendor figures.

The one benchmark number Black Forest Labs has attached to it is an internal Elo of 1,135 for text-to-video, which is a vendor-run evaluation and has not been reproduced independently — FLUX 3 Video does not currently appear on the independent video leaderboards. Treat the ranking as a claim. What is not in question is the delivery: a priced, documented, per-second API that returns a file, today.

Odyssey-3, specifically

Odyssey-3 is a preview with an architecture and no price. It is an autoregressive diffusion transformer trained on visual observation, and its distinguishing feature is the action decoder — "a learned output component attached to the world model," in Odyssey's words, trained on observation-action pairs to translate the model's internal representation into the actions a specific piece of hardware needs. Odyssey reports it driving robot arms from tens of hours of demonstrations, humanoid policies built with Flexion in real time, closed-loop driving from 20 hours of simulated data, indoor obstacle-avoiding drone flight, and extended play in Grand Theft Auto V with transfer to Red Dead Redemption 2 and Sleeping Dogs without extra policy training in those titles. The company also reports one comparative figure: simulation-trained driving policies traveled about 77% as far between safety-driver interventions as policies trained on real footage. There are no independent benchmarks, no technical report for Odyssey-3, and no release date beyond "the coming weeks." Every figure is Odyssey's own.

The contrast, dimension by dimension

What it outputs — FLUX 3 Video: an MP4 file of pixels. Odyssey-3: a simulated world state, plus motor actions through the action decoder.

Who consumes it — FLUX 3 Video: a person watching, or an editor finishing a cut. Odyssey-3: a robot controller, a driving policy, or an RL agent being trained.

Max output — FLUX 3 Video: up to 20s single pass, up to 1080p SDR, native audio. Odyssey-3: continuous rollout as long as the simulation holds, at whatever frame rate the deployment needs.

Price — FLUX 3 Video: $0.06/s draft 720p, $0.17/s 720p, $0.29/s 1080p (vendor). Odyssey-3: none published.

Availability — FLUX 3 Video: GA since 2026-08-04. Odyssey-3: preview, publicly "in the coming weeks."

Benchmarks — FLUX 3 Video: vendor Elo 1,135 T2V, unreproduced, absent from independent boards. Odyssey-3: no benchmark table, one vendor sim-to-real figure.

The price asymmetry is the honest headline

One of these has a price and the other does not, and that single fact tells you more about choosing between them than any capability claim. FLUX 3 Video has a metered cost per second of output, so a team can compute the cost of a hundred clips before generating the first one. Odyssey-3 has no published price, no API, and no date, so the cost of using it is currently unknowable — and so is the shape of the bill. World models of the kind Odyssey describes are typically billed very differently from video generation: not per second of beautiful pixels but per hour of simulation, because the consumer is a training loop that may run for millions of steps. Comparing $0.29 per second against nothing is not a comparison, and any page that presents it as one is filling a gap with a guess.

Where they would actually meet

The two are not substitutes, and the more interesting question is whether they compose. FLUX 3 Video is the better tool for anything where the deliverable is a file a human watches: a product shot, an ad variant, a social cutdown, a previsualization. Odyssey-3, if it delivers, is the better tool for anything where the deliverable is a policy that has to work in the physical world, or a scene a program has to reason about. In a film or game pipeline the plausible arrangement is not either-or — a world model generates an interactive environment, and a renderer produces the polished frames that leave the building.

There is a real limit on that composition today, and it is the SDR one. FLUX 3 Video outputs SDR MP4. Odyssey-3 outputs a simulated environment, not a graded file. Neither is an HDR finishing tool, so a pipeline that ends in a broadcast or theatrical deliverable still needs a mastering step after both of them.

Routing, and what it does and does not cover here

Neither model is routable through OrcaRouter today. FLUX 3 Video ships from Black Forest Labs' own API and several third-party platforms; Odyssey-3 is not available from anyone, because it is not released. Where a routing layer earns its place in a pipeline like this is in the layer above the video model — the language model that drafts the shot list, the vision model that checks the generated frame against the brief, the transcription model that turns a recorded voiceover into captions. Those are ordinary routed workloads, and they sit on one API across 200-plus models at provider list price with 0% markup, so a vendor price cut reaches your bill the same day rather than at the next contract renewal. The routing DSL is the piece that matters for a multi-stage pipeline: several models composed into a single call, with automatic failover when a provider degrades, so a render step does not fail because the prompt model's endpoint is having a bad afternoon. None of that hosts FLUX 3 Video or Odyssey-3, and this article does not claim otherwise.

A screenshot of Odyssey's own announcement page for Odyssey-3, dated September 15th 2026, headed "Introducing Odyssey-3: A General-Purpose Physical Intelligence", with the summary line that Odyssey-3 is a foundation world model that can power robots, drive cars, train AIs, pilot drones, and even play video games.

Which one, for whom

If your deliverable is a file — a clip, a cutdown, a previsualization — FLUX 3 Video is the one you can buy today, at a published per-second price, with native audio and twenty-second single-pass output, and the vendor's Elo claim is a reasonable bonus rather than the reason to choose it. If your deliverable is a controller that has to work on hardware, Odyssey-3 is the one aimed at your problem, and the honest advice is to wait: there is no price, no date, and no independent number yet, and the 77% sim-to-real figure is Odyssey's own and unreproduced.

The thing to watch is not which model wins, because they are not racing. It is whether Odyssey-3's action decoder generalizes across the six embodiments the announcement lists once outside groups can run it. That is the claim with the most upside if it holds, and the current evidence for it is a list of demonstrations rather than a measurement.

A screenshot of Artificial Analysis's Video Model Comparisons page, showing the Video Arena Quality Elo chart and the representative price-per-minute chart across 15 of 37 video models, including Kling 3.0 at 720p and 1080p, Wan 3.0, MiniMax H3 Max and Gemini Omni Flash.