A generated two-card infographic titled "A World to Enter or a File to Grade". The left card, headed Luma Ray 3.2 and captioned "a finished deliverable", lists ACES2065-1 EXR at 10/12/16-bit, up to 1080p with 5s and 10s clips, and no native audio. The right card, headed Odyssey-3 and captioned "a running environment", lists simulated state and continuous rollout, motor actions for hardware, and price not published.
Guides & Insights

Odyssey-3 vs Luma Ray 3.2: A World to Enter or a File to Grade

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Luma Ray 3.2 does something almost nothing else in video generation does: it hands you 10-, 12-, and 16-bit HDR frames in the ACES2065-1 EXR standard, which is the format a colourist actually wants and the reason the model exists in its current form. Odyssey-3 does something no video model does at all: it takes an action as input and predicts what the world does next, then converts that prediction into motor commands for a robot arm, a humanoid, a car, or a drone. Luma Ray 3.2, released by Luma AI on June 9, 2026, is a finishing-adjacent generation tool with a developer API. Odyssey-3, previewed by Odyssey on September 15, 2026, is an unreleased foundation world model with no price and no date. Comparing them is less a bake-off than a way of asking which of two very different outputs your pipeline is actually short of — a graded frame, or a world that responds to control.

Ray 3.2 is built around the deliverable

The Ray 3.2 release is notable less for its generation quality than for where it stops. It generates in 540p, 720p, and 1080p, in 5- or 10-second clips, in six aspect ratios covering vertical, square, and widescreen, and it accepts text, a start frame, an end frame, or an existing video for modification. The HDR path is the differentiator, and it comes with real constraints worth stating precisely: 10-, 12-, and 16-bit HDR output in ACES2065-1 EXR is available only at 720p and 1080p, and only for 5-second clips. EXR export has to be requested by enabling HDR. The standard output is MP4. Seamless looping applies to 5-second clips. Up to 16 keyframes per clip give shot-level control that sits closer to animation than to prompting, and facial performance tracking covers up to eight faces.

Two other facts shape the buying decision. Ray 3.2 is the first model in the Ray family with a full developer API — earlier Ray models were subscription-only — split into a pay-as-you-go Build tier and a reserved-throughput Scale tier with an SLA. And it does not generate synchronized audio for text-to-video or image-to-video work; audio survives only in the modify and reframe paths. Luma has run its own evaluation claiming parity with Google's Veo 3 and a lead over several rivals, but those are vendor-run comparisons with no published numeric scores, and should be read as positioning rather than measurement.

A generated two-column scoreboard for Odyssey-3 vs Luma Ray 3.2 sharing six dimension labels: output, max resolution, clip length, audio, price, availability. Odyssey-3 reads world state + actions, not a renderer, continuous rollout, none, not published, preview only. Luma Ray 3.2 reads MP4 + ACES2065-1 EXR, 1080p, 5s or 10s, no audio for text-to-video, about $0.30 per 5s at 720p, released 2026-06-09. Footer: "Odyssey-3 vendor-reported and unreleased; Luma Ray 3.2 figures per Luma and third-party listings."

Odyssey-3 is built around the controller

Odyssey-3's design centre is the opposite end of the pipeline. It is an autoregressive diffusion transformer trained on visual observation of the world, and its distinguishing component is a learned action decoder — in Odyssey's phrasing, "a learned output component attached to the world model" that translates the model's internal representation into the actions a given piece of hardware requires. Odyssey reports the same backbone controlling robot arms from tens of hours of demonstrations, driving humanoid policies built with Flexion in real time, producing closed-loop driving waypoints from a frozen backbone trained on 20 hours of simulated data, flying drones indoors around obstacles, and playing Grand Theft Auto V with transfer to Red Dead Redemption 2 and Sleeping Dogs without additional policy training there.

The single comparative number Odyssey published is that simulation-trained driving policies traveled about 77% as far between safety-driver interventions as policies trained on real footage. There is no benchmark table, no Odyssey-3 technical report, no independent evaluation, and no availability date beyond "the coming weeks." It is a preview, and it is priced at nothing because it cannot be bought.

The contrast, dimension by dimension

What comes out — Luma Ray 3.2: MP4, plus 10/12/16-bit ACES2065-1 EXR when HDR is enabled. Odyssey-3: a simulated world state plus motor actions.

Format ceiling — Luma Ray 3.2: ACES2065-1 EXR at 720p and 1080p, 5s clips only. Odyssey-3: not a file format; output is a rollout.

Resolution and length — Luma Ray 3.2: up to 1080p, 5s or 10s (10s text-to-video only). Odyssey-3: continuous, at whatever rate the deployment needs.

Audio — Luma Ray 3.2: none for T2V/I2V; preserved in modify and reframe. Odyssey-3: none; it is not a media generator.

Price — Luma Ray 3.2: per video by resolution, dynamic range, and duration, with HDR and EXR export multiplying the clip cost; third-party listings put a 5s 720p clip near $0.30 and 1080p near $1.20. Odyssey-3: unpublished.

Availability — Luma Ray 3.2: released 2026-06-09, Build and Scale API tiers. Odyssey-3: preview, "in the coming weeks."

Deliverable versus environment

The reason this pairing is instructive rather than arbitrary is that the two models terminate in different places. Ray 3.2 terminates in a file that goes to a colourist, a compositor, or a finishing house — the EXR path exists precisely so the generated image can survive being graded, composited, and re-graded without falling apart. That is a workflow decision, and it is why Ray 3.2 competes more directly with post-production tooling than with other generators.

Odyssey-3 does not terminate at all in the file sense. Its output is a simulated state that keeps running, and its consumer is a program. The test of whether it worked is not whether the frames look right but whether a policy trained inside it transfers to hardware. Those are different success criteria measured by different people, and no amount of resolution on either side makes them comparable.

Where they could touch is the previsualization-to-finishing path. A world model is the natural tool for exploring a space interactively — changing a camera move, re-blocking a scene, seeing the consequence immediately, which is exactly what a fixed generated clip cannot do. A model like Ray 3.2 is the natural tool for turning the chosen result into something deliverable in HDR. Today that path has a hole in the middle, because Odyssey-3 is not available and produces no file format a finishing pipeline consumes.

Routing around both of them

Neither Luma Ray 3.2 nor Odyssey-3 is served by OrcaRouter, and this article makes no claim that either is. What is worth noting is how much of a pipeline like this is neither model. The prompt model that drafts a shot description, the vision model that checks a frame against the brief, the transcription model that recovers dialogue from a reference track — those run constantly, at a fraction of the per-clip cost of the video model, and they are the parts that break at inconvenient hours. Running them on one API across 200-plus models at provider list price with 0% markup keeps that layer on one key and one bill, and automatic failover means a degraded prompt endpoint reroutes rather than stalling the queue behind it. When a vendor cuts a video-model price, the pass-through means that cut is live here the same day — but for the models we actually serve, which today does not include either of these two.

A screenshot of Odyssey's own announcement page for Odyssey-3, dated September 15th 2026, headed "Introducing Odyssey-3: A General-Purpose Physical Intelligence", with the summary line that Odyssey-3 is a foundation world model that can power robots, drive cars, train AIs, pilot drones, and even play video games.

What each is worth waiting for

Ray 3.2 is available now, and the reason to pick it is specific: if your deliverable passes through a colour pipeline, the ACES2065-1 EXR path is a capability most generators simply do not offer, and the resolution and duration constraints — 1080p ceiling, 5-second clips for HDR, no native audio — are the price of admission. If you need synchronized sound, this is the wrong model and you should look at one that generates it.

Odyssey-3 is worth watching for one reason and it is not any of Ray 3.2's dimensions. If one set of weights genuinely drives arms, humanoids, a car, a drone, and three games from tens of hours of data each, that is a claim about generality that no renderer makes, and it is the claim with the most upside in the current field. The problem is that the evidence for it, today, is Odyssey's own demonstration reel and a single sim-to-real figure. The thing that would change the read is not a better demo — it is a price, a date, and one number measured by somebody who does not work there.

A screenshot of Artificial Analysis's Video Model Comparisons page, showing the Video Arena Quality Elo chart and the representative price-per-minute chart across 15 of 37 video models, including Kling 3.0 at 720p and 1080p, Wan 3.0, MiniMax H3 Max and Gemini Omni Flash.