Hero card for Runway Solaris headed with the model name and the line "The interface is the video, and there is no code underneath", above a strip of five video frames each holding one abstract interface element, joined by an arrow labelled "each interaction conditions the next frame" and footed with "Research preview - no public API, no published price".
Engineering & Research

Runway Solaris: The Interface World Model That Replaces Code With Video

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Runway Solaris and GWM Worlds 2 are the two models Runway shipped at the turn of September 2026, and together they make one specific bet: that the next useful thing a video model does is not render a clip you watch, but hold a world you act inside. Solaris, announced on August 31, generates software interfaces frame by frame — a click or a drag conditions the next frame, and the picture is the application, with no HTML, CSS, JavaScript or DOM underneath it. GWM Worlds 2, announced September 3, runs the same trick on environments: continuous 720p at 24 frames per second with 48 kHz audio, in sessions that do not have to end. On September 10 Runway published the research behind both, and that post is more useful than the launch coverage, because it explains how a diffusion video model gets fast enough to feel live. What it does not explain is how to get access. Neither model has a public API, a published price, or downloadable weights.

Where Solaris and GWM Worlds 2 sit in Runway's line

Both models descend from the same place. Runway's Gen-4.5 video model is the base artefact; the GWM-1 general world model programme announced in December 2025 supplied the "world model" framing; Runway Characters, earlier in 2026, was the first public outing of the idea that you could prompt a video model and get something that responded rather than something that played.

Solaris takes that lineage somewhere narrower and stranger. Runway calls it the first interface world model, and the claim is not that it generates a picture of a UI but that the generation is the UI. A language model decides what should change when you interact; the world model renders the consequence. Runway's launch post frames conventional app development as "lossy compression" — a designer's intent compressed into code, then decompressed by a browser, with everything you can do limited to what a developer pre-programmed.

GWM Worlds 2 is the broader sibling. It generates interactive environments rather than interfaces, and its control surface is a format Runway calls WorldPrompt: a persistent "genesis" prompt that fixes environment, layout, materials, lighting, ambience, subjects, physical rules and camera perspective, plus a first frame to anchor the look; then a timestamped event stream of free-form text actions addressed to a subject or the scene, with start and end times so actions can overlap.

Solaris — interface world model, built on Gen-4.5, first in Runway's Interface World Models series, announced 31 August 2026

GWM Worlds 2 — second general world model, 720p at 24 fps, 48 kHz audio including generated speech with synchronised lip movement, announced 3 September 2026

Sessions — open-ended for GWM Worlds 2 rather than fixed-length clips; 720p sessions for Solaris

Access — Solaris: early-access request form. GWM Worlds 2: contact form. Neither has published pricing or a public endpoint

Stated latency — Solaris is described as running at interactive speeds, under 500 ms per frame

Independent benchmarks — none. Everything published so far originates with Runway

Single-model scoreboard for Runway Solaris with six rows reading Base model Gen-4.5 video model, Output 720p and under 500 ms per frame (stated), Access Early-access request form, Price Not published, Independent benchmarks None as of 11 Sept 2026, and Text rendering Not solved, with a footer noting all figures are stated by Runway and that no independent benchmarks or latency numbers have been published.

The three-stage conversion that buys the half-second

The September 10 research post, "Towards Instant Video Generation", is the first time Runway has described the machinery in any detail, and the shape of it is familiar from the diffusion-distillation literature even if the application is not. The starting point is a standard bidirectional video model — Gen-4.5 — where each generation step conditions on an initial first frame and a caption, and previously generated latents are kept in context. That model is accurate and slow: flow matching wants many denoising steps per frame, and a frame you have to wait for is not an interface.

The conversion happens in three moves. First, the architecture is turned into a temporally causal, frame-by-frame autoregressive generator — each frame depending only on what came before it, which is what makes streaming possible at all. Second, distribution matching distillation collapses the many denoising steps down to a few, which is where the speed comes from. This runs in two phases, and Runway is unusually clear about why both are needed. The off-policy phase has the student predict next states from ground-truth context while a frozen bidirectional teacher demonstrates good generations and a critic pushes the student toward the teacher's distribution — cheap, because there are no rollouts, but insufficient on its own, since in video an early mistake becomes the context for the next frame and pushes the model out of distribution fast. The on-policy phase therefore performs rollouts during training, so each generated latent becomes the context for the next, exactly as it will at inference. Runway says most of the real gains come from this phase, and that a curriculum of increasing sequence length beats training on long sequences from the start.

Third, the model is trained on its own outputs to stop visual drift over a long session. The three stages together are what let Runway use the word "interactive"; no latency figure, frame rate or throughput number is attached to any of them. "Good latency" is as specific as the post gets. What it does say is that the bottleneck moves from training to inference, because every frame has to leave the model fast enough to play back on shared hardware — and that cost per output at a given quality is what decides which use cases are viable at all.

Screenshot of Runway's research post "Towards Instant Video Generation", dated September 10, 2026, showing the opening paragraph that names Solaris and GWM Worlds 2 and a timeline diagram contrasting traditional staged video models with autoregressive causal diffusion.

The 61-to-24 study, read carefully

The number that travelled furthest from the Solaris announcement was a preference score, so it is worth being precise about what was measured. Runway ran a study with 250 participants and roughly 7,500 pairwise judgments across 30 interaction examples. Participants preferred Solaris to a coded interface generated by Claude Opus 5 61 percent of the time for following interaction instructions, against 24 percent for the coded version; on naturalness of behaviour the split was 71 to 21.

Three things that number is not. It is not an independent result — Runway designed it, ran it and reported it, and no third party has reproduced it. It is not a measure of whether the interface worked: instruction-following and behavioural naturalness are judgements about how it looked and felt, not whether a task was completed correctly. And it is not a comparison against Claude Opus 5 itself, which is a language model; it is a comparison against interfaces that a language model wrote as code. That distinction is the whole bet in one line — the claim is not that Solaris is smarter than an LLM, but that skipping the code layer produces a more natural result.

What the model still cannot do

Runway is more candid about the gaps than launch posts usually are, and the gaps are structural rather than cosmetic.

Text rendering — legible, stable text is genuinely hard for a model built on video-generation technology, and interfaces are made of text. Runway has floated a hybrid: let a conventional image model render text-heavy views during pauses in interaction

Accessibility — a generated interface has no DOM, so screen readers and accessibility APIs have nothing to attach to. This is not a bug to be filed; it is a property of the approach

Long sessions — coherence degrades over extended open-ended interaction, which is precisely the regime an interface lives in

Trust and grounding — a confident, wrong rendering is worse than no rendering, and Solaris depends on its starting frame being anchored in verified reference material

Resolution — 720p is generous for a video clip and cramped for data-dense desktop software

For GWM Worlds 2 the acknowledged limits are a different set: quick camera rotations degrade detail, texture and geometry; long-term memory is incomplete, so a room you leave and re-enter may not hold the same objects; image references stop at the first frame or a prefilled clip; and stateful dialogue usually needs an external harness to keep track of anything.

Why Runway thinks this is agent-training infrastructure

The most interesting argument around Solaris is not about interfaces at all. Because it can generate an interface that has never existed before — rather than retrieving a familiar layout — it could be used to train computer-use agents that have to generalise instead of memorise. That targets a real weakness: agents trained on a handful of well-known app layouts tend to be brittle the moment the layout changes. GWM Worlds 2 makes the same pitch for embodied work, and ships an explicit agent-control mode where an AI agent plans and executes tasks inside the world.

Screenshot of Runway's research page "Introducing GWM Worlds 2", carrying a Research Preview label, dated September 3, 2026, and describing interactive worlds generated in real time at continuous 720p and 24 fps with audio at 48,000 Hz.

It is a good argument and an incomplete one. An environment that cannot reliably remember that the lamp it switched on is still on is a poor place to train anything that depends on state, which is most of what agents do. Runway says as much in its own limitations section. The honest reading is that this is a promising direction with a state-tracking problem sitting directly across it.

What you can actually call today

Solaris is a research preview. Access runs through an early-access request form aimed at partners, and Runway has not announced a public launch date. There is no Solaris endpoint in Runway's own API reference, no pricing, and no released weights. GWM Worlds 2 is gated behind a contact form while Runway validates it with selected partners. If you are waiting to build on either one, you are waiting.

That matters for how you read the rest of the coverage, because a great deal of it describes Solaris in the present tense of a shipped product. It is not one yet, and Runway does not claim otherwise.

None of this is something OrcaRouter can help you call. Our catalogue is text and multimodal models — 195 of them at the time of writing, and nothing world-model-shaped among them. What we do route is the other half of the stack. If the interesting future here is an agent driving a generated interface, something still has to be the model deciding what to do, and that side is available today: one key across a catalogue of more than 200 models, with 0% markup on provider list price, automatic failover when a provider degrades, a routing DSL for composing several models into one call, and model fusion when you want a panel answering together instead of one. You can try a new model on a production path without betting the path on it.

What to watch

The question that decides whether Solaris becomes a product is not whether generated interfaces look impressive — they demonstrably do — but whether Runway can publish a latency figure and hold it. Everything downstream depends on it: the agent-training story, the cost-per-output argument, and the claim that video is where interfaces are going.

Three signals are worth tracking. A number, any number, attached to the under-500 ms claim. A second, independent reproduction of the preference study. And an endpoint — because until Solaris appears in Runway's API reference with a price next to it, the right way to treat it is as a research direction with an unusually good demo reel, which is a real thing to be and also not yet a tool.