A title card reading 'Utopai X Debuts at #2 in Text-to-Video', subtitled 'A film studio's post-train of MiniMax H3 - Elo 1,150, second only to Wan 3.0', with pill labels reading 'Elo 1,150', '5,455 votes' and 'No API'.
Guides & Insights

Utopai X Debuts at #2 in Text-to-Video: Inside a Film Studio's Post-Train of MiniMax H3

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Utopai X is a video model that did not exist last week, was not trained from scratch by a frontier lab, and is not sold through an API — and it is currently ranked second in the world at text-to-video. The model comes from Utopai Studios, an AI-native film and television studio based in Mountain View, and it is a post-train of MiniMax H3, MiniMax's omni-modal video model. On the Artificial Analysis text-to-video leaderboard with audio (the board Artificial Analysis now labels AA-Video-T2V v2.0, September 29, 2026 results), Utopai X sits at Elo 1,150 — behind only Aliba​sba's Wan 3.0 at 1,157, and ahead of ByteDance's Dreamina Seedance 2.5 at 1,143 and the MiniMax H3 (768p) base model it was tuned from at 1,138. It arrived on 2026-09-30, the same day the leaderboard was refreshed, and it is available in exactly one place: inside PAI, Utopai's own production platform. There is no public API, and Artificial Analysis' own pricing column for the row reads "No API available."

That combination — second in the world, one distribution channel, no rate card per second in the usual sense — is the whole story. A studio that makes films took someone else's foundation model, tuned it against its own production judgement, and jumped the queue. Whether that is a durable result or a snapshot of a brand-new arena is the question worth asking, and the answer depends on reading the leaderboard rather than the press release.

What the board actually shows

Artificial Analysis ranks video models using Elo derived from blind pairwise human preference voting. Voters see two clips generated from the same prompt and pick the one they prefer, without knowing which model produced either. The board changes as votes accumulate, and every row carries a confidence interval that says how much the position is allowed to move. Captured on 2026-10-01, the with-audio text-to-video board reads:

A scoreboard card reading 'Utopai X (based on MiniMax H3) - the scoreboard', with rows for text-to-video rank #2 (board range #1-3), Elo 1,150 (+/-11), votes 5,455, parent model MiniMax H3 (768p), parent model Elo 1,138 (+/-10), and availability PAI subscribers only with no API.

• #1 Wan 3.0 — Elo 1,157 (±10), 6,391 samples, released Aug 2026, $12.00/min.

• #2 Utopai X (based on MiniMax H3) — Elo 1,150 (±11), 5,455 samples, released Sep 2026, no API available.

• #3 Dreamina Seedance 2.5 — Elo 1,143 (±10), 6,375 samples, released Jul 2026, $34.12/min.

• #4 MiniMax H3 (768p) — Elo 1,138 (±10), 6,343 samples, released Jul 2026, $4.80/min.

• #5 FLUX 3 — Elo 1,129 (±10), 5,628 samples, released Jul 2026, $17.40/min.

Two details matter more than the ordering. First, the gap between first and second is seven Elo points, and the two intervals explicitly overlap — Artificial Analysis prints Utopai X's range as 1,139 to 1,161, which contains Wan 3.0's point estimate of 1,157. The board labels Utopai X's rank as a range, #1–3, not a fixed position. Second, the gap between Utopai X and the model it was post-trained from is twelve points with the same overlap problem, so "the post-train beats its parent" is directionally true and statistically soft. Utopai's own write-up is more careful than most launch posts about this: it says the Elo result "provides an independent assessment of how people respond to its generated footage," which is a fair reading of what a human-preference arena measures and not a claim about filmmaking.

The vendor-reported layer, by contrast, is confident. Utopai describes strengths in reflections and caustics — the concentrated patterns formed when light reflects or refracts — and in spatial consistency within a clip, both of which are development targets rather than benchmark results. Those are the vendor's words, not Artificial Analysis' measurements.

A studio instead of a lab

Utopai Studios develops, finances and produces its own film and television projects, and it builds its models inside that business rather than beside it. The stated mechanism is feedback: productions generate screenplays, approved character and location designs, shot plans and successive takes, and the studio's directors, cinematographers and artists record not only which takes were kept but why the others were rejected. That record becomes training and evaluation signal. The research team also runs an internal Elo system that combines human preference assessment with automated voting by vision-language models, with artists participating in the human side.

It is a plausible advantage and an unverifiable one from outside. Post-training can move a general video model toward the preferences of film and advertising users without replacing the foundation model — that much is consistent with the leaderboard. How much of the twelve-point move over MiniMax H3 comes from training data, preference optimisation, prompt handling, inference settings or serving-time changes is not something the launch materials say. Utopai has published no post-training data disclosure, no optimisation method, no compute budget and no safety evaluation. Independent researchers cannot reproduce the gain, and there is no technical report to argue with.

The distribution layer is where the studio thesis becomes concrete. PAI — Production Assistive Intelligence — is the platform Utopai launched alongside the model, and it holds project context: script breakdown into scenes and shots, a production library of characters, locations and objects, version history, shared approvals for enterprise teams, and a timeline that exports to Premiere Pro or DaVinci Resolve. The pitch is that generating a clip is the cheap part and keeping a shot consistent with the rest of the film is the expensive part. Utopai X generates footage inside that workspace "when selected by the filmmaker," in the company's own phrasing — the model is one option among several image, video and audio models available in PAI, not the only one.

What Utopai X costs, and how to read it

The Artificial Analysis Video Arena text-to-video leaderboard, last updated September 29, 2026, listing Wan 3.0 at Elo 1,157, Utopai X at 1,150, Dreamina Seedance 2.5 at 1,143, MiniMax H3 (768p) at 1,138 and FLUX 3 at 1,129, each with sample counts, release months and price per minute.

PAI is sold as a subscription with credits, not as a per-second API. The published tiers are $15/month for 1,000 credits, $49/month for 3,500 credits, $129/month for 10,000 credits and $379/month for 35,000 credits, with annual billing discounting roughly 8 to 10 percent, plus top-ups at $15 per 1,000 credits. Utopai X has its own line in PAI's credit table: 93 credits per second of video at 1080p. The platform's other video models sit lower — 50 credits per second at 1080p on the standard tier and 76 on the pro tier.

Converting that to a per-minute figure is arithmetic anyone can do and almost nobody should quote as a price. At 93 credits per second, one minute of Utopai X output costs 5,580 credits. The entry $15 plan buys 1,000 credits a month, which is roughly 10.8 seconds of footage if every credit went to this model. Scaling the entry credit rate linearly, the implied cost is on the order of $80-plus per minute of generated video. That number is worth stating only with its caveats attached: it assumes a strictly linear credit-to-dollar conversion, it ignores that credits are pooled across every model in PAI and expire or go unused, it ignores annual discounts and top-up bundles, and it says nothing about resolution or length limits, which Utopai has not specified separately for Utopai X. It is a useful order of magnitude, not a rate card.

For comparison, on the same Artificial Analysis board MiniMax H3 (768p) is listed at $4.80 per minute and Wan 3.0 at $12.00, while Dreamina Seedance 2.5 is listed at $34.12. Those are automated pricing feeds from the vendors' own APIs, not subscription-derived estimates, so the comparison is directional at best — but the direction is clear: at the entry tier, generation inside a credits-based studio platform is not competing with per-second API pricing. What a buyer is paying for is the workspace around the model, not the model's marginal second.

The part that is not public

Three gaps are worth naming, because each one blocks a different decision.

• No API, and no integration path. Utopai X currently has no direct integration route. A product team that wants to call it in a pipeline cannot. Artificial Analysis lists it as "No API available," and the launch materials describe no developer access.

• No separate specification. MiniMax H3 generates clips of 4 to 15 seconds at up to 2K with native stereo audio. Utopai has not published its own maximum duration or resolution for Utopai X, which means the 1080p figure in the credit table is the only resolution statement available and the clip length is unknown.

• No evaluation beyond the arena. The Elo result is an independent measurement of human preference on short clips. It is not a measure of whether footage serves an intended shot, cuts into an edit, or survives a revision. Utopai's own post acknowledges this directly — "evaluation also continues after a clip is generated" — which is an unusually honest framing for a launch and also a warning about how far the ranking travels.

On the studio side there is more history to point at. Utopai launched PAI 2.0 on 2026-06-02 and said at the time that the platform had reached $11 million in annual recurring revenue less than two months after its April debut, driven by commercial licences sold to production companies. The company says it has three theatrical films and two series scheduled for 2027 and is applying PAI and Utopai X to its own productions, including The Most Serious Fart, an animated feature written and directed by Mike Bender. Those are the studio's figures and the studio's schedule, and they are the reason a post-train with no API is worth writing about at all: the model is being used to make films the company intends to release, which is a harder test than a leaderboard.

Where the base model is still the practical choice

The OrcaRouter model page for MiniMax H3, showing the 0% markup badge, a description of omni-modal 4-to-15-second video generation at up to 2K with native stereo audio, a $0.08 per second price panel, a performance panel with p50 latency, and a release date of 7/31/2026.

For anyone who actually needs to generate video this week, the routable half of this family is the base model. MiniMax H3 is available on OrcaRouter at the provider's list price — $0.08 per second at 768p, which is $4.80 per minute, the same figure Artificial Analysis lists — with 0% markup, so any vendor price cut lands the same day rather than after a repricing cycle, and automatic failover across providers so a single endpoint outage does not take your pipeline down. It accepts text, image, video and audio references as one unified context and returns 4 to 15 second clips at 768P or 2K with native stereo audio, billed per second of generated output.

Utopai X is not routed, and nothing in this article should be read as a claim that it is. What the base model gives you is a way to test the H3 family's aesthetic and motion characteristics — the same foundation Utopai started from, before whatever post-training it applied — through one API key alongside more than 200 other models, without a PAI subscription. If Utopai X ever reaches a routable provider it would appear at list price the same day, with the vendor's rate passed through rather than marked up. Until then, the honest split is: Utopai X for studios inside PAI, MiniMax H3 for everyone assembling a pipeline.

What decides whether this was a debut or a moment

Three things are worth watching over the next quarter. First, whether the seven-point gap at the top survives another few thousand votes — the intervals overlap today, and a board that reshuffles when a new batch of comparisons lands has not settled anything. Second, whether Utopai opens any developer access at all; a model ranked second in the world that only a PAI subscriber can call is a product decision, and it is reversible. Third, whether the competition moves in the other direction. Wan 3.0 already generates up to 30 seconds of 1080p with native audio and accepts text, images, video, audio, documents and web pages as references, and it leads the board on lighting, materials and camera control in the use-case breakdown. A twelve-point post-training gain over a base model is suggestive; holding second place against a foundation model with a wider input envelope is a different job.

What is genuinely new here is not the Elo. It is the shape of the claim: a studio with a production pipeline argues that owning the decisions behind the footage — which take, which reference, which revision — is worth more than owning the pre-training run. That argument is testable, and the box office for the 2027 slate is one of the tests.