A generated hero card for HiDream-O1-Video-1.0 with the kicker 'OrcaRouter · model radar — video' above the headline 'Sixth on the board, one input in the API', a left card 'The board' listing AA image-to-video 6th of 80+, Elo 1,175 ±10, 3,231 samples and .80 per minute, a right card 'The documented API' listing one reference image, an optional prompt, force_10s only, and video, resolution and audio not exposed, and four chips reading Released Aug 2026, Announced 17 Sep 2026, Weights: none published, and no entry on the T2V or editing boards.
Engineering & Research

HiDream-O1-Video-1.0 Debuts Sixth on Artificial Analysis — and Its Public API Is Much Smaller Than Its Launch

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

HiDream-O1-Video-1.0 is sixth on the Artificial Analysis Image-to-Video leaderboard, with an Elo of 1,175 over 3,231 votes at $5.80 per minute of generated video. That is the position its maker announced — except HiDream.ai's launch release, dated 17 September 2026, said number four. Both numbers are true and they describe different things: the release quotes the rank the model held on the day it was announced, and the board has moved since. Read on 10 October 2026 with the audio filter on, HiDream V1 sits behind MiniMax H3 Max, MiniMax H3, Vidu Q4 Preview, Omni Flash and Dreamina Seedance 2.0 720p, in that order.

The more interesting story is not the rank. It is that HiDream-O1-Video-1.0 was announced as a model that takes text, images and video and returns five-to-twenty-second 1080p footage with natively synchronised audio — and that what you can actually call today takes exactly one reference image and an optional text prompt, and nothing else. Both of those statements are documented. The gap between them is the thing worth knowing before you build on this model.

The board, read properly

Artificial Analysis runs its video leaderboards in tiers, each with its own Elo scale and its own vote pool, so a score does not transfer between them. On AA-Video-I2V v1.0 with audio kept in, the top six are separated by twenty Elo points and the confidence intervals are between seven and ten points wide. That top six, as of 10 October 2026:

• MiniMax H3 Max — 1,195 ±9 over 5,894 samples, August 2026, $4.80/min

• MiniMax H3 — 1,181 ±8 over 7,284 samples, July 2026, $7.80/min

• Vidu Q4 Preview — 1,179 ±10 over 5,543 samples, October 2026, $5.04/min

• Gemini Omni Flash — 1,178 ±7 over 12,327 samples, May 2026, $6.00/min

• Dreamina Seedance 2.0 720p — 1,176 ±7 over 15,721 samples, March 2026, $9.07/min

• HiDream-O1-Video-1.0 — 1,175 ±10 over 3,231 samples, August 2026, $5.80/min

Read those intervals as a group and the honest summary is that positions one through six are, statistically, one contest with a very long tail of ties. The gap that is not inside the noise is the one between HiDream V1 and the rest of the field: seventh place, Alibaba's Wan 3.0, is at 1,164 — eleven points and a full interval below sixth, and nearly three hundred votes ahead of it. HiDream's own sample count is the smallest in that top six, which means its interval is among the widest and its position the least settled.

A generated scoreboard titled 'HiDream-O1-Video-1.0 — announced versus served', contrasting the vendor's launch claims (inputs: text, image and video; output: 1080p, 5 to 20 seconds; audio: natively synchronised; duration planned from narrative; generation as one joint multimodal process; physics of gravity, collision and inertia) against the HiHarness request surface (one reference image; aspect ratio preserved; no audio parameter; force_10s boolean only; no resolution selector documented; physics not exposed as a control), with a footer stating the launch claims are vendor-reported.

Two things Artificial Analysis states about the board itself are worth carrying forward. It records a single model as added in the last 30 days, and it says AA-Video-I2V v2.0 is coming, aligned to the AA-Video-T2V v2.0 taxonomy with 1080p throughout. That matters for a model whose launch claims are about 1080p and about text and video conditioning: the board it debuts on is explicitly a transitional one, and the replacement changes what gets measured.

Where the model is not

HiDream-O1-Video-1.0 appears on exactly one Artificial Analysis video board. It is absent from AA-Video-T2V v2.0, the text-to-video board, and from the video-editing board. I checked both on 10 October 2026 — Wan 3.0 leads text-to-video at 1,156 and video editing at 1,157, FLUX 3 is sixth on text-to-video at 1,128, and HiDream appears nowhere in either ranking.

That absence is a signal rather than a disqualification. A model that only consumes a reference image cannot be entered in a text-to-video vote, and one that only animates cannot enter an editing comparison. The practical consequence is that there is no independent measurement of HiDream V1 doing text-to-video or video editing, which is precisely two of the three input modes its launch release advertises.

Two dates, and the ninety days between them

There is a wrinkle in the dating that anyone quoting these numbers should handle carefully. Artificial Analysis records HiDream-O1-Video-1.0's release date as 28 August 2026 and the date it was introduced to the board as 10 September 2026. HiDream's own launch announcement is dated 17 September. The third-party model references that have since appeared describe it as published on 15 September and verified on 17 September.

What that pattern describes is a model that reached the evaluation pipeline before it reached the press release — normal enough, and it means the release date on the board is the earlier of the two, not an error. It does not describe a March model being re-announced in August, which is the failure mode this blog has been burned by before. HiDream-O1-Video-1.0 is genuinely new: the leaderboard infrastructure itself dates its arrival inside the last two months.

What the launch claims, and who says so

The architecture story comes from HiDream and has not been independently reproduced. Per the vendor: HiDream-O1-Video-1.0 — HiDream V1, or HD-V1 — jointly models text, video and audio inside one generation process rather than generating picture and adding sound afterwards, so that lip movement, ambient sound and physical action share a timeline. Duration is planned rather than fixed, with the model choosing a length between five and twenty seconds from the event being described. Physical behaviour — gravity, collision, inertia, deformation, light decay — is trained in as a generation constraint. Post-training uses diffusion reinforcement learning against a multimodal reward model aligned to human perception. Generation runs in three stages the company describes as planning first, joint generation second, alignment third.

HiDream positions the model inside a four-family portfolio built on a shared unified-transformer foundation — HiDream-O1-Image, HiDream-O1-Video, HiDream-O1-World and HiDream-O1-Embodied — and cites prior benchmark placements for the siblings, including HiDream-O1-Image 1.5 at second globally on Artificial Analysis's text-to-image board and HiDream-O1-World at first on WBench. Those are the vendor's figures for other models in the family, stated without independent confirmation here. The company's corporate material also describes a foundation-model technology stack of "more than 200 billion" parameters; that is a platform-level statement, not a specification of HiDream-O1-Video-1.0, and no source publishes a parameter count for this checkpoint. No open weights have been released for it.

A screenshot of the Artificial Analysis AA-Video-I2V v1.0 image-to-video leaderboard with the audio filter applied, captured 10 October 2026, headed by MiniMax H3 Max at Elo 1,195 and .80 per minute, with HiDream-O1-Video-1.0 sixth at Elo 1,175 over 3,231 samples, released August 2026 at .80 per minute, and a banner noting that AA-Video-I2V v2.0 is coming soon with 1080p videos throughout.

What you can actually call

Here the documentation is unusually clear, and it is narrower than the announcement. HiDream serves the model through HiHarness, its MaaS platform. The video endpoint is a POST to /api/maas/gw/v1/videos/generations with the model identifier HiDream-O1-Video-1.0. The required input is exactly one reference image — a public URL, or Base64 without the data-URL prefix, between 40 KB and 20 MB. The text prompt is optional and defaults to empty. The generation is asynchronous: you submit, receive a task_id, and poll /api/maas/gw/v1/videos/generations/results until result.status reads 1.

Three omissions in that schema are worth naming explicitly, because each one is a capability the launch release advertises.

• Reference video — the announcement lists video among the model's inputs; the documented request has no field for it.

• Explicit resolution — the announcement claims 1080p; the documented request has no resolution selector, and the result preserves the reference image's aspect ratio instead.

• Audio control — the announcement claims native synchronised audio; the documented request exposes no audio input or audio-generation parameter.

What the schema does expose is force_10s, a boolean defaulting to false. Left false, the model picks the duration. Set true, you get ten seconds. There is no numeric duration field, so the advertised five-to-twenty-second range is a model-level claim that the public request surface does not let you steer except by choosing one of two modes.

A screenshot of the HiHarness documentation page for the HiDream-O Series video models, captured 10 October 2026, describing HiDream-O1-Video-1.0 as a single-image-to-video model that generates a video from one reference image and a text prompt, with Model ID HiDream-O1-Video-1.0, input as one reference image by public URL or Base64, output of one video, resolution that preserves the input aspect ratio, a dynamic or fixed ten-second duration, and an asynchronous workflow of submitting a task and polling the result endpoint until result.status is 1.

There is one implementation detail in that response contract that will bite an integration written by pattern-matching other video APIs. result.status = 1 means every subtask has reached a terminal state — not that generation succeeded. You still have to read sub_task_results[].task_status, where 1 is success, 3 is generation failure and 4 is a safety-review failure. A client that treats terminal as successful will hand you a video URL that does not exist.

The price, in context

At the $5.80 per minute Artificial Analysis records, HiDream V1 is mid-priced for its bracket rather than cheap. Within the board's top six it is cheaper than MiniMax H3 ($7.80) and Dreamina Seedance 2.0 ($9.07) and dearer than MiniMax H3 Max and Vidu Q4 Preview. Against the models it will be compared with most often, the spread is wider: Google Veo 3.1 is $24.00 per minute, Kling 3.0 1080p Pro $20.16, Alibaba Wan 2.7 $9.00 and Wan 3.0 $12.00, while Grok Imagine Video 1.5 sits at $8.40.

Those per-minute rates are not directly comparable, because a minute of output at which resolution and with which audio is exactly the thing the API documentation leaves unspecified for this model. Treat the number as a starting point for your own arithmetic rather than a rate card.

Trying an unproven route without betting a pipeline on it

The case for caution here is not that the model is bad. It is that a preview-era serving surface — one input type, one duration switch, a two-state terminal flag — is the wrong place to hard-code a production path. HiDream-O1-Video-1.0 is not a model we serve at OrcaRouter today, so nothing below is a claim about hosting it: the place our product genuinely helps is on the other side of the decision, where you want to test a new video route against an incumbent and be able to walk away from it at any point. One OpenAI-compatible endpoint over 200-plus models, automatic failover so a route that starts failing quietly is not the route you ship, provider list prices passed through with no markup added, and a routing DSL that lets you compare two models inside one request rather than two integrations. Nothing about that requires HiDream's model specifically.

What would change this piece

Three developments would move the picture, in descending order of likelihood.

First, a wider serving contract. If HiHarness adds a reference-video field, a resolution selector or documented audio controls, the gap between the announcement and the API closes and the 1080p and text-to-video claims become testable. That is a documentation change, not a model change, and it could land any week.

Second, the AA-Video-I2V v2.0 board. Artificial Analysis has said it is coming, aligned to the T2V v2.0 taxonomy with 1080p throughout. A 1080p-only board would be the first independent environment in which HiDream V1's resolution claim means anything — and if the model still cannot be entered on the text-to-video board under the new taxonomy, its absence there becomes a substantive finding rather than an artefact of the old categories.

Third, sample count. 3,231 votes is a real measurement and a young one, the thinnest in the top six. The interval is ±10 now and will narrow in whichever direction the voters push it. Anyone quoting "sixth" as a settled fact should say the date they read it, because sixth is one tie in a six-way tie.

The question worth keeping is not whether HiDream-O1-Video-1.0 is competitive — on the one board where it can be measured, it plainly is, and at a defensible per-minute price. It is whether the model that eventually reaches a wide API is the one that was announced. On current documentation, those are still two different products.

HiDream-O1-Video-1.0 is not one of the routes we serve today, but the same kind of surface sits behind 200-plus models with provider list prices passed through and no markup added.