
HiDream-O1-Video-1.0 Debuts Sixth on Artificial Analysis — and Its Public API Is Much Smaller Than Its Launch
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 72 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 116 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 48 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 478 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 59 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 371 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
HiDream-O1-Video-1.0 is sixth on the Artificial Analysis Image-to-Video leaderboard, with an Elo of 1,175 over 3,231 votes at $5.80 per minute of generated video. That is the position its maker announced — except HiDream.ai's launch release, dated 17 September 2026, said number four. Both numbers are true and they describe different things: the release quotes the rank the model held on the day it was announced, and the board has moved since. Read on 10 October 2026 with the audio filter on, HiDream V1 sits behind MiniMax H3 Max, MiniMax H3, Vidu Q4 Preview, Omni Flash and Dreamina Seedance 2.0 720p, in that order.
The more interesting story is not the rank. It is that HiDream-O1-Video-1.0 was announced as a model that takes text, images and video and returns five-to-twenty-second 1080p footage with natively synchronised audio — and that what you can actually call today takes exactly one reference image and an optional text prompt, and nothing else. Both of those statements are documented. The gap between them is the thing worth knowing before you build on this model.
The board, read properly
Artificial Analysis runs its video leaderboards in tiers, each with its own Elo scale and its own vote pool, so a score does not transfer between them. On AA-Video-I2V v1.0 with audio kept in, the top six are separated by twenty Elo points and the confidence intervals are between seven and ten points wide. That top six, as of 10 October 2026:
• MiniMax H3 Max — 1,195 ±9 over 5,894 samples, August 2026, $4.80/min
• MiniMax H3 — 1,181 ±8 over 7,284 samples, July 2026, $7.80/min
• Vidu Q4 Preview — 1,179 ±10 over 5,543 samples, October 2026, $5.04/min
• Gemini Omni Flash — 1,178 ±7 over 12,327 samples, May 2026, $6.00/min
• Dreamina Seedance 2.0 720p — 1,176 ±7 over 15,721 samples, March 2026, $9.07/min
• HiDream-O1-Video-1.0 — 1,175 ±10 over 3,231 samples, August 2026, $5.80/min
Read those intervals as a group and the honest summary is that positions one through six are, statistically, one contest with a very long tail of ties. The gap that is not inside the noise is the one between HiDream V1 and the rest of the field: seventh place, Alibaba's Wan 3.0, is at 1,164 — eleven points and a full interval below sixth, and nearly three hundred votes ahead of it. HiDream's own sample count is the smallest in that top six, which means its interval is among the widest and its position the least settled.

Two things Artificial Analysis states about the board itself are worth carrying forward. It records a single model as added in the last 30 days, and it says AA-Video-I2V v2.0 is coming, aligned to the AA-Video-T2V v2.0 taxonomy with 1080p throughout. That matters for a model whose launch claims are about 1080p and about text and video conditioning: the board it debuts on is explicitly a transitional one, and the replacement changes what gets measured.
Where the model is not
HiDream-O1-Video-1.0 appears on exactly one Artificial Analysis video board. It is absent from AA-Video-T2V v2.0, the text-to-video board, and from the video-editing board. I checked both on 10 October 2026 — Wan 3.0 leads text-to-video at 1,156 and video editing at 1,157, FLUX 3 is sixth on text-to-video at 1,128, and HiDream appears nowhere in either ranking.
That absence is a signal rather than a disqualification. A model that only consumes a reference image cannot be entered in a text-to-video vote, and one that only animates cannot enter an editing comparison. The practical consequence is that there is no independent measurement of HiDream V1 doing text-to-video or video editing, which is precisely two of the three input modes its launch release advertises.
Two dates, and the ninety days between them
There is a wrinkle in the dating that anyone quoting these numbers should handle carefully. Artificial Analysis records HiDream-O1-Video-1.0's release date as 28 August 2026 and the date it was introduced to the board as 10 September 2026. HiDream's own launch announcement is dated 17 September. The third-party model references that have since appeared describe it as published on 15 September and verified on 17 September.
What that pattern describes is a model that reached the evaluation pipeline before it reached the press release — normal enough, and it means the release date on the board is the earlier of the two, not an error. It does not describe a March model being re-announced in August, which is the failure mode this blog has been burned by before. HiDream-O1-Video-1.0 is genuinely new: the leaderboard infrastructure itself dates its arrival inside the last two months.
What the launch claims, and who says so
The architecture story comes from HiDream and has not been independently reproduced. Per the vendor: HiDream-O1-Video-1.0 — HiDream V1, or HD-V1 — jointly models text, video and audio inside one generation process rather than generating picture and adding sound afterwards, so that lip movement, ambient sound and physical action share a timeline. Duration is planned rather than fixed, with the model choosing a length between five and twenty seconds from the event being described. Physical behaviour — gravity, collision, inertia, deformation, light decay — is trained in as a generation constraint. Post-training uses diffusion reinforcement learning against a multimodal reward model aligned to human perception. Generation runs in three stages the company describes as planning first, joint generation second, alignment third.
HiDream positions the model inside a four-family portfolio built on a shared unified-transformer foundation — HiDream-O1-Image, HiDream-O1-Video, HiDream-O1-World and HiDream-O1-Embodied — and cites prior benchmark placements for the siblings, including HiDream-O1-Image 1.5 at second globally on Artificial Analysis's text-to-image board and HiDream-O1-World at first on WBench. Those are the vendor's figures for other models in the family, stated without independent confirmation here. The company's corporate material also describes a foundation-model technology stack of "more than 200 billion" parameters; that is a platform-level statement, not a specification of HiDream-O1-Video-1.0, and no source publishes a parameter count for this checkpoint. No open weights have been released for it.

What you can actually call
Here the documentation is unusually clear, and it is narrower than the announcement. HiDream serves the model through HiHarness, its MaaS platform. The video endpoint is a POST to /api/maas/gw/v1/videos/generations with the model identifier HiDream-O1-Video-1.0. The required input is exactly one reference image — a public URL, or Base64 without the data-URL prefix, between 40 KB and 20 MB. The text prompt is optional and defaults to empty. The generation is asynchronous: you submit, receive a task_id, and poll /api/maas/gw/v1/videos/generations/results until result.status reads 1.
Three omissions in that schema are worth naming explicitly, because each one is a capability the launch release advertises.
• Reference video — the announcement lists video among the model's inputs; the documented request has no field for it.
• Explicit resolution — the announcement claims 1080p; the documented request has no resolution selector, and the result preserves the reference image's aspect ratio instead.
• Audio control — the announcement claims native synchronised audio; the documented request exposes no audio input or audio-generation parameter.
What the schema does expose is force_10s, a boolean defaulting to false. Left false, the model picks the duration. Set true, you get ten seconds. There is no numeric duration field, so the advertised five-to-twenty-second range is a model-level claim that the public request surface does not let you steer except by choosing one of two modes.

There is one implementation detail in that response contract that will bite an integration written by pattern-matching other video APIs. result.status = 1 means every subtask has reached a terminal state — not that generation succeeded. You still have to read sub_task_results[].task_status, where 1 is success, 3 is generation failure and 4 is a safety-review failure. A client that treats terminal as successful will hand you a video URL that does not exist.
The price, in context
At the $5.80 per minute Artificial Analysis records, HiDream V1 is mid-priced for its bracket rather than cheap. Within the board's top six it is cheaper than MiniMax H3 ($7.80) and Dreamina Seedance 2.0 ($9.07) and dearer than MiniMax H3 Max and Vidu Q4 Preview. Against the models it will be compared with most often, the spread is wider: Google Veo 3.1 is $24.00 per minute, Kling 3.0 1080p Pro $20.16, Alibaba Wan 2.7 $9.00 and Wan 3.0 $12.00, while Grok Imagine Video 1.5 sits at $8.40.
Those per-minute rates are not directly comparable, because a minute of output at which resolution and with which audio is exactly the thing the API documentation leaves unspecified for this model. Treat the number as a starting point for your own arithmetic rather than a rate card.
Trying an unproven route without betting a pipeline on it
The case for caution here is not that the model is bad. It is that a preview-era serving surface — one input type, one duration switch, a two-state terminal flag — is the wrong place to hard-code a production path. HiDream-O1-Video-1.0 is not a model we serve at OrcaRouter today, so nothing below is a claim about hosting it: the place our product genuinely helps is on the other side of the decision, where you want to test a new video route against an incumbent and be able to walk away from it at any point. One OpenAI-compatible endpoint over 200-plus models, automatic failover so a route that starts failing quietly is not the route you ship, provider list prices passed through with no markup added, and a routing DSL that lets you compare two models inside one request rather than two integrations. Nothing about that requires HiDream's model specifically.
What would change this piece
Three developments would move the picture, in descending order of likelihood.
First, a wider serving contract. If HiHarness adds a reference-video field, a resolution selector or documented audio controls, the gap between the announcement and the API closes and the 1080p and text-to-video claims become testable. That is a documentation change, not a model change, and it could land any week.
Second, the AA-Video-I2V v2.0 board. Artificial Analysis has said it is coming, aligned to the T2V v2.0 taxonomy with 1080p throughout. A 1080p-only board would be the first independent environment in which HiDream V1's resolution claim means anything — and if the model still cannot be entered on the text-to-video board under the new taxonomy, its absence there becomes a substantive finding rather than an artefact of the old categories.
Third, sample count. 3,231 votes is a real measurement and a young one, the thinnest in the top six. The interval is ±10 now and will narrow in whichever direction the voters push it. Anyone quoting "sixth" as a settled fact should say the date they read it, because sixth is one tie in a six-way tie.
The question worth keeping is not whether HiDream-O1-Video-1.0 is competitive — on the one board where it can be measured, it plainly is, and at a defensible per-minute price. It is whether the model that eventually reaches a wide API is the one that was announced. On current documentation, those are still two different products.
HiDream-O1-Video-1.0 is not one of the routes we serve today, but the same kind of surface sits behind 200-plus models with provider list prices passed through and no markup added.
