
MiniMax H3 (Hailuo 3.0): The 2K AI Video Model with Omni-Reference, Explained
- qwenNEWQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaNEWOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicNEWAnthropic: Claude Opus 52026-07-2461Intelligence78Coding
- googleNEWGoogle: Gemini 3.6 Flash2026-07-2150Intelligence69Coding
- googleNEWGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
- metaMeta: Muse Spark 1.12026-07-1651Intelligence71Coding
- kimiMoonshotAI: Kimi K32026-07-1557Intelligence76Coding
- openaiOpenAI: GPT-5.6 Luna2026-07-0951Intelligence71Coding
- openaiOpenAI: GPT-5.6 Terra2026-07-0955Intelligence77Coding
- openaiOpenAI: GPT-5.6 Sol2026-07-0959Intelligence77Coding
- grokxAI: Grok 4.52026-07-0854Intelligence72Coding
- tencentTencent: Hy32026-07-0641Intelligence59Coding
- obsidianQwen3.6 35B A3B Uncensored (Aggressive)2026-07-0232Intelligence42Coding
- obsidianGemma4 26B A4B Uncensored (Balanced)2026-07-0226Intelligence39Coding
- anthropicAnthropic: Claude Sonnet 52026-06-3053Intelligence72Coding
- klingKling: Kling 3.0 Turbo2026-06-1757Intelligence52Coding57Math
- z-aiZ.ai: GLM 5.22026-06-1651Intelligence69Coding60Math
- kimiMoonshotAI: Kimi K2.7 Code2026-06-1242Intelligence61Coding61Math
- anthropicAnthropic: Claude Fable 52026-06-0960Intelligence77Coding
- qwenQwen: Qwen3.7 Plus2026-06-0139Intelligence56Coding59Math
MiniMax H3 — better known as Hailuo 3.0 — is MiniMax's newest AI video-generation model, unveiled at WAIC 2026 (July 17, 2026) and now rolling out across Hailuo and partner platforms. It pushes MiniMax's video line to native 2K resolution, adds synchronized audio in a single pass, and introduces an unusually rich "omni-reference" control system. This guide explains what H3 actually is, what's genuinely new, and where it sits among the 2026 frontier video models — with capabilities sourced and quality claims clearly labeled.
Accuracy note: H3's capabilities below come from MiniMax's announcement and platform documentation. Video quality is largely subjective and judged by human-preference arenas; as a brand-new model, H3 does not yet have an established independent arena score (public leaderboards still list its predecessor, Hailuo 2.3). Treat comparative "which looks best" claims as provisional. Specs and pricing change — verify before you build.
TL;DR. H3 is a native-2K, 24fps text-to-video and image-to-video model that generates 5–15 second clips (extendable to ~30s) with synchronized dialogue, sound effects, and ambience in one pass. Its standout features are omni-reference (up to 9 reference images, 3 video clips, and 3 audio clips for consistency) and instruction-based editing (change characters, objects, scenes, sound, or pacing without regenerating). It's a big step over Hailuo 2.3 (which was 1080p, ~10s, no omni-reference). It competes with Kling 3.0, Google Veo 3.1, ByteDance Seedance 2.0, and others — strong on control and audio, with independent quality still to be benchmarked.
Key takeaways
• H3 = Hailuo 3.0, MiniMax's latest video model, unveiled at WAIC 2026 and available now via Hailuo/partner platforms.
• Native 2K at 24fps, 5–15s clips (extendable to ~30s), multiple aspect ratios, text-to-video and image-to-video.
• Omni-reference: up to 9 images, 3 video clips, and 3 audio clips as references for style, character, motion, and voice consistency.
• Synchronized dialogue, SFX, and ambience generated in a single pass; instruction-based editing without full regeneration.
• Big upgrade over Hailuo 2.3 (1080p, ~10s); competes with Kling 3.0, Veo 3.1, Seedance 2.0, Wan 2.7, and more.
What MiniMax H3 actually is
MiniMax is a major Chinese AI company (it completed a Hong Kong IPO in January 2026, reportedly raising about $619 million at a roughly $4 billion valuation, with backers including Alibaba and Tencent). Its consumer video brand is Hailuo, known for category-leading physics simulation, fast generation, and accessible pricing. H3 is the third-generation Hailuo model, unveiled alongside MiniMax's M3 text model and other multimodal solutions at WAIC 2026. Where M3 is a text/agentic LLM, H3 is squarely a video-generation model.
What's new: 2K, omni-reference, and one-pass audio
Three things define H3. First, native 2K at 24fps — a step up from Hailuo 2.3's 1080p — at film-standard cadence, with clips from 5 to 15 seconds (extendable to about 30 seconds via an Extend tool) and aspect ratios from 21:9 to 9:16. Second, omni-reference: you can feed up to 9 reference images, 3 video clips, and 3 audio clips simultaneously to lock in style, character identity, motion, or voice — an unusually generous control surface for keeping a character or look consistent across shots. Third, synchronized audio in a single pass: H3 generates dialogue, sound effects, and ambient atmosphere timed to on-screen action, rather than requiring a separate audio step.

Instruction-based editing
A quietly important feature is instruction-based editing: instead of regenerating a clip from scratch to make a change, you can describe an edit — swap a character, alter an object, change the scene, adjust the sound, or re-pace the shot — and H3 applies it. For iterative creative work, editing-in-place rather than re-rolling the dice on a fresh generation is a real workflow advantage.
How it compares to Hailuo 2.3
The generational jump is clear: H3 moves from 1080p to native 2K, extends maximum single-generation length from around 10 seconds to 15 (with a 30-second extend), and adds both omni-reference inputs and instruction-based editing, which 2.3 lacked. If you used Hailuo 2.3, H3 is a straightforward upgrade on resolution, length, control, and audio.
Where H3 sits in the 2026 frontier
The frontier video field in mid-2026 is crowded and strong. On human-preference arenas, Kling 3.0 has led text-to-video, ByteDance Seedance 2.0 has topped audio-inclusive rankings, and Google Veo 3.1 is noted for high-fidelity 48kHz synchronized dialogue; Alibaba Wan 2.7, xAI Grok Imagine Video 1.5, Runway Gen-4.5, and Luma Ray 3.2 are all serious options. H3's pitch isn't necessarily "highest raw fidelity" — several rivals push native 4K — but rather control and audio: omni-reference, one-pass synchronized sound, and instruction-based editing, at MiniMax's characteristically fast, accessible price point. Because H3 is brand new, it doesn't yet carry an independent arena score, so judge it on your own test prompts rather than on any single claim.
A note on OpenAI Sora
If you're mapping the field: OpenAI discontinued the Sora web and app experiences in April 2026, with its API scheduled to end in September 2026 — so Sora is not a model to start new video workflows on. The live frontier is the set above, and H3 joins it.
How to access H3 and the rest of the field
H3 is available through MiniMax's Hailuo platform and partner tools. Because no single provider hosts every model, teams that use multiple modalities often keep a vendor-neutral setup: a href="https://www.orcarouter.ai/">OrcaRouter/a> provides one OpenAI-compatible endpoint across a large catalog — primarily text and multimodal LLMs (including MiniMax's own M3), plus a limited set of media models such as Kling's turbo tier — while dedicated video models like H3 are reached through their native platforms. The practical pattern: run your LLM and agent traffic through one endpoint, and use the best video model per project directly.

Three scenarios where H3 shines
1. Character- and style-consistent short films
Omni-reference (up to 9 images, 3 clips, 3 audio) is built for keeping a character, look, and voice consistent across multiple shots — ideal for narrative shorts and episodic content.
2. Social and ad creative with sound
One-pass synchronized dialogue, SFX, and ambience plus flexible aspect ratios (9:16 to 21:9) suit fast social and advertising workflows that need finished audio, not just visuals.
3. Iterative creative editing
Instruction-based editing lets teams refine a clip descriptively instead of regenerating, which speeds up revision-heavy production.

FAQ
What is MiniMax H3?
H3 is Hailuo 3.0, MiniMax's latest AI video-generation model, unveiled at WAIC 2026 (July 17, 2026). It produces native 2K, 24fps video with synchronized audio, from text or images, with omni-reference control and instruction-based editing.
Is H3 the same as MiniMax M3?
No. M3 is MiniMax's text/agentic LLM; H3 (Hailuo 3.0) is its video-generation model. They were unveiled at the same event but do different jobs.
How long and what resolution are H3 clips?
Native 2K at 24fps, 5–15 seconds per generation, extendable to about 30 seconds, in aspect ratios from 21:9 to 9:16.
What is omni-reference?
A control system that accepts up to 9 reference images, 3 video clips, and 3 audio clips at once to keep style, character, motion, and voice consistent.
Is H3 better than Kling 3.0 or Veo 3.1?
It depends on the axis. Kling 3.0 offers native 4K and leads some arenas; Veo 3.1 is strong on 48kHz dialogue. H3 emphasizes omni-reference control, one-pass audio, editing, and price. As a new model, H3 has no established independent arena score yet — test on your own prompts.
Is H3 open-weight?
H3/Hailuo is offered as a platform model rather than an announced open-weight release (unlike MiniMax's open-weight M3 text model). Check MiniMax's terms for the latest.
Bottom line
MiniMax H3 (Hailuo 3.0) is a serious step up for MiniMax's video line: native 2K, one-pass synchronized audio, generous omni-reference control, and instruction-based editing, at an accessible price. It won't automatically out-resolve rivals pushing native 4K, and its independent quality is still to be benchmarked — but on control, audio, and iteration speed it's compelling. Evaluate H3 against Kling 3.0, Veo 3.1, Seedance 2.0, and the rest on your own footage, and keep your broader model stack vendor-neutral through one endpoint like OrcaRouter.
