
MiniMax H3 (Hailuo 3.0): The 2K AI Video Model with Omni-Reference, Explained
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0240Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0340Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2451Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2134Intelligence69Coding
MiniMax H3 — better known as Hailuo 3.0 — is MiniMax's newest AI video-generation model, unveiled at WAIC 2026 (July 17, 2026), released to the public on July 31, and now available four ways at once: MiniMax's Hailuo platform, Luma Agents, an open-weight download, and — since August 30 — Vercel's AI Gateway, where the new speed-first MiniMax H3 Max launched alongside it. It pushes MiniMax's video line to native 2K resolution, adds synchronized audio in a single pass, and introduces an unusually rich “omni-reference” control system. The newest development is on the self-hosting side: since September 1, the open H3-Base weights can be served through vLLM-Omni with FastVideo's FastH3 fast enough to render a complete 10.1-second video-and-audio MP4 in about 8.7 seconds — faster than its playback duration. This guide explains what H3 actually is, what's genuinely new, and where it sits among the 2026 frontier video models — with capabilities sourced and quality claims clearly labeled.
Accuracy note: H3's capabilities come from MiniMax's announcement and platform documentation. The Vercel AI Gateway availability is confirmed — Vercel's model catalog lists both MiniMax H3 and MiniMax H3 Max, and Vercel's changelog documents the launch discount — though the promo terms themselves are vendor-announced. The Luma Agents integration and the open-weight release remain vendor-announced as of this writing; Luma's own product documentation does not yet list H3, and the integration's access terms are unclear. Video quality is largely subjective and judged by human-preference arenas; H3 now has early Artificial Analysis arena scores — around 1,184 image-to-video and 1,226 text-to-video with audio, captured August 27, 2026 — though arena rankings shift as votes accumulate. The September 1 faster-than-playback serving figure — a complete 10.1-second MP4 with synchronized audio rendered in about 8.7 seconds — is reported by the vLLM-Omni team in their own blog post, measured on an 8× NVIDIA B300 server; it is a team benchmark that has not yet been independently reproduced, and it measures FastH3, FastVideo's four-step distilled adapter, not the full base model. Treat comparative “which looks best” claims as provisional. Specs and pricing change — verify before you build.
TL;DR. H3 is a native-2K, 24fps text-to-video and image-to-video model that generates 5–15 second clips (extendable to ~30s) with synchronized dialogue, sound effects, and ambience in one pass. Its standout features are omni-reference (up to 9 reference images, 3 video clips, and 3 audio clips for consistency) and instruction-based editing (change characters, objects, scenes, sound, or pacing without regenerating). Since launch it has spread fast: the base weights went open on August 3, MiniMax announced H3 in Luma Agents on August 6 (up to 15 seconds of 2K video with native stereo sound, though access terms aren't public yet), and on August 30 MiniMax H3 and MiniMax H3 Max landed on Vercel's AI Gateway with a two-week 50% launch discount. Since September 1, the open base weights can also be served for real-time generation: through vLLM-Omni with FastVideo's FastH3, a complete 10.1-second video-and-audio MP4 renders in about 8.7 seconds — faster than playback, per the vLLM-Omni team's benchmark. It's a big step over Hailuo 2.3 (which was 1080p, ~10s, no omni-reference), and it competes with Kling 3.0, Google Veo 3.1, ByteDance Seedance 2.0, and others — strong on control and audio, with early independent arena scores still to firm up.
Key takeaways
• H3 = Hailuo 3.0, MiniMax's latest video model, unveiled at WAIC 2026 and released July 31; available via Hailuo, Luma Agents, open weights, and Vercel's AI Gateway.
• Native 2K at 24fps, 5–15s clips (extendable to ~30s), multiple aspect ratios, text-to-video and image-to-video.
• Omni-reference: up to 9 images, 3 video clips, and 3 audio clips as references for style, character, motion, and voice consistency.
• Synchronized dialogue, SFX, and ambience generated in a single pass; instruction-based editing without full regeneration.
• Base weights open-sourced August 3 under a community license; the Context-IR and 2K-regeneration modules remain API-only.
• On August 30, MiniMax H3 and MiniMax H3 Max launched on Vercel's AI Gateway with a 50% launch discount through September 13.
• Since September 1, the open H3-Base weights can be served for real-time generation — a 10.1-second video-and-audio MP4 rendered in about 8.7 seconds, faster than playback — via vLLM-Omni and FastVideo's FastH3 (team-reported benchmark, not yet independently reproduced).
• Big upgrade over Hailuo 2.3 (1080p, ~10s); competes with Kling 3.0, Veo 3.1, Seedance 2.0, Wan 2.7, and more.
What MiniMax H3 actually is
MiniMax is a major Chinese AI company (it completed a Hong Kong IPO in January 2026, reportedly raising about $619 million at a roughly $4 billion valuation, with backers including Alibaba and Tencent). Its consumer video brand is Hailuo, known for category-leading physics simulation, fast generation, and accessible pricing. H3 is the third-generation Hailuo model, unveiled alongside MiniMax's M3 text model and other multimodal solutions at WAIC 2026 and released to the public on July 31, 2026. Where M3 is a text/agentic LLM, H3 is squarely a video-generation model.
What's new: 2K, omni-reference, and one-pass audio
Three things define H3. First, native 2K at 24fps — a step up from Hailuo 2.3's 1080p — at film-standard cadence, with clips from 5 to 15 seconds (extendable to about 30 seconds via an Extend tool) and aspect ratios from 21:9 to 9:16. Second, omni-reference: you can feed up to 9 reference images, 3 video clips, and 3 audio clips simultaneously to lock in style, character identity, motion, or voice — an unusually generous control surface for keeping a character or look consistent across shots. Third, synchronized audio in a single pass: H3 generates dialogue, sound effects, and ambient atmosphere timed to on-screen action, rather than requiring a separate audio step.

Instruction-based editing
A quietly important feature is instruction-based editing: instead of regenerating a clip from scratch to make a change, you can describe an edit — swap a character, alter an object, change the scene, adjust the sound, or re-pace the shot — and H3 applies it. For iterative creative work, editing-in-place rather than re-rolling the dice on a fresh generation is a real workflow advantage.
How it compares to Hailuo 2.3
The generational jump is clear: H3 moves from 1080p to native 2K, extends maximum single-generation length from around 10 seconds to 15 (with a 30-second extend), and adds both omni-reference inputs and instruction-based editing, which 2.3 lacked. If you used Hailuo 2.3, H3 is a straightforward upgrade on resolution, length, control, and audio.
Where H3 sits in the 2026 frontier
H3 now carries early Artificial Analysis arena scores — roughly 1,184 in image-to-video and 1,226 in text-to-video with audio, captured August 27, 2026 — though arena rankings shift as votes accumulate, so judge it on your own test prompts rather than on any single claim.
A note on OpenAI Sora
If you're mapping the field: OpenAI discontinued the Sora web and app experiences in April 2026, with its API scheduled to end in September 2026 — so Sora is not a model to start new video workflows on. The live frontier is the set above, and H3 joins it.
How to access H3 and the rest of the field
H3 is now reachable four ways, and that list has grown since launch.
• Hailuo (MiniMax's own platform): the full product, including instruction-based editing and the H3-Regenerate-2K upscale to 2K.
• Luma Agents: MiniMax announced on August 6, 2026 that H3 is now available in Luma's multi-model Agents product, generating up to 15 seconds of 2K video with native stereo sound. This is vendor-announced, and access terms aren't public yet — Luma's own product docs don't list H3 at this writing — so expect pricing and eligibility details to firm up shortly. Luma hosting a rival's video model inside its own Agents product is a sign of how interchangeable the 2026 video-model market has become.
• Vercel AI Gateway: on August 30, 2026, MiniMax and Vercel announced that both MiniMax H3 and MiniMax H3 Max are available through Vercel's AI Gateway, with requests billed through the gateway at 50% off from August 30 through September 13, 2026. The model IDs — minimax/minimax-h3 and minimax/minimax-h3-max — are unchanged, so existing gateway integrations pick up the discounted rate with no code change. Availability is confirmed in Vercel's model catalog and changelog; the promo terms are vendor-announced.
• Open weights: on August 3, 2026, MiniMax released H3's base weights — a 33B dense transformer, H3-Base — on Hugging Face and ModelScope under the MiniMax H3 Community License, as two checkpoints: H3-Base-FL2VA for text-to-video/audio and keyframe generation, and H3-Base-Ref2VA for reference-based generation. Two of the three system modules — H3-Context-IR (the prompt preprocessor) and H3-Regenerate-2K (the 2K upscale) — are not in the open release and remain API-only, so self-hosted runs cap below 2K. And since September 1, the same base weights can be served for real-time generation through vLLM-Omni and FastVideo's FastH3 — a four-step distilled adapter that renders a complete 10.1-second video-and-audio MP4 in about 8.7 seconds on an 8× NVIDIA B300 server, faster than its playback duration (team-reported, not yet independently reproduced).
MiniMax H3 is on OrcaRouter at the provider's list price — $0.08 per second at 768p and $0.13 per second at 2K — passed through with 0% markup and automatic failover, so you can call the H3 family through one OpenAI-compatible API instead of wiring up a vendor endpoint that is days old. Because no single provider hosts every model, the practical pattern is to run your LLM and agent traffic through one endpoint (OrcaRouter also carries MiniMax's own M3 text model) and use the best video model per project directly. MiniMax H3 Max isn't routed on OrcaRouter yet — it's served through third-party channels for now — so if you're prototyping against it, the failover habit pays off: try a days-old model without betting a production pipeline on it.

Real-time serving: video faster than playback
The development behind this update is on the self-hosting side. MiniMax H3's open base checkpoint is a 33B dense transformer, and generating a clip the standard way means running roughly 49 diffusion-transformer forwards over a long denoising schedule. FastVideo, the open-source video training-and-inference project, published a distilled student of H3-Base called FastH3: a four-step (DMD2-style) model that reuses H3's text encoder, video VAE, audio VAE, tokenizers, and schedulers but cuts the denoising loop to four transformer forwards over five sigma positions. vLLM-Omni, the multimodal serving runtime, fuses the FastH3 adapter at load time (pull request #6714) and exposes it through an OpenAI-compatible /v1/videos endpoint, alongside its earlier base-H3 support.
The headline number is measured on large hardware. On an 8× NVIDIA B300 server at 1344×768 and 24fps, the vLLM-Omni team reports a complete 10.1-second MP4 — video and synchronized audio — rendered end-to-end in about 8.7 seconds, faster than its playback duration. That is a team-reported benchmark published September 1, 2026, and not yet independently reproduced; the team itself notes the raw benchmark bundle is still pending publication and makes no base-vs-FastH3 quality-parity claim. On more modest hardware the numbers are less dramatic — the community recipe also covers 2× RTX 5090 and single-GPU setups — and FastH3 v1 covers text-to-video-audio (T2VA) only, at 1344×768 rather than the 2K that remains API-side.
For a reader, this shifts what “self-hosting H3” can mean. Before, the open base weights ran but slowly — fine for batch work, not for anything interactive. With FastH3 on vLLM-Omni, an on-premises H3 pipeline can turn around a ten-second clip in well under the clip's length, which is the threshold where real-time product loops — live previsualization, rapid shot iteration, agent-driven video — become practical. If you'd rather not run an eight-GPU server, the managed path is unchanged: MiniMax H3 is on OrcaRouter at the provider's list price, passed through with 0% markup and automatic failover, so you can call the H3 family through one OpenAI-compatible API without owning the hardware.
Three scenarios where H3 shines
1. Character- and style-consistent short films
Omni-reference (up to 9 images, 3 clips, 3 audio) is built for keeping a character, look, and voice consistent across multiple shots — ideal for narrative shorts and episodic content.
2. Social and ad creative with sound
One-pass synchronized dialogue, SFX, and ambience plus flexible aspect ratios (9:16 to 21:9) suit fast social and advertising workflows that need finished audio, not just visuals.
3. Iterative creative editing
Instruction-based editing lets teams refine a clip descriptively instead of regenerating, which speeds up revision-heavy production.

FAQ
What is MiniMax H3?
H3 is Hailuo 3.0, MiniMax's latest AI video-generation model, unveiled at WAIC 2026 (July 17, 2026) and released on July 31, 2026. It produces native 2K, 24fps video with synchronized audio, from text or images, with omni-reference control and instruction-based editing.
Is H3 the same as MiniMax M3?
No. M3 is MiniMax's text/agentic LLM; H3 (Hailuo 3.0) is its video-generation model. They were unveiled at the same event but do different jobs.
Where can I use H3 now?
Four places: MiniMax's Hailuo platform (the full product), Vercel's AI Gateway (both H3 and MiniMax H3 Max, with a 50% launch discount through September 13), Luma Agents (announced August 6, 2026 — up to 15 seconds of 2K video with native stereo sound, access terms still unclear), and self-hosted via the open H3-Base weights on Hugging Face and ModelScope — which, since September 1, can be served at faster-than-playback speed through vLLM-Omni with FastVideo's FastH3 (a 10.1-second video-and-audio MP4 in about 8.7 seconds on 8× B300, per the vLLM-Omni team). The base model is also routed on OrcaRouter at provider list price.
How long and what resolution are H3 clips?
Native 2K at 24fps, 5–15 seconds per generation, extendable to about 30 seconds, in aspect ratios from 21:9 to 9:16.
What is omni-reference?
A control system that accepts up to 9 reference images, 3 video clips, and 3 audio clips at once to keep style, character, motion, and voice consistent.
Is H3 better than Kling 3.0 or Veo 3.1?
It depends on the axis. Kling 3.0 offers native 4K and leads some arenas; Veo 3.1 is strong on 48kHz dialogue. H3 emphasizes omni-reference control, one-pass audio, editing, and price. H3 has early Artificial Analysis arena scores (around 1,184 image-to-video and 1,226 text-to-video with audio as of late August) that will shift as votes accumulate — test on your own prompts.
Is H3 open-weight?
Partly. MiniMax open-sourced H3's base weights on August 3, 2026 — a 33B transformer released on Hugging Face and ModelScope under the MiniMax H3 Community License, with FL2VA and Ref2VA checkpoints. Per MiniMax's license terms, the release restricts deployment regions and extends to generated outputs, with attribution required. The two modules that round out the flagship experience — H3-Context-IR (prompt preprocessing) and H3-Regenerate-2K (the 2K upscale) — are not in the open release and remain API-only. Since September 1, the open base weights can also be served for real-time generation through vLLM-Omni and FastVideo's FastH3 (FastH3 v1 covers text-to-video-audio only). So yes for the base model, with caveats.
Bottom line
MiniMax H3 (Hailuo 3.0) is a serious step up for MiniMax's video line: native 2K, one-pass synchronized audio, generous omni-reference control, and instruction-based editing, at an accessible price. Since launch it has also become dramatically easier to reach — it's routed on OrcaRouter at provider list price, it runs inside Luma Agents, its base weights are open for self-hosting (and since September 1 fast enough to render a 10.1-second clip faster than it plays, via vLLM-Omni and FastVideo's FastH3), and it and MiniMax H3 Max now sit on Vercel's AI Gateway with a two-week 50% launch discount. It won't automatically out-resolve rivals pushing native 4K, and its early arena scores are still firming up — but on control, audio, and iteration speed it's compelling. Evaluate H3 against Kling 3.0, Veo 3.1, Seedance 2.0, and the rest on your own footage, and keep your broader model stack vendor-neutral through one endpoint like OrcaRouter.
