MiniMax-H3

minimax/minimax-h3
VisionAudio
by MiniMax · 2026-07-31

MiniMax-H3 is MiniMax's omni-modal video generation model (the Hailuo 3 generation), released July 31, 2026. It reads text, image, video and audio references as one unified context and generates 4-15 second videos at 768P or 2K with native stereo audio. Billed per second of generated output: $0.08/s at 768P, $0.13/s at 2K.

Inputtext + image + video + audio
Outputvideo
p50 TTFT431 ms
PRICE$0.08/ per second
p50 TTFT431 ms7d
p95 TTFT1.32 s7d
TRAFFIC738tokens / 7d

MiniMax-H3 is the Hailuo 3 generation omni-modal video model from MiniMax, released July 31, 2026. It treats text, images, video and audio references as one unified context and generates 4-15 second…

What is MiniMax-H3?

Code samples

Call from any SDK

OpenAI-compatible — keep the SDK you already use

  • OpenAI SDKhttps://api.orcarouter.ai/v1

Pricing

Pricing
Per second$0.0800
CurrencyUSD
Billed per second of generated video plus metered reference inputs, not per API call

Performance

p50 TTFT
431 ms
Output speed
Collecting…
p95 TTFT
1.32 s
Error rate
0%

Public benchmarks

AA I2V Arena
1193.0 / 1500
AA T2V Arena
1237.0 / 1500
Source: artificial_analysis_arena

Community buzz

What developers are saying this week

Hacker News2 mentions · 7ddown 1 vs the previous week

FAQ

How is MiniMax-H3 billed?
Per second of generated video: $0.08 per second at 768P and $0.13 per second at 2K, for the duration you declare (4-15 seconds, integers). Input materials add to that: audio references are free, images beyond the first five cost $0.04 each, and a reference video is billed by its own duration at the same per-second rate.
How do I call MiniMax-H3?
Submit a task with POST /v1/video/generations using model minimax/minimax-h3, a prompt, and optionally duration (4-15), size (768P or 2K) and reference inputs via metadata. Poll GET /v1/video/generations/{task_id} until it succeeds, then download the video from the returned URL.
What inputs does it accept?
A text prompt (required, up to 7000 characters), plus optional image references (including first/last frame), a reference video, and a reference audio track.
What does it output?
A single video of 4 to 15 seconds at 768P or 2K, with stereo audio generated natively alongside the visuals — no separate audio pass needed.

Embed this badge

MiniMax: MiniMax-H3•pricing pending•431ms p50•via OrcaRouter
HTML <a href="https://www.orcarouter.ai/models/minimax/minimax-h3" target="_blank"> <img src="https://www.orcarouter.ai/embed/minimax/minimax-h3.svg" alt="MiniMax: MiniMax-H3 on OrcaRouter" /> </a>
Markdown [![MiniMax: MiniMax-H3](https://www.orcarouter.ai/embed/minimax/minimax-h3.svg)](https://www.orcarouter.ai/models/minimax/minimax-h3)

Model card as data

GET /api/public/models/minimax/minimax-h3Open
Machine-readable:/llms.txt/llms-full.txt