Hero title card for fal's H3 Max, reading 'H3 Max — 5-Second Clip in Under 3 Seconds', with stat cards showing 2.8s inference for a 5s clip, #1 Image-to-Video on Artificial Analysis, and a 768p max-resolution note.
Guides & Insights

fal's H3 Max Renders a 5-Second Clip in Under 3 Seconds — Now #1 in Image-to-Video

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The most surprising number in the new MiniMax H3 Max release is not the leaderboard rank, though that is notable too: the video model that fal post-trained from MiniMax H3 just took the #1 spot on Artificial Analysis' image-to-video leaderboard. It is the timing figure on fal's own model page, where a sample 5-second clip reports 2.77 seconds of inference. That is sub-real-time generation from a public API — the full clip renders before a real-time playback of the same clip would finish — and it is the reason this drop is drawing attention beyond the usual model-release crowd. MiniMax H3 Max generates 5-to-15-second clips with native audio at up to 768p, and right now it carries a 50% launch discount on the price that expires September 1.

Why "under 3 seconds" matters more than the podium

Every video model claims some version of fast, so it is worth being precise about what fal's number actually says. The 2.77 seconds comes from the inference timing on the fal model page's own sample output for a five-second clip — a vendor-reported figure, not an independently measured one, and not yet reproduced anywhere the rest of us can run. Even taken with that caveat, it is a striking number, because 2.77 seconds for five seconds of video is a real-time factor around 0.55. A model that generates a clip faster than the clip plays back turns video generation from a batch job into something you can put inside a loop: generate, judge, regenerate, all within the time a person would otherwise be waiting for a single render.

The tradeoff is the same one the tweet's "speed vs. quality" framing points at. A 0.55x real-time factor comes from fal co-optimizing the post-train with its own inference stack — fal's description, unreproduced by any independent benchmark beyond the leaderboards — and it is easier to hit at 768p than at higher resolutions. There is no free lunch in the other direction either: MiniMax H3 Max tops out at 768p, so the speed is partly a function of the ceiling. You are buying the fastest tier of the H3 family, and the resolution ladder stops where the speed starts.

Where it landed on the boards

Artificial Analysis' with-audio leaderboards, captured 2026-08-27, put MiniMax H3 Max at #1 in image-to-video with Elo 1,204 — a 13-point gap over ByteDance's Dreamina Seedance 2.0 720p at 1,191 and a 20-point gap over the base MiniMax H3 at 1,184. In text-to-video it sits at #3 with Elo 1,235, inside a six-point photo finish at the top: Wan 3.0 at 1,240, Gemini Omni Flash at 1,237, MiniMax H3 Max at 1,235, and the base MiniMax H3 at 1,226. These are blind human-preference Elo scores from Artificial Analysis' own arenas, not vendor numbers, and like every arena they drift as new models and new votes arrive.

A scoreboard card for MiniMax H3 Max listing Image-to-Video #1 (Elo 1,204), Text-to-Video #3 (Elo 1,235), generation speed of a 5s clip in ~2.8s (vendor-reported), max resolution 768p, clip length 5-15s with native audio, and price $0.04/s ($2.40/min) promo ending September 1.

Image-to-Video (with audio) — MiniMax H3 Max #1, Elo 1,204. Dreamina Seedance 2.0 720p #2, 1,191. MiniMax H3 #3, 1,184.

Text-to-Video (with audio) — Wan 3.0 #1, 1,240. Gemini Omni Flash #2, 1,237. MiniMax H3 Max #3, 1,235. MiniMax H3 #4, 1,226.

Capture date — 2026-08-27; both boards shift as votes accumulate.

The launch price, and the clock on it

The pricing has a deadline, which is the part most coverage will miss. fal is running a 50% launch discount on MiniMax H3 Max that its own model page says ends September 1. At the promotional rate, 768p costs $0.04 per second — $2.40 per minute — and 480p costs $0.025 per second — $1.50 per minute. After the discount, the standard rates are $0.08 per second at 768p ($4.80 per minute) and $0.05 per second at 480p ($3.00 per minute).

Put that next to the base MiniMax H3 and the promo is doing a lot of the work. On the same Artificial Analysis board, the base model is tracked at $7.80 per minute — but that is the 2K rate. At the 768p resolution where MiniMax H3 Max actually competes, the base MiniMax H3 lists at $0.08 per second, or $4.80 per minute — exactly what the post-train's standard rate will be after September 1. In other words: today's headline gap of $2.40 versus $7.80 is really a comparison between a discounted 768p tier and an undiscounted 2K tier. The clean comparison is at 768p, and there the price edge lasts until the discount ends.

Screenshot of the Artificial Analysis image-to-video leaderboard with audio, captured August 27 2026, showing MiniMax H3 Max at #1 with Elo 1,204 and a price of $2.40/min, ahead of Dreamina Seedance 2.0 720p and MiniMax H3.

What fal actually changed

fal describes MiniMax H3 Max as post-trained from MiniMax H3 for stronger prompt adherence and better aesthetics, then co-optimized with fal's custom inference stack for higher throughput — the throughput part being where the speed comes from. What is measurable from the outside is the output envelope: 5-to-15-second clips with native audio, at up to 768p, with the same family's text-to-video and reference-to-video endpoints available around it. The 768p ceiling is the real constraint to remember. It means "Max" is a claim about the leaderboard and the speed, not about the resolution ladder: the base MiniMax H3's open weights also top out at 768p, while its hosted API reaches 2K through an API-only regeneration module. A post-train derived from the open release inherits the same cap.

The open-weights question

fal has stated its intent to release the MiniMax H3 Max weights, which would make the post-train the highest-ranked open-weights model on both with-audio boards — ahead of the base MiniMax H3, which holds that title today. Whether that lands as genuinely open is the open question. The base MiniMax H3 opened on August 3 under a custom community license that, per MiniMax's own terms, excludes use of the model and its outputs in the EU, the UK, South Korea, and the US, and requires prior written authorization for commercial use above $20 million in annual revenue. If H3 Max's weights ship under anything similar, "open" will carry the same territory drawn around it — with the extra wrinkle that a post-train built by a third party sits in a greyer licensing spot than the original.

Calling it from a router

For practical purposes, the routable half of this family is the base model. MiniMax H3 is on OrcaRouter today at the provider's list price — $0.08 per second at 768p, $0.13 per second at 2K — with 0% markup and automatic failover, so teams can call the H3 family through one API without integrating a vendor endpoint that is days old. MiniMax H3 Max itself is not routed yet; it is served only through its developer's own API. That is exactly the situation the failover habit is for: you can prototype against the current #1 image-to-video model without betting a production path on it, and if H3 Max reaches a routed provider it will appear at list price the same day — the vendor's promo included — rather than through a hand-rolled integration that takes your time and their markup.

Screenshot of the OrcaRouter model page for MiniMax H3, showing list pricing of $0.08 per second at 768P and $0.13 per second at 2K, plus performance figures.

What to watch in the next few weeks

Three dates and decisions decide whether this debut is a moment or a footnote. First, September 1: when the 50% discount ends, the 768p price doubles to parity with the base model, and the "$2.40 image-to-video leader" becomes a "$4.80 image-to-video leader" — which changes the value story for anyone planning on sustained volume. Second, the weights: whether fal actually ships them, and under what license. Third, the arena: whether the six-point photo finish at the top of text-to-video holds as votes accumulate. None of that is settled, which is exactly why a sub-real-time model that just took the top of a leaderboard at half price is the start of the story rather than the end.

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube