A generated title card reading 'Vidu Q4 Preview vs MiniMax H3' with the subtitle 'Sixteen Elo apart, three boards apart' and three stat tiles reading '1,195 MiniMax H3 Max', '1,181 MiniMax H3' and '1,179 Vidu Q4 Preview', each captioned 'image-to-video Elo'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Vidu Q4 Preview vs MiniMax H3: The Top of the Board Is a Tie, and Only One of Them Is On Three Boards

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The image-to-video board at Artificial Analysis, read 8 October 2026, puts MiniMax H3 Max first at 1,195 Elo, MiniMax H3 second at 1,181, and Vidu Q4 Preview third at 1,179. Sixteen points separate first from third, on intervals of roughly ±10 and ±11, across 5,894, 7,284 and 5,543 votes respectively. That is a statistical tie dressed as a podium, and any page that opens by declaring a winner on this board is reading noise.

So the interesting question is not which of these two models is better. It is what happens to the comparison once you step off the one board where both appear — because MiniMax H3 was released 31 July 2026 as an omni-modal model that reads text, image, video and audio references as a single context, and Vidu Q4 Preview arrived on 7 October 2026 with two input modes and two API endpoints. One of them is measured in three places. The other is measured in one.

What the three-way tie does and does not mean

Both MiniMax entries and the Vidu build sit inside each other's intervals, so the honest statement is that the top three image-to-video models are indistinguishable by human preference at the current sample sizes. Two details make that tie less flattering to everyone than it sounds.

The first is age. MiniMax H3's votes are from July 2026 and H3 Max's from August; Vidu Q4 Preview's are from October. Bigger, older vote pools are measured against a different competitive field and a different set of opponents, which cuts both ways: better measured, and about a slightly different question. A 16-point gap in either direction here is not evidence.

The second is the price column. Artificial Analysis records MiniMax H3 at $7.80 per minute and H3 Max at $4.80 per minute on this board. Vidu Q4 Preview is recorded at $7.20. Read those as a ranking and H3 Max looks like the value pick of the three — but the underlying vendor rates explain the label rather than contradicting it. We list MiniMax-H3 at $0.08 per second at 768P and $0.13 per second at 2K on our own model card, which is $4.80 and $7.80 per minute. The two MiniMax figures on the board are the same model priced at two resolution tiers, not a premium tier and a cheap tier. Per-minute columns on a preference board are useful for orders of magnitude and dangerous for anything finer.

Where MiniMax H3 shows up and Vidu Q4 Preview does not

This is the part of the matchup that survives the tie, and it is a structural difference rather than a quality one.

MiniMax H3 (768p) is fourth on the text-to-video board at 1,137 Elo over 8,315 votes, with H3 Max fifth at 1,130 over 5,719. It is fourth again on the video-editing board at 1,119 over 7,023 votes, on a board that ranks models on editing a video from a text instruction with the audio kept. Vidu Q4 Preview appears on neither.

That absence is not a measurement gap — it is what the product is. Vidu Q4 Preview's API documents two endpoints, /ent/v2/img2video and /ent/v2/reference2video, and the product page lists two modes, Image-to-Video and Reference-to-Video. There is no text-to-video path in the documentation. There is no video-editing path either, which means the model cannot take an existing clip and change something in it. MiniMax H3 does both, and its omni-modal framing — text, image, video and audio references in one context — is the reason it can.

Worth being precise about the audio side, because both models do native audio and the two are not the same thing. Vidu Q4 Preview generates audio and video together, with an audio parameter in the image-to-video endpoint that outputs a video with dialogue and sound effects, and it accepts up to three reference audio clips to keep a character's voice consistent. MiniMax H3 outputs native stereo audio and accepts audio as an input modality alongside images and video. The reference-voice use case is Vidu's specific strength; the reference-everything use case is MiniMax's.

The per-second arithmetic, tier by tier

Both models bill per second of generated output, which makes this the rare comparison where the rate cards can be laid against each other without a conversion.

• 540p — Vidu Q4 Preview only: 9 credits per second, about $0.045 at the platform's published $0.005 standard credit rate

• 720P / 768P — Vidu Q4 Preview 19 credits per second, about $0.095 vs MiniMax-H3 $0.08 per second

• 1080p — Vidu Q4 Preview 24 credits per second, about $0.12 vs MiniMax-H3 no published 1080p tier

• 2K — Vidu Q4 Preview 38 credits per second, about $0.19 vs MiniMax-H3 $0.13 per second

• 4K — Vidu Q4 Preview only: 78 credits per second, about $0.39

The pattern that falls out of that is not "one is cheaper." Vidu Q4 Preview's floor is genuinely lower — its 540p tier is a blocking tier priced for exactly what it is used for, proving a composition before paying for a full-resolution render, and nothing in MiniMax's rate card plays that role. MiniMax H3 is cheaper at 2K, the top tier both models share, by about a third. And Vidu's 4K tier is the only 4K generation tier in the comparison, at a price that reflects it.

The launch figure of $0.014 per second that accompanied Vidu's announcement is the 540p tier at a discounted credit rate, not a general rate. Shengshu's platform is currently running limited-time package discounts of up to 30% on credit bundles, which move every tier down together rather than changing the shape of the curve. Treat the list-rate curve as the comparison and the discount as a temporary offset.

On our side: we route MiniMax-H3. It is live on the public model API at minimax/minimax-h3, billed per second of generated output at the vendor's list rate passed through with nothing added, callable on the same OpenAI-compatible endpoint as every text model we serve. Vidu Q4 Preview is not an OrcaRouter route, and this page is not claiming otherwise — the video models we do serve sit on one key and one request shape next to the language models, which is the arrangement that makes comparing several of them cheap. One practical consequence of pass-through pricing: when a vendor moves a per-second rate, that movement is visible on the route the same day rather than at contract renewal.

What "omni-modal" buys in an actual request

MiniMax H3's unified context is easy to read as a feature list and miss as an engineering difference. A model that accepts video as an input reference can be pointed at an existing shot and asked to continue, extend or restyle it; a model with no video input cannot. That is the mechanism behind H3's presence on the editing board, and it is also why the model's 4–15 second output range is less of a limitation than the number suggests — multi-turn extension is the normal way to build a longer sequence out of it.

Vidu Q4 Preview's 16-second ceiling is per generation too, and its Reference-to-Video mode is built for multi-shot storytelling within that window, with the documentation describing intelligent camera switching and consistency across multiple camera positions. What it does not have is a route to accept a clip you already shot and modify it. If your workflow starts from footage rather than from a still or a reference set, one of these two models is simply not available to you, and no Elo number closes that.

Which one, by job

Take MiniMax H3 when the input is anything other than an image or a reference set. Text-to-video, video-to-video editing, audio-in, multi-turn extension past fifteen seconds, or a pipeline where the model needs to sit on three different boards' worth of use cases — that is H3's territory, and it is the only one of the two that has any of it. Its 2K rate is also the cheaper of the two at the top tier they share.

Take Vidu Q4 Preview when the job is consistency of cast and voice, or when you need a 4K generation tier, or when you want a cheap 540p blocking pass to iterate a shot before committing. Fifteen image references and three audio references are a larger published reference budget than MiniMax documents, and no other model in this comparison generates natively at 4K. The narrower mode set is the trade for that.

What would change this: a full Vidu Q4 release that adds a text-to-video endpoint would put the two models on the same boards and turn this into a straight contest. AA-Video-I2V v2.0, announced as coming with 1080p throughout and a revised capability taxonomy, would replace the votes at the top of the board rather than re-sort them — and it would settle whether the current three-way tie is a tie or a measurement limit. Until then, the number to keep is sixteen, with the board name and the date beside it.

The one thing worth taking away from the podium itself: a first place that sits sixteen Elo above third, on intervals that wide, is not a ranking. It is three models that human voters cannot separate, and the decision between them has to come from everything else on this page.

A generated two-column scoreboard titled 'Vidu Q4 Preview vs MiniMax H3 - the scoreboard'. The left column reads: Image-to-video Elo 1,179 (3rd); Text-to-video not present; Video editing not present; Max resolution 4K generation tier; Clip length 1-16 seconds; Price at 2K about 0.19 USD per second. The right column reads: Image-to-video Elo 1,181 (2nd); Text-to-video 4th, 1,137; Video editing 4th, 1,119; Max resolution 2K; Clip length 4-15 seconds; Price at 2K 0.13 USD per second. A footer reads 'All Elo and rank figures per Artificial Analysis, 8 Oct 2026; prices per vendor rate cards.'A screenshot of the Vidu API documentation for Vidu Q4 Preview, showing the left navigation with the Vidu Q4 Preview entry selected and the 'Image to Video' page open, including the Create Task endpoint, the Authorization header and a cURL request example whose body sets model to viduq4-preview with a duration and resolution and the audio parameter set true.A screenshot of the OrcaRouter model page for minimax/minimax-h3, showing MiniMax-H3 described as MiniMax's omni-modal video generation model and the Hailuo 3 generation, released July 31, 2026, reading text, image, video and audio references as one unified context and generating 4-15 second videos at 768P or 2K with native stereo audio, billed per second of generated output at $0.08/s at 768P and $0.13/s at 2K, with an OpenAI-compatible code sample and a per-second price of $0.0800.