
HeyGen Video Launches at $0.01 a Second: What HeyGen's Post-Train of MiniMax H3 Actually Changes
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 118 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
HeyGen Video went live on 30 September at $0.01 per second through October, and that number is the whole story. A ten-second clip costs a dime. On HeyGen's own comparison chart the base generation it was tuned from — MiniMax H3, released 31 July 2026 — bills at $0.08 per second at the same 768p with audio, and the chart's other reference points run up to $0.40 per second for Google's Veo 3.1. Nothing about the launch is exotic: one model, three prompting modes, sound generated in the same call. What is unusual is a company taking an open-weights base model, post-training it for its own vertical, and pricing the result below the model it was derived from by nearly an order of magnitude — while saying in the same breath that blind testers ranked it ahead of that base. Both claims are HeyGen's, and the arithmetic behind them is where the buyer's decision actually lives.
What HeyGen Video actually is
HeyGen Video is a general-purpose video model, not an avatar engine, and the distinction matters if you have used HeyGen before. The company's own documentation draws the line explicitly: its avatar rendering engines animate a look you already own from a script you already wrote, while this one invents the subject, the setting, the light and the sound at once from a description. It ships under the identifier heygen-video-1 on POST /v3/models/videos, authenticated with an x-api-key header, and it runs only on paid API keys.
Three modes cover the conditioning:
• text_to_video — prompt only, output defaults to 16:9.
• image_to_video — prompt plus one image, which becomes the literal first frame including its EXIF orientation; the output follows that image's proportions.
• reference_to_video — the default mode. Prompt plus up to nine reference images, three reference videos and three reference audio clips, twelve items in total. List order becomes the label you address in the prompt: the first entry of reference_images is <Picture 1>, the first video is <Video 1>, and images, videos and audio are numbered independently.
The output envelope is tight, and worth reading before the price:
• Duration — any whole number of seconds from 5 to 15, defaulting to 5.
• Resolution — 480p or 768p, defaulting to 768p. At 16:9 that is 1344 × 768; at 9:16, 768 × 1344.
• Frame rate and audio — 24 fps in an MP4/H.264 container, with AAC audio at 32 kHz stereo carrying generated dialogue, ambience and effects.
• Aspect ratios — 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, plus adaptive for reference mode, which inherits the first reference image's shape.
• Prompt enhancement — turbo by default, quality for a longer pass, disabled to send your text through untouched.
The schema is strict — unknown fields are rejected rather than ignored — and seed is an unsigned 32-bit integer that makes a shot repeatable when you hold it and change one clause at a time. Download links are signed and expire, so polling refreshes them; callbacks are attempted once per job and HeyGen itself tells you to treat them as a latency optimisation rather than a guarantee.
The vendor's own numbers, and the caveats built into them
HeyGen put a comparison chart on the launch page, and it is the only quantitative claim in the piece, so it is worth taking apart rather than repeating.
The quality chart is an Elo table with HeyGen Video at the top — 1000 — then Dreamina Seedance 2.0 at 955, the H3 Max family at 952 down to 860, Kling 3.0 Pro at 857 and Veo 3.1 at 744. The footnote is doing a lot of work: the scores come from an internal eval set of Artificial Analysis Arena queries with 4,800 votes, and HeyGen's own score is normalised to 1000 for simpler comparison. A vendor normalising itself to the top of its own chart is not evidence that it leads; it is a presentation choice. The relative spacing below it is the informative part.
More useful is the head-to-head. HeyGen says blind testers preferred its model to fal's H3 Max in 55.8% of comparisons across 650 votes, with a range of 49–62%. That range is the honest detail: at the low end it straddles 50%, so on HeyGen's own numbers a preference over H3 Max is not clearly separated from noise. The per-axis breakdown is also mixed — 65% on staying true to the source image and 66% on timing, against 43% on consistency and 45% on music — which reads like a model that holds a single contained shot well and drifts over longer ones.

Speed is the least contestable figure. HeyGen reports its diffusion inference at 3.7 seconds for one ten-second image-to-video clip, against 4.1 seconds for H3 Max Turbo and 8.3 seconds for H3 Max, with captioner time excluded because it varies by mode. Anything under the length of the clip it produced is faster than real time.
The price, and a discrepancy in HeyGen's own copy
Two HeyGen sources give two different standard rates. The launch copy says pricing starts at $0.01 per second through October, described as 50% off a standard $0.02 — which makes October the promotional rate and $0.02 the list. The cost chart on the same page, footnoted as list price per second at 768p with audio before launch promotions, puts HeyGen Video at $0.03.
That is not a rounding difference; it is a 50% gap between what HeyGen's prose says its standard rate is and what HeyGen's chart says. The chart's other rows are consistent with outside pricing — MiniMax bills the base MiniMax H3 at $0.08 per second, which matches the $4.80 per minute Artificial Analysis records for it — so the H3-family figures are not the odd ones out. Until HeyGen publishes an API pricing page, treat the $0.01 as confirmed because it is the live promotional rate, and treat the post-October number as somewhere between $0.02 and $0.03 per second.

What a thousand seconds costs
Here is the comparison that matters, at 768p with audio, for one thousand seconds of generated video — a hundred ten-second clips:
• HeyGen Video in October — $10 at $0.01 per second.
• HeyGen Video after October — $20 at $0.02, or $30 if the chart's $0.03 is right.
• MiniMax H3 on OrcaRouter — $80 at the provider's list price of $0.08 per second, with 0% markup, so any vendor price move is live here the same day.
• fal's H3 Max — $80 at the charted $0.08 per second.
• Kling 3.0 Pro — $168 at $0.168 per second.
• Veo 3.1 — $400 at $0.40 per second.
For teams generating in bulk and discarding most of what they make — social cutdowns, ad variants, product loops — the discard rate is the real cost line, and at a dime a ten-second attempt the economics of iterating change shape. The catch is the ceiling: 768p at 24 fps is not a 2K delivery format, and "production-quality" is HeyGen's phrase, not a specification. The base MiniMax H3 reaches 2K through a separate regeneration module on MiniMax's hosted API at $0.13 per second, and HeyGen Video has no equivalent — so if your delivery target is above 768p, the cheap second is not the second you need.
For teams that want to compare the two, MiniMax H3 is on OrcaRouter as minimax/minimax-h3 at $0.08 per second, so the base model and the post-train can be put behind one key instead of two contracts. HeyGen Video itself is not a routed model here today — it is served through HeyGen's own API and several third-party platforms — which is exactly the case automatic failover exists for: you can prototype against a days-old model without betting a production path on it, and if it reaches a routed provider it appears at list price rather than a marked-up one.

What HeyGen says it is bad at
The most useful paragraph in HeyGen's documentation is the one listing failure modes, because vendors rarely write them. HeyGen Video is, in the company's own words, least reliable on long on-screen text, on soft organic motion such as petals, paper and hair, and on close hand work like assembling or operating equipment. It is at its best on short, contained shots: one subject, one place, one action, a locked or barely moving camera.
That profile lines up with the per-axis blind numbers above — good adherence to a supplied image, weak consistency — and it should shape how you test it. The capability the launch films lean on hardest is lettering generated inside the shot: every title in them was painted, stamped or pressed into the scene rather than composited. HeyGen's own guidance is that short phrases spelled out exactly render reliably, which is a narrow claim, and the same docs warn off long on-screen text. If your workflow depends on readable typography, that is the line to probe first.
The second MiniMax H3 post-train in five weeks
The interesting structural fact is that HeyGen Video is not the first commercial product built by fine-tuning MiniMax H3. fal shipped H3 Max on 27 August, a post-train of the same base co-optimised with fal's inference stack, and it debuted at the top of Artificial Analysis's with-audio image-to-video board. Now HeyGen has done the same thing for a different vertical — business video, product demos, training, property walkthroughs — and priced it below the base.
That pattern is the real signal. MiniMax's decision to release H3's weights under its own community licence turned the model into a substrate other companies can specialise and resell, and the second such product in five weeks suggests the conversion rate is real. It also means "MiniMax H3" is becoming a family rather than a model: the base, fal's H3 Max, and HeyGen Video now compete with each other at different prices and different resolutions, and a buyer comparing "H3" to anything else has to say which H3.
Artificial Analysis's text-to-video board shows how strongly that can land. Its second-ranked model with audio, at Elo 1150, is Utopai X — footnoted by Artificial Analysis as based on MiniMax H3 — ahead of the base MiniMax H3 itself at 1138 in fourth. Whatever the base model's own score, derivatives built on it are outperforming it in blind human preference on at least one independent board.
What is still unverified
Three things a reader should hold open.
First, HeyGen Video has no independent score. As of 1 October it appears on none of Artificial Analysis's three video boards — text-to-video, image-to-video or video-editing — so there is no blind, third-party Elo for it yet. Every quality number in this piece is HeyGen's own, on an eval set of Arena queries that HeyGen assembled with 4,800 votes. That is a defensible methodology and a vendor result, and it should be read as one.
Second, HeyGen has said nothing about releasing weights. The model it was tuned from is open-weights, but a post-train is a new artefact and HeyGen's own licences are its own; nothing in the launch material commits it either way.
Third, the documentation disagrees with itself. The API reference states the prompt accepts at most 5,000 Unicode characters, while the parameter table on the models page gives the range as 1 to 32,000 characters — a 6× spread on a hard input limit. The lower figure is the one that appears in the reference and in the changelog summary, so plan against 5,000 until HeyGen reconciles the two.
Who should act, and who should wait
If your video work fits inside a contained shot — a product on a bench, a room with a slow dolly, a single subject and a locked camera — and you deliver at 768p or below, the through-October rate is the cheapest credible way to generate narrated, scored footage at scale that anyone is currently publishing. Test it this month while the promotion holds, and design the test around the three things HeyGen admits it is weak at rather than around the launch film.
If your delivery target is above 768p, if you need long on-screen text, or if readable consistency across shots is the point, wait. The base MiniMax H3 remains the routable H3 at $0.08 per second for 768p, and its hosted 2K path sits at $0.13 per second — both behind one key on OrcaRouter, at the provider's list price with no markup, so if HeyGen's post-October price turns out to be $0.03 rather than $0.02 you will see the correction here the day it happens rather than the month after. What would change the verdict fastest is an independent Elo: once HeyGen Video shows up on a blind third-party board, the vendor's 55.8% either holds up or it does not.
