A title card for the comparison, showing "Grok Imagine Video 1.5 Lite vs Grok Imagine Video" above the line "52 Elo for $6.60 a minute" and "$8.40 / min vs $15.00 / min", with two flat outlined card shapes split by a divider.
Guides & Insights

Grok Imagine Video 1.5 Lite vs Grok Imagine Video: 52 Elo for Six Dollars a Minute

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Here is the whole decision in one sentence. G​rok Imagine Video 1.5 Lite costs $8.40 a minute and scores Elo 993 on Artificial Analysis's AA-Video-T2V v2.0 board; G​rok Imagine Video 1.5 costs $15.00 a minute and scores Elo 1045. Same family, same interface, same 1-to-15-second clip length, same seven aspect ratios — and a $6.60-per-minute gap that buys you a 52-point preference score and one hidden difference that does not appear in either headline number. G​rok Imagine Video 1.5 is the original: x​AI's fast generation model, first released January 29, 2026, stable at version 1.5 since June 16, 2026, with native audio and native 1080p. G​rok Imagine Video 1.5 Lite is the distilled sibling, dated August 2026 on the board, quicker and cheaper by construction. The question this page answers is not which model is better — the board already decided that — but which one your output actually needs.

The one difference the price sheet hides

Before the bullets, the sentence that decides more real deployments than the Elo gap does. On G​rok Imagine Video 1.5 Lite, 1080p output is rendered at 720p and upscaled. On G​rok Imagine Video 1.5, 1080p is 1080p.

That is the whole of it. A Lite clip requested at 1080p is a 720p render with an upscale pass on top. On a phone feed, at social bitrates, through a compression step that is going to flatten the difference anyway, this is invisible. The moment the clip enters a pipeline that crops, keys, composites, reframes, or grades, the missing detail is the first thing an editor notices and the last thing they can fix. The full model does not have this problem, and it is the main reason its per-minute price is nearly double.

Dimension by dimension

• Price — $15.00 per minute for G​rok Imagine Video 1.5 against $8.40 per minute for G​rok Imagine Video 1.5 Lite, both quoted by Artificial Analysis at 1080p default settings. Published per-second rates for the full model run $0.096 at 480p, $0.168 at 720p and $0.300 at 1080p, plus $0.05 per image input.

• Board rank — #11 at Elo 1045 for the full model against #17 at Elo 993 for Lite, on the with-audio AA-Video-T2V v2.0 board. Lite's interval is 983 to 1003, a ±10 band that does not overlap the parent's position seven rows up.

• Effective resolution — native 1080p against a 720p render upscaled to a 1080p container. The single most consequential row in this list.

• Clip length — 1 to 15 seconds in whole seconds for both. In the chat interface the full model defaults to 10 seconds; the API is where the longer durations live.

• Aspect ratios — the same seven for both: 16:9, 9:16, 1:1, 4:3, 3:4, 3:2 and 2:3.

• Audio — the full G​rok Imagine Video 1.5 ships native audio with dialogue, effects and music. Lite's own listing does not document an audio channel separately, and the announcement places it on the board's no-audio variant as well, so if a soundtrack is load-bearing for your deliverable, verify it on Lite before you commit.

• Inputs and conditioning — the family API takes a first-frame image and additional reference images, with the full model documented for reference-conditioned generation from up to seven images for character consistency. Lite's listing describes text-to-video and image-to-video support.

• Maturity — the full model has been stable at 1.5 since June 16, 2026, topped the Artificial Analysis arena in both text-to-video and image-to-video at launch, and held the #1 image-to-video slot on LMArena at roughly 1,400 Elo. Lite has an August 2026 board date and 5,715 blind comparisons behind its score.

A two-column scoreboard comparing Grok Imagine Video 1.5 Lite ($8.40/min, rank #17, Elo 993, 720p upscaled to 1080p, 1–15 second clips, audio not documented) with Grok Imagine Video 1.5 ($15.00/min, rank #11, Elo 1045, native 1080p, 1–15 second clips, native audio), with a footer separating the Artificial Analysis figures from the vendor-reported details.

The bill, worked out

Abstract per-minute pricing is hard to act on, so put it against a real brief. Suppose you need twenty finished ten-second clips a month.

• On G​rok Imagine Video 1.5 Lite at $8.40 a minute, ten seconds is a sixth of that, so $1.40 per generated clip. Twenty of them is $28.00 a month, before retries.

• On G​rok Imagine Video 1.5 at $15.00 a minute, the same ten seconds is $2.50, and twenty clips is $50.00 a month.

• The difference is $22.00 a month, or roughly $264 a year, for a hundred and fifty or so clips a year.

That is the number that makes the cheap tier look like a rounding error in the wrong direction, and it is worth saying plainly: at this volume the price gap is not the reason to pick Lite. The reason to pick Lite is speed of iteration. If Lite returns a draft faster, the twenty clips above are really sixty or eighty generations with two thirds of them thrown away, and the same $22 math turns into a genuinely different month. The per-second price is the visible lever; the wall-clock per generation is the lever that moves the bill.

Which reframes the choice as a workflow question rather than a price question. Failover is the other practical use for a cheaper sibling: routing a first pass at Lite and re-running only the keepers at full quality is a two-line change in which model the request names, and it is the cheapest way to find out whether Lite is good enough for your content without betting a delivery date on the answer. Neither model is served by OrcaRouter — we host neither G​rok Imagine Video 1.5 nor G​rok Imagine Video 1.5 Lite, and both are reached through the vendor's own API and several third-party platforms — but the layer that decides what to generate, writes the prompts, and files the output is text work, and text work is what OrcaRouter covers: one API across 200+ models at provider list price with 0% markup, automatic failover, and a routing DSL that composes several models into one call.

When the cheaper model is the wrong call

The failure mode is specific and worth naming, because it is easy to walk into. Teams adopt Lite for cost, build a month of work on top of it, then discover at the first deliverable that the 1080p footage does not survive the crop.

Lite is the right model when the clip's destination is a feed, a story, a social post, or an internal review — anywhere the output is watched at the size it was generated. It is the right model when you are iterating and the keeper will be regenerated at full quality anyway. It is the right model when volume is high and per-unit cost is the constraint that matters.

It is the wrong model when the clip will be cut into a longer piece at a different aspect ratio, when a compositor will pull a matte or a key, when the footage will be graded, or when the deliverable has a resolution specification attached to a contract. In all of those cases the $6.60 a minute you saved is paid back at the point where someone has to regenerate everything, and the Elo gap is beside the point.

A capture of the LMArena text-to-video leaderboard showing grok-imagine-video-1.5-720p in seventh place, behind wan-3.0-720p-audio, gemini-omni-flash, flux-3-video-20260811 and dreamina-seedance-2.5-720p.

Reading the board honestly

Two caveats on the numbers above. The first is the usual one for arena scores: 993 and 1045 are blind pairwise preference votes, not measurements of whether a clip matches a brief. A model can win a vote on color and lose the shot. The second is specific to this pairing — the gap between the two models is small enough that a reader should treat it as an ordering rather than a distance. Fifty-two Elo means Lite loses more comparisons than it wins against its parent; it does not mean Lite is 5% worse at your shot, and it certainly does not mean a Lite clip that happens to nail the brief is worse than a full-model clip that missed it.

What the board does establish is the direction and the size of the trade. Lite is neither the cheapest nor the best model in its band of the table: PixVerse V6 sits one row below it at Elo 985 for $6.90 a minute, and Kling 3.0 1080p (Pro) sits one row above at Elo 1000 for $10.08. What is unusual is the gap to its own parent — 52 Elo for 44% off is a smaller quality giveaway than the tier-below-a-flagship discount usually costs, and that ratio is what the frontier claim in the announcement rests on. It is also why this is a real choice rather than a marketing tier.

A capture of the OrcaRouter models page, showing the model catalogue with vendor groupings and model entries.

Which one, by output

If your clips are watched where they are generated — social, mobile, review, iteration — G​rok Imagine Video 1.5 Lite is the version to start with, and the $8.40 minute rate means a hundred experiments cost less than twenty finished deliveries. Regenerate the survivors on the full model if any of them graduate.

If your clips are finished rather than viewed — cropped, keyed, graded, delivered to a spec — skip Lite entirely and pay the $15.00. G​rok Imagine Video 1.5 is the only model of the two that returns native 1080p, and no amount of saved per-second cost survives a deliverable that has to be regenerated.

Almost nobody needs both at once. Everybody who generates at volume will end up using both, in that order — Lite to find the shot, the full model to keep it.