Hero title card for Grok Imagine Image 2.0 reading 'Grok Imagine Image 2.0 — #2 on the text-to-image Arena, precision editing, now live in Grok apps' with the caption line 'Announced August 7, 2026 by xAI' and a flat picture-frame-with-pen editing icon.
Guides & Insights

Grok Imagine Image 2.0: #4 on Artificial Analysis, at a Third of the Price

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Grok Imagine Image 2.0 has landed at #4 on Artificial Analysis's Text to Image Leaderboard, and the number that matters is not the rank — it is the price next to it. The model sits at 1,153.66 Elo, behind three entries: GPT Image 2.5 Flare (max), GPT Image 2.5 Sunburst (max) and GPT Image 2 (high). Ahead of every other lab's model, including Microsoft's MAI-Image-2.6 and the Nano Banana 2 model. That leaves the model as the highest-ranked image generator outside that trio, and the cheapest thing in that top four by a wide margin: Artificial Analysis prices it at $60 per 1,000 images against roughly $211 for the three entries above it. The model itself is from August 7, 2026 — what changed is where the independent board puts it, and what that does to the "world #2" framing the maker used at launch.

What the #4 placement actually measures

Artificial Analysis's Text to Image Leaderboard is a blind human-preference arena: voters see outputs from the same prompt without model labels and pick the better one, and the resulting Elo is published independently of any vendor. Grok Imagine Image 2.0's entry reads #4 at 1,153.66 Elo with a 95% confidence interval of 1,141.66 to 1,165.66. None of that number comes from xAI. (One housekeeping detail worth knowing when you read the board: Artificial Analysis lists the model's creator as SpaceXAI.)

The price half of the story is the more interesting half. The same page plots quality against price with a Pareto line marking the models that return the most Elo per dollar, and on the board's own published figures Grok Imagine Image 2.0 is the highest-scoring model priced under $100 per 1,000 images. The next model above it on quality, GPT Image 2 (high), costs $211 per 1,000 — roughly three and a half times as much for about 17 more Elo. That is a price-performance claim, not a claim about being the best model, and it is the version of the story the independent data actually supports.

The frontier position is real but narrow, and it is worth saying so plainly. Microsoft's MAI-Image-2.6 sits one rank below at $38.90 per 1,000, and Meta's Muse Image is $10 per 1,000 at 1,111 Elo. Elo intervals on this board run about ±12 points, so #4 and #5 — separated by 6.5 Elo — are inside each other's error bars. Treat the placement as "top five," not as a settled slot.

The 14-place climb is the gap to the tier Grok Imagine Image 2.0 replaced. xAI's previous quality tier, grok-imagine-image-quality, sits at #18 on the same board at 1,043.21 Elo, and the original grok-imagine-image is #29 at 1,017.86. That is roughly 110 Elo and 14 rank places between the model xAI now documents as its image default and the one it supersedes.

Where the "world #2" claim stands now

At launch, xAI cited a #2 finish on both the text-to-image and image-editing Arena boards, at roughly 1,320 Elo and marked Preliminary. Those were figures xAI highlighted in its own announcement — not an independent evaluation it commissioned — and they have moved since: the August 10 snapshot put the model at #3 with 1,316 Elo, behind MAI-Image-2.6 and GPT Image 2. The image-editing #2 remains xAI's own citation.

Artificial Analysis is a different board with a different method and a different crowd of voters, and it now places the model at #4. The two do not really contradict each other — they sample human preference differently, and both are snapshots of boards that move week to week. The honest summary of six weeks of independent ranking is that Grok Imagine Image 2.0 is consistently top-five and never quite the #2 xAI announced.

What Grok Imagine Image 2.0 actually ships

xAI positions Grok Imagine Image 2.0 as an editing environment rather than another one-shot generator. The feature set:

Magic Wand — region-level edits that leave the rest of the frame untouched

Segmentation — precise region selection for color grading, replacement, or adjustment

Background removal — transparent-background subject export

Multi-image input — up to five source images per edit; xAI's documentation now lists five, up from the three the API capped at launch

Smart Resize — recomposes one image into 9 aspect ratios (1:2 tall to 2:1 wide), inventing new frame content rather than cropping

Templates — preset workflows for product photography, headshots, e-commerce listings, game assets, and marketing posters

The API exposes controls the consumer app does not surface: aspect ratio — including 21:9 and 5:2, both new with 2.0 — resolution (1K by default, 2K optional), and a quality parameter taking low, medium or auto. Consumer-side, the model ships as the default Quality Mode on grok.com/imagine, iOS, and Android, and since August 13 it has also been listed on Runway's platform, where it joined the image and video models Runway already hosts. That listing is a vendor-and-partner statement rather than an independent evaluation, but it was the first sign of third-party distribution after launch.

On headline text rendering, xAI says the model "plans typography and layout the way a designer would," keeping dense multi-part compositions intact and rendering small text crisply. It also preserves a subject's identity and settings across multi-turn iterative edits.

What a first independent test found

A six-prompt, single-seed, no-cherry-picking comparison against gpt-image-2 found a clear split, and nothing in the new leaderboard placement overturns it.

Grok wins: photorealism and identity retention. Portrait hand anatomy was correct, with pore-level skin detail, subsurface scattering, and natural catchlight and rim-light. In a 45-degree geometric rotation task, Grok held facial identity above the ArcFace 0.5 threshold while gpt-image-2 drifted more.

Grok loses: production-critical areas. Asked for a Simplified-Chinese film-poster headline, it rendered Traditional Chinese characters — called a "hard disqualifier" for localization and publishing pipelines. A watercolor style-fusion request returned a photorealistic image with only a light painterly texture pass — a "category-level disqualification." Local edits completed all three requests but failed to preserve the window view and mis-anchored a highlight.

Summary: Grok Imagine Image 2.0 "leads on photorealism and identity, and trails on editing reliability, multilingual text, and true style transfer." A leaderboard averages preference across thousands of blind comparisons; it cannot tell you whether the model will get your Chinese headline right. One six-prompt test is not a body of evidence either — but between the two, the specific failure is the more actionable signal for a team shipping localized assets.

The API and the pricing math

The developer story has changed since launch. grok-imagine-image-2.0 is now the model xAI documents for image generation, single-image editing and multi-image editing, and it is what xAI's own model page recommends for images. The launch-era "developer API access coming soon" line is gone from the vendor's model and pricing pages.

• grok-imagine-image-2.0: from $0.04 per image, billed at the quality actually served

• grok-imagine-image (1.0): $0.02 per image, unaffected by the change below

• grok-imagine-image-quality: $0.05 per image, retiring November 2, 2026

Quality is the lever on cost. Auto — the default when the parameter is omitted — currently serves low for generation and medium for editing, so a generation request bills at the low rate and an edit at the medium one. Pinning low or medium makes the bill predictable rather than dependent on which path the service picks.

That last bullet carries a deadline. xAI's migration guide says the grok-imagine-image-quality slug is retired on November 2, 2026, after a 60-day notice that began September 2. From that date the slug keeps resolving, but requests are served by grok-imagine-image-2.0 with quality set to low — a compatibility redirect, not a hard shutdown. Because the low tier is $0.01 per image cheaper than the old quality tier at every resolution, teams that change nothing will see a price drop and a silent change in output quality. Migrating deliberately, by naming grok-imagine-image-2.0 and pinning quality where stable output matters, is the version of this you want. That date and the redirect behaviour are xAI's own documentation — vendor-stated, not independently tested.

Against the OpenAI options, the comparison is a pricing-model difference as much as a price difference. gpt-image-2 is token-priced — $8 per million image-input tokens and $30 per million image-output tokens on OpenAI's published rates — so per-image cost swings with resolution and detail. Artificial Analysis's representative figure of $211 per 1,000 images for GPT Image 2 (high) against Grok's $60 is that swing showing up as a single number. Flat per-image pricing is far easier to forecast at scale, which is most of why a model in fourth place is worth a look at all.

Trying it without betting the pipeline

The reasonable reaction to Grok Imagine Image 2.0 is still "try it, don't commit to it yet." Its position sits inside overlapping confidence intervals, one independent test flagged real weaknesses on localized text and style transfer, and xAI is retiring a slug underneath it in November. That is exactly the case for routing rather than wiring a direct integration.

On OrcaRouter, grok-imagine-image — xAI's image model in our catalogue — sits on the same OpenAI-compatible endpoint as gpt-image-2, so moving between xAI and OpenAI image models is a model-string change: no second contract, no new SDK. OrcaRouter passes provider list prices through with no markup, so a vendor price cut is live here the same day, and automatic failover can send errors, rate limits or refusals to a fallback model rather than into your production path. One honest caveat, unchanged from the original review: xAI's newest quality tier is a separate entry from the Grok image model already in our catalogue, so check the current model list rather than assuming a given Grok variant is routable today.

What to watch next

1. The November 2 migration. Whether teams move to grok-imagine-image-2.0 explicitly or drift onto the redirect and its low-quality default is the largest thing xAI can still get wrong with this model — and the redirect makes the mistake invisible unless you read the model field in your responses.

2. Whether #4 holds. The gap to MAI-Image-2.6 is 6.5 Elo with ±12-point intervals, so a few thousand more votes could reorder the board in either direction. The 110-Elo gain over the previous tier looks durable; the exact rank does not.

3. How much of the consumer surface reaches the API. Generation, editing and multi-image editing are documented. Templates and the named Smart Resize workflow remain app features, and that gap is where the remaining "consumer-only" caveat lives.

Bottom line: Grok Imagine Image 2.0 is a real, capable model that has now been independently placed — #4 on Artificial Analysis's Text to Image Leaderboard, the best non-OpenAI entry there, roughly 110 Elo clear of the tier it replaced, and on the board's price-quality frontier at $60 per 1,000 images against roughly $211 for the OpenAI models directly above it. The "world #2" framing was always xAI's launch snapshot, and the independent boards have never quite said that. Its two clearest independent weaknesses — Simplified-Chinese compliance and style-transfer reliability — sit on features teams most often need from an image API. Trial it immediately for photorealistic, identity-preserving work; wait on it for pipelines that ship localized text or stylized assets, and if you are already calling the retiring quality slug, migrate before November 2.