Hero title card for Grok Imagine Image 2.0 reading 'Grok Imagine Image 2.0 — #2 on the text-to-image Arena, precision editing, now live in Grok apps' with the caption line 'Announced August 7, 2026 by xAI' and a flat picture-frame-with-pen editing icon.
Guides & Insights

Grok Imagine Image 2.0: xAI's 'World #2' Image Model, Dissected

Penulis

Jim Song

Tanggal Terbit

Model terbaru · 20Lihat semua model
Benchmark: Artificial Analysis · diperbarui setiap hari
Kembali ke semua artikel

Grok Imagine Image 2.0, the image model x​AI shipped on August 7 as the new default Quality Mode inside G​rok, is the second-best image generator in the world right now — behind O​penAI's gpt-image-2 — if you trust the public leaderboard, which is exactly the thing to check before believing it. The text-to-image Arena snapshot x​AI points to does put grok-imagine-image-2.0 at number two, and the same announcement claims the number-two spot on the image-editing board. What follows is what actually shipped, how solid the second-place claim is, what a first independent test turned up, and what it will cost to call the model from code.

The release itself is real and recent: x​AI rolled Grok Imagine Image 2.0 out to the consumer apps — grok.com/imagine, iOS, and Android — on August 7, 2026, as the "Quality Mode" experience. Developer API access is listed as coming soon in most of the launch coverage, even though the quality-tier model behind the experience is already documented in x​AI's Imagine API. That gap between the consumer launch and the developer story is worth understanding, because the tools that make Image 2.0 interesting are editing tools, and they are exactly what you would want to call from an API.

What Grok Imagine Image 2.0 actually ships

x​AI is positioning this as an editing environment, not another one-shot generator. The feature set is built around touching part of an image and leaving the rest alone:

• Magic Wand — region-level edits where only the area you point at changes, and the rest of the frame is left untouched.

• Segmentation — select a precise region for separate color grading, replacement, or adjustment.

• Background removal — export a subject with a transparent background for use in other software.

• Multi-image input — up to 5 reference images in a single generation on the consumer side, so you can fuse a product, a style, and a layout without manual compositing. The current API documentation caps it at 3 reference images per request.

• Smart Resize — recompose one image into 9 aspect ratios, from 1:2 tall banners to 2:1 wide banners, with the model inventing the content that fills the new frame rather than cropping.

• Templates — preset workflows for product photography, professional headshots, e-commerce listings, game assets, and marketing posters.

The other headline is text rendering. x​AI says the model plans typography and layout the way a designer would, keeping dense multi-part compositions intact and rendering small text crisply — historically the weak spot of diffusion image models. The model was trained for fidelity across photography, graphic design, and illustration, and it holds a subject's identity and settings across multi-turn iterative edits, which is what makes the editing tools usable in a real workflow rather than as a party trick.

Scoreboard for Grok Imagine Image 2.0 listing six rows: Text-to-Image Arena #2, Elo 1,320 (preliminary); Image Editing Arena #2, Elo 1,439 (xAI-cited); Region editing Magic Wand + segmentation; Text rendering crisp, designer-planned; Multi-image input up to 5 references; API price from /usr/bin/bash.02/image (quality tier /usr/bin/bash.05), with footer 'Text-to-Image figure per arena.ai, Aug 2026; editing figure as cited by xAI.'

The 'second-best in the world' claim, and what it covers

Start with the leaderboard itself. On the August 2026 text-to-image Arena snapshot, across 76 models, O​penAI's gpt-image-2 (medium) leads with an Elo around 1,380, and Grok Imagine Image 2.0 — listed under the name grok-imagine-image-2.0, now branded "SpaceXAI" after the parent company's takeover of x​AI — sits second at around 1,320. That is the source of the "second-best image model in the world" headline, and it is real: the model genuinely appears at #2 on the public board.

LMArena Text-to-Image Arena leaderboard, August 2026, showing gpt-image-2 (medium) first at Elo 1,380 and grok-imagine-image-2.0 by SpaceXAI second at Elo 1,320 across 76 models, the entry marked Preliminary.

The asterisk is labeling, not existence. These are public-leaderboard figures that x​AI is highlighting in its own announcement, not an independent evaluation it commissioned. The image-editing #2 is x​AI's own citation of the editing board, and the gap to gpt-image-2 on either board is not broken out in the launch material. Arena standings also move week to week as new models and new votes land. None of that makes the second-place claim false; it makes it a snapshot, which is a different thing from a settled verdict.

What a first independent test found

The most useful thing that happened in the days after launch was a six-prompt, single-seed, no-cherry-picking comparison that ran Grok Imagine Image 2.0 and gpt-image-2 through identical tasks: compositional semantics, photorealistic anatomy, a multilingual poster, a geometric transform, a local edit, and multi-reference style fusion. It is one test, not a verdict, but it draws the boundary of the "#2" claim much more sharply than the leaderboard does.

Where G​rok wins is exactly where the leaderboard suggests it would. In the photorealistic portrait task, hand anatomy was correct, skin showed genuine pore-level detail with convincing subsurface scattering, and catchlight and rim-light setup read naturally. In the geometric-transform task — rotating a full-body subject 45 degrees with no pose listed in the prompt — G​rok held facial identity above the ArcFace 0.5 threshold while gpt-image-2 drifted more on identity. For raw photographic quality and identity retention, the model is genuinely strong.

Where it loses is the part that matters for production. Given a prompt for a Simplified-Chinese film-poster headline, G​rok rendered Traditional Chinese characters instead — the test called it a "hard disqualifier" for any localization or publishing pipeline. Asked to fuse a portrait, a watercolor style, and a scene layout into a watercolor painting, it returned a photorealistic image with only a light painterly texture pass — the test's "category-level disqualification." On the local-edit task, all three requested edits were completed, but the model failed to preserve the window view in the scene and mis-anchored a highlight on the new object, where gpt-image-2 kept the scene intact. The summary line from that test: G​rok leads on photorealism and identity, and trails on editing reliability, multilingual text, and true style transfer.

The API and the pricing math

Here is where the launch and the developer story separate. The consumer Quality Mode shipped August 7, and x​AI's launch materials list developer API access as coming soon. But the model that powers the experience — grok-imagine-image-quality — is already live in x​AI's Imagine API, documented for both image generation and image editing with natural-language prompts.

Screenshot of xAI's Imagine API documentation showing grok-imagine-image-quality for image generation and editing at /usr/bin/bash.05 per image for 1K and 2K resolution, the note that image edits are billed for both the input image and the generated output, and grok-imagine-video-1.5 at /usr/bin/bash.08 per second.

Pricing is where G​rok's image stack is genuinely different from O​penAI's. x​AI charges a flat per-image fee regardless of prompt length: $0.02 per image for the standard grok-imagine-image model, and $0.05 per image for the quality tier — the docs list $0.05 at both 1K and 2K resolution. Image edits bill for both the input image and the generated output. Compare that to gpt-image-2, which is token-priced — $8 per million image-input tokens and $30 per million image-output tokens — and whose cost per image swings roughly 35× from about $0.006 at low quality to about $0.21 at high quality for a 1024×1024 image. Flat per-image pricing is far easier to forecast at scale than a token bill that moves with prompt length, quality tier, and output resolution.

Trying a 48-hour-old model without betting the pipeline

For a model this new, the practical question is not whether it is good — it is how you run it next to the model you already use without making a two-day-old release the single point of failure in a production path. That is the case for calling it through a router rather than wiring a direct vendor integration.

G​rok's image models are callable through OrcaRouter's O​penAI-compatible API today, and O​penAI's gpt-image-2 sits on the same endpoint — one API key, and switching an image call from one to the other is a model-string change, not a new contract. OrcaRouter passes the provider's list price through with no markup, so when x​AI changes the per-image rate, the change is live here the same day, and a flat per-image fee means the cost math is the same whether you call G​rok directly or through the router. Automatic failover is the piece that matters most for a brand-new model: if a G​rok image call errors, rate-limits, or refuses, the request can land on a fallback like gpt-image-2 without the pipeline stalling — which is how you trial a model that shipped yesterday without betting a production path on it. One honest caveat: the newest quality-tier experience is rolling out, so check the current model list on the platform rather than assuming every G​rok variant is routed.

Tontonan Berikutnya

Three signals will tell you whether Grok Imagine Image 2.0 keeps its #2 seat. First, API availability for the full consumer experience — multi-image input, Smart Resize, and the templates are the differentiating features, and they only matter to developers once they are exposed outside the apps. Second, more independent evaluation: one six-prompt test is not a body of evidence, and the editing-reliability and Simplified-Chinese findings deserve a second look from a different lab. Third, the leaderboard itself — Arena standings move with votes, and a model that debuts at #2 can slide just as fast as it rose.

The shortest honest summary: Grok Imagine Image 2.0 is a real, fresh, capable image model whose #2 ranking is a genuine leaderboard snapshot, and whose editing-first toolset is ahead of anything x​AI has shipped before. It is also, right now, a model whose two clearest independent data points — weak Simplified-Chinese compliance and unreliable style transfer — sit squarely on the features teams most often need from an image API. If you generate photorealistic imagery and need to keep a subject's identity across edits, it is worth trialing immediately. If your pipeline ships localized text or stylized assets, wait for the next iteration or another independent run before betting on it.

© 2026 OrcaRouter

Untuk Penyedia

Mengoperasikan platform inferensi? Hadirkan model Anda di OrcaRouter.

Hubungi kami

Gabung komunitas kami

DiscordEmailXGitHubYouTube