
Ideogram 4.5 vs GPT Image 2.5: Ideogram Picked This Fight Itself
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 262 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 114 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 969 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 49 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Most model comparisons are assembled by whoever is writing them. This one was assembled by Ideogram. The launch page for Ideogram 4.5, live since September 30, 2026, runs a side-by-side editing demo and names its opponents in the caption: "the same edits on Ideogram 4.5 vs. GPT Image 2.5 Sunburst, Nano Banana Pro, and Nano Banana 2." Then it makes a claim about what the viewer will see — "GPT Image and Nano Banana outputs become unusable within a few edits, while Ideogram 4.5 stays clean edit after edit."
That is unusually direct. A vendor choosing the benchmark target and telling you the result in advance is worth taking seriously in one sense and sceptically in another. GPT Image 2.5 is the strongest thing on the independent boards in both image categories, so it is the right opponent to have picked. And it is a two-identifier release — Flare and Sunburst, both shipped September 8, 2026 — with public arena scores that Ideogram 4.5 does not currently match. The interesting question is not who wins the arena. It is whether Ideogram's claim is measuring something the arena does not.
Where both models stand, precisely
• Ideogram 4.5 — shipped September 30, 2026. Precise-edit model; generate plus edit at four rendering speeds; native 2K; edits demonstrated at source resolution up to 4,016 × 6,016 (24.2 MP). Text-to-image 34th of 167 at 1,014 ± 11 Elo (2,247 votes); image editing 23rd at 1,064 ± 11 Elo (2,382 votes). $0.03–$0.22 per image by speed setting.
• GPT Image 2.5 Sunburst — shipped September 8, 2026 as the precision-first half of OpenAI's split image upgrade, API identifier gpt-image-2.5-sunburst. Text-to-image No. 1 at 1,197 ± 9 Elo (14,023 votes); image editing No. 1 at 1,182 ± 8 Elo (17,899 votes). Priced at $210.70 per 1,000 images on the board's conversion.
• GPT Image 2.5 Flare — the fast half of the same release, identifier gpt-image-2.5-flare. No. 2 on both boards: 1,190 text-to-image, 1,162 editing. Same $210.70 per 1,000.
Put the two editing numbers side by side and the gap is 118 Elo — 1,182 against 1,064 — with confidence intervals of ±8 and ±11. The intervals do not overlap. On the board that measures single-shot editing preference, Sunburst wins, and it wins with roughly seven and a half times the vote count behind it.
What the leaderboard is and is not measuring
This is the part that decides whether Ideogram's claim survives contact with the evidence.
Artificial Analysis's editing arena is blind pairwise human preference: voters see the same source image and the same instruction, get two edited results with no labels, and pick the better one. It is a strong method for "which of these two outputs looks more like what the instruction asked for."
It cannot measure the thing Ideogram is advertising. A voter judging one edit cannot see whether the model quietly altered the wallpaper, shifted a face by two millimetres, or re-rendered grain across the rest of the frame. The failure Ideogram is claiming to fix only becomes visible across a sequence — turn four, turn seven, turn ten — and no mainstream arena presents sequences to voters. A 1,064 Elo says Ideogram's individual edits are good. It says nothing at all about whether they stay good.
That cuts both ways for Ideogram, and the honest version is this: the leaderboard does not confirm the drift claim, and it also does not refute it. What it does refute is the implicit stronger claim a reader might take from the launch page — that 4.5 is the better editor, full stop. Against the board, it is the lower-ranked editor by a margin wider than its error bars.

What each one does that the other does not
• Source-resolution editing — Ideogram 4.5's distinguishing engineering claim. Its page shows an edit applied to a 4,016 × 6,016 source without downscaling, with the promise that the edit boundary is preserved well enough to stitch the crop back into the original. For print work, that is the difference between a deliverable and a compromise.
• Breadth of parameter surface — GPT Image 2.5's. OpenAI added xhigh and max to the quality ladder and made transparent backgrounds a first-class parameter, with background set to transparent, opaque or auto and PNG or WebP output. Transparent output was a gap on GPT Image 2 and became a headline feature on the 2.5 pair.
• Speed lane — Flare exists specifically to be the fast default. OpenAI claims up to 50% lower latency than the previous generation; Sunburst is deliberately slower and aimed at edits that have to survive scrutiny.
• Price shape — Ideogram meters per image by rendering speed ($0.03 to $0.22). OpenAI meters by tokens on the same rate card as GPT Image 2 — $8 per million image-input tokens, $30 per million image-output tokens — which is why the independent per-image conversion lands near $0.21 for a high-quality 1024 × 1024.
The price lines cross in an awkward place. At Ideogram 4.5's Low and Medium settings it is dramatically cheaper per image than the OpenAI pair. At High — which is where the drift demonstration lives, and where you would run a 24-megapixel source — the two are close to the same number. The comparison most buyers will actually run is High-quality Ideogram 4.5 against high-quality Sunburst, and on that comparison the price is roughly a tie with a 118-Elo deficit on the measured board.
Which is fine, if the drift claim is real. That is the entire bet.


The accessibility difference nobody markets
One practical asymmetry runs the other way from the scoreboard. GPT Image 2.5's two identifiers are newer than anything on OrcaRouter's catalogue — our OpenAI image line today is GPT-Image-1.5 at $8/$32 per million tokens and GPT-Image-2 at $8/$30, both at provider list price with no markup added. The 2.5 pair arrived on OpenAI's rate card at the same $8/$30 as GPT-Image-2, which means the moment they are routable they land at the same pass-through price rather than a new one, and the cost question between them becomes about token efficiency rather than about which vendor you call.
Ideogram 4.5 is on Ideogram's own API and on partner platforms. It is not on our catalogue. That matters for the comparison in a boring way: one of these two models is a drop-in for an OpenAI-compatible client and the other is a new integration. If you are already calling GPT-Image-2 through an OpenAI-shaped endpoint, the cost of trying Sunburst is a model-ID change; the cost of trying Ideogram 4.5 is a second contract and a second client, before you even get to evaluating quality.
A routing layer does not remove that difference — it only removes the plumbing from the side that is already compatible, and lets you put a slice of traffic on a new entrant without committing a production path to it. Failover matters more than usual here, because the thing you would be trialling is a model whose central claim has no independent verification and whose sample sizes are a fifth of its competitor's.
How to settle it
Do not run this comparison on single images. That is the arena's method and it is the wrong instrument — it will tell you Sunburst wins, which you already knew.
Run it on sequences. Take five or ten images from your own work, write the same chain of six or eight progressive edits, and execute the identical chain on Ideogram 4.5 (High) and on Sunburst (max). Then diff turn one against turn eight for each model, on the regions you never asked the model to touch. Count the pixels that moved. That number — unchanged-region divergence across a sequence — is what Ideogram is claiming to win on, and to my knowledge nobody publishes it for either model.
Then price both runs per finished sequence. At that point you will have an answer that is worth more than 118 Elo, because it will be about your images rather than a voter panel's.
What this matchup actually is
Ideogram 4.5 did not enter a contest it wins on the boards, and it appears to know that — which is why the launch page demos a behaviour rather than printing a rank. The behaviour it demos is real enough to be worth testing and specific enough to be falsifiable: either repeated edits hold together or they do not, and a ten-turn test on your own files settles it in an afternoon.
The reasonable read today is that GPT Image 2.5 Sunburst is the better editor by the only independent measure that exists, by a margin outside the error bars, with far more evidence behind the number — and that Ideogram 4.5 is selling a narrower, unmeasured property at a price that is cheaper at low settings and roughly equal at the setting its claim depends on. If multi-turn fidelity is your production bottleneck, the test is cheap and the model deserves it. If it is not, the board has already answered.
