
Ming-Image-0.1-Design vs GPT Image 2.5: A Lab's Own Number Against a Board of Strangers' Votes
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
Put Ming-Image-0.1-Design and GPT Image 2.5 in the same sentence and the obvious comparison is quality. The useful comparison is provenance. GPT Image 2.5 holds first and second place on the live Artificial Analysis UI/UX Design board, at 1,227 and 1,218 Elo, on 1,287 and 1,303 blind human votes respectively — that is, a ranking produced by people who did not build the model and were not told which output was whose. Ming-Image-0.1-Design has exactly one score attached to it, 1,082 Elo, and that score appears on a leaderboard graphic shipped inside inclusionAI's own Hugging Face repository, on a board we could not find anywhere else, for a model with zero downloads. One of these numbers is evidence. The other is a claim with a typeface. That gap, not the 145 Elo between them, is what should shape a decision about either model.
The two numbers side by side, and the trap in comparing them
The numbers are close enough to look comparable and different enough in kind that they are not. Here is the set, stated precisely.
• GPT Image 2.5, on the live board — first place at 1,227.0 Elo for the Flare (max) configuration, second at 1,218.1 for Sunburst (max), with sample counts of 1,287 and 1,303 and a tracked price of $210.72 per 1,000 images. The model was listed on that board as released 8 September 2026.
• Ming-Image-0.1-Design, on its own card — first place at 1,082 Elo, ahead of Ideogram 4.0 (Quality) at 1,052 and the FLUX.2 dev variants between 994 and 1,000, on a graphic headed "Text to Image Leaderboard: UI/UX Design" with the footer "Elo scores from blind preference votes in our Image Arena". No sample count is shown. No methodology is published. The board itself is not on the public web as far as we can find.
• The trap — Elo is computed per board. A score is only meaningful against the other scores on the same board, which is why the same FLUX.2 release reads 1,026 on the general text-to-image board and 1,065 on the UI/UX Design category view. Subtract 1,082 from 1,227 and you have not measured a quality gap between two models. You have subtracted a number from a different number.

What makes a blind vote a blind vote
Artificial Analysis publishes enough of its method to be checked. Its image arena pairs two generations from anonymous models, asks a voter to pick a winner, and converts the outcomes into an Elo rating — the sample counts on the board are the number of such comparisons each model has accumulated. That structure is what makes a score resistant to the model's own authors: the people who trained GPT Image 2.5 do not appear in the loop, and a model cannot be ranked first by being the one its maker prefers. It is not a perfect instrument — the pool of voters is self-selected and the prompts are not a controlled benchmark suite — but it is an instrument someone other than the vendor is holding.
The graphic inside the Ming-Image-0.1-Design repository borrows that format without the parts that make it verifiable. "Blind preference votes" is the claim. "Our Image Arena" is the operator, and the operator is inclusionAI. There is no published vote count, no visible prompt distribution, and no way to inspect the pairings. It is entirely possible that the evaluation behind that graphic is honest and rigorous — the model's team would know. It is also entirely uncheckable from outside, which is the only thing that matters if you are the one deciding whether to build on it.
Worth adding, because it cuts against the easy narrative: Ming-Image-0.1-Design is not on the live Artificial Analysis board at all. Not in the UI/UX Design category, not on the general text-to-image board, not on the open-weights view. Its absence is not evidence that it is worse — a model with zero downloads and no hosted endpoint has no way to accumulate votes. It is evidence that no independent number exists yet, which is a different statement.
Price is the part that is not ambiguous
On cost, the two models are not close and there is nothing to argue about.
• GPT Image 2.5 is a commercial endpoint. Artificial Analysis tracks it at $210.72 per 1,000 images in the UI/UX category — roughly 21 cents an image, which puts it at the top of the price range for the field. For context on the same board, Grok Imagine Image 2.0 is tracked at $60 per 1,000, MAI-Image-2.6 at $38.9, and Muse Image at $10.
• Ming-Image-0.1-Design has no price because it has no endpoint. The weights are MIT-licensed and free to download; running them costs whatever one 80 GiB CUDA GPU costs you per hour. The model card's validated configuration is a single card of that class, at 2048 x 2048 resolution and 12 sampling steps.
• Neither figure is stable. Both vendors revise rates, and a per-image comparison written today has a shelf life measured in weeks. This is the one place where OrcaRouter's economics are worth stating plainly: we pass provider list prices through at 0% markup, so a vendor rate change is live on our side the same day rather than after a repricing cycle of our own — which matters when the delta between two candidates is 21 cents against 6 cents an image and both can move.
What you can actually call, today, and where we sit in it
Availability is where this comparison stops being a thought experiment.
GPT Image 2.5 is a real product with a real API behind it, sold by OpenAI on its own terms. It is not routable through OrcaRouter — our catalogue carries no GPT Image 2.5 entry as of 23 September 2026 — and we are not going to describe a route we do not have. What the catalogue does carry from the same line is the previous generation, openai/gpt-image-2 at $8.00 per million input tokens and $30.00 per million output tokens, alongside openai/gpt-image-1.5, and those sit behind the same key as everything else we route.
Ming-Image-0.1-Design is not routable anywhere. The repository has inference disabled, Hugging Face reports no inference provider deployment, and the Quick Start points at a companion GitHub repository that returns 404. No inclusionAI model of any kind is in our catalogue. The only way to run it is to download the shards and serve them yourself.
So the shape of the choice is: one option you can buy today at the top of the market price, and one option you can run today for the price of a GPU and your own engineering time, with no independent evidence that it performs. Anyone who tells you Ming-Image-0.1-Design is a cheaper alternative to GPT Image 2.5 has not priced the second option.

How to hold both without betting on either
The practical advice here is narrower than a verdict, because the two models are not substitutes yet.
If you need UI or design-quality image generation in production, GPT Image 2.5 is the defensible choice and the price is the thing to negotiate with yourself about, not the quality — it leads the independent board by a margin wide enough that the ranking is not in question. If you want to know whether Ming-Image-0.1-Design is good, you have two options and only one of them is free of guesswork: pull the MIT weights onto an 80 GiB card and run your own evaluation set, or wait for someone else to. Do not treat the 1,082 as a result in the meantime. And if your real requirement is optionality rather than either model specifically, the discipline that pays is keeping the generation call behind a single interface — one key across 200-plus models, a model-ID change instead of a rewrite, automatic failover when an endpoint misbehaves — so that whichever of these two ships an independent number first, adopting it costs you a configuration line and not a sprint.
One closing note on what would change this piece. A Ming-Image-0.1-Design row appearing on the Artificial Analysis UI/UX Design board, with a sample count next to it, would convert the vendor's 1,082 from a claim into a data point — and it would be the first time anyone outside inclusionAI has said anything measurable about the model. That is the event to watch for, and it is a more useful thing to monitor than the download counter.

