Hero card for the comparison between Qwen-Image 2.1, an open-weights model with no independent score, and MAI-Image-2.5-Pro, the measured leader of the Artificial Analysis image editing arena at Elo 1,272
Guides & Insights

Qwen-Image 2.1 vs MAI-Image-2.5-Pro: 6,100 Blind Votes Against Zero

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

One of these two models has been ranked by roughly 6,100 blind human comparisons. The other has been ranked by nobody outside its own lab. MAI-Image-2.5-Pro, Microsoft's premium image model, sits at number one on the Artificial Analysis image editing arena with an Elo of 1,272 — the most-voted result in the category. Qwen-Image 2.1, which the Qwen-Image team open-sourced on 20 September 2026, has no independent score at all, and will not have one until enough people run it in public to generate votes. That gap is the actual subject of this comparison, because it is the only difference between the two models that is not a vendor claim. Everything else — the architecture, the resolution, the price — is knowable but unproven, and the two models are close enough on paper that the measurement gap is what should drive the decision.

What the votes actually say, and what they do not

Artificial Analysis runs blind pairwise comparisons: two models get the same prompt, voters pick the better output, Elo falls out of the results. It is not a perfect instrument, but it is the closest thing this category has to a shared scoreboard, and the sample size matters as much as the number.

MAI-Image-2.5-Pro — image editing #1 at Elo 1,272 on roughly 6,100 comparisons. In text-to-image it sits at #7 with Elo 1,292, behind GPT Image 2 (high) at 1,369, Reve 2.1 at 1,322, Nano Banana 2 at 1,320, GPT Image 1.5 (high) at 1,310 and MAI-Image-2.5 at 1,304.

Qwen-Image 2.1 — no entry on either board. Not a low score; no score. The release is one day old at the time of writing.

The shape of those numbers is worth reading carefully, because it is not the shape the marketing implies. Microsoft's model is genuinely first in editing and genuinely seventh in generation. That is a specific, defensible position — an editing specialist that leads its category — and it is a different thing from being the best image model, which it is not. Meanwhile the model that beats it on text-to-image is GPT Image 2 (high), which is not the newest OpenAI model in this field either. Independent leaderboards reward the model that was measured most, and in this matchup that is unambiguously the Microsoft one.

The one spec where the open model wins outright

There is a concrete, non-subjective difference between these two, and it is resolution.

Qwen-Image 2.1 generates natively at 2K, with 2048×2048 as the documented default at 40 inference steps, seven aspect-ratio presets, and RGBA output with a real alpha channel rather than a matte.

MAI-Image-2.5-Pro caps output at 1,048,576 pixels — about one megapixel. Its headline spec numbers are elsewhere: 32,000 text input tokens, a 131k context window and 4,096 output tokens, which is a statement about how much instruction it can absorb, not how much image it can emit.

For most social and product work, a one-megapixel cap is not a constraint. For print, for large-format compositing, for anything where a crop has to survive, it is. This is the one dimension in the whole comparison where you can point at a number and say the open model is ahead — and note that it is an architectural fact from the model card, not a quality claim. Qwen-Image 2.1 will render at 2048×2048 whether or not it renders anything worth looking at.

Price, which is where the comparison gets uncomfortable

MAI-Image-2.5-Pro is priced as a premium model and the numbers are not subtle: $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output tokens. At a 1024-pixel output that resolves to roughly $108.50 per thousand images. For scale, that is about 2.3× the cost of MAI-Image-2.5 and roughly 5.4× MAI-Image-2.5-Flash, so Microsoft has deliberately segmented its own line rather than pricing the Pro tier as an incremental upgrade.

Qwen-Image 2.1 has no hosted price, because there is no hosted endpoint — it is a weights download under a research licence. Its cost is your GPU time, and its constraint is that you cannot legally sell what it produces. Comparing $108.50 per thousand images against "free weights" is comparing a service against a machine, and the honest version of that comparison is: the metered price buys you a measured, supported, commercially licensed model, and the weights buy you a 33 GB download and a licence file that says non-commercial only.

That trade is exactly where a routing layer earns its place. OrcaRouter fronts 200+ models behind one OpenAI-compatible endpoint at provider list price with zero markup, with automatic failover across providers — so a team that wants to evaluate an unmeasured open model does not have to move production onto it to do so, and because list price is passed through, any vendor price change on a routed model lands on our side the same day rather than at the next contract renewal. Neither MAI-Image-2.5-Pro nor Qwen-Image 2.1 is among the models we route today, and this piece is not going to suggest otherwise. The image line we do front is OpenAI's GPT-Image family, Google's Imagen 4 tiers and Gemini image previews, and xAI's Grok Imagine image endpoint.

Two-column scoreboard for Qwen-Image 2.1 and MAI-Image-2.5-Pro: Qwen-Image 2.1 rows read Released 20 Sept 2026 / Licence research-only, non-commercial / Editing score none / Text-to-image score none / Output 2048x2048 native, RGBA / Price self-host, no hosted endpoint; MAI-Image-2.5-Pro rows read Announced 23 July 2026 / Licence commercial, via Microsoft / Editing score AA #1, Elo 1,272, about 6,100 votes / Text-to-image score AA #7, Elo 1,292 / Output capped at 1,048,576 pixels / Price about $108.50 per 1,000 images; footer reads 'MAI figures per Artificial Analysis; Qwen-Image 2.1 has no independent score yet.'

Where the vendor claims sit, and how much weight to give them

Both vendors publish numbers that no third party has reproduced, and they should be read with the same scepticism in both directions.

Microsoft's claims — 96.8% text-rendering accuracy across English, Chinese, Japanese, Korean and Spanish; 27% better prompt adherence than GPT-Image-1.5; 94% multi-image character consistency; and up to 84–89% lower GPU cost in specific Microsoft-internal scenarios. The last one is the one to be careful with: it describes Microsoft's own serving configuration, not anything a customer can verify.

Alibaba's claims — the Qwen-Image 2.1 blog post carries a Qwen-Image-Bench comparison chart, and the numbers in it are the vendor's own. There is also a figure circulating in third-party coverage describing the model as a 20-layer transformer, which contradicts the model card's 32-layer single-stream DiT; the model card is the primary source and the 32-layer figure is the one to use.

The difference between the two situations is not the honesty of either vendor. It is that Microsoft's model has been independently measured and Alibaba's has not, so Microsoft's claims can be checked against something and Alibaba's cannot be checked against anything yet.

What each one is for

Read together, these are not competing for the same slot.

MAI-Image-2.5-Pro is the editing instrument. It leads the editing board on a large vote base, it accepts a long instruction (32k text input tokens, 131k context), it is commercially licensed, and it is available now as a public preview on Microsoft Foundry and the MAI Playground. If your work is iterative image correction and you need to know before you commit that it is the best available at that job, this is the measured answer, and it costs about $108.50 per thousand images.

Qwen-Image 2.1 is the open research artifact. It renders larger, it ships an inspectable prompt-rewriting checkpoint alongside the image model, it accepts up to 10 reference images with local edits by circle, painted annotation or mask, and it is free to download and illegal to sell with. If your work is evaluation, benchmarking, or building on weights you can read, it is the more interesting object — and you are doing that work without a leaderboard to tell you whether it is any good.

There is also a timing asymmetry worth naming. MAI-Image-2.6 entered private preview on 19 August 2026, which means the model at the top of the editing board is already one generation behind what Microsoft is preparing to ship. Qwen-Image 2.1, by contrast, is on day one, with an early-access programme that asked its testers to publish by 29 September. Both scoreboards in this comparison are going to look different within a month.

Screenshot of the Artificial Analysis text-to-image leaderboard captured September 9 2026, showing GPT Image 2 (high) first at Elo 1,178 with MAI-Image-2.5-Pro listed at No. 7 with Elo 1,292, and no Qwen-Image 2.1 entry anywhere on the boardScreenshot of the Qwen vendor blog page for Qwen-Image-2.1, captured in English, headlined 'Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation', dated 2026/09/20, with download links to GitHub, Hugging Face and ModelScope and the four headline improvements listed as compact and efficient, native transparency with unified creation and editing, versatile editing with up to 10 reference images, and realistic textures

What would settle it

One thing: Qwen-Image 2.1 appearing on a leaderboard. Everything in this article that is not the resolution cap or the price is a claim by one lab or the other, and the entire reason MAI-Image-2.5-Pro can be recommended with a straight face is that 6,100 people who did not work for Microsoft looked at its output next to something else and preferred it. That is the standard the open model has not met, and the standard it should be held to.

Until then, the practical read is simple. If you need an editing model today, under a commercial licence, with a scoreboard behind it, the answer is Microsoft's and it costs about $108.50 per thousand images. If you want to find out whether a 7B open-weights model with native 2K output and an inspectable prompt rewriter is better, you have to do the measuring yourself — and the most useful thing you can do with the result is publish it, because the whole category is currently short one data point.