A generated hero title card headlined Qwen-Image-2.1 Tops Both Image Arenas, subtitled What the #1 open-source label hides, showing two leaderboard cards labelled Image Edit Arena with a rank chip reading 16 of 56 and a score chip reading 1367, and Text-to-Image Arena with a rank chip reading 17 of 79 and a score chip reading 1228, above a tag reading qwen-research licence.
Guides & Insights

Qwen-Image-2.1 Tops Both Image Arenas: What the #1 Open-Source Label Hides

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The team behind Qwen-Image-2.1 thanked the Arena for the recognition, and the recognition is real. Qwen-Image-2.1 is the highest-scoring model on both Arena image boards that the boards do not label Proprietary: 1,367 on Image Edit and 1,228 on Text-to-Image, in the September 22, 2026 snapshot. Every row above it on both boards carries a Proprietary tag. Three facts the announcement does not carry are just as checkable as the claim itself — the overall rank those scores buy, the Preliminary flag both rows ship with, and a licence string reading qwen-research rather than the Apache 2.0 on its earlier image models. The claim survives contact with the boards. It just says less than the headline implies.

Reading the two boards side by side

Arena scores are Elo-style ratings built from blind pairwise votes: a person sees two outputs, does not know which model produced which, and picks. That is a categorically different evidence class from a vendor benchmark, which is why a claim resting on it is worth verifying rather than repeating. Both boards were captured on September 22, 2026, and both were read row by row rather than skimmed from the ranking graphic.

A screenshot of the LMArena Image Edit Arena leaderboard, captured September 22 2026, showing the board header reading 29,902,508 votes and 56 models, with ranks 13 to 22 visible: gemini-3-pro-image-preview at 1385, reve-2.1 at 1375, gpt-image-1.5-high-fidelity at 1370, qwen-image-2.1 at rank 16 with 1367 and a confidence interval of plus or minus 9 and 4,841 votes marked Preliminary, reve-2.0 at 1358, uni-1.1-max at 1334, grok-imagine-image at 1329, uni-1.1 at 1315, gemini-3.1-flash-lite-image at 1314 and qwen-image-2.0-pro-2026-06-22 at 1304, with every row above qwen-image-2.1 labelled Proprietary.

The Image Edit board lists 56 models and 29,902,508 votes. Qwen-Image-2.1 sits at rank 16 with a score of 1,367, a confidence interval of ±9, and 4,841 votes. The fifteen rows above it — from gpt-image-2.5-sunburst at 1,526 down to gpt-image-1.5-high-fidelity at 1,370 — are all Proprietary. So are the two rows immediately below it, reve-2.0 at 1,358 and uni-1.1-max at 1,334.

A screenshot of the LMArena Text-to-Image Arena leaderboard, captured September 22 2026, showing the board header reading 6,258,152 votes and 79 models, with ranks 14 to 21 visible: gemini-3-pro-image-2k at 1246, gpt-image-1.5-high-fidelity at 1239, gemini-3-pro-image-preview at 1232, qwen-image-2.1 at rank 17 with 1228 and a confidence interval of plus or minus 11 and 2,843 votes marked Preliminary, ideogram-4.0-quality at 1204 labelled Ideogram Open Model, qwen-image-2.0-pro-2026-06-22 at 1191, uni-1.1-max at 1188 and mai-image-2 at 1183, with Proprietary labelling on every row above qwen-image-2.1.

The Text-to-Image board the same day lists 79 models and 6,258,152 votes. Qwen-Image-2.1 is rank 17 at 1,228 with a confidence interval of ±11 and 2,843 votes. Same pattern: the sixteen rows above it are Proprietary, and the first non-Proprietary row below it is ideogram-4.0-quality at 1,204, labelled Ideogram Open Model.

Two details from those rows are worth pulling out, because they cut against the way the claim is usually repeated.

Rank and score tell different stories — 1,367 is the best non-Proprietary number on the image-editing board, and it is also sixteenth place, 159 points behind the leader and 3 points off fifteenth.
On one board, Alibaba's own closed model wins — qwen-image-3.0-pro, tagged Proprietary, sits at 1,254 on Text-to-Image, twenty-six points above qwen-image-2.1's 1,228. On the editing board the open model beats the earlier proprietary qwen-image-2.0-pro-2026-06-22 at 1,304. The house champion is not the same model on both boards.
The gap to the next open model is wide on one board and narrow on the other — on Image Edit the next non-Proprietary row is hunyuan-image-3.0-instruct at 1,302, sixty-five points back; on Text-to-Image it is ideogram-4.0-quality at 1,204, twenty-four points back.

The licence string is doing more work than the ranking

Every Arena row carries an organisation and licence label. Qwen-Image-2.1's reads Alibaba · qwen-research on both boards. That is not the Apache 2.0 label on Alibaba's earlier image models: qwen-image-edit at 1,241 and qwen-image-edit-2511 at 1,235 on the editing board are both tagged Apache 2.0, as are qwen-image-2512 at 1,125 and qwen-image at 1,057 on Text-to-Image. The vendor's blog post closes the same loop, publishing the weights under the Qwen Research License Agreement dated 20 September 2026, which permits research and evaluation and not commercial use.

So the honest reading of "#1 open-source model" is narrower than it sounds. On the board's own two-way split — Proprietary, or not — Qwen-Image-2.1 is first. On the OSI definition of open source, which does not accept a field-of-use restriction, it is not open source at all, and neither are most of the models it is being compared against: the same boards carry rows labelled flux-non-commercial-license, tencent-hunyuan-community, krea-2-community-license and Ideogram Open Model. The Arena's "Open Source" filter is a coarse proprietary-versus-not switch. It is a useful one. It is not a licence review.

That distinction matters more here than the two-point margins, because it is the one thing that cannot change with another week of votes. If you want a permissively licensed image model from this family today, the Apache 2.0 rows are qwen-image-edit and qwen-image-edit-2511 — both older, and both more than a hundred and twenty points lower on the same editing board (1,241 and 1,235 against 1,367). The best-scoring open-labelled Qwen image model and the commercially usable Qwen image model are not the same download.

Preliminary means the number can still move

Both of Qwen-Image-2.1's rows carry the board's Preliminary marker, which the Arena applies when a row has not accumulated enough votes for the score to be treated as settled. The vote counts are the reason: 4,841 votes against a board total of 29,902,508 on Image Edit, and 2,843 against 6,258,152 on Text-to-Image. By comparison, the rows around it carry far more — gpt-image-1.5-high-fidelity has 589,013 votes on the editing board, and qwen-image-2.0-pro-2026-06-22 has 307,705 on the Text-to-Image board.

A ±9 and a ±11 confidence interval on a model three days old is not a warning sign; it is what a new release looks like. But it does mean the specific ordering at the top of the open tier can shift as votes accumulate, particularly on Text-to-Image where the margin over the next open row is twenty-four points.

What the model is, on the vendor's own sheet

Alibaba's own write-up describes a unified text-to-image and editing model with a 7B-parameter visual generation component built from 32 Single-Stream DiT layers, native generation and editing of transparent images, support for up to 10 reference images, and a 2048×2048 native output ceiling — the full pipeline download is roughly 33 GB. The transparency support is the least common item on that list; the Arena's editing board is not obviously measuring it, so a board rank is not evidence either way on that specific capability.

The vendor's headline benchmark is Qwen-Image-Bench, where the model is reported at 60.28, ahead of Nano Banana 2.0 at 59.82 and GPT Image 1.5 at 59.65. That is a vendor-run evaluation with no third-party reproduction, and the margins are under a point — a spread that would not survive a change of prompt set. It is worth stating plainly: the Arena numbers above are other people's votes, and the Qwen-Image-Bench numbers are Alibaba's own. They are not the same kind of evidence and should not be averaged into one impression of the model.

Deployment support arrived with the release rather than after it: day-zero integrations for the model exist in ComfyUI, SGLang, vLLM-Omni, LightX2V and the Diffusers library, which means the practical question for most teams is not whether the model can be served but whether the serving cost fits.

Where a claim like this lands if you need an image API this week

Being precise about our own position matters here. OrcaRouter does not route Qwen-Image-2.1, or any Qwen-Image model of any version. If the Arena scores are what convinced you, the model is reachable through the vendor's own API and several third-party platforms, or as a self-hosted download — not through us. We have no Qwen image route to sell you, and the board rank is not a reason to pretend otherwise.

What we do route is the rest of the image tier those boards rank, and several of the rows above Qwen-Image-2.1 are on it. OpenAI's GPT-Image-2, GPT-Image-1.5 and GPT-Image-1-mini are available, as are Google's Imagen 4 tiers and the Gemini image preview endpoints, and xAI's Grok Imagine image endpoint. That mix is the practical answer to a claim like this one. If the Arena had put an open model at the top outright, switching would be a one-line code change on a single key. It did not — the top of both boards is proprietary, and the interesting decision is between a research-licensed download and a hosted endpoint you can call today.

Three properties of the platform matter for that decision. Routing runs on one OpenAI-compatible endpoint across 200+ models, so comparing the hosted rows against each other does not mean a second contract or a second SDK. Provider list price is passed through at zero markup, which means a vendor price change reaches your invoice the same day it is announced rather than at the next contract renewal. And automatic failover across providers means a new, thinly-voted model can sit behind an established one in a fallback chain instead of becoming a production dependency on day three — which is roughly where Qwen-Image-2.1 is right now.

The claim is true, and thinner than it reads

What Alibaba posted is verifiable: on the September 22, 2026 snapshot of both Arena image boards, Qwen-Image-2.1 is the best-scoring model not tagged Proprietary. What the post compresses is everything that qualifies it — sixteenth and seventeenth overall, a preliminary score on a few thousand votes, and a research licence standing in for the open-source licence the phrase brings to mind.

A generated scoreboard titled Qwen-Image-2.1 - the scoreboard, with rows reading Image Edit Arena score 1367 at rank 16 of 56, Text-to-Image Arena score 1228 at rank 17 of 79, row status Preliminary on both boards, licence qwen-research non-commercial, Apache 2.0 sibling qwen-image-edit at 1241, and vendor benchmark Qwen-Image-Bench 60.28 unreproduced, above a footer reading Arena figures per the LMArena boards captured September 22 2026; Qwen-Image-Bench figure vendor-reported, no third-party reproduction.

None of this makes the result small. Being the first non-proprietary row on both boards, three days after release, is a real change in where the open tier sits — the permissively licensed Qwen image models that precede it trail by more than a hundred and twenty points on the same editing board. It is just not the same statement as "the best image model that is open source," and the difference is a licence line and a rank column that take about a minute to check.