
Qwen-Image-2.1 Is the Top Open-Weights Image Model on Both Artificial Analysis Boards
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 125 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 103 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 212 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Alibaba's Qwen team released Qwen-Image-2.1 with open weights on 20 September 2026. Twelve days later, the independent scoreboards have caught up with it, and the result is more interesting than the launch. On the AA-Image-T2I v2.0 leaderboard, Qwen-Image-2.1 sits at rank 18 of 166 with an Elo of 1,034, from 5,286 blind comparisons. On AA-Image-Editing v2.0 it sits at rank 18 of 97 with an Elo of 1,074, from 5,536. Both of those rows are marked Open Weights on a board that marks every model individually. No other model carrying that marker on either board is placed above them. That is the whole finding: on the only two public image boards that score models by blind pairwise comparison, the released Qwen-Image-2.1 is the strongest set of downloadable weights in the field, and it is Qwen-Image-2.1 specifically, not a preview of it, that holds both positions.
What follows is where those two numbers come from, what the rows around them look like, and the three details on the board that most of the launch coverage left out — starting with the fact that the board records no API anywhere for the model.
Where the two rows sit
Artificial Analysis builds both boards out of blind pairwise comparisons: two models answer the same prompt, a voter picks the better output, and an Elo falls out of the aggregate. The two boards share a house style and nothing else. They score different tasks, they have different populations, and their Elo values are not comparable to each other. Qwen-Image-2.1 happens to land on the same rank in both, which is a coincidence of two different distributions.
Text-to-image, 166 models on the board:
• The board's top row is GPT Image 2.5 Sunburst (max) at 1,107 Elo from 14,023 appearances — a proprietary model with more than twice Qwen-Image-2.1's sample.
• Qwen-Image-2.1 sits at rank 18 with 1,034 Elo and a 95% confidence interval of 1,024 to 1,044, from 5,286 appearances.
• Read the row counts before the ranks. Qwen-Image-2.1 sits at rank 18 on both boards, but those ranks come from different populations — 166 models on text-to-image, 97 on editing — and the two Elo values belong to two separate scales. The identical rank is arithmetic, not evidence that the two boards agree about anything.
• Alibaba's own two Pro-tier models are above it on this board: Qwen-Image-3.0-Pro at 1,089 Elo, rank 14, and Qwen-Image-3.0 at 1,075, rank 16. Neither carries the Open Weights marker. The board's September snapshot lists both as July 2026 releases, which is the version of the story where the newest Qwen image model is not the highest-scoring one — and on text-to-image it is not.
Image editing, 97 models on the board:
• GPT Image 2.5 Sunburst (max) tops this board too, at 1,182 Elo from 17,899 appearances.
• Qwen-Image-2.1 is at rank 18 with 1,074 Elo, interval 1,065 to 1,083, from 5,536 appearances — a slightly larger sample than its text-to-image row, which makes sense for a model whose editing side got the louder launch copy.
• The rows around it are crowded. grok-imagine-image sits one place below at 1,071 and Luma UNI 1 Max at 1,069. The nearest open-weights row beneath it on this board is HunyuanImage 3.0 Instruct at 1,066 on 5,221 appearances — 8 Elo adrift. Seven rows sit inside eighteen Elo points.
Reading the "top open-weights model" claim honestly
The phrase in the headline is doing specific work, and it deserves to be checked rather than repeated.
• It is a rank among a marker, not an overall rank. Qwen-Image-2.1 is 18th on both boards overall. The Open Weights label is what puts it first on each, and the label is applied per model on the board's own authority — 47 text-to-image rows carry the marker; 36 rows carry it on the editing board, which lists 97.
• The marker describes weights, not terms. Qwen-Image-2.1 ships under the Qwen Research License Agreement dated 20 September 2026, which grants rights for non-commercial purposes only. The board's Open Weights label is accurate about downloadability and silent about what you may do with the download. Anyone reading "top open-weights image model" as "free to use in a product" has read the label past its meaning.
• It is a snapshot, not a settled ranking. The board moves daily as votes accrue. Qwen-Image-2.1's confidence intervals are ±10 Elo on text-to-image and ±9 on editing; the rows below it sit on overlapping intervals. Treat the position as current rather than permanent.
The detail the launch coverage did not have: no API
Both Qwen-Image-2.1 rows carry an explicit No API available note. That is a real cost to the ranking, and it tells you something about who generated the 5,286 and 5,536 votes — if there is no endpoint to point a prompt at, the comparisons came from people running the weights themselves and submitting output. Independent Elo for a self-hosted-only model is a genuinely harder measurement than Elo for something anyone can call, and it is why the samples here are a third the size of the top rows.
For a developer, the shape of this is familiar: the strongest downloadable image model in the field is not the one you can call today. Practically, that splits the work. You either serve Qwen-Image-2.1 yourself on your own hardware, or you call one of the models around it on this board. The second option is where routing earns its keep — one API over 200+ models with provider list price passed through at zero markup, automatic failover when a provider degrades, and a routing DSL for composing several models into a single call. OrcaRouter does not serve Qwen-Image-2.1 today; the image lines on the platform are Google's Imagen tiers and Gemini image previews, xAI's Grok Imagine image endpoint and OpenAI's GPT-Image family. On this board that is the GPT Image 2.5 Sunburst row at the very top, and it is a legitimate answer to "I want the best score and I want it callable this afternoon."

What to watch on the board
Three things will move this story, and all three are visible on the boards already.
The first is whether an API appears. If Alibaba or a serving partner puts Qwen-Image-2.1 behind an endpoint, the No API available note disappears, vote volume should climb toward the size of the neighbouring rows, and the confidence interval tightens. A tighter interval on a 1,034 or a 1,074 is a much stronger claim than the current one.
The second is Alibaba's own stack. Qwen-Image-3.0-Pro at 1,089 and Qwen-Image-3.0 at 1,075 both outrank Qwen-Image-2.1 on text-to-image while carrying no Open Weights marker. If Qwen-Image-2.1 is meant to be the downloadable counterpart to that line rather than its successor, the interesting question is whether the gap narrows — and on the editing board it already has, where Qwen-Image-3.0-Pro leads Qwen-Image-2.1 by two Elo points.
The third is the licence. Every Open Weights row on both boards is equally eligible for the label, and the licences behind them are not remotely equal. The board will never tell you that. It is the one column the scoreboard does not have.

FAQ
Is Qwen-Image-2.1 the best open-weights image model?
On these two boards, yes — it is the highest-Elo row carrying the Open Weights marker on both, at 1,034 on text-to-image and 1,074 on editing. Overall it ranks 18th on each — and the 18 is a rank within 166 rows on one board and within 97 on the other, standing in for two Elo scales that cannot be compared to each other. The claim only holds if both qualifiers stay attached.
Why do the two boards show different Elo values for the same model?
They score different tasks against different model populations. A text-to-image Elo and an editing Elo are two separate scales; comparing 1,034 to 1,074 measures the boards, not the model.
Can I call Qwen-Image-2.1 through an API?
Both board rows record No API available. The weights are downloadable from Alibaba's repositories, so the model runs — it just does not have a hosted endpoint that the board has credited with serving it.
If you want a number today rather than a download, the top of both boards is callable through the vendors' own APIs, and the GPT-Image line, Google's Imagen tiers and xAI's Grok Imagine are the image models OrcaRouter fronts directly. Keep the weight file for the experiment and the endpoint for the deadline.
