
Qwen-Image-2.1 Tops Both Image Arenas: What the #1 Open-Source Label Hides
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
The team behind Qwen-Image-2.1 thanked the Arena for the recognition, and the recognition is real. Qwen-Image-2.1 is the highest-scoring model on both Arena image boards that the boards do not label Proprietary: 1,367 on Image Edit and 1,228 on Text-to-Image, in the September 22, 2026 snapshot. Every row above it on both boards carries a Proprietary tag. Three facts the announcement does not carry are just as checkable as the claim itself — the overall rank those scores buy, the Preliminary flag both rows ship with, and a licence string reading qwen-research rather than the Apache 2.0 on its earlier image models. The claim survives contact with the boards. It just says less than the headline implies.
Reading the two boards side by side
Arena scores are Elo-style ratings built from blind pairwise votes: a person sees two outputs, does not know which model produced which, and picks. That is a categorically different evidence class from a vendor benchmark, which is why a claim resting on it is worth verifying rather than repeating. Both boards were captured on September 22, 2026, and both were read row by row rather than skimmed from the ranking graphic.

The Image Edit board lists 56 models and 29,902,508 votes. Qwen-Image-2.1 sits at rank 16 with a score of 1,367, a confidence interval of ±9, and 4,841 votes. The fifteen rows above it — from gpt-image-2.5-sunburst at 1,526 down to gpt-image-1.5-high-fidelity at 1,370 — are all Proprietary. So are the two rows immediately below it, reve-2.0 at 1,358 and uni-1.1-max at 1,334.

The Text-to-Image board the same day lists 79 models and 6,258,152 votes. Qwen-Image-2.1 is rank 17 at 1,228 with a confidence interval of ±11 and 2,843 votes. Same pattern: the sixteen rows above it are Proprietary, and the first non-Proprietary row below it is ideogram-4.0-quality at 1,204, labelled Ideogram Open Model.
Two details from those rows are worth pulling out, because they cut against the way the claim is usually repeated.
• Rank and score tell different stories — 1,367 is the best non-Proprietary number on the image-editing board, and it is also sixteenth place, 159 points behind the leader and 3 points off fifteenth.
• On one board, Alibaba's own closed model wins — qwen-image-3.0-pro, tagged Proprietary, sits at 1,254 on Text-to-Image, twenty-six points above qwen-image-2.1's 1,228. On the editing board the open model beats the earlier proprietary qwen-image-2.0-pro-2026-06-22 at 1,304. The house champion is not the same model on both boards.
• The gap to the next open model is wide on one board and narrow on the other — on Image Edit the next non-Proprietary row is hunyuan-image-3.0-instruct at 1,302, sixty-five points back; on Text-to-Image it is ideogram-4.0-quality at 1,204, twenty-four points back.
The licence string is doing more work than the ranking
Every Arena row carries an organisation and licence label. Qwen-Image-2.1's reads Alibaba · qwen-research on both boards. That is not the Apache 2.0 label on Alibaba's earlier image models: qwen-image-edit at 1,241 and qwen-image-edit-2511 at 1,235 on the editing board are both tagged Apache 2.0, as are qwen-image-2512 at 1,125 and qwen-image at 1,057 on Text-to-Image. The vendor's blog post closes the same loop, publishing the weights under the Qwen Research License Agreement dated 20 September 2026, which permits research and evaluation and not commercial use.
So the honest reading of "#1 open-source model" is narrower than it sounds. On the board's own two-way split — Proprietary, or not — Qwen-Image-2.1 is first. On the OSI definition of open source, which does not accept a field-of-use restriction, it is not open source at all, and neither are most of the models it is being compared against: the same boards carry rows labelled flux-non-commercial-license, tencent-hunyuan-community, krea-2-community-license and Ideogram Open Model. The Arena's "Open Source" filter is a coarse proprietary-versus-not switch. It is a useful one. It is not a licence review.
That distinction matters more here than the two-point margins, because it is the one thing that cannot change with another week of votes. If you want a permissively licensed image model from this family today, the Apache 2.0 rows are qwen-image-edit and qwen-image-edit-2511 — both older, and both more than a hundred and twenty points lower on the same editing board (1,241 and 1,235 against 1,367). The best-scoring open-labelled Qwen image model and the commercially usable Qwen image model are not the same download.
Preliminary means the number can still move
Both of Qwen-Image-2.1's rows carry the board's Preliminary marker, which the Arena applies when a row has not accumulated enough votes for the score to be treated as settled. The vote counts are the reason: 4,841 votes against a board total of 29,902,508 on Image Edit, and 2,843 against 6,258,152 on Text-to-Image. By comparison, the rows around it carry far more — gpt-image-1.5-high-fidelity has 589,013 votes on the editing board, and qwen-image-2.0-pro-2026-06-22 has 307,705 on the Text-to-Image board.
A ±9 and a ±11 confidence interval on a model three days old is not a warning sign; it is what a new release looks like. But it does mean the specific ordering at the top of the open tier can shift as votes accumulate, particularly on Text-to-Image where the margin over the next open row is twenty-four points.
What the model is, on the vendor's own sheet
Alibaba's own write-up describes a unified text-to-image and editing model with a 7B-parameter visual generation component built from 32 Single-Stream DiT layers, native generation and editing of transparent images, support for up to 10 reference images, and a 2048×2048 native output ceiling — the full pipeline download is roughly 33 GB. The transparency support is the least common item on that list; the Arena's editing board is not obviously measuring it, so a board rank is not evidence either way on that specific capability.
The vendor's headline benchmark is Qwen-Image-Bench, where the model is reported at 60.28, ahead of Nano Banana 2.0 at 59.82 and GPT Image 1.5 at 59.65. That is a vendor-run evaluation with no third-party reproduction, and the margins are under a point — a spread that would not survive a change of prompt set. It is worth stating plainly: the Arena numbers above are other people's votes, and the Qwen-Image-Bench numbers are Alibaba's own. They are not the same kind of evidence and should not be averaged into one impression of the model.
Deployment support arrived with the release rather than after it: day-zero integrations for the model exist in ComfyUI, SGLang, vLLM-Omni, LightX2V and the Diffusers library, which means the practical question for most teams is not whether the model can be served but whether the serving cost fits.
Where a claim like this lands if you need an image API this week
Being precise about our own position matters here. OrcaRouter does not route Qwen-Image-2.1, or any Qwen-Image model of any version. If the Arena scores are what convinced you, the model is reachable through the vendor's own API and several third-party platforms, or as a self-hosted download — not through us. We have no Qwen image route to sell you, and the board rank is not a reason to pretend otherwise.
What we do route is the rest of the image tier those boards rank, and several of the rows above Qwen-Image-2.1 are on it. OpenAI's GPT-Image-2, GPT-Image-1.5 and GPT-Image-1-mini are available, as are Google's Imagen 4 tiers and the Gemini image preview endpoints, and xAI's Grok Imagine image endpoint. That mix is the practical answer to a claim like this one. If the Arena had put an open model at the top outright, switching would be a one-line code change on a single key. It did not — the top of both boards is proprietary, and the interesting decision is between a research-licensed download and a hosted endpoint you can call today.
Three properties of the platform matter for that decision. Routing runs on one OpenAI-compatible endpoint across 200+ models, so comparing the hosted rows against each other does not mean a second contract or a second SDK. Provider list price is passed through at zero markup, which means a vendor price change reaches your invoice the same day it is announced rather than at the next contract renewal. And automatic failover across providers means a new, thinly-voted model can sit behind an established one in a fallback chain instead of becoming a production dependency on day three — which is roughly where Qwen-Image-2.1 is right now.
The claim is true, and thinner than it reads
What Alibaba posted is verifiable: on the September 22, 2026 snapshot of both Arena image boards, Qwen-Image-2.1 is the best-scoring model not tagged Proprietary. What the post compresses is everything that qualifies it — sixteenth and seventeenth overall, a preliminary score on a few thousand votes, and a research licence standing in for the open-source licence the phrase brings to mind.

None of this makes the result small. Being the first non-proprietary row on both boards, three days after release, is a real change in where the open tier sits — the permissively licensed Qwen image models that precede it trail by more than a hundred and twenty points on the same editing board. It is just not the same statement as "the best image model that is open source," and the difference is a licence line and a rank column that take about a minute to check.
