Generated hero card for the Qwen-Image-2.1 versus Hy Image3.5 article, with the headline Qwen-Image-2.1 vs Hy Image3.5, the subtitle The Boards Rank Them in Opposite Orders, chips reading Editing: Hy Image3.5 leads 1,088 to 1,074 and Text-to-image: Qwen-Image-2.1 leads 1,034 to 975, and the footer Both models scored per Artificial Analysis, September 2026, with the OrcaRouter logo composited bottom-right
Guides & Insights

Qwen-Image-2.1 vs Hy Image3.5: The Independent Boards Rank Them in Opposite Orders

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put Qwen-Image-2.1 and Hy Image3.5 head to head on the two Artificial Analysis image boards and the boards disagree about which one is better. On text-to-image, Qwen-Image-2.1 is 59 Elo clear of Hy Image3.5 — 1,034 against 975, rank 18 against rank 66. On image editing, the same two models swap places: Hy Image3.5 sits at 1,088 Elo and rank 14, Qwen-Image-2.1 at 1,074 and rank 18. A 14-point gap the other way. The two models went public two days apart — Qwen-Image-2.1 with open weights on 20 September 2026, Hy Image3.5 in public preview on 22 September 2026 — and this is the first time either has had a score on the same instrument, which is what makes the disagreement worth reading rather than reporting.

The obvious move is to say the editing board is noisy and move on. That is half right. The other half is that the two rows are not equally solid, the price on one of them is not the price on the other, and one of these models carries rights to the weights while the other never will.

Two boards, two orders

Artificial Analysis runs blind pairwise comparisons — same prompt to two models, a voter picks the winner, Elo accumulates — and publishes a separate board per task. The text-to-image board carries 166 rows; the editing board carries 97. Both rows for both models come from the same provider's snapshot, so this is not a version mismatch.

• Text-to-image — Qwen-Image-2.1 at 1,034 Elo (interval 1,024 to 1,044) from 5,286 comparisons, rank 18 of 166. Hy Image3.5 at 975 Elo (interval 966 to 984) from 5,334 comparisons, rank 66 of 166. The intervals do not touch. A 59-point separation on samples of similar size is a real ordering, not a rounding artefact.

• Image editing — Hy Image3.5 at 1,088 Elo (interval 1,079 to 1,097) from 5,550 comparisons, rank 14 of 97. Qwen-Image-2.1 at 1,074 Elo (interval 1,065 to 1,083) from 5,536 comparisons, rank 18 of 97. Those intervals overlap across a four-point band at 1,079 to 1,083. The ordering is directional, not established.

• Sample sizes are close on both boards — 5,286 against 5,334 on text-to-image, 5,536 against 5,550 on editing — so the difference in confidence is not a matter of one model being under-voted.

• The board labels the models differently, and that matters: the Qwen-Image-2.1 rows carry the Open Weights marker, the Hy Image3.5 rows do not.

So the honest reading is asymmetric. Text-to-image: Qwen-Image-2.1 wins clearly. Editing: Hy Image3.5 is probably ahead, by a margin inside the noise of the measurement.

Where the asymmetry comes from

Hy Image3.5's editing strength is not a surprise once you look at what Tencent shipped it with. The preview went public with up to five reference images per request, text-to-image and image-to-image in the same model, and multi-turn editing — the editing surface got the launch emphasis. Qwen-Image-2.1 was built with a unified generation-and-editing stack that also handles reference images, but its board rows show the editing side finishing two Elo below Alibaba's own Qwen-Image-3.0-Pro, which is a narrower result than the launch framing implied.

Tencent's own claim is not what the board shows either. The vendor quotes the preview at roughly a 30% blind-evaluation win rate over Hy Image3.0. That number has not been reproduced by anyone outside Tencent, and the independent board does not corroborate it on text-to-image — where Hy Image3.5 preview at 975 is 20 Elo below the open-weights HunyuanImage 3.0 Instruct at 995. On editing the same comparison goes the vendor's way: Hy Image3.5 preview at 1,088 is 22 Elo above HunyuanImage 3.0 Instruct at 1,066. One generation of progress on one board, and a step back on the other, from Tencent's own succession.

Screenshot of the Artificial Analysis image editing leaderboard, AA-Image-Editing v2.0, captured in English and cropped from the top row down through the Ideogram 4.5 row, showing GPT Image 2.5 Sunburst (max) first at 1,182 Elo from 17,899 samples, the Hunyuanimage 3.5 (Preview) row at rank 14 with 1,088 Elo and 5,550 appearances, the Qwen-Image-2.1 row at rank 18 with 1,074 Elo and 5,536 appearances and a No API available note, and the Hunyuanimage 3.0 Instruct row at 1,066 with No API available

The price axis, where the two are not comparable at all

Hy Image3.5 preview is a metered API. Tencent Cloud bills ¥0.15 for a 2K image on the mainland endpoint and US$0.024 on the international one, charged only for images the API returns. The Artificial Analysis price column records $24.00 per 1,000 images, which is the same figure in the units the board uses — and against the $210.70 the top row of the text-to-image board carries for the same thousand images, it is less than a ninth of the price.

Qwen-Image-2.1 has no price line, because it has no endpoint. Both of its board rows carry an explicit No API available note. You do not buy this model; you download it, and you pay for it in GPU hours and in the licence question below. The 5,286 and 5,536 votes on its rows came from people running the weights themselves. That fact also gives the ranking a caveat worth keeping: an independent Elo for a self-hosted-only model measures a different population than an Elo for something anyone can call.

If the plan is to put either of these into a product, the practical split is now clear, and it is the same split the price implies: the model with the better editing score is the one you rent, and the model with the better text-to-image score is the one you run. OrcaRouter is the renting side of that — one API over 200+ models with provider list price passed through at zero markup, automatic failover across providers, and a routing DSL for composing several models into one call. We do not route Hy Image3.5 preview or Qwen-Image-2.1; the image lines on the platform are Google's Imagen tiers and Gemini image previews, xAI's Grok Imagine image endpoint and OpenAI's GPT-Image family. Neither of these two is a first pick from that list on the strength of a board row alone, which is exactly why the licence question below is the one that decides it.

The axis neither board has a column for

Generated comparison scoreboard for Qwen-Image-2.1 and Hy Image3.5, listing editing Elo 1,074 against 1,088 with ranks 18 and 14, editing intervals 1,065 to 1,083 against 1,079 to 1,097, text-to-image Elo 1,034 against 975 with ranks 18 and 66, and the board labels Open Weights, no API for Qwen-Image-2.1 against no Open Weights and $24.00 per 1,000 images for Hy Image3.5

Qwen-Image-2.1 is released under the Qwen Research License Agreement, dated 20 September 2026, which grants rights for non-commercial purposes only and requires a separate licence from the vendor for commercial use. The Hugging Face repositories report the licence as other for that reason — it is not an OSI licence and should not be read as one. The Open Weights marker on both board rows is accurate and says nothing about this.

Hy Image3.5 preview has no weights to license. Tencent has not published them; the preview is served from Tencent's own platform and there is no downloadable release of this generation. What you get is a rate card, a reference-image limit and a service you can start and stop. Nothing to audit, nothing to host, nothing to negotiate — and nothing you keep if Tencent changes the price or retires the preview.

The comparison that actually resolves the open-weights question is Tencent's own predecessor rather than Qwen's model. HunyuanImage 3.0 Instruct is on both boards with the Open Weights marker — 1,066 on editing, 995 on text-to-image. Qwen-Image-2.1 beats it on both. So if the requirement is "downloadable weights I can run on my own hardware," the best option in this matchup is Alibaba's, on either board, under a licence that still forbids selling the output.

There is no version of this where the paperwork resolves itself. You can have the better editing score with no weights and a metered bill, or the better text-to-image score with weights you may not commercialise. Pick which constraint you can live with, then go and measure both on your own prompts — the editing gap between them is four points wide inside overlapping intervals, which is the kind of margin a single held-out prompt set can erase.