A generated hero card headlined Qwen-Image-2.1 vs GPT-Image-2.5, showing a card labelled Qwen-Image-2.1 with a download icon and the text 33 GB weights beside a card labelled GPT-Image-2.5 with a metered dial icon and the text token billed, captioned one you download, one you meter.
Guides & Insights

Qwen-Image-2.1 vs GPT-Image-2.5: The Gap Is One Number, and It Is Not Quality

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put Qwen-Image-2.1 and GPT-Image-2.5 side by side on the specs and they look like siblings: both unified generation-and-editing image models, both with transparent-background output, both shipping in the last five weeks, both capable of 2K-class stills. The difference that decides the choice between them is a single number, and it is not a benchmark score. Qwen-Image-2.1 is free to download and impossible to sell. GPT-Image-2.5 is impossible to download and trivial to sell. Everything else — the architecture, the latency, the transparency implementation — is downstream of that. Qwen-Image-2.1 was open-sourced on 20 September 2026; GPT-Image-2.5 was announced with its API models live on 8 September 2026.

Two ways to add transparency, one month apart

That both models gained transparent output within weeks of each other is a coincidence worth understanding, because the two implementations are not equivalent.

Qwen-Image-2.1 puts alpha inside the model. It denoises in a 64-channel RGBA latent space with 16× spatial compression, so the alpha channel is a first-class component of the thing the sampler produces. There is no segmentation step and no matting pass. Alongside that sits an unusual extra: because the model can decide per prompt whether to emit RGB or an alpha-channel image, it can also pull a subject out of an ordinary photograph and hand back a transparent layer rather than a binary mask.

GPT-Image-2.5 takes the parameter route. OpenAI's API accepts a background setting of auto, opaque, or transparent, and the model responds accordingly. That is a capability toggle on a hosted endpoint rather than a property of the latent space, and it comes with the advantage that you do not think about it: set the flag, get a cutout, move on. The trade is that you get what the endpoint gives you, at the resolution and cost the endpoint decides.

Which one is better is genuinely unresolved. No public comparison of Qwen-Image-2.1's alpha edges against a dedicated matting model existed at the time of writing, and the only published check on the Qwen side is a functional one — transparent PNG output passed, alpha retained across the full 0–255 range, RGBA PSNR of 60.69 dB in a single sample. That is a correctness check, not a quality verdict. It tells you the channel is real.

What you can actually measure

Parameters — Qwen-Image-2.1's generation component is a 7B, 32-layer single-stream diffusion transformer, paired with a Qwen3-VL 8B text encoder; roughly 33 GB of weights in total, with the DiT accounting for about 14 GB of that. OpenAI publishes no parameter count for GPT-Image-2.5.
Resolution — Qwen-Image-2.1 outputs natively up to 2048×2048 at 40 inference steps with seven aspect-ratio presets. GPT-Image-2.5 accepts sizes up to 3840×2160 and 2160×3840, with each edge capped at 3840 px and multiples of 16 required.
Reference images — Qwen-Image-2.1 takes up to 10 for multi-subject composition and virtual try-on. OpenAI's image models accept reference inputs for editing, but the practical limit and the fidelity behaviour differ per model and per quality tier.
Latency — SGLang-Diffusion publishes 2.748 s for a 1024×1024, 40-step generation on a single B200, and 4.74 s on a 24 GB RTX 4090 with Cache-DiT and INT8 kernels, denoise only. OpenAI reports GPT-Image-2.5's Flare variant at up to 50% lower latency than GPT-Image-2, with one launch partner measuring 2–4× — vendor-reported and partner-reported respectively, not a controlled comparison.
Price — Qwen-Image-2.1 has no hosted price because there is no hosted endpoint; the cost is your own GPU time. GPT-Image-2.5 bills at the same token rates as GPT-Image-2: $8.00 per 1M image-input tokens, $30.00 per 1M image-output tokens, $5.00 per 1M text-input tokens.
Licence — Qwen-Image-2.1 ships under the Qwen Research License Agreement dated 20 September 2026, non-commercial use only. GPT-Image-2.5 is a commercial API with C2PA provenance and an invisible watermark on outputs.

One row in that list is thinner than it looks. OpenAI does not publish per-image prices for GPT-Image-2.5, only token rates, and image-token consumption varies with resolution and quality tier. Third-party measurements put the range from roughly $0.0059 to $0.712 per image across 1K, 2K and 4K at low, medium and high settings at a 1:1 aspect ratio — useful as an order of magnitude, not as a quote. If you are forecasting, build the estimate from token counts and your own prompt distribution, not from a per-image figure someone else measured.

A generated two-column scoreboard titled Qwen-Image-2.1 vs GPT-Image-2.5 - the scoreboard. Left column Qwen-Image-2.1 rows: Generation component 7B DiT; Transparency native RGBA; Resolution 2048 x 2048 native; Reference images up to 10; Price self-hosted GPU time; Licence non-commercial research only. Right column GPT-Image-2.5 rows: Generation component undisclosed; Transparency API parameter; Resolution up to 3840 x 2160; Reference images per the editing API; Price 8 and 30 dollars per 1M tokens; Licence commercial, C2PA and watermark. Footer: Qwen figures per the vendor model card, unaudited; OpenAI figures per the vendor API reference.

The two costs nobody puts in the same column

The reason this comparison resists a simple answer is that the two models bill you in different currencies, and the exchange rate is your own situation.

GPT-Image-2.5 bills per token, which means the meter is visible, predictable, and scales linearly with volume. That is a genuine advantage for a product with known traffic: you can put it in a budget, and a price cut from the vendor moves your cost the same day. It is also a tax on exploration, because the tenth iteration of a prompt costs the same as the first. The input_fidelity behaviour makes this sharper than the headline rates suggest — the API can automatically apply high fidelity on image inputs, which consumes more input tokens than a casual read of the price list implies, so edit loops cost more than the per-image figures people quote.

Qwen-Image-2.1 inverts both properties. The marginal cost of the hundredth generation is electricity. That makes it the better instrument for iteration, for batch backfills, and for anything where the prompt is still being discovered. What it does not give you is a number for the finance sheet — you are paying for a GPU whether it is busy or idle — and it does not give you the right to sell the output. The Qwen Research License Agreement permits non-commercial research and evaluation, and states that commercial deployment requires a separate licence from Hangzhou Tongyi Laboratory Technology. There is no published price for that licence and no stated terms. That is not a footnote; for a commercial pipeline it is the entire decision.

Running both without a second contract

The realistic answer for most teams is neither-or, it is both. GPT-Image-2.5 handles the production path where the output ships to a customer and the per-token meter is a line item. Qwen-Image-2.1 handles the exploration path where you need two hundred variants of a transparent product shot and no vendor will sell you that volume at a sensible price.

The friction is that the two arrive by completely different routes. One is an API key. The other is a 33 GB download, a Python environment, and a GPU you own. OrcaRouter covers the hosted half: 200+ models behind one OpenAI-compatible endpoint, provider list price passed through at zero markup, automatic failover across providers, a routing DSL for expressing fallbacks, and model fusion for combining several models into one call. Being precise about this model matters — we do not route GPT-Image-2.5 today; our OpenAI image routes are GPT-Image-2, GPT-Image-1.5 and GPT-Image-1-mini, all at OpenAI list price, and the day an upstream provider adds the gpt-image-2.5 identifiers, the same key calls them with no code change. We route no Qwen-Image model of any version either, so Qwen-Image-2.1 is not a route on our side at all. What the pass-through pricing buys you in the meantime is that when OpenAI moves a rate, the number on our side moves with it rather than on a lag.

The verdict, stated as a condition rather than a winner

If the output has to be sold, this is not close: GPT-Image-2.5 is the only one of the two you are licensed to use, and the token meter is the price of that. If the output is research, internal tooling, or evaluation, Qwen-Image-2.1 is the stronger instrument — native 2K, alpha out of the sampler, ten reference images, and a marginal cost of zero once the hardware is paid for. If you need both, run both, and treat the licence as the boundary that decides which pipeline a given job goes into.

Two things would change this. A published commercial term sheet for Qwen-Image-2.1 would collapse the gap on the licence row, which is currently doing most of the work in this comparison. And an independent image leaderboard entry for Qwen-Image-2.1 would tell us whether the vendor's Qwen-Image-Bench result — 60.28, ahead of Nano Banana 2.0 at 59.82 and GPT Image 1.5 at 59.65, first among open models — holds up outside the lab that produced it. Neither exists yet. Until they do, the honest recommendation is to pick on the licence and the meter, not on the scoreboard.

A screenshot of OrcaRouter's own model page for openai/gpt-image-2, captured September 21 2026, showing the model description as OpenAI's token-billed image model reached through the Images API, the /v1/images/generations endpoint, and the published list rates of $8.00 per 1M image input tokens and $30.00 per 1M image output tokens.A screenshot of the Qwen vendor blog page for Qwen-Image 2.1, captured September 21 2026, showing the 2026/09/20 publication date, the headline Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation, the Now open weights banner, and the GitHub, Hugging Face and ModelScope buttons.