
Qwen-Image-2.1 vs GPT-Image-2.5: The Gap Is One Number, and It Is Not Quality
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Put Qwen-Image-2.1 and GPT-Image-2.5 side by side on the specs and they look like siblings: both unified generation-and-editing image models, both with transparent-background output, both shipping in the last five weeks, both capable of 2K-class stills. The difference that decides the choice between them is a single number, and it is not a benchmark score. Qwen-Image-2.1 is free to download and impossible to sell. GPT-Image-2.5 is impossible to download and trivial to sell. Everything else — the architecture, the latency, the transparency implementation — is downstream of that. Qwen-Image-2.1 was open-sourced on 20 September 2026; GPT-Image-2.5 was announced with its API models live on 8 September 2026.
Two ways to add transparency, one month apart
That both models gained transparent output within weeks of each other is a coincidence worth understanding, because the two implementations are not equivalent.
Qwen-Image-2.1 puts alpha inside the model. It denoises in a 64-channel RGBA latent space with 16× spatial compression, so the alpha channel is a first-class component of the thing the sampler produces. There is no segmentation step and no matting pass. Alongside that sits an unusual extra: because the model can decide per prompt whether to emit RGB or an alpha-channel image, it can also pull a subject out of an ordinary photograph and hand back a transparent layer rather than a binary mask.
GPT-Image-2.5 takes the parameter route. OpenAI's API accepts a background setting of auto, opaque, or transparent, and the model responds accordingly. That is a capability toggle on a hosted endpoint rather than a property of the latent space, and it comes with the advantage that you do not think about it: set the flag, get a cutout, move on. The trade is that you get what the endpoint gives you, at the resolution and cost the endpoint decides.
Which one is better is genuinely unresolved. No public comparison of Qwen-Image-2.1's alpha edges against a dedicated matting model existed at the time of writing, and the only published check on the Qwen side is a functional one — transparent PNG output passed, alpha retained across the full 0–255 range, RGBA PSNR of 60.69 dB in a single sample. That is a correctness check, not a quality verdict. It tells you the channel is real.
What you can actually measure
• Parameters — Qwen-Image-2.1's generation component is a 7B, 32-layer single-stream diffusion transformer, paired with a Qwen3-VL 8B text encoder; roughly 33 GB of weights in total, with the DiT accounting for about 14 GB of that. OpenAI publishes no parameter count for GPT-Image-2.5.
• Resolution — Qwen-Image-2.1 outputs natively up to 2048×2048 at 40 inference steps with seven aspect-ratio presets. GPT-Image-2.5 accepts sizes up to 3840×2160 and 2160×3840, with each edge capped at 3840 px and multiples of 16 required.
• Reference images — Qwen-Image-2.1 takes up to 10 for multi-subject composition and virtual try-on. OpenAI's image models accept reference inputs for editing, but the practical limit and the fidelity behaviour differ per model and per quality tier.
• Latency — SGLang-Diffusion publishes 2.748 s for a 1024×1024, 40-step generation on a single B200, and 4.74 s on a 24 GB RTX 4090 with Cache-DiT and INT8 kernels, denoise only. OpenAI reports GPT-Image-2.5's Flare variant at up to 50% lower latency than GPT-Image-2, with one launch partner measuring 2–4× — vendor-reported and partner-reported respectively, not a controlled comparison.
• Price — Qwen-Image-2.1 has no hosted price because there is no hosted endpoint; the cost is your own GPU time. GPT-Image-2.5 bills at the same token rates as GPT-Image-2: $8.00 per 1M image-input tokens, $30.00 per 1M image-output tokens, $5.00 per 1M text-input tokens.
• Licence — Qwen-Image-2.1 ships under the Qwen Research License Agreement dated 20 September 2026, non-commercial use only. GPT-Image-2.5 is a commercial API with C2PA provenance and an invisible watermark on outputs.
One row in that list is thinner than it looks. OpenAI does not publish per-image prices for GPT-Image-2.5, only token rates, and image-token consumption varies with resolution and quality tier. Third-party measurements put the range from roughly $0.0059 to $0.712 per image across 1K, 2K and 4K at low, medium and high settings at a 1:1 aspect ratio — useful as an order of magnitude, not as a quote. If you are forecasting, build the estimate from token counts and your own prompt distribution, not from a per-image figure someone else measured.

The two costs nobody puts in the same column
The reason this comparison resists a simple answer is that the two models bill you in different currencies, and the exchange rate is your own situation.
GPT-Image-2.5 bills per token, which means the meter is visible, predictable, and scales linearly with volume. That is a genuine advantage for a product with known traffic: you can put it in a budget, and a price cut from the vendor moves your cost the same day. It is also a tax on exploration, because the tenth iteration of a prompt costs the same as the first. The input_fidelity behaviour makes this sharper than the headline rates suggest — the API can automatically apply high fidelity on image inputs, which consumes more input tokens than a casual read of the price list implies, so edit loops cost more than the per-image figures people quote.
Qwen-Image-2.1 inverts both properties. The marginal cost of the hundredth generation is electricity. That makes it the better instrument for iteration, for batch backfills, and for anything where the prompt is still being discovered. What it does not give you is a number for the finance sheet — you are paying for a GPU whether it is busy or idle — and it does not give you the right to sell the output. The Qwen Research License Agreement permits non-commercial research and evaluation, and states that commercial deployment requires a separate licence from Hangzhou Tongyi Laboratory Technology. There is no published price for that licence and no stated terms. That is not a footnote; for a commercial pipeline it is the entire decision.
Running both without a second contract
The realistic answer for most teams is neither-or, it is both. GPT-Image-2.5 handles the production path where the output ships to a customer and the per-token meter is a line item. Qwen-Image-2.1 handles the exploration path where you need two hundred variants of a transparent product shot and no vendor will sell you that volume at a sensible price.
The friction is that the two arrive by completely different routes. One is an API key. The other is a 33 GB download, a Python environment, and a GPU you own. OrcaRouter covers the hosted half: 200+ models behind one OpenAI-compatible endpoint, provider list price passed through at zero markup, automatic failover across providers, a routing DSL for expressing fallbacks, and model fusion for combining several models into one call. Being precise about this model matters — we do not route GPT-Image-2.5 today; our OpenAI image routes are GPT-Image-2, GPT-Image-1.5 and GPT-Image-1-mini, all at OpenAI list price, and the day an upstream provider adds the gpt-image-2.5 identifiers, the same key calls them with no code change. We route no Qwen-Image model of any version either, so Qwen-Image-2.1 is not a route on our side at all. What the pass-through pricing buys you in the meantime is that when OpenAI moves a rate, the number on our side moves with it rather than on a lag.
The verdict, stated as a condition rather than a winner
If the output has to be sold, this is not close: GPT-Image-2.5 is the only one of the two you are licensed to use, and the token meter is the price of that. If the output is research, internal tooling, or evaluation, Qwen-Image-2.1 is the stronger instrument — native 2K, alpha out of the sampler, ten reference images, and a marginal cost of zero once the hardware is paid for. If you need both, run both, and treat the licence as the boundary that decides which pipeline a given job goes into.
Two things would change this. A published commercial term sheet for Qwen-Image-2.1 would collapse the gap on the licence row, which is currently doing most of the work in this comparison. And an independent image leaderboard entry for Qwen-Image-2.1 would tell us whether the vendor's Qwen-Image-Bench result — 60.28, ahead of Nano Banana 2.0 at 59.82 and GPT Image 1.5 at 59.65, first among open models — holds up outside the lab that produced it. Neither exists yet. Until they do, the honest recommendation is to pick on the licence and the meter, not on the scoreboard.


