A generated hero card for the article 'Qwen-Image 2.1 vs FLUX 3' showing a two-column matchup graphic with a lock icon between two rounded model cards labelled 'Qwen-Image 2.1' and 'FLUX 3', and a ribbon reading 'One you can run and cannot sell; one you could sell and cannot run'.
Guides & Insights

Qwen-Image 2.1 vs FLUX 3: Two Image Models You Cannot Quite Use

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Qwen-Image 2.1 and FLUX 3 are the two most interesting image releases of the second half of 2026, and neither one is straightforwardly usable. Qwen-Image 2.1, which the Qwen-Image team open-sourced on 20 September 2026, is a 7B-parameter unified generation-and-editing model you can download this afternoon and cannot legally put in a commercial product. FLUX 3, which Black Forest Labs announced on 23 July 2026, is a unified multimodal foundation model whose image tier — the part you would actually compare — still has no public endpoint, no model ID, no price and no benchmarks two months after the announcement. The comparison is therefore not "which one is better." It is which of the two blockers you can work around.

Two different kinds of unavailable

It is tempting to file both under "not ready yet" and move on. That would hide the only decision that matters, because the two blockers have opposite shapes.

Qwen-Image 2.1's blocker is legal, and it is a line in a file you can read before you download anything. The weights are on Hugging Face and ModelScope, the model card is written, the inference code is published, and four separate serving frameworks had support merged before or on release day. Nothing technical stands between you and a running model. What stands between you and shipping it is the Qwen Research License Agreement, dated 20 September 2026, whose grant-of-rights clause permits use "FOR NON-COMMERCIAL PURPOSES ONLY" and directs anyone wanting commercial rights to request a separate licence from the vendor.

FLUX 3's blocker is physical. There is no image endpoint to call, no model identifier to pass, and no price card to budget against. BFL's own pricing page covers the FLUX 3 Video tiers and the older FLUX.2 image line; it does not cover a FLUX 3 image tier, because there is nothing to price. Access to the image model is request-only rather than a get-started button: BFL's own FLUX 3 page lists the image tier as "coming soon", and it is the action tier that carries the early-access label.

So one of these is a model you can run tonight and cannot sell, and the other is a model you could sell and cannot run. Which one is more useful to you depends entirely on which side of that sentence you are standing on — an internal research team and a commercial product team should reach opposite conclusions from the same two facts.

The dimension-by-dimension read

Status — Qwen-Image 2.1 weights, licence and model card published 20 September 2026, with a vendor blog post the same day. FLUX 3 announced 23 July 2026; only the video tier has shipped, reaching general availability on 4 August 2026.

What you can actually call today — Qwen-Image 2.1 runs locally through Diffusers, ComfyUI, vLLM-Omni, SGLang and LightX2V, or through the vendor's own API and several third-party platforms. FLUX 3 lists its image tier as "coming soon" with no endpoint to call, and offers a priced, generally available video endpoint.

Architecture — Qwen-Image 2.1 is a 32-layer single-stream DiT with 7B parameters in the visual generation component, a Qwen3-VL 8B text encoder, and a 64-channel RGBA VAE at 16× spatial compression. FLUX 3 is a single unified flow model trained jointly across image, video and audio, built on a self-supervised flow-matching method BFL calls Self-Flow, with an action-prediction tier developed alongside Mimic Robotics.

Native transparency — Qwen-Image 2.1 generates and edits RGBA images, edits text inside transparent layers, and extracts subjects from RGB photographs into transparent layers, folding in the capability Alibaba shipped separately as Qwen-Image-Layered in December 2025. BFL has published no transparency claim for FLUX 3 Image.

Reference images — Qwen-Image 2.1 accepts up to 10 input references, with local editing by circle, painted annotation or a separate mask. No published figure exists for FLUX 3 Image.

Resolution — Qwen-Image 2.1 is native 2K, defaulting to 2048×2048 at 40 inference steps, with seven documented aspect-ratio presets. FLUX 3 Image's resolutions are described in the announcement only as "broad," with nothing published.

Price — no published price for either. Qwen-Image 2.1 has no listed hosted rate at the time of writing; FLUX 3 Image has no price card because it has no endpoint. The only FLUX 3 pricing that exists is video, at $0.06 per second draft, $0.17 per second HD and $0.29 per second FHD.

Licence — Qwen Research License Agreement, non-commercial only, commercial use by separate request. FLUX 3's licence terms for the image tier have not been published at all.

One asymmetry deserves its own line, because it is the trap in this matchup. FLUX 3 does have published benchmark numbers — an internal Elo of 1,135 for text-to-video and human-preference win rates reported against Grok Imagine Video, Seedance 2.0, Gemini Omni Flash, Runway Gen-4.5 and Luma Ray 3.2. Every one of those figures describes the video tier. None of them transfers to image generation, and BFL itself has published no image benchmarks or win rates. Quoting a FLUX 3 video Elo as evidence about FLUX 3's image quality is the single easiest mistake to make in this comparison.

A generated two-column comparison scoreboard titled 'Qwen-Image 2.1 vs FLUX 3 — the scoreboard'. Left column 'Qwen-Image 2.1': rows reading Status: released 20 Sep 2026; Callable today: local, five frameworks; Architecture: 7B, 32-layer DiT; Transparency: native RGBA; Reference images: up to 10; Price: none published; Licence: research, non-commercial. Right column 'FLUX 3': rows reading Status: announced 23 Jul 2026; Callable today: video only; Architecture: unified flow model; Transparency: none claimed; Reference images: not published; Price: video $0.06-$0.29 per second; Licence: not published for image. Footer reads 'Qwen-Image 2.1 figures per the vendor model card, unaudited; FLUX 3 image tier has published no figures at all.'

What each side can actually show

Put the evidence side by side and the asymmetry inverts from the previous section.

FLUX 3 has a shipped, priced, generally available product — just not the one in the title of this article. FLUX 3 Video generates clips up to 20 seconds with synchronised audio in a single pass, supports text-to-video, image-to-video, video-to-video, keyframe-to-video and generative continuation, and is already being tested by Canva, Burda, Magnific, Krea and Picsart. That is a real, deployed system with a real customer list. The image tier is the part that has not arrived, and BFL's original guidance that it would follow "in the coming weeks" has now outlived two full months without a public endpoint.

Qwen-Image 2.1 has the opposite profile. Its quality claims come from the vendor's own Qwen-Image-Bench comparison chart, and no third party has reproduced them. But the thing itself is in your hands: 33 GB of weights, an Apache-style repository layout with a licence file that is honest about being research-only, and a day-zero serving path across five frameworks. There is also one early-access hands-on review, from a tester who ran the final weights through a ModelScope Studio interface and reported roughly 10–15 seconds for text-to-image and 18–23 seconds for editing, strong camera-preset and multi-character role-assignment behaviour, and multi-reference consistency starting to degrade from about three input images. One reviewer, no timer, no published prompt set. Directionally useful, formally unverified.

The honest summary of the evidence: Qwen-Image 2.1 has a runnable model and no independent quality data; FLUX 3 Image has vendor claims and neither a runnable model nor any image data at all. If your question is "which of these should I plan a pipeline around," the answer today is neither, and the reason differs.

A screenshot of the Black Forest Labs FLUX 3 model page, captured September 20 2026, showing the FLUX 3 family overview with the Video tier marked as generally available and the Image tier listed as coming soon, alongside the FLUX 3 Video per-second pricing tiers.

Where the workaround actually is

For a commercial team, the Qwen-Image 2.1 blocker has a shape you can plan around. The licence is explicit, the commercial-licence contact is named in the repository, and the model is small enough — 7B in the generation component — that the cost of evaluating it on your own hardware is measured in an afternoon, not a quarter. Evaluate first, negotiate second, and keep the research build off any customer-facing path in the meantime. That is a normal procurement problem.

The FLUX 3 Image blocker is not a procurement problem, because there is nothing to procure. You can join an early-access queue, and you can watch for the endpoint to appear, and that is the entire available action. If you need FLUX-family image generation in production today, the callable options are the older FLUX.2 and FLUX Pro/Dev lines, which are priced and available — and which are a different generation of model from the one this article is about.

For a team that wants to test both candidates without committing either to a production path, this is the case the routing layer exists for. OrcaRouter puts 200+ models behind one OpenAI-compatible endpoint at provider list price with zero markup, with automatic failover across providers, so an unproven or unreleased model sits behind a route next to something you already trust instead of replacing it — and because list price is passed through, any vendor price change on a routed model is live the same day. Neither Qwen-Image 2.1 nor FLUX 3 Image is routable at the time of writing, and we are not going to pretend otherwise; what we do route is the image line around them, including OpenAI's GPT-Image family, Google's Imagen 4 tiers and Gemini image previews, and xAI's Grok Imagine image endpoint. Those are the models that can carry a workload while these two sort themselves out.

One more thing worth watching on the Qwen side specifically. Alibaba shipped Qwen-Image 1.0 and 2.0 as Apache-2.0 open weights, then shipped Qwen-Image 3.0 in July 2026 as a fully closed hosted model with no weights, no licence and no technical report, then shipped Qwen-Image 2.1 in September as open weights under a research-only licence. The direction of travel is not a straight line, and the licence terms on the next release are a genuine open question rather than a safe assumption in either direction.

A screenshot of the Qwen vendor blog page for Qwen-Image 2.1, captured September 20 2026, showing the 2026/09/20 date, the headline 'Compact, Efficient, and Unified Image Creation', the 'Now open weights' hero banner, GitHub, Hugging Face and ModelScope buttons, and the 7B parameter figure.

The verdict, such as it is

If you are a researcher or an internal evaluation team, Qwen-Image 2.1 is the more useful of the two right now, by a wide margin — it is downloadable, documented, framework-supported and honest about its terms. If you are a commercial team, both are blocked and the blocks are not equivalent: one is a contract you can start a conversation about, the other is a product that does not exist yet.

If you need to ship image generation this month, neither of these is the answer, and the useful question is which of the models you can call gets you closest. That is a benchmarking exercise, and it is worth doing with your own prompts rather than anyone's launch chart — including the vendor's, and including this one.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily