A generated hero card headlined Qwen-Image-2.1 vs Gemini Omni 1.1 Flash, showing a card labelled Qwen-Image-2.1 with a picture-frame icon and a transparency checkerboard beside a card labelled Gemini Omni 1.1 Flash with a film-strip icon, captioned one returns a file, the other returns a clip.
Guides & Insights

Qwen-Image-2.1 vs Gemini Omni 1.1 Flash: One Returns a PNG, the Other Cannot

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The most useful thing to know about Qwen-Image-2.1 and Gemini Omni 1.1 Flash is that they do not compete. Ask Gemini Omni 1.1 Flash for a still image and you get an error: the vendor's own documentation lists image generation, image editing, and interleaved image-and-text output as unsupported capabilities for the model. Ask Qwen-Image-2.1 for a video and you get a 2048×2048 still with an alpha channel. Any comparison table that puts them in two columns is comparing a camera to a projector. The real question — the one that actually costs money — is what happens when you chain them, and whether the chain survives contact with the rate card. Qwen-Image-2.1 went public on 20 September 2026; Gemini Omni 1.1 Flash has been generally available since 27 August 2026.

What each model physically returns

This is the whole comparison, and everything else follows from it.

Output format — Qwen-Image-2.1 returns a still image, up to 2048×2048 natively at 40 inference steps, optionally with a real alpha channel; Gemini Omni 1.1 Flash returns video, up to 40 seconds assembled from 10-second generation and extension calls, with no still-image mode at all.
Transparency — Qwen-Image-2.1 denoises in a 64-channel RGBA latent space, so alpha comes out of the sampler; Gemini Omni 1.1 Flash has no transparent-background output, and requesting one fails.
Reference input — Qwen-Image-2.1 accepts up to 10 reference images for multi-subject composition; Gemini Omni 1.1 Flash accepts up to 10 images per prompt as *input* and up to three 3-second reference video clips.
Resolution ceiling — Qwen-Image-2.1 is native 2K; Gemini Omni 1.1 Flash offers 360p, 720p, 1080p and 4K, but 1080p and 4K are upscales of a native 720p render rather than native generations.
What it costs — Qwen-Image-2.1 has no hosted price because there is no hosted endpoint, so the cost is your own GPU time; Gemini Omni 1.1 Flash bills $0.10 per second at 720p, with a 360p draft tier around $0.03 per second.
Licence — Qwen-Image-2.1 ships under the Qwen Research License Agreement dated 20 September 2026, non-commercial only; Gemini Omni 1.1 Flash is a commercial API whose output carries SynthID and C2PA provenance by default.

The last row is the one people skim past. On one side you have a model you can download and cannot sell. On the other, a model you cannot download and can sell freely — with an invisible watermark in every frame.

Why the "versus" framing keeps appearing anyway

Search for the two names together and you will find pages ranking them head to head. The reason is that the underlying question is real even when the framing is wrong: both models sit in the same creative pipeline, and both are the newest thing in it. What people are actually trying to decide is where the still ends and the motion begins.

Gemini Omni 1.1 Flash earned its place on that question with independent numbers rather than vendor ones. On Artificial Analysis it held the top spot on both the text-to-video and image-to-video arenas in the no-audio category as of 2 September 2026, at Elo 1324 and 1364. With audio enabled it drops to Elo 1237 and 1180 — second and fourth. That audio gap is the sort of detail a spec sheet hides and a leaderboard exposes.

Qwen-Image-2.1's evidence is thinner and of a different kind. The vendor's own Qwen-Image-Bench puts it at 60.28, ahead of Nano Banana 2.0 at 59.82 and GPT Image 1.5 at 59.65, and first among open models — but that is a vendor-run benchmark with margins under a point, and it is not a third-party reproduction. Independent testing so far consists of a GenAI Showdown human-scored run at 7/15, up from 4/15 for the original Qwen-Image and behind Ideogram 4's 8/15. The model had no entry on the Artificial Analysis image boards at the time of writing. Not a low score — no score.

A generated two-column scoreboard titled Qwen-Image-2.1 vs Gemini Omni 1.1 Flash - the scoreboard. Left column Qwen-Image-2.1 rows: Output still image, 2048 x 2048; Transparency native RGBA; Reference input up to 10 images; Resolution native 2K; Price self-hosted only; Licence non-commercial research only. Right column Gemini Omni 1.1 Flash rows: Output video, up to 40 seconds; Transparency none; Reference input 10 images plus 3 clips; Resolution native 720p, upscaled; Price 0.10 dollars per second at 720p; Licence commercial, SynthID and C2PA. Footer: Qwen figures per the vendor model card, unaudited; Gemini figures per Google documentation and Artificial Analysis.A screenshot of the Qwen vendor blog page for Qwen-Image 2.1, captured September 21 2026, showing the 2026/09/20 publication date, the headline Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation, the Now open weights banner, and the GitHub, Hugging Face and ModelScope buttons.

The handoff, and what it actually costs

If you are building something that needs both a still and a clip, the pipeline is one direction only: Qwen-Image-2.1 makes the frame, Gemini Omni 1.1 Flash animates it. The reverse does not exist.

The arithmetic is unforgiving once you run it. A 10-second 720p clip on Gemini Omni 1.1 Flash bills about $1.00. At the 360p draft tier the same 10 seconds is roughly $0.30, and a full 40-second scene at 720p — one generation plus three extension calls — lands near $4.00. Meanwhile the still that seeds the whole thing costs nothing but the electricity on your own card. Which means the interesting cost decision is not "which model" but "how many takes." A still you can iterate on for free; a video you cannot. Teams that budget video generation per asset rather than per attempt are the ones that get surprised, because the first frame is free and the retries are not.

There is a quality trap in the handoff too. Gemini Omni 1.1 Flash renders natively at 720p and upscales to 1080p and 4K. If your still is a 2048×2048 Qwen-Image-2.1 frame with a clean alpha channel, feeding it into a video path that renders at 720p and upscales throws away most of what the image model gave you — and the alpha channel is dropped entirely, because there is no transparent video output. Transparency survives in the compositor, not in the model chain. Plan the composite after the clip comes back, not before.

Where the routing layer fits

The awkward part of this pairing is that neither model is reachable the way the rest of your stack probably is. Gemini Omni 1.1 Flash is a vendor API with its own key and its own per-second meter. Qwen-Image-2.1 is a roughly 33 GB download — a 7B diffusion transformer plus a Qwen3-VL 8B text encoder — that you serve yourself, under a licence that permits non-commercial use only. Two billing models, two access paths, one project.

OrcaRouter exists for exactly the surrounding surface: one OpenAI-compatible endpoint across 200+ models, provider list price passed through at zero markup, automatic failover across providers, a routing DSL for composing fallbacks, and model fusion for panel-style calls. Being precise about what that covers here matters. We route no Qwen-Image model of any version, so Qwen-Image-2.1 is not a route you can call on our side — it runs on your hardware or not at all. We also do not route Gemini Omni 1.1 Flash. What we do carry on the motion side is Kling's video line, including kling-v3, kling-v3-omni and kling-video-o1, plus MiniMax H3 — so the animation half of this pipeline has a hosted route that does not require a second vendor contract, and on the image side we carry the OpenAI GPT-Image family, Google's Imagen 4 tiers and Gemini image previews, and xAI's Grok Imagine image endpoint. Because list price is passed through rather than marked up, a vendor price cut on any of them is live here the same day.

Which one you actually need

You need assets, not footage — Qwen-Image-2.1, and the alpha channel is the reason. Transparent product cutouts, icons, sprites, and layered composites are where a sampler that emits RGBA saves an entire matting pipeline.
You need motion and you need it measured — Gemini Omni 1.1 Flash. It is the only one of the two with independent arena results at the top of a board, and the only one with a published per-second price you can put in a forecast.
You need both — budget the video half per attempt, not per asset, and do not assume the alpha channel makes it through.
You need to sell the output — Gemini Omni 1.1 Flash, and only Gemini Omni 1.1 Flash, until Qwen publishes commercial terms for Qwen-Image-2.1. There is no published price or term sheet for that licence today.

The honest summary is that the comparison that ranks them is the wrong artifact. What is worth watching next is whether Qwen attaches commercial terms to Qwen-Image-2.1, and whether Google closes the audio gap on the Omni leaderboards — those two facts, not another scoreboard, decide which half of this pipeline gets the budget.

A screenshot of the Google Gemini API documentation models index, captured September 21 2026, showing the left navigation listing Nano Banana, Veo and Gemini Omni Flash under the models section, and the main panel opening with the Gemini 3 family of stable models.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily