A generated hero card headlined Qwen-Image-2.1 vs Nano Banana 2 Lite, showing a receipt icon with the text 0.034 dollars per image beside a stopwatch icon with the text 4.74 s per generation, captioned the cheap one is metered, the fast one is not free.
Guides & Insights

Qwen-Image-2.1 vs Nano Banana 2 Lite: $340 for Ten Thousand Images, or 13 Hours of GPU

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ten thousand 1024-pixel images on Nano Banana 2 Lite costs about $340. Ten thousand images from Qwen-Image-2.1 costs roughly thirteen hours of a 24 GB GPU and a licence that forbids you to sell any of them. That is the trade in one line, and it is the only comparison in this family where the two models actually do the same job at the same moment. Qwen-Image-2.1 went open-weights on 20 September 2026. Nano Banana 2 Lite — the vendor's model ID google/gemini-3.1-flash-lite-image, officially Gemini 3.1 Flash-Lite Image — launched on 30 June 2026 and is the cheapest still generator in the Nano Banana line.

The two candidates, priced honestly

Nano Banana 2 Lite generates a 1K image in roughly four seconds, about 2.7× faster than Gemini 3.1 Flash Image, and bills a flat $0.034 per 1K image. Google's token rates behind that figure are $0.25 per 1M input tokens and $30 per 1M image-output tokens, which works out to about $0.0336 at 1K; batch mode halves it to roughly $0.0168. There is no free tier, and every output carries an invisible SynthID watermark.

Qwen-Image-2.1 does not have a price, because there is nothing to buy. It is a roughly 33 GB download — a 7B, 32-layer single-stream diffusion transformer paired with a Qwen3-VL 8B text encoder — that you serve yourself. The nearest thing to a rate card is SGLang-Diffusion's measured latency: 4.74 seconds for the denoising pass on a 24 GB RTX 4090 using Cache-DiT and INT8 kernels, 8.23 seconds for a full 1024×1024, 40-step generation on an RTX PRO 6000 Blackwell at a 40.1 GiB peak footprint, and 2.748 seconds on a single B200. Those are single-request latencies including the components noted, not throughput figures, so treat 13 hours for ten thousand images as an order of magnitude rather than a benchmark.

Put those next to each other and the arithmetic is not "$340 versus free". It is $340 versus thirteen hours of a card you either already own or rent, plus the engineering time to stand up a serving stack. The crossover volume is far lower than most teams assume — but only if the card is not sitting idle the rest of the month, and only if the output is something you are permitted to produce.

Three workloads, and which model takes each

Volume is the wrong axis on its own. The work type decides it faster.

High-volume on-screen assets — social crops, thumbnails, A/B variants. Nano Banana 2 Lite wins on almost every component. Four seconds per image, fourteen aspect ratios including 1:4 and 8:1, a flat per-image price you can forecast to the cent, batch mode at half. Nothing about this workload needs alpha or 2K, and the watermarked commercial licence is exactly what you want when the output ships.
Print, large-format compositing, or anything you will crop hard. Qwen-Image-2.1, and it is not close. Native 2048×2048 at 40 steps against a model that is hard-capped at roughly 1 megapixel, with no 2K mode, no upscaling, and no API parameter to raise it — Google directs 2K and 4K work to Nano Banana 2 or Nano Banana Pro instead. If you need to crop a third of the frame away and still have a usable asset, Lite cannot get you there.
Transparent cutouts, icons, sprites, layered composites. Qwen-Image-2.1. The alpha channel is a component of the 64-channel RGBA latent space the sampler denoises in, so there is no segmentation pass and no matting pass; the model can also pull a subject out of an ordinary photograph into a transparent layer. Nano Banana 2 Lite has no documented transparent-background output.
Anything that has to be commercially licensed. Nano Banana 2 Lite, without qualification. Qwen-Image-2.1 ships under the Qwen Research License Agreement dated 20 September 2026 and permits non-commercial research and evaluation only; commercial deployment requires a separate licence from Hangzhou Tongyi Laboratory Technology, and there is no published price or term sheet for one.

One more row belongs in that list, and it cuts against the open model. Nano Banana 2 Lite knows things — it ships with a knowledge cutoff of January 2025 and a prompt-adherence profile built on the Gemini 3.1 Flash Lite backbone, with strong in-image text rendering including Chinese, Japanese and Korean. Qwen-Image-2.1's independent evidence on instruction-following is thin and mixed. An AI-judge arena comparison that pitted Nano Banana 2 Lite against a Qwen image model — the entry does not pin a version, so read it as a signal about the family rather than about 2.1 specifically — gave Lite the win on the odd spatial prompts ("astronaut riding a horse", "capybara taxi driver") and on text rendering, where the Qwen output rendered a Halloween invitation title as "Hallofarty Invitation". Lite's own documented weak spots are small faces, fine text detail, dense infographics, and complex masked edits. Neither model is reliably better at following a strange instruction; they fail in different places.

A generated two-column scoreboard titled Qwen-Image-2.1 vs Nano Banana 2 Lite - the scoreboard. Left column Qwen-Image-2.1 rows: Speed 4.74 s denoise on an RTX 4090; Price self-hosted GPU time; Resolution 2048 x 2048 native; Transparency native RGBA; Reference images up to 10; Licence non-commercial research only. Right column Nano Banana 2 Lite rows: Speed about 4 s per 1K image; Price 0.034 dollars per 1K image; Resolution hard cap at 1K; Transparency not supported; Reference images one image for editing; Licence commercial, SynthID watermark. Footer: Qwen figures per the vendor model card and SGLang measurements, unaudited; Google figures per the vendor launch material.

The reference-image question, unresolved

Multi-image editing is where a volume pipeline stops being able to substitute one model for the other, and it is also the row where the public information conflicts.

Qwen-Image-2.1 accepts up to 10 reference images and supports virtual try-on, group portraits and wardrobe composition on top of that. For Nano Banana 2 Lite, sources disagree sharply: one provider's model page states support for up to 14 reference images for style transfer and multi-reference editing, while a comparison write-up states the opposite — that Lite accepts a single image for editing, and that the 14-reference capability belongs to the full Nano Banana 2, with max five people plus six objects. Given the conflict, the defensible position is that Lite supports image input and multi-image composition but that its reference capacity is the differentiator Google uses to separate it from Nano Banana 2. If your workflow depends on six consistent characters across a series, verify the limit against Google's own model card before you architect around it.

This is also where the models stop being interchangeable in a broader sense. Google positions Nano Banana 2 Lite and Gemini Omni Flash as a pair — the still model and the video model, built to chain. Qwen-Image-2.1 has no video counterpart, and its animation trick is a chained sequence of local region edits that builds a simple frame-by-frame loop, not a video model.

Where a routing layer earns its place

Neither of these models is on OrcaRouter. We route no Qwen-Image model of any version, so Qwen-Image-2.1 is self-hosted or nothing. Nano Banana 2 Lite — the gemini-3.1-flash-lite-image identifier — is also not on our catalogue; the Google image routes we do carry are google/gemini-3.1-flash-image-preview, google/gemini-2.5-flash-image, google/gemini-3-pro-image-preview and the three Imagen 4 tiers, plus OpenAI's GPT-Image-2, GPT-Image-1.5 and GPT-Image-1-mini, and xAI's Grok Imagine image endpoint.

What that buys a mixed pipeline is the thing this comparison keeps running into: a decision that changes per asset. Volume assets go to a cheap hosted route, transparent cutouts go to your own card, and a vendor price move on the hosted side lands the same day because we pass provider list price through at zero markup rather than pricing it ourselves. Add automatic failover across providers, a routing DSL for composing fallbacks, and model fusion for panel-style calls, and the hosted half stops being a vendor relationship you have to manage per model — 200+ models, one OpenAI-compatible endpoint, one key. The self-hosted half stays your problem, which is the honest cost of running a research-licensed checkpoint on hardware you own.

Which one to pick, and the condition that flips it

If you are producing volume on-screen assets for commercial use, Nano Banana 2 Lite is the answer and the licence question never arises: $0.034 an image, four seconds, forecastable, sellable. If you need 2K, native transparency, or ten reference images, Qwen-Image-2.1 is the only one of the two that has any of them, and you pay for that in hardware, engineering time, and a non-commercial licence that makes it a research instrument rather than a production dependency.

One fact would collapse the gap. A published commercial term sheet for Qwen-Image-2.1 — with a price and terms someone can actually sign — turns a 13-hour GPU job into a line item that competes with $340, and at that point the resolution and alpha advantages start winning workloads that currently go to Lite by default. Until that exists, the split is clean: Lite for anything that ships, Qwen-Image-2.1 for anything that stays inside the building, and a serving stack for the second one that you maintain yourself.

A screenshot of OrcaRouter's own model page for google/gemini-3.1-flash-image-preview, captured September 21 2026, showing the model title Nano Banana 2 (Gemini 3.1 Flash Image Preview), the Vision, JSON and Reasoning capability tags, the 2026-02-26 date, the generateContent endpoint and a $0.15 price.A screenshot of the Qwen vendor blog page for Qwen-Image 2.1, captured September 21 2026, showing the 2026/09/20 publication date, the headline Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation, the Now open weights banner, and the GitHub, Hugging Face and ModelScope buttons.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily