
Ming-Image-0.1-Design vs Nano Banana 2: The Layer You Get Back
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
Ask a designer what goes wrong after the image is generated and the answer is almost never the image. It is that the generated image is flat. Every element is welded to every other element, so the headline cannot be moved, the logo cannot be swapped, the background cannot be recoloured, and the whole thing has to be regenerated from a prompt to change one word. Ming-Image-0.1-Design, published quietly by Ant Group's inclusionAI team on 17 September 2026 under an MIT licence, is built around that problem: 6.15 billion parameters, native RGBA output, and a companion repository that takes a flattened design and returns it as separate layers. Nano Banana 2 — the vendor's Gemini 3.1 Flash Image, generally available since 28 May 2026 and routable today as google/gemini-3.1-flash-image-preview — solves a different half of the same problem, and the difference between the two halves is what decides which one belongs in your pipeline.
What Nano Banana 2 is actually good at, and what it costs
Google's model is the one with a published price and a production track record, and both are worth stating precisely because they are the reason people reach for it.
• Price — roughly $0.067 per 1024×1024 image, $0.101 at 2048×2048, and about $0.151 at the 4K end, with batch mode at approximately half those figures and text input billed separately at $0.25 per million tokens. That converts to about $101 per thousand 2K images, forecastable to the dollar before you generate anything.
• Output — up to 4K, with ratio presets including the 1:4 and 1:8 extremes that social and banner formats need, and Google-reported consistency across five characters and fourteen objects.
• Provenance — every output carries SynthID watermarking, which is a compliance feature rather than a quality one and is the kind of thing a legal review asks about before a design review does.
• Grounding — it can search the web during generation, which is how it handles a request that depends on a fact it does not have memorised.
None of that is a transparency story. Nano Banana 2 returns an image, and the image is the whole of the artifact.

What Ming-Image-0.1-Design returns instead
The model card is unusually specific for an unannounced release, and the specifics are all pointed at design work rather than photography. Recommended inference is 2048×2048 — or 1024 when you want speed — at 12 sampling steps, classifier-free guidance at 1.0, BF16 throughout, on a single 80 GiB CUDA GPU. Weights are all BF16 safetensors, the repository is roughly 71.5 GB across connector/, mllm/, mlp/, scheduler/, transformer/ and vae/, and there is no model_index.json — the documented path is git clone of the inference repo and python infer.py, with vLLM-Omni named for deployment and a separate vision-language model (Ling-3.0-flash-VL or qwen3.8-27B) named for prompt enhancement.
The RGBA output is the first half of the answer to the flat-image problem. The second half is Ming-Image-0.1-Design-Layer, a separate repository with its own pipeline tag (image-text-to-image), 6,154,908,736 parameters, the same MIT licence, and one job: take a flattened design image plus a layer plan and return N RGBA layers. It runs at 1024 in its default bucket, 512 for speed, 12 steps, guidance at 2.0, with flash_attention_2. Its card shows Crello test-set results as an image rather than a table, so there is no number to quote — which is the honest state of the evidence and worth saying plainly.
Put those two facts together and the shape of the trade is clear. Nano Banana 2 gives you a better final image with less work. Ming-Image-0.1-Design gives you an image you can still take apart afterwards, and it gives you the layer decomposition as a second model you run when you need it.
The board, read carefully
Ming-Image-0.1-Design sits first on the Artificial Analysis Text to Image Leaderboard for UI/UX Design in the open-weights view, at an Elo of 1,082, ahead of Ideogram 4.0 (Quality) at 1,052, Ideogram 4.0 at 1,015, HunyuanImage 3.0 Instruct at 1,005, FLUX.2 [dev] at 1,000, FLUX.2 [dev] Flash at 999 and FLUX.2 [dev] Turbo at 994. The copy we are reading is the one reproduced on the model's own card, and it carries Artificial Analysis branding.
Two caveats that change how much weight it carries. It is an open-weights-filtered board, so a hosted model like Nano Banana 2 is not eligible and its absence is not a defeat. And a blind-preference Elo is an average over a crowd, not a test of your prompt — it says a panel preferred these outputs, not that the model will set your headline in the right typeface.
![The Artificial Analysis Text to Image Leaderboard for UI/UX Design, open-weights view, with Ming-Image-0.1-Design ranked first at an Elo of 1,082 ahead of Ideogram 4.0 (Quality) at 1,052, Ideogram 4.0 at 1,015, HunyuanImage 3.0 Instruct at 1,005 and the FLUX.2 [dev] variants below them.](https://cms.orcarouter.ai/api/media/file/3-1125.png)
Nano Banana 2's standing is measured on the hosted boards and on Google's own published figures, and the two should not be blurred together. It has been on the boards since spring, which is the more useful fact here: it has had months of independent traffic, and the Ming release has had none.
Where the money and the risk actually land
The pricing difference is not a difference of degree. Nano Banana 2 is metered and therefore forecastable: a thousand 2K images is about $101, and complexity does not change the bill much because image output tokenises predictably. Ming-Image-0.1-Design has no hosted price at all, because there is no hosted endpoint — the cost is your own GPU time against a 71.5 GB download on an 80 GiB card, and the honest comparison is not $101 against free but $101 against an amortised figure you have to compute yourself.
This is the one place in this comparison where a routing layer does something concrete rather than rhetorical. The Gemini image preview that serves Nano Banana 2 is on our endpoint, and it is on it at provider list price with zero markup — so if Google changes the per-image rate, the number on our side moves the same day rather than at the next contract renewal. Automatic failover across providers is the second half of that: a route can carry production traffic on a model with months of board history while an unannounced open-weights release is evaluated beside it, which means the evaluation never sits on the critical path and a bad afternoon on either side does not become an outage. That is a real arrangement, and it is the arrangement this particular pair of models is asking for.

Neither model is the safe default, for different reasons
Ming-Image-0.1-Design was uploaded on 17 September 2026 with no announcement, no release note, no stated limitations and a download counter reading zero at the time of writing. The leaderboard position above is the only independent measurement that exists. Everything else — throughput on your card, how it behaves above a handful of reference images, whether the alpha channel is clean at the edges of a cutout — is unmeasured in public, and no third party has written about it.
Nano Banana 2 has the opposite profile. It has independent scores, a published price, a commercial licence and SynthID, and it will not give you a layer back. If your workflow ends at "the image looks right", that is not a cost. If your workflow continues past it into a layout tool, it is the whole cost.
• Need a transparent PNG or a design you can decompose — Ming-Image-0.1-Design is the only one of the two, and the Layer companion is the reason to look at it rather than at its leaderboard position.
• Need a forecastable per-image price, 4K output, unusual aspect ratios and provenance watermarking — Nano Banana 2, at about $0.067 per 1K image, and it is on the same endpoint as everything else you call.
• Need to sell what comes out — Ming-Image-0.1-Design's MIT licence permits it, which is unusual for a 2026 image release and is worth checking against your own legal position rather than assuming.
• Need either one to be verifiable before you commit — only one of them can be, and it is the one with no announcement.
The thing to watch is not a benchmark. It is whether the Layer repository turns out to be the useful half of this release. Layer decomposition is the step where generated design work stops being a picture and starts being an asset, and if that model does what its card implies, the flat-image problem is the part of this comparison that will age fastest.
