Hero title card reading 'Ming-Image-0.1-Design vs Gemini Omni 1.1 Flash', subtitle 'One returns a layered design, the other cannot return a still at all', with two minimal icon cards. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Ming-Image-0.1-Design vs Gemini Omni 1.1 Flash: A Design Deliverable Needs Both Halves

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Ask Ming-Image-0.1-Design for a video and you get a still. Ask Gemini Omni 1.1 Flash for a still and you get an error — its own documentation lists image generation, image editing and interleaved image-and-text output as unsupported capabilities for the model. So a page that puts these two in two columns is comparing a renderer to a projector. What is actually worth comparing is the thing design teams keep getting wrong: a shipped deliverable usually needs a screen and a moving asset, and these two models each own one half of that, with a handoff in the middle that quietly destroys work.

Ming-Image-0.1-Design is inclusionAI's 6B text-to-image model, released under MIT with weights published 17 September 2026, built for text-dense graphic design. Gemini Omni 1.1 Flash is Google's production video model, generally available since 27 August 2026, built for short-form motion with synchronized audio. Neither is a substitute for the other, and the interesting question is what happens between them.

The two halves, stated plainly

Start with what each one physically returns, because everything else follows.

• Output — Ming-Image-0.1-Design returns a still image, up to 2048 x 2048 at 12 sampling steps and CFG 1.0, or 1024 x 1024 for speed; Gemini Omni 1.1 Flash returns video, up to 40 seconds assembled from 10-second generation and extension calls

• Transparency — Ming-Image-0.1-Design emits native RGBA when the prompt opens with a documented transparency phrase, so alpha comes out of the sampler; Gemini Omni 1.1 Flash has no transparent output, and there is no transparent video format in the pipeline either

• Structure — Ming-Image-0.1-Design's companion Layer model decomposes a flattened design into 2–9 separate RGBA PNGs that recompose into the original; Gemini Omni 1.1 Flash returns a single flattened clip

• Reference input — Ming-Image-0.1-Design generates from a prompt alone with no reference image required; Gemini Omni 1.1 Flash accepts up to 10 images per prompt and up to three 3-second reference video clips, and reads up to 10 seconds of prior footage when extending a scene

• Resolution ceiling — Ming-Image-0.1-Design is native 2048; Gemini Omni 1.1 Flash renders natively at 720p and upscales to 1080p and 4K, with a 360p drafting tier

• Cost — Ming-Image-0.1-Design has no hosted price because there is no hosted endpoint, so the cost is your own GPU time on one 80 GiB card; Gemini Omni 1.1 Flash bills $0.10 per second at 720p, with the 360p draft tier around $0.03 per second

• Licence — MIT for Ming-Image-0.1-Design, commercial use permitted; a commercial API for Gemini Omni 1.1 Flash whose output carries SynthID watermarking

A rendered two-column scoreboard headed 'Two halves', titled 'Ming-Image-0.1-Design vs Gemini Omni 1.1 Flash'. The Ming-Image-0.1-Design column reads: returns a still to 2048 x 2048; 12 steps, CFG 1.0; transparency native RGBA; structure 2-9 editable RGBA layers; cost your own 80 GiB GPU; independent placement #1 open weights in the UI/UX slice. The Gemini Omni 1.1 Flash column reads: returns video to 40 seconds; 10-second calls with 10s lookback; transparency none and no alpha video; structure one flattened clip; cost $0.10 per second at 720p; independent placement top of the video arenas with no audio.

Where the handoff breaks

If you are building a design pipeline that produces both, the direction is one way only: Ming-Image-0.1-Design makes the frame, Gemini Omni 1.1 Flash animates it. The reverse does not exist. And the seam between them is where the design work gets lost.

The first casualty is the alpha channel. A UI asset generated with a transparent background is the reason to use a sampler that emits RGBA in the first place — it drops onto an existing layout without a cut-out pass. Feed that frame into a video path and the transparency is gone, because there is no transparent video output to preserve it in. Transparency survives in your compositor, applied after the clip comes back, not in the model chain. Teams that plan the composite before the clip exists end up redoing it.

The second casualty is resolution. Ming-Image-0.1-Design renders natively at 2048. Gemini Omni 1.1 Flash renders natively at 720p and upscales upward. A 2048-pixel design frame fed into a 720p render path is mostly discarded detail, and the upscale that follows is a vendor-side step you are not controlling.

The third is text. The entire premise of Ming-Image-0.1-Design is that rendered text stays legible at the size it will be read — that is the axis its 1,084 Elo in the Artificial Analysis UI/UX Design slice measures. A video model animating that frame does not re-render the copy; it moves the pixels it was given. Any motion that changes the perspective, scale or lighting of the frame will move the typography with it, and small type that was legible in a still will not survive the transform.

None of that is a defect in either model. It is what happens when a deliverable is split across two systems that were never designed to hand off to each other, and the fix is a production decision, not a model choice: decide which layer of the deliverable is allowed to be flattened, and flatten it deliberately.

What each half is actually good at

Ming-Image-0.1-Design's evidence is a third-party arena. It sits at rank 45 of 104 on the Artificial Analysis Text to Image leaderboard with an Elo of 995 across 21,342 votes, and first among open-weights models in that board's UI/UX Design slice at 1,084 Elo on 2,100 votes in the slice. The gap between those two numbers is the honest description of the model: a measured specialist, not a generalist. Its Layer companion's figures — RGB L1 error of 0.0574 and alpha soft IoU of 0.8923 at 1024 on the Crello test set — are vendor-reported and unreproduced, and should be labelled as such.

Gemini Omni 1.1 Flash's evidence is also independent, but on a different board. On Artificial Analysis it held the top spot on both the text-to-video and image-to-video arenas in the no-audio category as of 2 September 2026, at Elo 1,324 and 1,364. With audio enabled it drops to 1,237 and 1,180 — second and fourth. That audio gap is a detail a spec sheet hides and a leaderboard exposes, and it is worth knowing before you budget a campaign around a top-of-board claim.

The two boards do not overlap, which is the point. There is no arena where a design still and a ten-second clip are judged against each other, and any article that implies otherwise is inventing a scoreboard.

Screenshot of Google's Gemini API models documentation at ai.google.dev/gemini-api/docs/models, captured 24 September 2026 in English. The left navigation lists Gemini 3, Nano Banana, Veo, Gemini Omni Flash, Lyria 3.5 & Lyria Clip, Text-to-speech, Transcribe and other model families. The body shows the Models guide heading, a Gemini 3 section with Gemini 3.8 Flash, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking cards marked New and Stable, and a Generative media models entry in the on-this-page rail.

What the split costs, per attempt

The economics of the two halves are not comparable and should not be averaged into one budget line.

On the still side, Ming-Image-0.1-Design costs nothing per image once the hardware exists. That makes iteration free in the way that matters: you can generate twenty dashboard variants in an afternoon and throw away nineteen. On the motion side, a 10-second 720p clip on Gemini Omni 1.1 Flash bills about $1.00; the same 10 seconds at the 360p draft tier is roughly $0.30; a full 40-second scene at 720p — one generation plus three extension calls — lands near $4.00. Google has not published per-second rates for the 1080p and 4K tiers, and reseller listings quoting $0.15 and $0.30 per second are marketplace estimates rather than vendor numbers, so any high-resolution production budget stays provisional.

The practical consequence is that the two halves want opposite working styles. Iterate on the still, where retries are free. Commit on the video, where every retry is metered. A team that iterates on the video half is the team that gets surprised by the invoice, because the first frame costs nothing and the retakes cost everything.

Where a routing layer fits here is narrower than a pitch would suggest, and worth stating precisely. Ming-Image-0.1-Design is an MIT download — OrcaRouter routes no inclusionAI model, so it runs on your hardware or not at all. Gemini Omni 1.1 Flash is a vendor API with its own key and its own per-second meter, and we do not route it either; Google has not opened the Omni video API to third-party routing. What we do carry on the motion side is Kling's video line, including kling-v3, kling-v3-omni and kling-video-o1, plus MiniMax H3 — so if the animated half of a deliverable needs a hosted route rather than a second vendor contract, that exists on our side. On the image side we carry the OpenAI GPT-Image family, Google's Imagen 4 tiers and Gemini image previews, and xAI's Grok Imagine image endpoint. Because provider list price is passed through at 0% markup rather than marked up, a vendor price cut on any of them is live on our side the same day.

A rendered summary card headed 'The handoff', titled 'What a still loses on the way to motion', listing four losses: the alpha channel (Ming-Image-0.1-Design emits it from the sampler, Gemini Omni 1.1 Flash has no transparent video output); resolution (2048 x 2048 in against a path that renders natively at 720p and upscales); typography (the video model moves the pixels it is given and does not re-render the copy); and working style (stills iterate for free, video bills about $1.00 per 10 seconds at 720p and about $0.30 at the 360p draft tier).

The decision, in one pass

• You need editable design assets — Ming-Image-0.1-Design, and the layer decomposition is the reason. Transparent cutouts, icons, sprites and layered composites are where a sampler that emits RGBA removes a whole stage from the pipeline.

• You need motion, measured — Gemini Omni 1.1 Flash. It is the only one of the two with independent arena results at the top of a board, and the only one with a published per-second price you can put in a forecast.

• You need both — budget the video half per attempt rather than per asset, do not assume the alpha channel survives, and plan the composite after the clip returns.

• You need to sell the output — Gemini Omni 1.1 Flash is unambiguously the safer half, since MIT covers the Ming-Image-0.1-Design weights and Google's API terms cover the video. The watermarks are the trade: every Omni clip carries SynthID, and no Ming-Image-0.1-Design output does.

What is worth watching next is not another head-to-head. It is whether inclusionAI attaches a hosted endpoint to Ming-Image-0.1-Design — which would remove the 80 GiB entry cost and change this comparison entirely — and whether Google closes the audio gap on the Omni video boards. Those two facts, not another scoreboard, decide which half of a design deliverable gets the budget.