
Qwen3.8-Omni-Flash vs Gemini Omni 1.1 Flash: One Watches the Video, One Makes It
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Short answer: these two are not substitutes, and the comparison only makes sense once you stop treating "omni" as a category. Qwen3.8-Omni-Flash, which the vendor opened to API customers on 18 September 2026, takes text, image, audio and video in and returns text — it is a model for reading an hour of meeting or footage and telling you what happened. Gemini Omni 1.1 Flash, which the vendor shipped on 27 August 2026, takes a text or image prompt and returns video with synchronized audio — it is a model for making the footage. One consumes media, the other produces it. If your pipeline needs both, you need both, and the honest question is not which wins but which one you are actually short of.
Neither model is open-weight, neither is routable through OrcaRouter, and both are priced on axes that cannot be converted into each other — tokens on one side, seconds of generated video on the other. Where they genuinely overlap is narrower than the "omni vs omni" framing suggests, and that overlap is worth pinning down before the price arithmetic, because the price arithmetic is where the two get compared wrongly most often.
The category mistake worth clearing up first
Alibaba's launch materials benchmark Qwen3.8-Omni-Flash against Gemini 3.8 Flash, and a great deal of the coverage that followed compressed that into "beats Gemini." That is a different Google model from Gemini Omni 1.1 Flash, and the two are not interchangeable reference points: Gemini 3.8 Flash is the general-purpose multimodal model, and Gemini Omni 1.1 Flash is the video-generation model with a separate price list and a separate API surface. Every Qwen-versus-Gemini number circulating from the launch — WildClawBench-MM 71.0 against 58.9, SpotSoundBench 67.2 against 39.7, AgenticVBench 36.8 against 45.0 — is against Gemini 3.8 Flash, not against the Omni video model, and none of it is independent in any case. If you came here holding a scoreboard that puts the two models in this article on the same axis, it is almost certainly the wrong Google model in the right-hand column.
What each one actually does
• What it takes in — Qwen3.8-Omni-Flash accepts text, image, audio and video, including stereo and four-channel spatial audio; Gemini Omni 1.1 Flash accepts text and image prompts plus up to three seconds of reference video.
• What it gives back — Qwen3.8-Omni-Flash returns text only; Gemini Omni 1.1 Flash returns video with natively synchronized audio, in scenes extendable to 40 seconds in ten-second increments.
• Control surface — Qwen3.8-Omni-Flash exposes context, sampling and an agentic mode that decides for itself what to look at; Gemini Omni 1.1 Flash exposes first-frame and last-frame control, a ten-second lookback for continuity, and a 360p drafting tier that Google describes as roughly one-third the cost and about 60% faster.
• Resolution and finish — Qwen3.8-Omni-Flash has no output resolution because it produces no media; Gemini Omni 1.1 Flash outputs at 720p, 1080p and 4K.
• Context and duration — Qwen3.8-Omni-Flash carries one million tokens and up to one hour of continuous audio-video in a single call; Gemini Omni 1.1 Flash is bounded by the clip it is generating, not by an input window.
• How you pay — Qwen3.8-Omni-Flash is billed per token: $0.15 per million input, $0.016 cached, $0.47 output on the international list; Gemini Omni 1.1 Flash is billed per second of generated video, with Google's published rate at $0.10 per second for 720p.
• Openness — closed on both sides. No weights for either model, and neither is available to route through our API today.

The overlap is video understanding, and it is asymmetric
The one place a buyer could reasonably put these side by side is a pipeline that has to both understand footage and produce it — say, an editing assistant that reads a long recording, finds the moments worth keeping, and generates a cut. On the understanding half, Qwen3.8-Omni-Flash is built for exactly this: an hour of continuous audio and video per call, speaker separation, and an agentic mode that raises accuracy on the vendor's OmniVideoBench from 63.4 to 67.8 while cutting tokens per query from 145,736 to 79,117. On the production half, Gemini Omni 1.1 Flash is the one that can generate the resulting clip with synchronized audio, and Qwen's model cannot produce a frame of video at all.
The asymmetry runs the other way too, and it is easier to miss. A video-generation model's grip on understanding is instrumental — it needs enough of a scene to keep a shot consistent across a ten-second extension, which is a different requirement from answering a question about what was said in the third hour of a meeting. Google's model has the longer track record on generation quality and the shorter one on long-form analysis; Alibaba's has the reverse. If your task is "find the three minutes that matter in this recording," the generation model is not the tool. If your task is "produce a thirty-second clip matching this brief," the understanding model is not the tool either.
Why the price comparison keeps going wrong
A per-second rate and a per-token rate do not convert, and the attempt to convert them produces the two errors that dominate this matchup. The first is comparing a whole generated clip against a per-token input rate and concluding the video model is hundreds of times more expensive — which is true and meaningless, because they are not selling the same thing. The second is comparing only input prices and missing that Alibaba's 98% audio cut applies to input audio hours, while Google's rate applies to output video seconds, so the two "cuts" are not even on the same side of the meter.
What can be said honestly is this. Under Alibaba's own methodology — a two-minute sample multiplied by thirty, video at 720p and one frame per second — an hour of combined audio and video input to Qwen3.8-Omni-Flash is quoted at about $0.20, down from $3.27 against the previous generation. Generating forty seconds of 720p video with Gemini Omni 1.1 Flash at Google's published $0.10 per second costs $4.00, and drafting the same forty seconds at 360p lands near $1.20 on Google's one-third-cost framing. Both sets of figures are vendor-reported and neither has been independently reproduced. The useful conclusion is not which number is smaller but that analysing an hour of footage and generating forty seconds of it sit in the same order of magnitude — which is the real reason a production pipeline that does both needs a cost model, not a winner.
That is also the argument for keeping the two on one key rather than two contracts. Neither omni model is routable through OrcaRouter today: Qwen3.8-Omni-Flash runs only on Alibaba's own platforms, and Google has not opened the Omni video API to third-party routing, so we do not host either and will not pretend to. What is live in our catalogue is each family's sibling — Qwen3.8-Flash and Gemini 3.8 Flash — both callable today through one endpoint at provider list price with 0% markup, which is what makes an A/B against either vendor's own numbers a matter of a config change rather than a second integration.

Where each one wins, stated plainly
Qwen3.8-Omni-Flash wins on anything measured in hours of input: meeting transcription and diarisation, long-recording summarisation, audio-heavy agentic tool use, and cost per hour of media processed. Its vendor-reported audio results are strong and its weakness is where the model was never aimed — the reasoning-heavy boards, where the first third-party runs put it in the 28th percentile on reasoning with a speed figure in the 21st, on a sample small enough that the numbers should be treated as a prompt to test rather than a verdict.
Gemini Omni 1.1 Flash wins on output: shot generation with synchronized audio, continuity across extended scenes, frame-level control, and the 4K finish that a delivery pipeline may simply require. Its advantage is not a benchmark score — it is that the other model cannot do the job at all. Where it is weaker is anything that requires holding an hour of context, because it is not built to.
The decision, and the thing to watch
If you are choosing between them, the question is which half of the pipeline is unserved. A team that already generates video and needs to understand long recordings should be testing Qwen3.8-Omni-Flash on its own audio this week; a team that analyses media and needs to produce it should be testing Gemini Omni 1.1 Flash. A team that needs both is looking at two vendors, two billing models and two integrations, which is the situation where a single routing layer stops being a convenience and starts being the thing that keeps the cost model legible.
Two things would change this read. An independent evaluation of Qwen3.8-Omni-Flash would replace every vendor number above with a comparable one — none exists yet, and its absence is the largest single gap in this comparison. And a published rate for 1080p and 4K output from Google would make the high-end arithmetic checkable; today those figures circulate only as marketplace estimates rather than as Google's own list.

One closing note on framing. "Omni" has stopped meaning anything precise — it now covers a model that reads an hour of meeting audio and one that renders a forty-second clip with sound, and treating the word as a category is how buyers end up comparing a per-token rate to a per-second rate and drawing a conclusion from it. The models are complements with a narrow overlap, and the correct posture toward both, five days and four weeks after their respective releases, is the same: measure them on your own material, and keep a second path open.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
