A hero title card for the comparison 'Qwen3.8-Omni-Flash vs Qwen 3.8' with the eyebrow 'Head to head - same family, opposite briefs', the subtitle 'A specialist built for hours of media against the 2.4-trillion-parameter open core built for depth per token', and three stat chips reading 'Price per M tokens: 13x apart', 'Context: 1M on both', and 'Open weights: one side only'.
Guides & Insights

Qwen3.8-Omni-Flash vs Qwen 3.8: A Specialist Bill vs a Flagship One

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

If you want the short version: Qwen3.8-Omni-Flash is roughly thirteen times cheaper per token than Qwen 3.8 and is the only one of the two built to sit inside an hour of audio-video without choking — while Qwen 3.8, the 2.4-trillion-parameter open core at the top of the Qwen3.8 line, is the one you use when the answer has to be right rather than cheap, and the only one of the two whose weights you can download. They share a context length, a family name and an August-to-September launch window. They are not substitutes, and the entire value of this comparison is in knowing which question you are actually asking.

To be precise about the names, because the family is crowded. Qwen 3.8 here means the 2.4T-A95B mixture-of-experts open core — the model behind the Qwen3.8-Max endpoint, which Alibaba unveiled on 3 August 2026 and published weights for on 12 August, and which Artificial Analysis indexes at 40 on its Intelligence Index. Qwen3.8-Omni-Flash is the native omni-modal model Alibaba put into API service on 18 September 2026, built on the Qwen3.8-Flash architecture line, taking text, image, audio and video in and returning text. Everything about the flagship is vendor-reported or independently indexed, as noted; everything about the omni model is vendor-reported alone, because it is hours old and nobody outside Alibaba has measured it.

Two models, one family, opposite design briefs

The temptation is to read this as a big model and a small model from the same generation, the way you would read Qwen3.8-Max against Qwen3.8-Flash. That reading is wrong in a way that costs money. The Qwen 3.8 core is a reasoning engine that happens to accept images and video; its throughput is measured in tokens per second on hard problems, its price is set to reflect the compute a 95-billion-active-parameter forward pass costs, and its evaluation sheet is a list of reasoning and coding boards. Qwen3.8-Omni-Flash is a perception-and-action engine that happens to reason; its price is set against the cost of ingesting hours of media, and its evaluation sheet is a list of audio, video and tool-calling boards. One is optimised for depth per token. The other is optimised for tokens per hour of content.

That difference shows up most sharply in a place the marketing does not mention. Both models accept roughly a million tokens and both accept video. But Qwen3.8-Omni-Flash is explicitly engineered around the duration problem — up to one hour of continuous audio-video in a single call, with an Agentic Understanding mode in which the model chooses what to sample instead of processing every frame, cutting tokens per query from 145,736 to 79,117 while raising accuracy on OmniVideoBench from 63.4 to 67.8. Qwen 3.8 has no equivalent published mechanism. Feeding it two hours of video is possible on paper; doing it at a price that survives contact with a finance team is a different question, and it is the question the omni model was built to answer.

The dimensions that actually decide it

• Price — Qwen3.8-Omni-Flash $0.15 per million input and $0.47 per million output, against Qwen 3.8 at $2.00 and $6.00. Cache read: $0.016 against $0.25.

• Open weights — Qwen 3.8 published weights on 12 August 2026 under a custom licence; Qwen3.8-Omni-Flash is API-only and Alibaba has not said it will be released.

• Context — one million tokens on both sides; 991K maximum input on the omni model, with 131K maximum output and a 262K reasoning budget.

• Media duration — up to one hour of continuous audio-video in a single omni call; the flagship is documented for video input with no per-hour cost story attached.

• Output modality — text on both sides. Neither generates speech, despite the omni model's lineage.

• Independent score — Qwen 3.8 sits at 40 on the Artificial Analysis Intelligence Index. Qwen3.8-Omni-Flash has no independent score of any kind yet.

A two-column scoreboard, 'Qwen3.8-Omni-Flash vs Qwen 3.8 - the scoreboard'. Left column Qwen3.8-Omni-Flash: Price per M $0.15 in / $0.47 out, Context 1M tokens, Media per call 1 hour continuous A/V, Output text only, Open weights not released, Independent score none yet. Right column Qwen 3.8: Price per M $2.00 in / $6.00 out, Context 1M tokens, Media per call video input with no per-hour price, Output text only, Open weights published 12 Aug 2026, Independent score AA Index 40. The footer reads 'Qwen 3.8 pricing and Index score per Artificial Analysis and Alibaba; Qwen3.8-Omni-Flash figures vendor-reported and unreproduced.'

Why the benchmark tables cannot settle this

There is no head-to-head board, and there is not going to be one soon. Alibaba's launch materials for Qwen3.8-Omni-Flash benchmark it exclusively against Qwen3.5-Omni-Plus, its direct predecessor, and against Gemini 3.8 Flash; the flagship's numbers come from a different set of reasoning and coding evaluations entirely. The two sheets do not share a single row. Anyone publishing a table that puts them side by side on one axis has built that axis themselves.

What can be said honestly is directional. On the vendor's own figures the omni model is a large step over the omni line it replaces — an average gain above 26% across roughly thirty evaluations, with WildClawBench-MM at 71.0, UniClawBench at 69.6, LongAudioSpan at 82.7 — and a genuine but partial one against Gemini 3.8 Flash, where it leads on audio-centric boards such as SpotSoundBench (67.2 against 39.7) and MMAU (81.8 against 76.9) and trails on the pure video-reasoning AgenticVBench (36.8 against 45.0). On the flagship side, Qwen 3.8's Index score of 40 is an independent measurement and worth more than any of the above. Those are different kinds of evidence and should not be averaged together.

A screenshot of the Qwen3.8-Omni-Flash launch post on qwen.ai (captured 18 September 2026, English UI), showing the 'Conducting Deep Research with Audio and Video' section with an embedded demo video, a MODEL WORKFLOW panel listing steps including Write report, Grab key frames, Dispatch Web research and Screenshot check, and a TIMELINE panel with chapter timestamps such as 0:00 Intro + a 5-minute composite and 10:02 Arrangement & Matching.

A worked cost comparison, with the assumption stated

Take a long-document-and-video analysis job: 300,000 input tokens and 20,000 output tokens, a plausible shape for a ninety-minute recorded meeting with slides. On Qwen3.8-Omni-Flash that is 300,000 × $0.15 per million = $0.045 of input and 20,000 × $0.47 per million = $0.0094 of output, so about 5.4 cents. On Qwen 3.8 the same call is 300,000 × $2.00 per million = $0.60 and 20,000 × $6.00 per million = $0.12, or about 72 cents. Slightly over thirteen times the cost for the same token count.

The number that changes the decision is not the ratio but the volume. At a hundred such jobs a day, that is roughly $5 a day against roughly $72 a day — the difference between a feature you ship and a feature you scope down. But if the task is one hard extraction from a 300K-token contract where a wrong answer costs an hour of human review, the 13x premium is irrelevant and the stronger reasoning model wins on the first mistake it does not make. Set the assumption before you look at the price line, not after.

Where each one actually wins

Qwen3.8-Omni-Flash wins on anything measured in hours. Meeting transcription with diarisation and follow-up task extraction, long-video summarisation, audio-visual research reports, captioning pipelines, and any agent that has to decide for itself which minute of a recording matters. The vendor-reported diarisation improvement — AliMeeting speaker error rate falling from 88.11 to 3.35 — is the kind of delta that turns a manually-reviewed pipeline into an automated one, though a Chinese technical analysis of the launch argues much of that gain comes from front-end and diarisation engineering wrapped around the model rather than from the model itself, so benchmark it on your own audio before you retire the human review step. The omni model also wins straightforwardly on price for any high-volume text workload where its text quality is adequate, which Alibaba claims is comparable to text-only models of the same size — an unreproduced claim, and worth testing rather than assuming.

Qwen 3.8 wins wherever the output is judged rather than counted: hard reasoning, long-horizon agentic work, large-scale code work, million-token document analysis where the synthesis is the product. It also wins on a dimension that has nothing to do with benchmarks — it is open-weight. If your constraint is data residency, on-premise deployment, fine-tuning, or simply not wanting a vendor in the inference path, the omni model is not a candidate at all, and no price advantage closes that gap. That single constraint decides more real deployments than every number above.

Running them without two contracts

The practical friction in a comparison like this is rarely the model choice; it is that you end up integrating two vendors, two billing relationships and two sets of client code to find out which one you wanted. Qwen3.8-Max — the endpoint carrying the 2.4-trillion-parameter open core — is live on OrcaRouter under the model id qwen/qwen3.8-max, at $2.00 per million input tokens and $6.00 per million output, with a one-million-token context and text, image and video input, alongside more than 200 other models on one key at provider list price with 0% markup and automatic failover across providers if an endpoint degrades. Qwen3.8-Omni-Flash is not in that catalogue — it runs on Alibaba's own platforms — so the honest configuration today is the flagship on our route and the omni model called directly, with the routing layer giving you a place to put the loser of the test once you have run it.

A screenshot of the OrcaRouter model page for Qwen3.8 Max at www.orcarouter.ai/models/qwen/qwen3.8-max (captured 18 September 2026, English UI), showing the model id qwen/qwen3.8-max with a FEATURED badge and a byline date of 2026-08-03, a spec panel reading ctx 1M tokens, Input text + image + video, Output text, p50 TTFT 10.00s, and a metrics row reading INPUT $2.00 / 1M tokens, OUTPUT $6.00 / 1M tokens, TRAFFIC 84.0M tokens / 7d, above OpenAI- and Anthropic-compatible SDK code samples.

The verdict

Choose Qwen3.8-Omni-Flash when the input is long, the output is short, and volume makes unit price the binding constraint. Choose Qwen 3.8 when the input is dense, the output is the product, and either correctness or deployment control is non-negotiable. The overlap zone — video understanding — is the only place a genuine head-to-head exists, and in that zone the omni model holds the cost and duration argument while the flagship holds the reasoning and open-weights argument. Run both on your own data before you commit; the omni model has no independent evaluation yet, and a launch-day price advantage is exactly the kind of thing that is worth confirming against a bill rather than an announcement.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily