
Qwen3.8-Omni-Flash vs Qwen 3.8: A Specialist Bill vs a Flagship One
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 984 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
If you want the short version: Qwen3.8-Omni-Flash is roughly thirteen times cheaper per token than Qwen 3.8 and is the only one of the two built to sit inside an hour of audio-video without choking — while Qwen 3.8, the 2.4-trillion-parameter open core at the top of the Qwen3.8 line, is the one you use when the answer has to be right rather than cheap, and the only one of the two whose weights you can download. They share a context length, a family name and an August-to-September launch window. They are not substitutes, and the entire value of this comparison is in knowing which question you are actually asking.
To be precise about the names, because the family is crowded. Qwen 3.8 here means the 2.4T-A95B mixture-of-experts open core — the model behind the Qwen3.8-Max endpoint, which Alibaba unveiled on 3 August 2026 and published weights for on 12 August, and which Artificial Analysis indexes at 40 on its Intelligence Index. Qwen3.8-Omni-Flash is the native omni-modal model Alibaba put into API service on 18 September 2026, built on the Qwen3.8-Flash architecture line, taking text, image, audio and video in and returning text. Everything about the flagship is vendor-reported or independently indexed, as noted; everything about the omni model is vendor-reported alone, because it is hours old and nobody outside Alibaba has measured it.
Two models, one family, opposite design briefs
The temptation is to read this as a big model and a small model from the same generation, the way you would read Qwen3.8-Max against Qwen3.8-Flash. That reading is wrong in a way that costs money. The Qwen 3.8 core is a reasoning engine that happens to accept images and video; its throughput is measured in tokens per second on hard problems, its price is set to reflect the compute a 95-billion-active-parameter forward pass costs, and its evaluation sheet is a list of reasoning and coding boards. Qwen3.8-Omni-Flash is a perception-and-action engine that happens to reason; its price is set against the cost of ingesting hours of media, and its evaluation sheet is a list of audio, video and tool-calling boards. One is optimised for depth per token. The other is optimised for tokens per hour of content.
That difference shows up most sharply in a place the marketing does not mention. Both models accept roughly a million tokens and both accept video. But Qwen3.8-Omni-Flash is explicitly engineered around the duration problem — up to one hour of continuous audio-video in a single call, with an Agentic Understanding mode in which the model chooses what to sample instead of processing every frame, cutting tokens per query from 145,736 to 79,117 while raising accuracy on OmniVideoBench from 63.4 to 67.8. Qwen 3.8 has no equivalent published mechanism. Feeding it two hours of video is possible on paper; doing it at a price that survives contact with a finance team is a different question, and it is the question the omni model was built to answer.
The dimensions that actually decide it
• Price — Qwen3.8-Omni-Flash $0.15 per million input and $0.47 per million output, against Qwen 3.8 at $2.00 and $6.00. Cache read: $0.016 against $0.25.
• Open weights — Qwen 3.8 published weights on 12 August 2026 under a custom licence; Qwen3.8-Omni-Flash is API-only and Alibaba has not said it will be released.
• Context — one million tokens on both sides; 991K maximum input on the omni model, with 131K maximum output and a 262K reasoning budget.
• Media duration — up to one hour of continuous audio-video in a single omni call; the flagship is documented for video input with no per-hour cost story attached.
• Output modality — text on both sides. Neither generates speech, despite the omni model's lineage.
• Independent score — Qwen 3.8 sits at 40 on the Artificial Analysis Intelligence Index. Qwen3.8-Omni-Flash has no independent score of any kind yet.

Why the benchmark tables cannot settle this
There is no head-to-head board, and there is not going to be one soon. Alibaba's launch materials for Qwen3.8-Omni-Flash benchmark it exclusively against Qwen3.5-Omni-Plus, its direct predecessor, and against Gemini 3.8 Flash; the flagship's numbers come from a different set of reasoning and coding evaluations entirely. The two sheets do not share a single row. Anyone publishing a table that puts them side by side on one axis has built that axis themselves.
What can be said honestly is directional. On the vendor's own figures the omni model is a large step over the omni line it replaces — an average gain above 26% across roughly thirty evaluations, with WildClawBench-MM at 71.0, UniClawBench at 69.6, LongAudioSpan at 82.7 — and a genuine but partial one against Gemini 3.8 Flash, where it leads on audio-centric boards such as SpotSoundBench (67.2 against 39.7) and MMAU (81.8 against 76.9) and trails on the pure video-reasoning AgenticVBench (36.8 against 45.0). On the flagship side, Qwen 3.8's Index score of 40 is an independent measurement and worth more than any of the above. Those are different kinds of evidence and should not be averaged together.

A worked cost comparison, with the assumption stated
Take a long-document-and-video analysis job: 300,000 input tokens and 20,000 output tokens, a plausible shape for a ninety-minute recorded meeting with slides. On Qwen3.8-Omni-Flash that is 300,000 × $0.15 per million = $0.045 of input and 20,000 × $0.47 per million = $0.0094 of output, so about 5.4 cents. On Qwen 3.8 the same call is 300,000 × $2.00 per million = $0.60 and 20,000 × $6.00 per million = $0.12, or about 72 cents. Slightly over thirteen times the cost for the same token count.
The number that changes the decision is not the ratio but the volume. At a hundred such jobs a day, that is roughly $5 a day against roughly $72 a day — the difference between a feature you ship and a feature you scope down. But if the task is one hard extraction from a 300K-token contract where a wrong answer costs an hour of human review, the 13x premium is irrelevant and the stronger reasoning model wins on the first mistake it does not make. Set the assumption before you look at the price line, not after.
Where each one actually wins
Qwen3.8-Omni-Flash wins on anything measured in hours. Meeting transcription with diarisation and follow-up task extraction, long-video summarisation, audio-visual research reports, captioning pipelines, and any agent that has to decide for itself which minute of a recording matters. The vendor-reported diarisation improvement — AliMeeting speaker error rate falling from 88.11 to 3.35 — is the kind of delta that turns a manually-reviewed pipeline into an automated one, though a Chinese technical analysis of the launch argues much of that gain comes from front-end and diarisation engineering wrapped around the model rather than from the model itself, so benchmark it on your own audio before you retire the human review step. The omni model also wins straightforwardly on price for any high-volume text workload where its text quality is adequate, which Alibaba claims is comparable to text-only models of the same size — an unreproduced claim, and worth testing rather than assuming.
Qwen 3.8 wins wherever the output is judged rather than counted: hard reasoning, long-horizon agentic work, large-scale code work, million-token document analysis where the synthesis is the product. It also wins on a dimension that has nothing to do with benchmarks — it is open-weight. If your constraint is data residency, on-premise deployment, fine-tuning, or simply not wanting a vendor in the inference path, the omni model is not a candidate at all, and no price advantage closes that gap. That single constraint decides more real deployments than every number above.
Running them without two contracts
The practical friction in a comparison like this is rarely the model choice; it is that you end up integrating two vendors, two billing relationships and two sets of client code to find out which one you wanted. Qwen3.8-Max — the endpoint carrying the 2.4-trillion-parameter open core — is live on OrcaRouter under the model id qwen/qwen3.8-max, at $2.00 per million input tokens and $6.00 per million output, with a one-million-token context and text, image and video input, alongside more than 200 other models on one key at provider list price with 0% markup and automatic failover across providers if an endpoint degrades. Qwen3.8-Omni-Flash is not in that catalogue — it runs on Alibaba's own platforms — so the honest configuration today is the flagship on our route and the omni model called directly, with the routing layer giving you a place to put the loser of the test once you have run it.

The verdict
Choose Qwen3.8-Omni-Flash when the input is long, the output is short, and volume makes unit price the binding constraint. Choose Qwen 3.8 when the input is dense, the output is the product, and either correctness or deployment control is non-negotiable. The overlap zone — video understanding — is the only place a genuine head-to-head exists, and in that zone the omni model holds the cost and duration argument while the flagship holds the reasoning and open-weights argument. Run both on your own data before you commit; the omni model has no independent evaluation yet, and a launch-day price advantage is exactly the kind of thing that is worth confirming against a bill rather than an announcement.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
