
Qwen3.8-Omni-Flash Is Here: 1M-Token Context, 98% Cheaper Audio, Text Out Only
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiNEWOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3135Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
The launch opened Qwen3.8-Omni-Flash to API customers today, 18 September 2026, and the number it led with is a bill rather than a score. Per hour of audio input, Qwen says the new model costs 98% less than its predecessor Qwen3.5-Omni-Plus; per hour of combined audio and video, 93% less. Everything else in the announcement follows from that framing. Qwen3.8-Omni-Flash is a native omni-modal model — text, image, audio and video in, one million tokens of context, up to a full hour of continuous audio-video inside a single call — and Alibaba is selling it not as a better describer of media but as the first Qwen model built to act on it. The tagline is "Omni Senses. Agentic Delivery." The catch, stated plainly in the same materials, is that the model answers only in text: it hears and watches, it does not speak.
Nothing below has been independently reproduced. The model is hours old, no third-party evaluation of it exists yet, and every benchmark and price figure in the launch materials is Alibaba's own. Where a number is vendor-reported this piece says so in the sentence that carries it; where a reading comes from a critic rather than the vendor, it is attributed. The one thing that is independently checkable today — that the model is live on Alibaba's platforms and not in anyone else's catalogue — is the thing the launch is least interested in talking about.
What actually shipped
The spec sheet is unusually clean for an omni-modal launch, which is itself the story: the previous generation of omni models in this family were assembled from separate perception encoders bolted onto a text LLM, and this one is built directly on the Qwen3.8-Flash architecture line, with modalities handled natively rather than grafted on.
• Input — text, image, audio and video, with stereo and four-channel spatial audio accepted; output is text only.
• Context — one million tokens, with maximum input at roughly 991K (about 983K in thinking mode), maximum output at 131K, and a reasoning budget that reaches 262K.
• Duration — up to one hour of continuous audio or audio-video per call, which is the number that makes the meeting and long-video use cases possible without chunking.
• Speech coverage — recognition across 74 languages and speech generation across 29 in the surrounding tooling; one Chinese-language account of the release states 113 languages and dialects for audio input.
• Availability — the Qianwen AI platform and Alibaba Cloud Model Studio (Bailian), under the API id qwen3-8-omni-flash. Weights are not being released. A Realtime variant is described as the first omni-modal model with sound-source localisation.

The 98% is a cost claim, and the arithmetic behind it is worth seeing
A 98% reduction in the cost of an hour of audio is a much larger number than a 98% reduction in the price of a token, and the gap between those two things is where the announcement does its work. The per-hour figure is derived by pricing a two-minute sample and multiplying by thirty, with video sampled at 720p and one frame per second. That is a legitimate way to quote a production workload, and it is also a description of what the model is being asked to look at: one frame per second is enough to know what was said and roughly what happened on screen, and nowhere near enough to inspect frame-level operations, judge edit quality, or catch a single-frame artefact. Two changes are being multiplied together — a genuine unit-price cut, and a much more aggressive sampling policy — and only the first is a model improvement.
The unit prices themselves are concrete. International pricing is quoted at $0.15 per million input tokens, $0.016 per million cached input tokens, and $0.47 per million output tokens. Domestic Bailian pricing is quoted at ¥0.8 per million for multimodal input, down from roughly ¥18. Note that the international and mainland billing systems are separate and the figures are not interchangeable; anyone comparing them directly will get a wrong answer.
This is the part of the release where routing matters more than usual. OrcaRouter passes provider list price through at 0% markup, so a vendor price cut of this kind reaches our customers the day it lands rather than at the next contract renewal. One genuinely verifiable comparison already sits in our own catalogue: Qwen3.8-Flash, the multimodal model this one is built alongside, is routable today at $0.15 per million input tokens and $0.47 per million output — the same list price Alibaba is quoting for the new omni model — and it carried 2.47 billion tokens of traffic through our API in the seven days before this was written, which is a fair measure of how much appetite the tier already has. The omni model itself is not in our catalogue yet, and until it is, it runs on Alibaba's own platforms and nowhere else. We are not going to pretend otherwise.
Every benchmark in the launch compares one thing
The headline is an average improvement of more than 26% across roughly 30 evaluations — and every one of those evaluations is against Qwen3.5-Omni-Plus. That is the right baseline for a launch and a narrow one: it tells you the family moved, not where it landed relative to anything you might already be paying for.
The individual figures, all vendor-reported: WildClawBench-MM 71.0, up 36.5 points and the single largest move in the set; UniClawBench 69.6; AgenticVBench up 22.3 points to 36.8; LongAudioSpan 82.7, up 8.3; OmniVideoBench 63.4, up 9.6; OmniCap-IF 28.2. The most striking pair is speaker diarisation on AliMeeting, where the error rate drops from 88.11 to 3.35 and cpWER from 89.61 to 17.18 — a jump large enough that it deserves scrutiny rather than applause. A Chinese technical analysis of the launch argues exactly that, attributing the improvement substantially to front-end and diarisation engineering wrapped around the model rather than to the model's own reasoning, and cautioning that the system reads as hearing-specialised with comparatively weaker deep video reasoning. That is a critic's read, not a measured replication, but the size of the delta is the reason to keep it in mind.

The other number worth holding onto is not a score at all. In what Alibaba calls Agentic Understanding mode, the model decides for itself what to look at and listen to rather than processing every frame, gathering evidence in coarse-to-fine rounds. On OmniVideoBench that mode raises accuracy from 63.4 to 67.8 while cutting tokens per query from 145,736 to 79,117 — a 45.7% reduction. Accuracy up and cost down simultaneously is the strongest claim in the release, and it is the one that most directly supports the agentic pitch.
The Gemini comparison is narrower than the coverage suggests
Alibaba's positioning is that audio-video capability "approaches" Gemini 3.8 Flash while overall audio performance surpasses it. Stripped of framing, the vendor-reported comparisons look like this. WildClawBench-MM: 71.0 against 58.9. SpotSoundBench: 67.2 against 39.7. MMAU: 81.8 against 76.9. DailyOmni: 85.1 against 84.0. Those are real gaps on audio and audio-centric tool use. They are also selective. On AgenticVBench, the pure video-reasoning board, the ordering reverses — Gemini at 45.0 against Qwen's 36.8 — and Alibaba's own account concedes Gemini still leads several video-understanding tests. A reasonable summary is that Qwen has taken the audio half of omni and is still behind on the video half, which is a different claim from the one the headlines ran with.
"Qwen's first omni-modal model" needs a qualifier
The framing that Qwen3.8-Omni-Flash is Qwen's first omni-modal model circulated widely today and is worth being precise about, because the family has an earlier entry with a fair claim to the title. Qwen2.5-Omni, a 7B open-weight model released in March 2025 under a Thinker-Talker architecture, also took text, image, audio and video as input — and it generated speech as well as text. What is genuinely first here is native omni-modality: one model built on the Qwen3.8-Flash foundation rather than a text LLM with perception encoders attached, plus an agent-first design brief. Both statements are true; only one of them is the one that got repeated. If you are evaluating the model on the assumption that no Qwen omni model existed before this week, you will mis-price the upgrade path from Qwen2.5-Omni, which is a comparison we have written up separately.
The tooling is where the agentic claim actually lives
Two things shipped alongside the model, and they matter more than the model card suggests, because they are what turns perception into delivery. Qwen-MM-Plugins is a multimodal plugin suite for agent harnesses that gives an external agent image, audio, video and document understanding plus long-video memory and video editing; it advertises support for Codex, Claude Code, Qwen Code, Gemini CLI and several others, and ships named tools including Video2Note, which turns hours of video into illustrated PDF notes. Qwen-Live Harness is an open-sourced runtime for continuous real-time omni-modal interaction, acting as a scheduling layer where a user converses live and delegates heavier work to background agents. One caveat from the launch-day coverage: the Qwen-Live Harness GitHub page was returning a 404 when Gigazine checked it, so treat the tooling as announced rather than verified until the repositories resolve.
The sentence the launch buries: it does not speak
For a model whose predecessor generated speech, the fact that Qwen3.8-Omni-Flash is multimodal-in and text-out is the most consequential line in the spec, and it appears in the materials largely as a footnote. It means the model cannot be dropped into a speech-to-speech product on its own. Any dubbing, voice agent or real-time conversation built on it needs an external text-to-speech stage and, for video, an external renderer — which is precisely why the plugin suite and the live harness exist. The practical consequence is a pipeline with more moving parts than "one omni model," and latency that has to be budgeted across all of them.
There is a neat demonstration of the gap, and of what the model is good for, buried in the release. In a "model optimising model" experiment, Qwen3.8-Omni-Flash autonomously improved the Sichuan-dialect recognition of Qwen2.5-Omni-3B, cutting character error rate from 25.79% to 15.30% inside twelve hours and building 3,413 training samples across four rounds. A model that can diagnose and repair a smaller sibling's dialect handling is doing something a transcription API cannot. That, rather than the benchmark table, is the clearest argument for what "agentic" is supposed to mean here.

What would change our read
The honest state of play is that this is a well-specified launch with a compelling price story, no independent evaluation, and at least one structural limitation that the announcement soft-pedals. Three things would settle it. An Artificial Analysis or LMArena entry would turn vendor numbers into comparable ones, and none exists yet. A third-party test of the 45.7% token reduction — the claim that accuracy rises while cost falls — would confirm the most useful and least glamorous result in the release. And independent measurement of AliMeeting diarisation would show whether the 88-point drop belongs to the model or to the pipeline around it.
Until then the sensible disposition is the one that applies to any model on its first day: worth testing, not worth betting a production path on. If you are already routing through OrcaRouter, the practical move is to keep the workload on the models you have measured, and treat Qwen3.8-Omni-Flash as a candidate to audition against them on your own audio and video rather than on either vendor's benchmark table. If your use case is long-form meeting or media analysis in Chinese and English, the token economics here are strong enough to justify the test now.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
