A generated hero title card for Qwen3.8-Omni-Flash reading 'One door, no weights', with the subtitle 'The model has a single host, no open weights, and its first non-vendor scores are thin' and four chips: 'Runs on: Alibaba only', 'Weights: not released', 'Independent scores: none', 'Context: 1M tokens'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Qwen3.8-Omni-Flash Five Days In: One Door, No Weights, and the First Scores That Aren't Alibaba's

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Five days after the vendor opened Qwen3.8-Omni-Flash to API customers, the model still has exactly one home. It runs on the vendor's own Qianwen platform and on its Bailian cloud service; the third-party listings that have appeared since launch are proxies in front of those same endpoints, not a second place the model lives. No weights have been released. The launch-day figures — a 98% cut on the cost of an hour of audio, a 26% average gain over Qwen3.5-Omni-Plus, one million tokens of context, text-only output — remain the vendor's own and unreproduced. What has actually changed in five days is not the model but the ground around it: a thin layer of third-party scores has appeared, framework support landed on day zero, and the comparison everyone is reaching for is against Gemini 3.8 Flash rather than the Qwen3.8-Flash this model was built alongside. Each of those deserves a closer look than the launch coverage gave it.

The door is a good door, and it is the only one

Availability has widened without becoming plural. The model answers an OpenAI-shaped request unchanged, which means existing client code mostly works; the thing that breaks is the endpoint, because the international and mainland hosts are separate services with separate billing. Code pointed at the default will either fail or bill in the wrong currency, and the per-hour price story only holds on the international side. Open-source client libraries shipped day-zero support for the model, which is genuinely useful and also not a second provider: a library that speaks to Bailian is a convenience wrapper around the same single source of supply.

That is the structural risk worth naming. A single-source model has no failover path. If the endpoint rate-limits, degrades, or a region goes dark, there is nowhere to go, and no amount of client-library support changes that. It is also why we are stating our own position plainly rather than implying otherwise: Qwen3.8-Omni-Flash is not in our catalogue. We do not host it, and a reader who signs up expecting to route it would be misled. What is routable is the rest of the family — Qwen3.8-Flash and Qwen3.8-Max are both live on our API — and OrcaRouter passes provider list price through at 0% markup, so when Alibaba's omni pricing moves, or the day the omni model becomes available to route, the change reaches customers the same day instead of at the next contract renewal.

A generated single-model spec card for Qwen3.8-Omni-Flash with six rows reading 'Released: 18 Sep 2026', 'Context: 1M tokens', 'Media per call: up to 1 hour', 'Output: text only', 'Price: $0.15 in / $0.47 out per M', 'Independent score: none yet', and the footer 'Vendor-reported; no independent evaluation exists.' The OrcaRouter logo is composited in the bottom-right corner.

The first scores that aren't Alibaba's are thin, and they are not flattering

Nothing on Artificial Analysis exists for this model, and that absence is the honest headline five days in: the leaderboards everyone quotes for adjacent models do not yet have a row for the omni one. The gap has been filled by something else. A third-party evaluation platform has begun publishing runs against the model, and its summary is worth reading precisely because of how much it does not say. It reports email classification at 99.0% accuracy (86th percentile), ethics at 95.0% (23rd percentile), and reasoning at 56.0%, which it places in the 28th percentile and calls a notable weakness; speed lands in the 21st percentile, cost in the 58th, with an 89% success rate overall.

Treat those with real caution, and not only because the sample is small. The same page dates the model's release to 20 September, two days after Alibaba shipped it, gives no run dates for any of the results, and renders an empty benchmark table beneath a summary that describes numbers the table does not contain. That is the profile of an early, lightly supervised harness rather than a measured evaluation, and the honest reading is that these are the first non-vendor data points rather than the first reliable ones. They are a reason to test the model yourself, not a reason to conclude anything. The direction, though, is consistent with what the vendor's own materials imply by omission: this is a perception and long-media model, and the reasoning-heavy boards are not where it was aimed.

There is also a trap in the search results that will catch people for the next few weeks. Qwen 3.8, the text model, carries an Artificial Analysis Intelligence Index in the forties, and that number has been quoted in more than one summary as though it described the omni model. It does not. They are different models with different jobs, and the omni model has no index at all yet.

A headless screenshot of Alibaba's own Qwen site (qwen.ai) showing the Qwen3.8-Omni-Flash 'Conducting Deep Research with Audio and Video' demo page, with the model workflow panel, the timeline of a 40-minute video, and the English UI navigation.

The benchmark table is still one vendor comparing to itself

The launch figure — an average improvement of more than 26% across roughly thirty evaluations — is measured entirely against Qwen3.5-Omni-Plus. That tells you the family moved. It does not tell you where the model sits next to anything you might already be paying for, and the selective comparisons that circulated alongside it do not close the gap either. The vendor-reported picture against Gemini 3.8 Flash is a genuine win on the audio boards and a loss on the ones that measure agentic video work: WildClawBench-MM 71.0 against 58.9 and SpotSoundBench 67.2 against 39.7 on one side, AgenticVBench 36.8 against 45.0 and OmniGAIA 74.0 against 78.6 on the other. Alibaba's own summary language — audio-video performance "approaching" Gemini, overall audio performance exceeding it — is more careful than the headlines it produced, and the careful version is the accurate one.

• Context — 1M tokens, with roughly 991K maximum input (about 983K in thinking mode) and 131K maximum output.

• Modality — text, image, audio and video in; text out only. No speech generation in the base model.

• Duration — up to one hour of continuous audio or audio-video per call, which is what makes meeting and long-media work possible without chunking.

• Language coverage — recognition across 74 languages plus 39 Chinese dialects in the launch materials, with wider figures (113 languages and dialects) quoted elsewhere for audio input.

• Input detail — stereo and four-channel spatial audio accepted, with the Realtime variant described as the first omni-modal model able to locate a target by its sound.

• Price — $0.15 per million input tokens, $0.016 cached, $0.47 output on the international side, and a separate mainland price list that is not comparable.

The 98% is real, and it is not your bill

By Alibaba's own methodology — a two-minute sample multiplied by thirty, video sampled at 720p and one frame per second — an hour of audio input falls from $0.28 to under $0.01, and an hour of combined audio and video from $3.27 to $0.20. Those are large, specific, vendor-reported reductions, and they are reductions on one line item. The arithmetic that matters is what share of your workload that line item is. A pipeline that spends most of its tokens on output, on cached context, or on video frames sampled more densely than one per second will see a much smaller percentage move on the invoice than the headline promises, and the headline is about input audio, not about the total. Add the text-only output and the gap widens again: anything that has to speak back needs a second model, and that model is billed separately.

This is the part of the picture where routing earns its keep. The value of a pass-through router is not a discount — it is that the vendor's own price list is what you pay, and that a workload can be spread across several models so that a single-source model is a component rather than a dependency. An omni model that is the only thing you can call is a single point of failure at both the availability and the pricing layer.

A headless screenshot of the OrcaRouter models catalogue at www.orcarouter.ai/models, showing the model grid and the 'one API key, one bill' framing over the English-language site navigation.

What five days of production would actually tell you

Three things would settle the open questions. Independent measurement of the AliMeeting diarisation result — an error rate falling from 88.11 to 3.35 and cpWER from 89.61 to 17.18 — would show whether that belongs to the model or to the diarisation and front-end engineering wrapped around it, which is where at least one Chinese technical analysis attributes most of the gain. A reproduced test of the agentic mode's token reduction, where accuracy rose from 63.4 to 67.8 on OmniVideoBench while tokens per query fell from 145,736 to 79,117, would confirm the single most useful claim in the release: that the model can get better and cheaper at the same time. And an Artificial Analysis or arena row would turn every vendor number above into something comparable, which is the precondition for the model being chosen rather than tried.

Until then the disposition that makes sense is the one that applies to any five-day-old model with one host and no independent evaluation: worth measuring on your own audio and video, not worth making the only path through your pipeline. If your workload is long-form meeting or media analysis in Chinese or English and your costs are dominated by input hours rather than output tokens, the economics here are strong enough to justify a test this week. If your workload needs speech back, or needs a fallback when a provider degrades, the model as it stands today cannot be the whole answer — and no amount of day-zero framework support changes that.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily