A generated hero card titled 'Qwen3.8-Omni-Flash vs InternLumina U2' under the eyebrow 'Rented senses vs owned weights', with the subtitle 'A closed long-media API against a 16B diffusion model you can download.' and four chips: 'Weights: closed vs Apache 2.0', 'Generates images: one side only', 'Context: 1M vs undisclosed', 'Routable here: neither'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Qwen3.8-Omni-Flash vs InternLumina U2: Rented Senses vs Owned Weights

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The useful way to compare Qwen3.8-Omni-Flash and InternLumina U2 is not by benchmark row, because they do not appear on the same ones. It is by asking who is allowed to run them. The vendor's Qwen3.8-Omni-Flash, open to API customers since 18 September 2026, is a closed model on the vendor's own platforms: one million tokens of context, up to an hour of continuous audio and video per call, text-only output, and no weights. Shanghai AI Laboratory's InternLumina U2 is a 16B-A1B mixture-of-experts diffusion model under Apache 2.0 whose Ascend checkpoints are downloadable from Hugging Face right now, with the NVIDIA checkpoints and the technical report still listed as coming. One is a service you rent by the token. The other is a file you can hold, with a caveat about which hardware it currently runs on. That difference decides more real deployments than any score in either vendor's table.

What InternLumina U2 actually is

The architecture is the interesting part and it is genuinely unusual. InternLumina U2 is a sparse MoE with 16 billion total parameters and 1 billion active, and it is a diffusion language model rather than an autoregressive one — it refines a whole sequence in parallel instead of emitting tokens left to right. Its visual representation is fully discrete: eight codebooks of visual tokens built on Apple's AToken, which is what lets one model both read an image and produce one without a separate decoder bolted on the end. Text question answering, text-to-image generation, image understanding, image editing, video understanding and 3D understanding are all the same model doing the same kind of work over the same vocabulary.

• Size and shape — InternLumina U2 is 16B total, 1B active, sparse MoE; Qwen3.8-Omni-Flash's parameter count is undisclosed.

• Generation — InternLumina U2 generates and edits images and handles 3D understanding; Qwen3.8-Omni-Flash generates no media at all and returns text.

• Decoding — InternLumina U2 is a diffusion LLM decoding in parallel; Qwen3.8-Omni-Flash is autoregressive.

• Context and duration — Qwen3.8-Omni-Flash carries a 1M-token window and up to one hour of continuous audio-video in a single call; InternLumina U2's published materials describe none of that, and long-form audio is not among its listed capabilities.

• Weights — InternLumina U2 is Apache 2.0 with Ascend checkpoints published and NVIDIA checkpoints pending; Qwen3.8-Omni-Flash has no weights and no announced plan for any.

• How you get it — InternLumina U2 is a download from Hugging Face with inference code on GitHub; Qwen3.8-Omni-Flash is an API call to Alibaba Cloud Model Studio or the Qianwen platform.

• Independent scores — none for either. InternLumina U2's own card labels its benchmark table "preliminary, partial results"; Qwen's are all against its own predecessor and unreproduced.

A generated two-column scoreboard titled 'Qwen3.8-Omni-Flash vs InternLumina U2 - the scoreboard'. Six shared dimension labels, left column Qwen3.8-Omni-Flash: 'Weights: not released', 'Size: undisclosed', 'Context: 1M tokens', 'Output: text only', 'Generates images: no', 'Access: vendor API only'; right column InternLumina U2: 'Weights: Apache 2.0 on Ascend', 'Size: 16B total, 1B active', 'Context: not published', 'Output: text and images', 'Generates images: yes', 'Access: download from Hugging Face'. The footer reads 'Qwen figures vendor-reported; InternLumina card says preliminary, partial results.' The OrcaRouter logo is composited in the bottom-right corner.

The weights question has moved since the preview coverage

If you read our earlier write-up of InternLumina U2 when the code first appeared, the state of play has changed in one direction and stayed put in another, and both halves matter. The Hugging Face repository is no longer a stub: the ascend/ folder now holds seven safetensors shards plus the configuration, tokenizer and modelling code, which means the model is genuinely downloadable and runnable by anyone with the right hardware. The nvidia/ folder is still marked coming soon, the technical report is still forthcoming, and the repository's own benchmark section still says "preliminary, partial results." So the accurate statement today is not "the weights are out" or "the weights are missing" — it is that one hardware stack's checkpoints are published and the other is not, which is a real constraint for anyone whose cluster is NVIDIA.

Two other things have moved quietly. The GitHub repository has continued to receive commits well past the initial September 1 drop, so the released code is being maintained rather than parked. And the download counter on the Hugging Face repository still reads zero, which says less about the model than about the audience: a 16B-A1B diffusion model with Ascend-first checkpoints is aimed at a lab, not at a product team looking for a hosted endpoint this quarter.

Where the two genuinely overlap, and where they do not

Image understanding is the intersection, and it is narrower than it looks. Both models can look at a picture and answer questions about it, which is the entire job for Qwen3.8-Omni-Flash and one of six jobs for InternLumina U2. Everything past that diverges hard. InternLumina U2 can then edit that image or generate a new one from a prompt, in the same model, using the same visual vocabulary it used to read it. Qwen3.8-Omni-Flash cannot produce a pixel, but it will accept an hour of audio and video in one call and tell you who spoke, when, and what was decided — which InternLumina U2's published materials do not claim to do at all.

The overlap that matters less than people assume is video. Both are described as handling video, but they are doing different things with it: one is understanding a long recording as a document, the other is handling video as a visual modality alongside 3D. A team that needs an hour of footage summarised is not served by the second, and a team that needs a frame generated is not served by the first.

A headless screenshot of the Hugging Face model page for InternLumina-U2 showing the Apache 2.0 licence badge, the tags for diffusion, multimodal, image-generation, image-editing and vision-language, the model card heading 'A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding, Image Generation and Editing', and the 'Technical Report (Coming Soon)' note.

Cost, control, and what you are actually buying

The pricing comparison is not close in kind. Qwen3.8-Omni-Flash bills per token — $0.15 per million input, $0.016 cached, $0.47 output on the international list, with a separate mainland price list that is not comparable — and its launch headline was a vendor-reported cut of more than 98% on the cost of an hour of audio input, computed by pricing a two-minute sample and multiplying by thirty with video sampled at 720p and one frame per second. InternLumina U2 bills nothing, because there is nothing to bill: the cost is your hardware and your time, and the model's own card does not publish a throughput figure that would let you turn that into a per-token number.

That is the trade in one line. The API gives you capability you cannot host and an hourly cost that scales with use; the open weights give you a fixed cost and full control over the data path, at the price of an inference stack you have to build and a checkpoint set that currently means Ascend hardware. For regulated workloads where media cannot leave the building, the second is not a preference — it is the only option on this page. For a team that needs long-audio analysis this month, the first is the only option on this page.

OrcaRouter does not host either model today, and we would rather say that plainly than imply a routing path that does not exist: Qwen3.8-Omni-Flash runs only on Alibaba's own platforms, and InternLumina U2 is a download rather than an endpoint. What we do run is the rest of the Qwen3.8 line — Qwen3.8-Flash and Qwen3.8-Max — through one API at provider list price with 0% markup, which is the practical way to hold a working baseline for the closed-model side of this comparison while the open-weight side is still getting its NVIDIA checkpoints in order.

A headless screenshot of the OrcaRouter model page for Qwen3.8 Flash at www.orcarouter.ai/models/qwen/qwen3.8-flash, showing the model header, the code sample, the input modalities and capabilities, the published pricing and the p50 TTFT figure.

Which one you should actually pick

Pick InternLumina U2 if you need one model that reads, generates and edits images and you are willing to own the inference stack — and if your accelerators are Ascend, or you can wait for the NVIDIA checkpoints. It is the only model of the two that can create an image, and the only one you can run without sending data to a vendor. Treat its benchmark numbers as unproven until the technical report lands, because the model card says exactly that.

Pick Qwen3.8-Omni-Flash if your problem is hours of audio and video rather than pixels out — meeting intelligence, long-recording analysis, audio-heavy agentic work in Chinese and English — and if you can accept a single-source dependency and text-only output. Its cost story is strong on input audio hours and much weaker on total invoice, its benchmarks are vendor-reported and unreproduced, and the first third-party runs to appear put it in the 28th percentile on reasoning, on a sample small enough that it should prompt a test rather than a decision.

If you need both, they are complements, not alternatives: an open-weight image model you host and a closed long-media model you call. Neither is routable through us today. The thing to watch on one side is the NVIDIA checkpoint and the technical report; on the other, the first independent evaluation, which does not exist yet and which is the single largest gap in everything written about Qwen3.8-Omni-Flash so far.