A generated title card in the OrcaRouter house style reading “PixelUMM vs InternLumina-U2”, with four statistic tiles: “41.67 / 57.33” (PixelUMM MMMU and Video-MME), “0.81 / 51.26” (InternLumina-U2 GenEval and Video-MME), “Apache-2.0” (InternLumina-U2 code and both checkpoints) and “1-way NC” (the PixelUMM checkpoint licence). A footer line reads “Vendor-reported figures; no independent rerun of either model.”. The OrcaRouter logo is composited into the padded strip at the bottom right.
Guides & Insights

PixelUMM vs InternLumina-U2: Two Encoder-Free Betas, and the Only Difference That Decides It

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Two models, both public, both unusually designed, and only one of them you can build a product on. PixelUMM is NVIDIA's encoder-free unified model — 15.2 billion parameters on a Qwen3-8B backbone, reading and writing raw pixels through 16×16 patches and 4-frame tubelets, with no VAE and no vision encoder anywhere in the stack. InternLumina-U2 is Shanghai AI Laboratory's 16B mixture-of-experts diffusion model that compresses every visual input into eight complementary discrete codes through a tokenizer called AToken. Both wagered that the visual interface is the bottleneck. The wager that matters more is the one in the license files: InternLumina-U2 ships Apache-2.0 weights, PixelUMM ships a checkpoint under a noncommercial licence. If you are choosing between them for anything commercial, that sentence is the whole comparison and the rest of this page is context for it.

Both sets of figures below come from the labs themselves. Neither model has an independent evaluation, an arena Elo, or a third-party reproduction. Neither is served by any hosted API. What differs is what you are permitted to do with the artifact once you download it, and how far along each release actually is.

Where each one stands right now

The release timelines have moved at different speeds and both of them quietly.

• PixelUMM — GitHub repository created 4 September 2026; arXiv preprint 2609.38597 dated 29 September 2026; Hugging Face model repository created 1 October 2026 at 21:40 UTC. No NVIDIA blog post, press release or launch thread surfaced for it.

• InternLumina-U2 — GitHub repository created 1 September 2026; Hugging Face stub registered the same day; Ascend-trained checkpoint weights added 10 September 2026; NVIDIA/CUDA checkpoint weights added 24 September 2026. Seven shards per stack, both downloadable today.

The InternLumina-U2 that existed in early September — code only, weights "coming soon", an empty model repository — is not the InternLumina-U2 that exists now. Anyone who read about it at the start of the month and concluded the weights were missing should re-check; both the Huawei Ascend and the NVIDIA/CUDA checkpoints are up, under Apache-2.0, with tokenizer, config, chat template and remote modelling code alongside them.

A headless-browser capture of the Hugging Face model card for the InternLumina-U2 repository, showing the model title, the Apache-2.0 licence tag, and the repository's download and like counters alongside the file listing for the checkpoint shards.

Two answers to the same question

Both models start from the same complaint: carrying a vision encoder for understanding and a VAE for generation means representing every image twice, roughly doubling visual context and forcing a training pipeline to maintain two interfaces.

PixelUMM's answer is to delete both. Raw pixels go straight in through single-layer linear projections — a 16×16 patch per image position, a 4-frame tubelet per video position — and the backbone is a decoder-only Transformer whose Mixture-of-Transformers design pairs shared attention with task-specific parameters. Text is predicted autoregressively; pixels are predicted by flow matching. The paper spends eight sections on design studies — patch size, patch artifacts, pixel-space versus VAE-space dynamics, compute scaling — which tells you the team's real question was whether the paradigm holds at all.

InternLumina-U2's answer is to keep a discrete interface but make each token carry far more. The AToken tokenizer describes every visual position with eight complementary codebooks covering texture, embedded text and geometry; the eight embeddings are concatenated per position and projected into one backbone token. Decoding runs the other way, with spatially parallel denoising and a multi-codebook autoregressive head predicting the eight codes in order. The backbone is a sparse diffusion LLM of the LLaDA-2.0 lineage, 16B total with 1B active.

• Visual interface — PixelUMM: raw pixel patches and tubelets, no tokenizer at all. InternLumina-U2: eight-codebook fully-discrete tokens from AToken, no continuous latent.

• Backbone — PixelUMM: Qwen3-8B (15.2B total), single dense decoder with understanding and generation experts. InternLumina-U2: 16B total / 1B active sparse MoE diffusion language model.

• Generation method — PixelUMM: pixel-space flow matching plus autoregressive text. InternLumina-U2: masked diffusion denoising over discrete codes.

• Tasks claimed — PixelUMM: text-to-image, text-to-video, image-to-text, video-to-text. InternLumina-U2: those plus image editing and 3D understanding.

• Weight format — PixelUMM: 128 .distcp shards (~30 GB) behind a hidden .metadata index. InternLumina-U2: seven .safetensors shards per hardware stack, standard Hugging Face layout.

A generated scoreboard card titled “PixelUMM vs InternLumina-U2 — the scoreboard”, with the two model names as column headings and six dimension rows spanning both columns: visual interface, backbone, generation method, video understanding, weight format and licence. PixelUMM's column shows raw pixel patches and tubelets, a Qwen3-8B backbone at 15.2B, pixel-space flow matching, Video-MME 57.33 with MVBench 70.53, 128 distributed-checkpoint shards, and a noncommercial checkpoint licence; InternLumina-U2's column shows eight-codebook discrete AToken tokens, a 16B-A1B sparse MoE diffusion backbone, masked diffusion denoising, Video-MME 51.26 with MVBench 59.74, seven safetensors shards per hardware stack, and Apache-2.0. A footer notes that the figures are vendor-reported with no independent rerun.

The scoreboard, with its provenance attached

Read every row as a lab's own claim. PixelUMM's come from its preprint; InternLumina-U2's are the lab's own description of them as "preliminary, partial", with the full comparison tables deferred to a technical report that has not appeared.

• MMMU (image reasoning) — PixelUMM 41.67 (authors' own). InternLumina-U2: not reported.

• ChartQA — PixelUMM 82.96. InternLumina-U2 86.52.

• DocVQA — PixelUMM 90.42. InternLumina-U2: not reported.

• Video-MME (no subtitles) — PixelUMM 57.33. InternLumina-U2 51.26.

• MVBench — PixelUMM 70.53. InternLumina-U2 59.74.

• GenEval overall — PixelUMM 0.83 with an LLM prompt rewriter, 0.77 without. InternLumina-U2 0.81.

• DPG-Bench overall — PixelUMM 85.74. InternLumina-U2 87.10.

• Video generation — PixelUMM VBench Part 1 quality 84.10 / semantic 79.80, Part 2 total 83.24. InternLumina-U2: not reported.

• 3D understanding — PixelUMM: no such task. InternLumina-U2 reports 3D MM-Vet 41.8.

The comparison is uneven because the two labs publish different tables, not because one is winning. PixelUMM has a full video-generation section and no editing or 3D; InternLumina-U2 claims six task families and reports a partial table across them. Cross-lab numbers on the same benchmark name are not the same measurement — splits, prompts and sampling differ — so treat the ChartQA and GenEval rows as roughly level and stop there.

Licensing is where this stops being a tie

This is the part no benchmark table shows. The PixelUMM repository is Apache-2.0 and the card says so, prominently. The checkpoint is a separate artifact under the NVIDIA One-Way Noncommercial License, limited to non-commercial research or evaluation, and one source file additionally retains a CC BY-NC 4.0 notice inherited from DiT. Research use: fine. Internal product feature: a conversation with legal, and possibly a short one.

InternLumina-U2 is Apache-2.0 end to end, code and both checkpoints, with no separate noncommercial carve-out. For a team that needs to ship, that is not a small edge — it is the entire decision. Both models are research-grade in maturity. Only one of them is unrestricted in use.

Which one to actually try, and how

If your work is video understanding or you want a model that generates video and images from one set of weights, PixelUMM is the more complete artifact and its paper is the more useful read — but it is a research project on a noncommercial licence, so it belongs on an evaluation bench, not in a pipeline. If your work is chart, document and image-generation breadth, or you need editing and 3D, InternLumina-U2 covers more ground and carries no licence ceiling; its cost is that the 16B-A1B diffusion backbone and the AToken tokenizer mean real serving engineering, and its published numbers are explicitly partial.

Neither is on any hosted API as of this writing, and that includes OrcaRouter — we route neither model. When that changes, the argument for reaching them through a routing layer rather than by direct integration is the same one that applies to any newly opened model: one key across 200+ models, provider list prices passed through at 0% markup, and automatic failover so a model that turns out to be worse than its vendor table suggested can be demoted by changing a route rather than rewriting a deployment. Until then the honest advice is the boring one — download whichever licence permits your use case, run your own evaluation on your own data, and do not plan around either lab's numbers until someone outside the lab reruns them.

What would change the answer

Three things, in order of impact. A commercial licence on PixelUMM's checkpoint would make the comparison genuinely close and would move it from "read the paper" to "evaluate the weights". A technical report from Shanghai AI Laboratory carrying the full comparison tables would turn InternLumina-U2's partial numbers into something checkable, and would also reveal how much of the six-task breadth is at parity versus at token depth. And the first independent reproduction of either model — a GenEval or MVBench rerun by a team with no stake in the result — is the moment either of these stops being a vendor table. Until then, one is an unusual architecture you can only study, and the other is an unusual architecture you can ship.