
PixelUMM vs Microsoft Mage-VL: One Deletes the Visual Tokenizer, the Other Rewrites It
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 123 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 103 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 211 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Both PixelUMM and Microsoft Mage-VL were built by teams convinced that the way vision gets turned into tokens is the wrong bottleneck — and they fixed it in opposite directions. Mage-VL, Microsoft's codec-native streaming model, keeps a tokenizer and makes it far more aggressive: it follows the structure of modern video codecs, keeps every anchor frame and only those predicted-frame patches where the codec spends bits, and reports cutting visual tokens by more than 75% while gaining up to 3.5× wall-clock speedup over uniform frame sampling. PixelUMM, NVIDIA's encoder-free unified model, removes the tokenizer entirely — every 16×16 patch of raw pixel reaches the backbone through a single linear projection, with no VAE and no vision transformer anywhere. One compresses harder. The other refuses to compress at all. And only one of them can generate anything.
That last difference is what makes this pairing more interesting than a spec fight. Mage-VL is a watcher — a streaming perception model with a proactive gate that decides when to speak. PixelUMM is a reader that also draws: the same weights answer questions about an image and generate new video from a text prompt. They are two incompatible answers to "what should a unified vision model be", and each lab's numbers are its own.
Where the compression happens
Mage-VL's premise is a modern Moravec's paradox: vision-language models are strong at hard offline reasoning yet slow and compute-hungry at simple real-time perception. Its fix is codec alignment. Instead of decoding a stream into uniformly sampled frames and pushing a dense grid through a frozen web-pretrained ViT, Mage-VL separates a stream into anchor (I) frames and predicted (P) frames, retains all anchor patches, and keeps only the predicted-frame patches carrying real motion or new detail. The encoder, Mage-ViT, is trained from scratch on a 16×16 patch grid with 3D rotary position encoding and is explicitly codec-agnostic — the same interface accepts H.264/AVC or HEVC motion vectors and residual energy, or the learned rate map of a neural codec, with no architecture change or retraining.
PixelUMM's premise is that the encoder itself is the problem, not its efficiency. Its paper argues that models like BAGEL carry two visual interfaces — a ViT for semantic features and a VAE for reconstruction latents — which roughly doubles visual context per conditioning image and forces vision-language pretraining pipelines to be rebuilt around a second stream. PixelUMM deletes both: images become 16×16 spatial patches, videos become 4-frame spatiotemporal tubelets, and raw pixels reach a decoder-only Transformer through single-layer linear projections. Understanding is autoregressive text; generation is pixel-space flow matching.
• Compression strategy — Mage-VL: temporal, codec-derived; keep anchors, sparsify predicted frames. PixelUMM: none; keep every patch of raw pixel.
• What is trained from scratch — Mage-VL: the entire visual stack, on about 100M unlabeled images and videos. PixelUMM: the pixel embedders and decoders, on top of a Qwen3-8B language backbone.
• Backbone — Mage-VL: Qwen3-4B-Instruct-2507, the only pretrained component, behind a two-layer MLP projector. PixelUMM: Qwen3-8B, about 15.2B parameters in total.
• Streaming behaviour — Mage-VL: a cognition gate scores each rolling window and stays silent until a response-worthy event completes, invoking the full model only then. PixelUMM: no streaming mode; requests are generation or understanding calls.

What each one does, and what it does not
Mage-VL is a single checkpoint that simultaneously provides image and video understanding and the proactive streaming gate — the same weights answer offline questions and drive event-gated commentary. It bundles the codec processor, the neural codec package and the gate. A second release, microsoft/Mage-ViT, is the standalone visual encoder from the from-scratch pretraining stage, offered as a drop-in front end for other multimodal training. What Mage-VL does not do is generate. It reads.
PixelUMM covers text-to-image, text-to-video at 96 frames and 24 fps, and image- and video-conditioned text. It ships four checkpoints — S8-F22-R05 as the default covering all four tasks, S8-F18-R01 tuned for better text-to-video at 480p and 720p but unable to do video understanding, and two intermediate stages. What it does not do is streaming, editing or 3D. If you need a model that watches a live feed and speaks when something happens, PixelUMM is the wrong shape entirely, and no benchmark column will tell you that.
The adoption gap is the first honest signal
The two releases are not equally mature, and the download counters say so more plainly than any launch post.
• Mage-VL — published on Hugging Face on 25 July 2026 under Apache-2.0, with a companion tech report and an arXiv identifier. By early October it was showing roughly 13,800 downloads and 414 likes, and a small ecosystem of community work had grown around it.
• PixelUMM — published on Hugging Face on 1 October 2026 under a noncommercial checkpoint licence. Its download count was zero and its like count was three when this was written. It is days old.
There is still no announcement from either lab for the respective release — Mage-VL shipped in July without one and PixelUMM shipped in October without one. That symmetry is real, but it would be a mistake to read it as the two being equivalent artifacts. Mage-VL has had ten weeks of community attention; PixelUMM has had days. A model nobody outside the lab has run is a different proposition from a model a few thousand people have downloaded, and only one of these two is in the second category.

The scoreboard, with provenance
Every figure is the developing lab's own. Mage-VL's come from Microsoft's card and tech report; PixelUMM's from its preprint. Neither has been independently reproduced.
• Visual tokens — Mage-VL: reported reduction of over 75% versus dense frame sampling. PixelUMM: no reduction; the argument is that raw patches remove a second encoding rather than shrink the first.
• Wall-clock speed — Mage-VL: up to 3.5× faster than uniform frame sampling at matched accuracy, Microsoft-reported. PixelUMM: no comparative speed claim in the paper.
• Encoder quality — Mage-VL: Mage-ViT reports 99.33% CIFAR-10 and 85.69% ImageNet at a 256-token budget, from about 100M unlabeled media. PixelUMM: no equivalent encoder benchmark, since there is no encoder to benchmark.
• Video understanding — Mage-VL: reported gains over Qwen3-VL-4B on every video and temporal-grounding benchmark it reports, including +22.5 on QVHighlight and +17.1 on ActivityNet. PixelUMM: MVBench 70.53, Video-MME 57.33 without subtitles, LongVideoBench 59.61, LVBench 40.41.
• Image understanding — Mage-VL: on par with Qwen3-VL-4B on static images, Microsoft-reported. PixelUMM: MMMU 41.67, AI2D 80.12, DocVQA 90.42, ChartQA 82.96.
• Generation — Mage-VL: none. PixelUMM: GenEval overall 0.83 with a prompt rewriter, DPG-Bench 85.74, VBench Part 1 quality 84.10.
• Streaming — Mage-VL: proactive event gating with reported top TimVal, F1, ROC-AUC and PR-AUC on SoccerNet streaming. PixelUMM: not supported.
• Licence — Mage-VL: Apache-2.0, code and weights. PixelUMM: Apache-2.0 code, NVIDIA One-Way Noncommercial License on the checkpoint.
The two columns do not overlap enough to rank. Mage-VL reports an efficient encoder beating a same-scale competitor; PixelUMM reports a 15B model landing in the same band as the specialist systems around it, and says outright in its own paper that differing training data mean the results "cannot establish which architecture is superior".
The practical split
Pick by job, not by score. If you are building real-time video perception — a monitor that watches a stream and comments when something happens, with token cost as the binding constraint — Mage-VL is the only one of the two that does it, it is Apache-2.0, and its 4B-class scale means a single node can serve it. If you need one model that reads images and video and also generates them, PixelUMM is the only one of the two that does that, but it is a research artifact on a noncommercial licence, its weights are 128 distributed-checkpoint shards behind a hidden index, and text-to-video runs Cosmos guardrails by default with a separate Python environment for the gate.
For a team that needs both capabilities, the argument for reaching them through one routing layer rather than two direct integrations is straightforward even though neither is hosted today: a single key, provider list prices passed through at 0% markup, and failover rules that let you put a newly opened Apache-2.0 model in front of a slice of traffic while the proven model holds the rest. OrcaRouter is built for that shape of problem — and to be explicit, we route neither PixelUMM nor Mage-VL right now, so this is a description of the workflow, not a listing.
What would settle it
Two events matter. The first is an independent rerun of Mage-ViT's encoder numbers or Mage-VL's video results by someone with no stake in them — ten weeks of downloads have not produced one yet. The second is any change to PixelUMM's checkpoint licence, which is the single fact that currently divides a research artifact from a deployable one. Until either lands, the useful thing you can say about this pair is that Microsoft bet on making the visual interface cheaper and NVIDIA bet on removing it, and neither bet has been graded by anyone outside the lab that placed it.
