
PixelUMM vs North Micro Vision Instruct: A Noncommercial Research Model Against a 2.4B Reader You Can Ship
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 123 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 103 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 211 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Put North Micro Vision Instruct and PixelUMM side by side and the size difference is almost a distraction: 2.4 billion parameters for Cohere's native-resolution reader against 15.2 billion for NVIDIA's unified model. The interesting axis is the encoder. North Micro Vision Instruct is built the conventional way — a 400M vision encoder initialised from SigLIP 2 SO400M, feeding a 2B language model, wrapped in a 262,144-token vocabulary and shipped under Apache-2.0. PixelUMM is built on the argument that this exact arrangement is the bottleneck, and deletes the encoder: raw 16×16 pixel patches go into a Qwen3-8B backbone through single-layer linear projections, and the same weights generate video as well as answer questions about it. One is a specialised reading appliance you can download and deploy today. The other is a research thesis with a download link and a licence that keeps it out of products.
Both models' numbers below are their own labs'. Neither has been independently reproduced, and the two publish different benchmark suites, so the overlap is narrower than it looks.
The case for the boring architecture
North Micro Vision Instruct was published on Hugging Face on 10 August 2026, has had eight weeks to accumulate users, and by early October was showing roughly 180,000 downloads and 150 likes — an order of magnitude more attention than PixelUMM's days-old repository. Its pitch is specific and unglamorous: most vision models downscale an image to a fixed grid and lose the fine detail documents live on, so this one processes images at native resolution, preserving aspect ratio and small text.
• Parameters — 2.4B total: a 2B language model plus a 400M vision encoder custom-trained from SigLIP 2 SO400M.
• Context — a 128K-token language backbone, but a validated multimodal operating range of about 8K tokens. The card is explicit that longer multimodal contexts rely on extrapolation and are unbenchmarked.
• Inputs and outputs — interleaved text and images in, text out. Multilingual across eleven-plus languages. No generation of any kind.
• Reported scores — DocVQA 0.921, ChartQA 0.808, OCRBench 0.792, all Cohere-reported and unreproduced. MMMU 0.329, which the card does not hide: this is a reading specialist, not a reasoning generalist.
• Deployment — Transformers 5.16.0 and an accelerator, a single modern GPU at 2.4B in bfloat16, Flash Attention 2 optional rather than required.
Nothing about that design is novel. It is also the reason it can be downloaded, fine-tuned and put in front of real documents this week.

The case against it, made by the other side
PixelUMM's preprint opens by naming the cost of that architecture. A frozen web-pretrained vision transformer supplies semantic features for understanding; a separate VAE supplies reconstruction-oriented latents for generation; carrying both roughly doubles the visual context per conditioning image and forces vision-language pretraining pipelines to be rebuilt around a second stream. North Micro Vision Instruct avoids the doubling only by refusing the second job entirely — it does not generate, so it needs one interface, not two.
PixelUMM takes the argument to its conclusion. No encoder, no VAE, no tokenizer: 16×16 pixel patches for images, 4-frame spatiotemporal tubelets for video, both reaching a decoder-only Transformer through single-layer linear projections, with understanding by autoregressive text and generation by pixel-space flow matching. Its eight empirical sections are studies of design choices rather than leaderboard wins — patch size, patch artifacts, pixel-space versus VAE-space training dynamics, compute scaling — which is a paper asking whether the paradigm holds, not one claiming it already won.
• Visual path — North Micro Vision Instruct: 400M pretrained vision encoder at native resolution. PixelUMM: no encoder; raw patches through a linear projection.
• What it can do — North Micro Vision Instruct: read images and documents, answer in text, ground and caption. PixelUMM: read images and video, and generate images and 4-second video from text.
• Scale — 2.4B dense versus about 15.2B total on a Qwen3-8B backbone, rendered as "8B MoT" in the paper's own tables.
• Licence — Apache-2.0 on code and weights versus Apache-2.0 on code with the NVIDIA One-Way Noncommercial License on the checkpoint.

Reading the two scoreboards honestly
The two models publish different benchmarks, and where they overlap the scales differ.
• DocVQA — North Micro Vision Instruct 0.921 as an accuracy fraction. PixelUMM 90.42 as a percentage. Same benchmark name, different splits and prompt setups, and only one of the two claims native-resolution handling as the reason.
• ChartQA — North Micro Vision Instruct 0.808. PixelUMM 82.96. Again roughly the same band once you stop comparing the raw strings.
• OCRBench — North Micro Vision Instruct 0.792. PixelUMM 78.00.
• MMMU — North Micro Vision Instruct 0.329, PixelUMM 41.67. Both low by the standards of large vision-language models, and both labs report them without decoration.
• PixelUMM-only — MVBench 70.53, Video-MME 57.33, LongVideoBench 59.61, GenEval 0.83 with a prompt rewriter, VBench Part 1 quality 84.10.
• North Micro Vision Instruct-only — grounding, captioning and multi-image tasks; a documented absence of tool calling and explicitly no agentic behaviour.
The striking thing about the overlap is how close a 2.4B specialist gets to a 15.2B generalist on document and chart reading. That is not evidence that the encoder-free approach fails — it is evidence that document reading is a task where a compact native-resolution encoder is a very good fit, and PixelUMM's architecture is aimed at a different argument. PixelUMM's own paper concedes the point in a sentence worth quoting: because training data differ across models, its results "cannot establish which architecture is superior".
Where the decision actually lands
If you have documents, charts, forms or screenshots and you need answers this month, North Micro Vision Instruct is the practical choice and it is not close: Apache-2.0, 2.4B, a standard Transformers load path, no guardrail environment, no distributed-checkpoint index, no second Python stack. Fine-tune it on your own document distribution and you have a specialist. The caveat is scope — point it at a reasoning task and the MMMU number will find you.
PixelUMM's offer is breadth and a research question. It reads and generates, image and video, from one set of weights, and it is the more interesting read by a wide margin. But its checkpoint is under a noncommercial licence, its weights are 128 distributed-checkpoint shards behind a hidden metadata index totalling around 30 GB, it wants a CUDA 13 toolkit with FlashAttention compiled from source, and its text-to-video path runs Cosmos guardrails that need a gated repository and a separate environment. That is a lab bench, not a deployment.
The routing angle arrives later, if it arrives at all. Neither model is available through any hosted API today — OrcaRouter routes neither — and neither will be until their licensing permits it. When a model like North Micro Vision Instruct does show up behind a shared endpoint, the reason to prefer that over a direct integration is the ordinary one: one key across 200+ models, provider list prices passed through at 0% markup, failover that demotes a model which underperforms its card, and no second contract to negotiate when you want to A/B a replacement. That is a description of the workflow, not a claim about availability.
The one test that would separate them
Run both on the same document set at the same resolution and score the small-text cases, then flip one variable: strip the vision encoder out of the comparison by feeding PixelUMM its native-resolution inputs. If the encoder-free model holds up on dense text at a fraction of the parameter cost, that is the strongest possible evidence for the thesis — and it is a test anyone with a spare card can run, because both checkpoints download. Until somebody does, the pragmatic summary stands: a small Apache-2.0 reader is the one you deploy, and a large noncommercial generalist is the one you study.
