
MiniCPM-V 4.7 vs Microsoft Mage-VL: Two Very Different Bets on What a Vision Model Should Cost Per Frame
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 150 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 98 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 248 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
MiniCPM-V 4.7 and Microsoft Mage-VL are both vision-language models that claim to save you tokens, and they get there by opposite routes. Mage-VL, released by Microsoft on July 25, 2026, attacks the video side: it borrows the structure of a video codec, keeps every anchor-frame patch and only the predicted-frame patches that carry real motion, and claims a visual-token reduction of more than 75% with up to a 3.5× wall-clock speedup over uniform frame sampling. MiniCPM-V 4.7, uploaded by OpenBMB on October 6, 2026, attacks the sequence side: a 35.2-billion-parameter sparse MoE with 30 of its 40 layers on a linear-attention path, which makes the KV cache grow slowly across a 256K context. One model compresses what it looks at. The other compresses what it remembers. And only one of them comes with anything you can evaluate.
What each one is
Microsoft Mage-VL is a codec-native, proactive-streaming multimodal foundation model at a 4B scale, released under Apache-2.0 with a technical report and a substantial evaluation section. It was built deliberately against what its authors call a Moravec's paradox for VLMs — strong at offline reasoning, slow at real-time perception. The visual encoder, Mage-ViT, is trained entirely from scratch on roughly 100 million unlabeled images and videos rather than initialised from a web-pretrained ViT, and the only pretrained component in the stack is the Qwen3-4B-Instruct-2507 language backbone. The full checkpoint is a single unified model that does image understanding, offline video reasoning, and event-gated streaming commentary without separate variants, and a companion microsoft/Mage-ViT release provides the encoder on its own. It has accumulated about 10,600 downloads and 420 likes since release.
MiniCPM-V 4.7 is openbmb/MiniCPM-V-4.7-35B-A3B, a 35,212,875,824-parameter BF16 checkpoint in sixteen shards totalling 70.4 GB, written by Transformers 5.2.0 and uploaded on October 6, 2026 with no model card.

Its text config is tagged qwen3_5_moe_text with 256 experts and 8 active per token; its layer_types array runs three linear-attention layers to every one full-attention layer across 40 layers; its vision tower is the in-house minicpmv4_7_vision at 27 layers with 16× downsampling and up to nine image slices. Context is 256K. The repository has zero downloads, three likes, no benchmarks, and no licence.
The six rows a reader can check
Same six dimensions, both sides. Where a number does not exist, the gap is left visible.
• Parameters — Mage-VL: 4.74B, dense, all active per token. MiniCPM-V 4.7: 35.2B total, sparse, 8 of 256 experts per token. The MiniCPM number is not commensurable with a dense one.
• Licence — Mage-VL: Apache-2.0, stated in the card and the repo tags. MiniCPM-V 4.7: nothing declared.
• Where the token savings come from — Mage-VL: visual-token sparsity at the encoder, codec-aligned, claimed over 75% reduction. MiniCPM-V 4.7: linear attention at the decoder, which shrinks KV-cache growth over long sequences but does not reduce the tokens you feed it.
• Context — Mage-VL: trained through a long-context stage on 350K videos as rolling codec windows of up to 384 or 768 frames. MiniCPM-V 4.7: max_position_embeddings: 262144.
• Vision encoder provenance — Mage-VL: from scratch, trained on ~100M unlabeled frames; the card reports 85.69% ImageNet with 256 tokens and monotonic improvement with token budget. MiniCPM-V 4.7: lineage tower reused from MiniCPM-V 4.6's design at the same 1152-hidden, 27-layer shape, with a 16x downsample default.
• Evidence — Mage-VL: a full card with DocVQA 95.14, InfoVQA 80.33, OCRBench 81.80, ChartQAPro 32.57, MMStar 67.32, CV-Bench-3D 94.75 and a +53.1 CrossPoint gap over Qwen3-VL-4B, all vendor-reported against a matched backbone. MiniCPM-V 4.7: none.
• Streaming behaviour — Mage-VL: a cognition gate that scores each rolling window and stays silent until an event completes, trained on ~3.3M streaming samples. MiniCPM-V 4.7: not described anywhere.

The part everyone gets wrong about Mage-VL's numbers
The Mage-VL card's benchmark tables are strong, and the strongest ones are also the most easily misread. The comparison rows pit Mage-VL-4B against Qwen3-VL-4B, Phi-4-Multimodal-Instruct and Phi-4-Reasoning-Vision, and the headline gains — +22.5 on QVHighlight, +24.5 on VideoEval-Pro, +53.1 on CrossPoint — are real in the sense that they appear in the vendor's tables. They are also the vendor's own measurements, produced by the team that trained the model, on the same harness. No independent party has reproduced them. That is normal for a four-month-old model and it is not a mark against it, but it does mean the correct verb is "reports," not "scores."
Two things are worth crediting beyond the tables. First, the matched-backbone design — holding the Qwen3-4B decoder fixed and swapping only the ViT — is a cleaner piece of evidence than a leaderboard position, because it isolates the visual stack rather than the whole system. Second, the training corpus for Mage-ViT is about 100M unlabeled frames against the billions of image-text pairs that web-pretrained encoders use, so matching SigLIP2-at-10B-class behaviour on cluster discrimination is a specific, checkable claim about data efficiency rather than a general boast.
MiniCPM-V 4.7 has none of this. Not a weaker version — none.

There is no table, vendor or otherwise, and the config file that describes the architecture explicitly excludes itself from saying anything about accuracy.
Architecture: two answers to the same question
Both models are trying to answer "how do you make multimodal inference cheap," and the contrast is instructive because the two answers compose rather than compete.
Mage-VL reduces the input. Its 16×16 patch grid is shared between anchor and predicted frames, and the predicted frames only contribute patches where the codec is spending bits — which is where motion and new detail are. The result is variable-length token streams per frame, which is why the stack needs a projector that can hand a variable sequence to a causal decoder, and why the same interface can accept H.264/HEVC motion vectors or a neural codec's learned rate map without retraining. Then a 3D rotary encoding keeps spatio-temporal positions coherent across the sparsity. The token budget is the quantity being managed.
MiniCPM-V 4.7 reduces the state. Its linear-attention layers keep a fixed-size recurrent state instead of a growing key-value cache, so the memory cost of a long conversation stops scaling linearly with token count. Its sequence budget is the quantity being managed, and the 256K context is the payoff. Notably, MiniCPM-V 4.7's config advertises the 16x visual downsample — it compresses the input too — but is silent on whether it retains the switchable 4x mode that MiniCPM-V 4.6 exposed.
Put simply: Mage-VL's trick pays off on video where most frames are nearly identical to their neighbours. MiniCPM-V 4.7's trick pays off on long documents and long multi-turn sessions where the tokens accumulate. A pipeline that ingests a live stream and then holds a long conversation about it would benefit from both, and neither model yet does the other's job.
Serving them into a real workflow
Mage-VL is practical to try today: one checkpoint, a documented architecture, Apache-2.0, and a card that tells you what the repository bundles. It has been out since July and the tooling around it has had time to settle.
MiniCPM-V 4.7 is not practical to try today, for the boring reasons. 70 GB in BF16 with no quantisation means accelerator memory you probably do not have idle, and a missing licence means the question of whether you may use it at all is open. The custom code paths require trust_remote_code, and whether the MoE-plus-linear-attention combination has kernels in your serving stack is untested.
What is true for both is that a self-hosted vision model is one component of a system that also has to call frontier models. That is the case for putting the hosted half behind a single router rather than a second contract and a second SDK: OrcaRouter reaches 200-plus models on one key, passes each provider's list price through with no markup added by us, and fails over automatically when a provider degrades. Neither Mage-VL nor MiniCPM-V 4.7 is a hosted model on that service — both are weights you serve yourself — so treat this as the plumbing on the far side of the handoff, not as an availability channel for either model.
Verdict
Mage-VL is a finished, licensed, benchmarked model with a specific and well-argued design philosophy, and the case for it rests on video streaming and spatial reasoning where its codec-native encoder does something structurally different from everyone else. MiniCPM-V 4.7 is a larger, unlicensed, unmeasured checkpoint whose design philosophy is legible but whose behaviour is a blank. On everything a reader can act on today, Mage-VL wins by default — not because it is better, but because it is the only one of the two that can be judged at all. Revisit this comparison when the MiniCPM-V 4.7 README appears; if that day brings benchmarks and a licence, the interesting question will be whether a 256K-context linear-attention MoE beats a codec-native sparsity encoder on long video, and that is a genuinely open fight.
