
MiniCPM-V 4.7 vs North Micro Vision Instruct: Cohere Shipped a Documented 2.4B Model; OpenBMB Shipped a 70 GB Blank
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 150 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 98 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 248 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
MiniCPM-V 4.7 and North Micro Vision Instruct are both small-lab answers to the same brief — a compact vision-language model you can fine-tune and deploy yourself — but only one of them is finished. Cohere's North Micro Vision Instruct arrived on August 10, 2026 as a 2.4-billion-parameter native-resolution model under Apache-2.0, with a technical blog post explaining the architecture and a card listing what it is for. MiniCPM-V 4.7 arrived on October 6, 2026 as a 35.2-billion-parameter sparse mixture-of-experts checkpoint with sixteen shards, no README, no licence field, and no benchmark of any kind. The interesting comparison is not "which is smarter." It is that the two releases represent opposite postures toward the community that has to use them.
The two releases, and what shipped with each
North Micro Vision Instruct is a 2.4B open-weight vision-language model from CohereLabs, released under Apache-2.0. It is explicitly positioned as a compact foundation for prototyping, task-specific fine-tuning and specialised multimodal work. Its defining feature is native-resolution image processing that preserves aspect ratios and fine visual detail rather than resampling into a fixed grid, and the card claims capability across VQA, captioning, grounding, OCR, charts and documents. It supports multiple images per conversation and eleven languages — English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese and Arabic. Model weights are 2,484,847,856 parameters in BF16, and since release the MLX community has produced 4-bit, 5-bit, 6-bit, 8-bit, mxfp8, mxfp4, nvfp4 and bf16 conversions, which is what a healthy small-model release looks like after a couple of months. It has about 166,500 downloads and 150 likes.
MiniCPM-V 4.7 is openbmb/MiniCPM-V-4.7-35B-A3B: 35,212,875,824 parameters in BF16, 70.4 GB across sixteen safetensors shards, class MiniCPMV4_7ForConditionalGeneration, saved by Transformers 5.2.0 and uploaded on October 6, 2026.

Its text backbone is a Qwen3.5-derived sparse MoE — 256 experts, 8 selected per token — with 40 layers running a fixed three-linear-to-one-full attention pattern, 256K context, and an in-house 27-layer vision tower at 16× downsampling. There is no model card. The repository holds three likes and no downloads.
Side by side, on the six things that matter
• Parameters — North Micro Vision Instruct: 2.48B dense, every parameter active per token. MiniCPM-V 4.7: 35.2B total, 8-of-256 sparse. Not a like-for-like number.
• Licence — North Micro Vision Instruct: Apache-2.0, tagged and stated. MiniCPM-V 4.7: none declared, which defaults to all rights reserved.
• Fine-tuning surface — North Micro Vision Instruct: explicitly designed as a base for task-specific fine-tuning, with community quantisations already available. MiniCPM-V 4.7: a 70 GB BF16 checkpoint with no quants and no stated training recipe.
• Resolution handling — North Micro Vision Instruct: native resolution, aspect-ratio preserving. MiniCPM-V 4.7: downsample_mode: "16x" with max_slice_nums: 9, a slicer-based high-resolution path.
• Language coverage — North Micro Vision Instruct: eleven languages listed on the card. MiniCPM-V 4.7: not stated anywhere.
• Context — North Micro Vision Instruct: sized for its 2.4B deployment class. MiniCPM-V 4.7: max_position_embeddings: 262144, 256K.
• Evidence of quality — North Micro Vision Instruct: a published card, a technical blog post, community conversions and heavy adoption. MiniCPM-V 4.7: nothing to cite.

The asymmetry that makes this comparison lopsided
There is a particular kind of comparison article that would be easy to write here: pull the MiniCPM-V 4.6 numbers off OpenBMB's GitHub, note that they beat the equivalent class of small models, and imply the 4.7 inherits them. That would be dishonest, and the reason is worth stating precisely. MiniCPM-V 4.7 is not a bigger version of 4.6. It drops the 1.3B dense Qwen3.5-0.8B backbone for a 35B sparse Qwen-MoE, changes the attention pattern to 3:1 linear-to-full, and switches the vision compression default to 16x. Every one of those is a change to the part of the model that determines accuracy. Family lineage is not a benchmark.
The other direction of dishonesty is subtler: treating the absence of a card as evidence of a problem. It is not. OpenBMB has shipped nine MiniCPM-V generations since 2024, several of them genuinely excellent for their size, and a twelve-minute window between repository creation and last commit is what a staged-then-published artifact looks like, not what a broken one does. The correct reading of MiniCPM-V 4.7 is "unknown," and unknown is a distinct thing from "bad."
Cohere's side has its own honesty requirements.

The North Micro Vision card's quality claims are Cohere's own measurements, and the technical deep dive is written by the same team. That the community built sixteen MLX quantisations within days is adoption evidence, not accuracy evidence — people convert models that are easy to convert, and a clean 2.4B Apache-2.0 release is easy to convert. Neither the download count nor the quant spread should be read as a score.
What each is actually for
North Micro Vision Instruct is a fine-tuning target. That is the whole design brief, stated in the card's own words: a compact foundation for prototyping and specialised applications. If you have a narrow document-parsing, chart-extraction or multilingual captioning task and a few thousand labelled examples, a 2.4B Apache-2.0 model with native-resolution input and eight-bit-plus quantisations is a sensible place to start, and you can move it to an edge device or a single mid-range GPU when it works.
MiniCPM-V 4.7, if it turns out to be usable, is not that. At 70 GB unquantised it is a server model, and its distinguishing features — a 256K window and linear attention that keeps the KV cache flat — point at long-context multimodal workloads: hour-long video transcripts, large document sets, extended multi-turn analysis with images in the history. Those are real workloads, and the architecture is a reasonable bet on them. The bet just has not been scored.
There is a nuance about the two compression strategies worth holding onto. Native resolution and 16× downsampling are not simply better and worse. A native-resolution encoder spends tokens to preserve fine detail, which is what OCR and grounding need. A 16× downsample spends fewer tokens and risks losing exactly that detail. Which is right depends on whether your task is reading a dense form or describing a scene, and MiniCPM-V 4.6's switchable 4x/16x mode existed precisely because the answer changes by task. Whether 4.7 kept that switch back is one of the things the config does not say.
Where a router fits, and where it does not
Both of these models are weights you serve yourself. Neither North Micro Vision Instruct nor MiniCPM-V 4.7 is a hosted model on OrcaRouter, and nothing here should be read as a claim that it is. The routing question enters one step later, when the small self-hosted model is doing the routine work and something larger has to handle the cases it fails on. That handoff is what a single-key router is for — 200-plus models across providers, each at its list price with no markup added on our side, with automatic failover when one provider degrades — and it is worth having the interface settled before you need it, because the day your 4.7 evaluation succeeds is the day you want a second endpoint without a second contract.
The verdict, and what would change it
Cohere released a model. OpenBMB released a checkpoint. On every axis a reader can act on — licence, licence clarity, documented behaviour, fine-tuning readiness, quantisation, language coverage — North Micro Vision Instruct is the answer today, and it is not close. That is not a claim about capability, because no capability comparison is possible. MiniCPM-V 4.7 is roughly ten times the parameter count of the model it is being compared to and has published nothing to justify or refute that size. The honest posture is to treat it as pending: a real artifact from a lab with a real track record, sitting in a repository with one file missing that would change everything — the README.
