A generated hero title card for 'MiniCPM-V 4.7 vs LFM2.5-VL 3B' with the subtitle '70 GB of weights against a 3B model you can run on a laptop', a left card reading 'MiniCPM-V 4.7: 35.2B sparse, no licence, no card', a right card reading 'LFM2.5-VL 3B: 3.1B dense, LFM Open License v1.0' and a pill reading 'One is measured, one is not', with the OrcaRouter logo in the bottom-right corner.
Guides & Insights

MiniCPM-V 4.7 vs LFM2.5-VL 3B: A 70 GB Mystery Against a 3B Model You Can Ship on a Laptop

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

MiniCPM-V 4.7 and LFM2.5-VL 3B are sold as the same category of thing — compact vision-language models you run yourself — and they could hardly be further apart in practice. LFM2.5-VL 3B is a finished product: a documented 3.1-billion-parameter checkpoint from Liquid AI with a licence, a model card, GGUF quants in the wild, and a declining number of good reasons not to just download it and try. MiniCPM-V 4.7 is a 35.2-billion-parameter sparse mixture-of-experts checkpoint that OpenBMB uploaded to Hugging Face on October 6, 2026 with no model card, no licence, and no benchmarks whatsoever. Comparing them is not comparing two models. It is comparing a model against an artifact that will probably become one.

The matchup, stated honestly up front

This page can give you a complete picture of LFM2.5-VL 3B, because Liquid AI published one. It can give you a complete picture of MiniCPM-V 4.7's shape and nothing at all about its behaviour, because OpenBMB published a config.json and stopped. Everything below about the MiniCPM side is read from that configuration file and the weight index; everything about the Liquid side is from the model card, and the card's own performance claims are vendor-reported and labelled as such. Where a number exists for one model and not the other, the honest thing is to leave the gap visible rather than paper over it.

What each one is

LFM2.5-VL-3B was released on August 11, 2026. It is a multimodal variant of the LFM2.5 family, built for on-device deployment, and it does not start from scratch: it takes the LFM2.5-2.6B language backbone and pairs it with a SigLIP2 NaFlex vision encoder. NaFlex is the part that matters for document work — it preserves native aspect ratio and resolution instead of forcing every image into a fixed square, which is why the card's headline improvements are grounding with natural-language queries and full-page OCR with layout annotation. It ships under the LFM Open License v1.0 and covers sixteen languages including English, Chinese, Japanese, Korean, Arabic, Hindi, French, German, Spanish, Portuguese, Italian, Polish, Russian, Thai, Vietnamese and Indonesian. Because it is 3B, GGUF builds appeared within a day of release and a DSpark variant followed in September, so the ecosystem around it is real: you can pull a quant and run it on a laptop today.

MiniCPM-V 4.7 is a 35,212,875,824-parameter BF16 checkpoint, 70.4 GB across sixteen shards, class MiniCPMV4_7ForConditionalGeneration, uploaded to openbmb/MiniCPM-V-4.7-35B-A3B on October 6, 2026.

A generated two-column comparison scoreboard titled 'MiniCPM-V 4.7 vs LFM2.5-VL 3B — the scoreboard'. Left column 'MiniCPM-V 4.7' lists Parameters 35.2B sparse, On disk 70.4 GB BF16, Licence none declared, Vision 16x downsample, Context 256K, Benchmarks none, Quants none. Right column 'LFM2.5-VL 3B' lists Parameters 3.12B dense, On disk 6 GB BF16, Licence LFM Open v1.0, Vision SigLIP2 NaFlex, Context on-device class, Benchmarks vendor card, Quants GGUF and MLX. The footer reads 'LFM figures vendor-reported; MiniCPM figures read from config.json.'

Its language backbone is tagged qwen3_5_moe_text — a Qwen3.5-derived sparse MoE with 256 experts and 8 selected per token — and its 40 layers run a fixed three-linear-to-one-full attention pattern, so 30 of the 40 layers use a Mamba-style linear path rather than conventional attention. Context is 256K. The vision tower is the MiniCPM-V lineage's own, at roughly the same shape MiniCPM-V 4.6 used. There is no README in the repository, which on Hugging Face means there is no licence.

The comparison that is actually possible

Six dimensions, both sides, and a note on each about how much weight the number carries.

• Total parameters — LFM2.5-VL 3B: 3.12B, all active per token. MiniCPM-V 4.7: 35.2B total, sparse, 8 of 256 experts active per token. The MiniCPM figure is not comparable to a dense parameter count and the actual active parameter budget is not published; treat "A3B" as a name, not a measurement.

• Effective size on disk — LFM2.5-VL 3B: a single ~6 GB safetensors file in BF16, or a fraction of that as GGUF. MiniCPM-V 4.7: 70.4 GB across sixteen shards, BF16 only.

• Licence — LFM2.5-VL 3B: LFM Open License v1.0, published and linked from the card. MiniCPM-V 4.7: none declared. Absent a license: tag the default position is all rights reserved, so this is a blocker, not a footnote.

• Vision approach — LFM2.5-VL 3B: SigLIP2 NaFlex, native-resolution with aspect-ratio preservation, card claims full-page OCR with layout annotation. MiniCPM-V 4.7: custom minicpmv4_7_vision tower with downsample_mode: "16x" and max_slice_nums: 9.

• Context — LFM2.5-VL 3B: the card's serving examples are modest compared with a long-context design, and the family targets on-device memory budgets. MiniCPM-V 4.7: max_position_embeddings: 262144, with tokenizer model_max_length agreeing at 256K.

• Evidence of quality — LFM2.5-VL 3B: a published card with vendor-reported improvements over LFM2-VL-3B, third-party quants, and a Hugging Face history of about 25,900 downloads. MiniCPM-V 4.7: zero downloads, three likes, no card, no benchmarks, no independent evaluation. Nothing to cite.

• Tooling — LFM2.5-VL 3B: GGUF and MLX builds from third parties, plus a DSpark variant. MiniCPM-V 4.7: none. No quantisation of any kind exists yet.

A screenshot of the Hugging Face model page for openbmb/MiniCPM-V-4.7-35B-A3B (captured 7 October 2026) showing the model header with three likes, the tags safetensors, minicpmv4_7 and region:us, and the file listing with no README or licence field.

The two-thirds of this page that is missing

Every comparison article in this space ships with a benchmark paragraph. This one cannot. There is no MMMU, no OCRBench, no DocVQA, no grounding number, no latency measurement and no VRAM figure for MiniCPM-V 4.7 — not because they were inconvenient to look up, but because nothing in the repository contains them and no third party has produced them. A model uploaded with no card is, from an evaluation standpoint, unmeasured. Anyone who tells you otherwise is inferring from the family name or from the parameter count, both of which are bad predictors of multimodal quality.

The same is true of a subtler question: whether the sparse MoE actually helps. A 35B-A3B design trades total memory for per-token compute, and the instruction-following quality of a sparse model depends entirely on how well the router was trained. The config shows a routing auxiliary loss coefficient of 0.001 and a shared expert alongside the 256 routed ones, which is a conventional setup — but conventional setup is not evidence of good routing.

On the Liquid side, the missing piece is different: the card's quality claims are the vendor's own.

A screenshot of the Hugging Face model page for LiquidAI/LFM2.5-VL-3B (captured 7 October 2026) showing the model header, the image-text-to-text pipeline tag, the lfm2_vl architecture tag, the multilingual language tags and the opening of the model card explaining that LFM2.5-VL-3B builds on LFM2-VL-3B with a SigLIP2 NaFlex vision encoder.

The OCR and grounding improvements are described by the people who trained the model, and while third-party quants and download volume show the community found it useful enough to convert, that is adoption evidence rather than quality evidence. Nobody has published an independent head-to-head.

Which one you would actually reach for

The decision tree here is unusually clean, because the two models are not competing for the same job.

If you need a vision-language model running on a laptop, a phone, a Jetson, or a single modest GPU, LFM2.5-VL 3B is the only one of the two that answers the question. It is small enough to quantise, it has a licence that permits the work, and someone has already done the conversion work for you. For document layout, screenshot parsing and grounded object description on a size and language footprint that fits an edge budget, the card's stated strengths line up with the deployment target.

MiniCPM-V 4.7 is not a candidate for that job at 70 GB in BF16. It is a candidate for the opposite job: a long-context, high-throughput multimodal model on server hardware, in workloads where the 256K window and the linear-attention KV-cache savings are the point. Whether it does that job well is unknown, and the licence question has to be answered before it does that job at all in anything commercial.

There is a third option worth naming, which is that you may not need to choose. If the architecture is "a small local model handles the bulk of the traffic and hands the hard cases to something hosted," the local half of that is a decision you can make today and the hosted half is a routing decision. OrcaRouter covers a couple of hundred hosted models behind a single key at each provider's list price with nothing added on top, with failover when a provider degrades — but be clear about what that is not: neither LFM2.5-VL 3B nor MiniCPM-V 4.7 is a hosted model on that router. Both are weights you serve yourself. The router is what sits behind them.

Where this stands, and what to refresh

Liquid AI shipped a finished small model. OpenBMB shipped an unfinished large one. The comparison that readers want — accuracy against latency, OCR against OCR — cannot be written yet for the MiniCPM side, and the piece that would be dishonest is the one that fills the space with parameter-count arithmetic. Watch the repository's README. The day it appears it will carry the licence and the first real numbers, and on that day this becomes a genuine head-to-head instead of a specification sheet next to a product.

The hosted half of that pattern is a routing decision rather than a second contract. each provider's list price with no markup added by us Neither LFM2.5-VL 3B nor MiniCPM-V 4.7 is a hosted model on OrcaRouter - both are weights you serve yourself.