A generated hero title card for 'MiniCPM-V 4.7' with the subtitle 'A 35B-A3B vision model uploaded with no model card, no licence and no benchmarks', a card reading 'Uploaded 6 Oct 2026', a checklist card reading 'Missing: model card, licence, benchmarks, quants', a pill reading '35.2B total parameters' and a pill reading '256K context', with the OrcaRouter logo in the bottom-right corner.
Engineering & Research

MiniCPM-V 4.7: OpenBMB Uploaded a 35B-A3B Vision-Language Model With No Model Card, No License, and No Benchmarks

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

MiniCPM-V 4.7 appeared on Hugging Face on the evening of October 6, 2026 as the repository openbmb/MiniCPM-V-4.7-35B-A3B — and it is one of the strangest things the MiniCPM series has ever shipped, because almost nothing shipped with it. There is no model card. There is no README. There is no declared licence. There are no benchmark tables, no launch post, no technical report, and no entry in OpenBMB's own GitHub documentation, which as of this week still calls MiniCPM-V 4.6 "the latest and most efficient model in the MiniCPM-V series." What does exist is 16 shards of bfloat16 weights, 70.4 GB of them, describing a 35.2-billion-parameter mixture-of-experts vision-language model with a 256K context window and an unusually aggressive hybrid-attention design. That is enough to say what the model is. It is not yet enough to say what it is good at, and this piece keeps those two things strictly apart.

What actually landed, byte for byte

The repository was created at 18:36 UTC on October 6, 2026 and had its last commit twelve minutes later at 18:48 UTC. In between, the uploader pushed the weight files and the configuration that a Transformers loader needs, and stopped. The complete file list is eighteen entries long:

• Weights — sixteen safetensors shards named model-00001-of-00016 through model-00016-of-00016, with the index file model.safetensors.index.json reporting a total tensor size of 70,425,751,648 bytes.

• Parameter count — the Hugging Face API reads 35,212,875,824 parameters, all in BF16. Note that this is the total count, which for a sparse MoE is not the number of parameters active on any given token.

• Config — config.json (2,984 bytes), generation_config.json (186 bytes), preprocessor_config.json, processor_config.json.

• Tokenizer — tokenizer.json (20 MB), tokenizer_config.json, and chat_template.jinja (7,250 bytes).

• Missing — README.md. A direct fetch of the raw file returns HTTP 404. On Hugging Face a missing README means no model card, which means no licence field, which means the repo carries no license: tag at all.

The repository has three likes, zero downloads, and an empty discussion tab. A community upload of a 70 GB checkpoint usually attracts comments within hours — questions about quants, about serving, about what the licence is. This one has not, which is consistent with it being spotted by a handful of people watching the openbmb organisation rather than announced to anyone.

A generated single-column scoreboard titled 'MiniCPM-V 4.7 — the scoreboard' with six rows: Total parameters 35.2B, 8 of 256 experts; Context 256K tokens; Text backbone Qwen3.5 sparse MoE; Attention 30 of 40 layers linear; Vision 16x downsample, 9 slices; Benchmarks none published, with the footer 'Figures read from config.json and the weight index; no vendor benchmarks exist.'

The architecture, read straight off config.json

Because there is no card to paraphrase, the configuration file is the primary source, and it is unusually informative. The class is MiniCPMV4_7ForConditionalGeneration, the model type is minicpmv4_7, and it was saved by Transformers 5.2.0 — all of which say this is a first-party OpenBMB checkpoint in the mainline MiniCPM-V lineage, not a community fine-tune wearing the name.

• Language backbone — model_type: qwen3_5_moe_text. A Qwen3.5-derived sparse MoE, the first time the MiniCPM-V line has used a Qwen-MoE text stack rather than the dense small Qwen variants that powered MiniCPM-V 4.5 and 4.6.

• Sparsity — 256 experts, 8 selected per token, with an expert intermediate size of 512 and a shared expert of the same size. A 35B total checkpoint on a 8-of-256 routing pattern activates a small fraction of that per forward pass, which is the entire point of the A3B name: the weights are large, the compute is not.

• Depth and width — 40 hidden layers, hidden size 2048, 16 attention heads with 2 key/value heads and a head dimension of 256.

• Hybrid attention — the layer_types array is 40 entries long and runs three linear_attention layers to every one full_attention layer on a fixed interval of 4. Ten of the forty layers do conventional attention; the other thirty use a Mamba-style linear path with a convolution kernel of 4. This is the single most consequential design choice in the file, and it is the same direction the wider field moved in through 2026.

• Positional encoding — RoPE with a theta of 10,000,000 and partial_rotary_factor: 0.25, meaning only a quarter of each head's dimensions are rotated. Multimodal RoPE is enabled with mrope_interleaved: true, a section split of [11, 11, 10], and mrope_mode: canvas.

• Context — max_position_embeddings: 262144, and the tokenizer's model_max_length agrees. 256K tokens.

• Multi-token prediction — mtp_num_hidden_layers: 1, a single speculative-decoding head, the same trick MiniCPM-V 4.6 carried.

• Vision tower — minicpmv4_7_vision, hidden size 1152, 27 layers, GELU-tanh activations, patch size 14 with an image_size of 980, and an insert_layer_id of 6, which is where the visual embeddings get spliced into the language stack. The tower's shape is close to the one in MiniCPM-V 4.6, so the vision side is an evolution rather than a rebuild.

• Vision compression — downsample_mode: "16x" and max_slice_nums: 9 in the image processor. MiniCPM-V 4.6 introduced a switchable 4x/16x token-compression scheme; the 4.7 config advertises the 16x setting as the default, with the slicer allowing up to nine sub-images for high-resolution inputs.

• Vocabulary — 248,144 tokens, with <|image_pad|> at id 248,056 and <|video_pad|> at 248,057. Video is a first-class input, exactly as it has been since MiniCPM-V 4.5.

• A curious leftover — the tokenizer config still declares <|audio_start|>, <|audio_end|> and <|audio_pad|>. This almost certainly means the vocabulary is shared with the omni MiniCPM-o branch rather than that MiniCPM-V 4.7 speaks audio. Reading it as an audio feature would be a mistake the config alone cannot rule out.

A screenshot of the Hugging Face model page for openbmb/MiniCPM-V-4.7-35B-A3B (captured 7 October 2026) showing the model header with three likes, the tag chips safetensors, minicpmv4_7 and region:us, and the repository file listing beginning with .gitattributes, chat_template.jinja, config.json and generation_config.json, with no README or model card rendered.

What the family tells us that this repository does not

The 4.6 generation is the reference point, and the contrast is the story. MiniCPM-V 4.6 was released on May 11, 2026 as a 1.3-billion-parameter model built on a SigLIP2-400M vision encoder and a Qwen3.5-0.8B language model, under Apache-2.0, with an explicit pitch: it scores 13 on the Artificial Analysis Intelligence Index — a vendor-reported figure — while using dramatically fewer tokens than comparable small models, and it runs on iOS, Android and HarmonyOS with the edge adaptation code open-sourced. Its whole identity was "small, efficient, on-device."

MiniCPM-V 4.7 is 1.3B multiplied by roughly 27. The 35B-A3B name puts it in the class occupied by sparse flagships, not phones. Whether OpenBMB intends it as a server-side companion to the edge line, as a teacher for a future small model, or as a capability ceiling test is not stated anywhere, and the repository contains no hint either way.

What is genuinely new, and worth flagging for anyone watching the series, is the Qwen-MoE backbone plus the 3:1 linear-to-full attention ratio. Both are departures. Everything OpenBMB publishes about the family is still written around 4.6, so this checkpoint is running ahead of its own documentation.

A screenshot of the GitHub repository OpenBMB/MiniCPM-V (captured 7 October 2026) showing the repository header and README, whose model table lists MiniCPM-V 4.6 as the latest and most efficient model in the MiniCPM-V series — with no mention of MiniCPM-V 4.7 anywhere on the page.

What is not knowable yet — and why that list matters

It is worth being blunt about the size of the hole here, because the temptation with a 70 GB checkpoint is to fill it with plausible inference.

• No licence. This is not a formality. MiniCPM-V 4.6 is Apache-2.0, and the community has come to expect that from this line. A repo with no license: tag is, by default, all rights reserved in most jurisdictions — you cannot safely build on it until the tag appears. That single missing file is the most consequential absence in the repository.

• No benchmarks, vendor-reported or otherwise. There is not a single number to argue about. Anyone quoting an MMMU or OCRBench score for MiniCPM-V 4.7 today is quoting something that does not exist in the source.

• No independent evaluation. Artificial Analysis and similar trackers index models by name; a checkpoint with no card and no announcement typically sits unmeasured for a while.

• No serving recipe. Whether the released weights run stock in vLLM, SGLang or llama.cpp is untested by anyone outside OpenBMB. The custom code paths (MiniCPMV4_7ForConditionalGeneration, MiniCPMV4_7Processor, MiniCPMV4_7ImageProcessor, MiniCPMV4_7VideoProcessor) all require trust_remote_code, and the MoE plus linear-attention combination is not a shape every inference engine has a kernel for.

• No quantisations. MiniCPM-V 4.6 shipped GGUF, AWQ, GPTQ and BNB variants, and reached Ollama's library in June 2026. None exist for 4.7. For a 35B model this is the difference between a laptop and a cluster.

• No relationship to 4.6 stated. OpenBMB could be replacing the edge line, extending it, or testing something orthogonal. The repository does not say, and neither does the GitHub README, which still lists 4.6 as the current release as of its September 8, 2026 commit.

There is also a plausible-but-unconfirmed reading of the timeline: the twelve-minute gap between repo creation and last commit, the absence of a card, and the absence of any post are what a checkpoint looks like when it is staged for a launch rather than after one. That is a hypothesis about intent, not a fact about the artifact, and it should be treated as one.

How you would actually run it, when there is something to run

Nothing here is quotable as a supported path yet, but the config does constrain the options. A 35B-parameter BF16 checkpoint needs roughly 70 GB of accelerator memory before KV cache, so single-GPU hobbyist deployment is off the table until quantisations land. The 256K context and the 3:1 linear-attention ratio mean the KV cache grows far more slowly than a conventional transformer of the same depth, which is precisely why that architecture is worth the complexity — long-context multimodal work is where the design pays for itself. When a card appears, the first things to check are the licence, whether an official vLLM or SGLang recipe is published, and whether the 16x downsample default can be traded for the 4x setting that 4.6 exposed.

For teams that want to evaluate a model like this the moment it becomes usable, the practical problem is not the weights, it is the plumbing around them. A model that lives on your own GPU still needs everything around it — a router that puts 200-plus hosted models behind one key, at each provider's list price with no markup of ours layered on top, and that fails over automatically when a provider degrades. That is what OrcaRouter is for, and it is worth saying plainly that MiniCPM-V 4.7 itself is not a hosted model: it is an open-weights checkpoint you serve yourself. The router matters here as the other half of the architecture — the frontier model your 4.7 box hands the hard cases to, reachable through the same client you already wrote.

The state of play, and the one file to refresh

MiniCPM-V 4.7 exists. Its weights are downloadable today, its weight count is exact, and its training-era design choices — sparse MoE, 3:1 linear attention, 256K context, 16x visual compression — are all legible from the configuration file. What does not exist is any statement from OpenBMB about what the model does, what it costs to run in practice, or what you are allowed to do with it. For a lab whose last vision model was a 1.3B Apache-2.0 edge release with a full comparison table, that is a jarring gap, and the sensible posture is to watch the repository rather than the discourse. The one file to keep refreshing is README.md: the moment it appears, it will carry the licence, the benchmarks, and probably the explanation of what a 35B MiniCPM-V is doing in a series built on small models.

Until then, the honest summary is the boring one. This is a real checkpoint from a real lab, uploaded quietly, and the interesting question — is it any good — has no published answer.

OrcaRouter reaches 200-plus hosted models through one key with each provider's list price passed straight through at 0% markup and automatic failover between providers. provider list price passed through at 0% markup MiniCPM-V 4.7 is not one of them - it is an open-weights checkpoint you serve yourself, and the router is what sits on the far side of the handoff.