
MiniCPM5-2B Weights Are Live: The Small Model That Claimed the Sub-4B Crown Goes Open-Source
- openaiNEWOpenAI: GPT-6 Astra2026-09-0455Intelligence77Coding
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0247Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0247Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0157Intelligence82Coding
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2646Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1849Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1541Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1242Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1251Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0547Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0347Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3141Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2454Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2140Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2128Intelligence49Coding
The model itself is seven weeks old. MiniCPM5-2B was unveiled at the 2026 World Artificial Intelligence Conference on July 19, where ModelBest and OpenBMB called it the top-ranked model under 4B parameters on Artificial Analysis's leaderboard and promised the weights would follow "in the near future." What happened this weekend is that the near future arrived without any of the launch fanfare: the MiniCPM5-2B repository went live on Hugging Face on September 6, and OpenBMB's changelog marks the release September 7, under an Apache-2.0 license, with the full package — BF16 weights, GGUF, MLX and GPTQ quantizations, base and SFT checkpoints, a DSpark draft model for speculative decoding, and the training datasets behind them.
No press release accompanied it and no launch event. It simply appeared: a Hugging Face repo one day, a changelog line the next. That makes this a what-we-know-so-far piece — with one early update. The repo tells us a great deal: the model is genuinely downloadable and runnable today, and the spec sheet is public. And an outside score has already landed: Artificial Analysis lists MiniCPM5-2B on its Intelligence Index v4.2 at 15, the top open-weights result under 4B total parameters on that index, so the sub-4B ranking no longer rests on the vendor's word alone. What is still not confirmed is the model card's headline figures, which OpenBMB reproduced in its own harness and nobody outside the lab has re-run. Here is what is knowable from the repo, what is now independently measured, and what still needs verification before you build on it.
What just happened
MiniCPM5-2B is the second model in the MiniCPM5 series, following MiniCPM5-1B, which OpenBMB open-sourced on May 19. When the 2B shipped as a product at WAIC, the story was a chip story as much as a model story — ModelBest framed it as an on-device agent base for phones, PCs and smart cockpits, and named nine chips with day-zero support through its FlagOS community, from Huawei Ascend to NVIDIA to a spread of domestic accelerators. The open-sourcing was framed as imminent. Seven weeks passed.
On September 6 the openbmb/MiniCPM5-2B repo appeared. As of this writing the model card is the release announcement, and it is unusually complete for a quiet drop:
• Weights — the final RL + OPD post-trained checkpoint in BF16, licensed Apache-2.0, and not gated — anyone can download it.
• Companion checkpoints — MiniCPM5-2B-SFT, -Midtrain and -Base for people who want the model before RL/OPD, plus -GGUF, -MLX, -GPTQ and a -DSpark draft model, all listed in the MiniCPM5 collection on Hugging Face.
• Data — the training corpora are open too, released as part of the UltraData family: Ultra-FineWeb, UltraX, UltraData-Code, UltraData-Math, UltraData-SFT-Agent-2609 and UltraData-RL-2609.
• Mirrors — the README points to both Hugging Face and ModelScope.
One line in the card matters more than it looks: "no custom kernels, no model-code fork." MiniCPM5-2B uses the standard LlamaForCausalLM architecture, so any engine that already serves Llama-family models can load it without waiting for a bespoke implementation.
Why a 2.5B model is worth a second look
The WAIC pitch was a rankings pitch. ModelBest said MiniCPM5-2B scored an Artificial Analysis Intelligence Index of 17, placing it first among models under 4B parameters on the AA leaderboard at launch. Two caveats belong next to that number. First, index scores are only comparable within a methodology version: MiniCPM5-1B's earlier 17.9 came from a previous AA snapshot, so it is no evidence that the smaller model "scores higher," and by the same rule a newer score on a newer methodology is not a regression. Second, AA's figures feed the model card, but the card's headline average is OpenBMB's own, reproduced in OpenBMB's own harness. That distinction matters because AA has since moved to a newer methodology and published a fresh score — the first outside check on the pitch since the open-weights release.
An outside number arrived in the same week as the weights. On September 4 Artificial Analysis shipped Intelligence Index v4.2, a methodology revision that raises the weight on private, held-out tests to 40% and adds two new components — AA-Briefcase, an agentic evaluation, and Surge's GDP.pdf, a long-context document-reasoning task. AA's index entry for MiniCPM5-2B now carries a v4.2 score of 15: the top result among open-weights models under 4B total parameters, and first of the 47 open models in that size class on AA's listing. That keeps the model where the July pitch put it, this time on a methodology that is harder to game and with weights anyone can download and re-check. The 15 is not comparable to the launch-era 17, which came from v4.1 — the methodology is stricter, not the model weaker — and the composite does not re-verify the card's individual rows, most of which OpenBMB reproduced in its own harness. What it does do is land on the two axes OpenBMB says it optimized for: v4.2 leans harder on agentic tasks, and AA's verbosity tracking shows MiniCPM5-2B generating 57M output tokens across the index tasks, the most concise model in its class against a 72M median. OpenBMB thanked AA for the run and framed it in exactly those terms — agentic performance and token efficiency at 2.6B, it said, are what the model was optimized for.
The substantive claim was agentic capability in a small package: native tool calling, deep search and code generation, trained in from the ground up rather than bolted on. The card describes a three-stage pipeline — base training, mid-training, then post-training of SFT on 400B tokens of deep-thinking data, specialized RL teachers, and On-Policy Distillation (OPD) that merges sixteen RL expert models, five of them agentic, back into one small dense network. OpenBMB credits the RL + OPD stage with an average gain of 10.96 points on reasoning and general tasks and 6.96 points on agentic tasks — internal deltas, but they explain the unusual shape of this model: a 2.5B parameter network built like a reasoning model and pointed at tool use.

What the repo actually contains
The spec sheet is unambiguous, because it is the config file. MiniCPM5-2B has 2,516,756,480 parameters, of which 1,981,982,720 are non-embedding — vendors count the non-embedding figure when they say "2B-class." It is a dense transformer with 42 layers at a hidden size of 2048, grouped-query attention with 16 query heads and just 2 KV heads, a vocabulary of 130,560, and BF16 tensors. The whole thing ships as a single safetensors shard, small enough to pull on a laptop connection and run on a single mid-range GPU or a phone-class NPU.
• Params — 2,516,756,480 total; 1,981,982,720 non-embedding.
• Architecture — dense LlamaForCausalLM, 42 layers, hidden 2048, GQA 16 query heads / 2 KV heads, vocab 130,560.
• Context — 131,072 tokens configured (max_position_embeddings), no RoPE scaling in the config.
• License — Apache-2.0, for both the model and the released training data.
• Format — BF16; GGUF, MLX and GPTQ companions published separately.

One number from the July deck does not show up in the checkpoint. The WAIC materials touted a 512K-token context window; the released config sets max_position_embeddings to 131,072 (128K) with no rope_scaling entry. It is possible the 512K figure is meant to be reached through inference-time extension the way several small models stretch their context — but as shipped, the repo documents 131,072, and that is the number to design against until OpenBMB publishes an extension recipe.
The numbers, sorted by who measured them
The model card is careful about provenance, and it repays reading the footnotes. Scores it marks with a dagger "come from the official Artificial Analysis release" — rows such as GPQA-Diamond 70.2, HLE 8.9, AA-LCR 59.0, Terminal-Bench v2.1 8.6 and GDPval-AA v2 19.6. Everything else is described as "reproduced internally," which means OpenBMB ran those evals in its own harness. That category includes the attention-getting figures: an internal average of 53.9 across the card's comparison set, AIME 2025/2026 at 86.5, MATH-500 at 94.6, LiveCodeBench v6 at 69.1, SWE-bench Verified at 46.4, BFCL v4 at 66.6, τ²-Bench Telecom at 97.1, GAIA (Text-103) at 88.7 and MMLU-Pro at 70.8.
Three things keep those numbers honest to read. The 53.9 average is a relative score, not an absolute one — the card normalizes each radar axis so the best model in the comparison scores 100. The comparison set is vendor-selected: other 2B-4B open models including Qwen3.5-2B, Qwen3.5-4B, Gemma-4-E2B-it, granite-4.2-3B, LFM2.5-2.6B and LFM2.5-8B-A1B. Within that set the card shows MiniCPM5-2B clearing the next-best, Qwen3.5-4B at 51.1 — a genuinely striking result for a model half the size, and precisely the kind of result that needs an independent reproduction before anyone builds on it. And a couple of the individual rows are hard to square with the marketing: an SWE-bench Verified of 46.4 from a 2.5B dense model would be remarkable if it reproduces, while an HLE of 8.9 is a reminder that "sub-4B leader" and "frontier" are different leagues. None of these are contradicted by the repo, and AA's v4.2 composite has since confirmed the model's overall placement at the top of the sub-4B open-weights class — but these specific rows are OpenBMB's own, still unverified by anyone outside the lab.
Running MiniCPM5-2B today
Because the architecture is plain Llama, mainstream engines load it directly. OpenBMB's README gives quickstarts for vLLM (0.21 or newer), SGLang (0.5.16 or newer) and Transformers (5.6 or newer), with recommended sampling at temperature 1.0 and top-p 0.95, and enable_thinking switched on when you want the reasoning pass. The GGUF and MLX companions cover llama.cpp, Ollama, LM Studio and Apple Silicon; the card also lists the ArcLight runtime and the usual fine-tuning stacks (TRL + PEFT, LLaMA-Factory, ms-swift, unsloth).
Two deployment details are worth knowing before you start. For tool and function calling, SGLang is the recommended backend: the model emits XML-style tool calls, and SGLang's built-in minicpm5 parser converts them to OpenAI-compatible tool_calls natively, which means standard tool-calling clients work without a custom shim. And the DSpark draft model exists to make the reasoning pass cheaper: launched in SGLang with the DSPARK speculative-decoding algorithm, it accelerates decoding while leaving the target model's outputs unchanged — the kind of thing that decides whether a 128K-context agent loop on a laptop is usable or merely possible.

If you would rather evaluate a batch of small open models than hand-tune one, the routing layer earns its keep here: when a provider lists MiniCPM5-2B, a router is the cheap way to A/B it against Qwen3.5-2B and the other sub-4B open checkpoints on the same task without a second integration. On OrcaRouter that means one API across 200+ models, provider list prices passed through at 0% markup, and automatic failover if a brand-new checkpoint stalls or loops on a production path. The repo is barely two days old and no hosted API we can point to lists MiniCPM5-2B yet — so the honest default today is to treat it as a self-host experiment, not a dependency.
Who should pull these weights
Pull them today if your work is on-device agents, edge tool-calling, or small-model coding and reasoning experiments, and you want the most aggressively-trained 2B-class option on the table. The Apache-2.0 license and the open training data remove the two reasons most teams hesitate on a Chinese-lab small model — no commercial-use restriction, and no black-box corpus. Pull them warily if you are choosing a production model on the strength of the benchmark table alone: the card's headline rows are OpenBMB's own reproductions, and AA's v4.2 composite — the one outside measurement so far — does not vouch for them. A 2.5B model that reportedly clears Qwen3.5-4B is a claim that deserves the two hours it takes to reproduce SWE-bench yourself.
What to watch next. The repo is barely two days old, so the near-term signals are still mechanical: whether llama.cpp and the Ollama library pick up the GGUF, whether the ModelScope mirror and FlagOS chip builds go live cleanly, and whether a hosted API lists MiniCPM5-2B — that is the moment it stops being self-host-only. The substantive signal this piece flagged as missing — an independent run by Artificial Analysis or a university lab — has now arrived in the form of the v4.2 index score, and it keeps MiniCPM5-2B first among open-weights models under 4B. What is still outstanding is row-level verification: nobody outside OpenBMB has reproduced the card's internal figures, most notably SWE-bench Verified at 46.4 and the 53.9 internal average. Until those specific rows are re-run, treat MiniCPM5-2B the way you would treat any quietly-released model with an impressive card: as a very promising hypothesis you can run locally tonight, and a model whose most striking numbers you should verify before you trust it with real work.
