
Liquid AI d1-3B vs MiniCPM5-2B: Probabilities or Prose, at Laptop Scale
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 60 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 356 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Two open-weight models, both small enough to run on the machine in front of you, and they disagree about what a model should hand back. Liquid AI d1-3B is Liquid AI's 3.12B decision model, open-weighted on October 7, 2026: it reads a state and a set of named questions and returns a label, a probability per option and a confidence, in one forward pass, generating nothing. MiniCPM5-2B is ModelBest and OpenBMB's 2.52B dense model, shown off at the World Artificial Intelligence Conference on July 19 and quietly published to Hugging Face in early September 2026 under Apache-2.0 — a reasoning, coding and tool-calling assistant that writes answers, calls functions in XML, and already has over a million downloads. The smaller one has four times the context window and a far more permissive license; the larger one cannot produce a token even if you want it to. Which of those two shapes belongs in your pipeline depends entirely on whether the thing downstream of the call wants a number or a sentence.
The reversal worth noticing first
Almost every intuitive expectation about this pairing is backwards, so start there.
MiniCPM5-2B is the smaller model and it holds the longer context: 131,072 tokens against 32,768 for Liquid AI d1-3B. It is also the more liberally licensed of the two — Apache-2.0, ungated, with the training datasets released alongside the weights as part of the UltraData family. Liquid AI d1-3B ships under Liquid's own lfm1.0 license, which is downloadable and fine-tunable, but is a vendor license rather than a standard permissive one.
Liquid AI d1-3B is the larger model — 3.12B against 2.52B total, though the comparison flips again on non-embedding parameters — and it is the one that cannot write. It has vision; the MiniCPM is text-only. It is post-trained for exactly one thing; the MiniCPM is post-trained for many. And where MiniCPM5-2B arrives with a comparison chart placing it at the top of its size class, Liquid AI d1-3B arrives with a single index score and a set of latency tables.
None of that makes one better. It means the two are aimed at different jobs, and the giveaway is the shape of the answer each one returns.
What comes back on the wire
MiniCPM5-2B is a generative language model. It reasons, it answers in prose, and it emits XML-style tool calls that SGLang's built-in minicpm5 parser converts into OpenAI-compatible tool_calls without a custom adapter. Its post-training is the interesting part of its story: 400 billion tokens of deep-thinking supervised fine-tuning, then reinforcement learning with a critic-based algorithm borrowed from JustRL II, then on-policy distillation folding sixteen specialised RL teachers — five of them agentic — back into one dense network. The card credits RL plus distillation with an average gain of 10.96 points on reasoning and general tasks and 6.96 on agentic ones, and reports an average of 53.9 across its own comparison set, above the largest model included in that set. Those are ModelBest's figures, internally produced, with no independent reproduction published yet.
Liquid AI d1-3B returns structure. A noul question comes back as P(yes) between 0 and 1. A choice question comes back as the label, a confidence and a probability per option. A score question comes back as an expected level on an ordered two-to-ten scale with its distribution and legend. The usage object reports output_tokens: 0, and there is a raw probabilities accessor if you want the unmassaged vectors. Its post-training was the opposite kind of work: average two checkpoints together, fine-tune several seeds over different data mixtures, merge them, then spend the effort on long-input training, shuffling answer-option order and stripping shortcuts out of the training data — because a decision model that learns "option B is usually right" instead of reading the state is worse than useless.
One asks a model to think and then say something. The other asks a model what it already believes.

Side by side, with provenance labelled
Every figure below is vendor-reported. MiniCPM5-2B's numbers come from ModelBest's own card and its own comparison set; Liquid AI d1-3B's come from Liquid running the official Decision Index scorer itself rather than submitting to the public leaderboard. Neither model has been independently benchmarked by a third party as of the open-weights release.
• Parameters — Liquid AI d1-3B 3.12B total with a 400M SigLIP2 vision encoder and 128,000-token vocabulary; MiniCPM5-2B 2,516,756,480 total, 1,981,982,720 non-embedding, 42 layers, 16 query heads to 2 KV heads, vocabulary 130,560.
• Context — 32,768 tokens for Liquid AI d1-3B; 131,072 for MiniCPM5-2B.
• Input modalities — text and images for Liquid AI d1-3B; text only for MiniCPM5-2B.
• Output — probabilities, labels, confidences and ordered scores, with zero output tokens, for Liquid AI d1-3B; generated text, reasoning and XML tool calls for MiniCPM5-2B.
• Headline score — Liquid AI d1-3B 48.57 on Decision Index 0.2.1, first among models under 10B in the vendor's table; MiniCPM5-2B an average of 53.9 across its own comparison set, which is a different set measuring different things.
• Speed — 8 ms per question on an RTX 4090, 26 ms on a Jetson AGX Orin 64 GB, 50 ms on a Jetson Orin Nano, and 64 packed states per pass at 475 per second for Liquid AI d1-3B. MiniCPM5-2B's speed is a function of how many tokens it generates, so there is no single number to put opposite it.
• License — Apache-2.0, ungated, training data released, for MiniCPM5-2B; Liquid's lfm1.0 for Liquid AI d1-3B.
• Evidence — over a million Hugging Face downloads and a month of community quantisation for MiniCPM5-2B; two days of existence and no third-party run for Liquid AI d1-3B.
The job each one is actually good at
There is a workflow both models can do, and the difference in how they do it is the whole argument.
Suppose you are triaging support tickets, or deciding whether a retrieved passage answers a question, or scoring an LLM output against a rubric. MiniCPM5-2B can do all three — it is a capable small reasoner with tool calling and a long context, and you would run it like any other assistant, asking for a label and then parsing whatever comes back. Sometimes it returns clean JSON. Sometimes it returns JSON with an explanation wrapped around it. Sometimes it returns the right answer in the wrong shape, and that miss is a bug in your parser rather than in the model.
Liquid AI d1-3B cannot do anything else, which is precisely the point. On the same three tasks it returns a probability and a label and nothing to parse, three questions over one state cost 1.3x the time of one, and the failure mode shifts from "the format broke" to "the confidence was miscalibrated" — which is at least measurable, because the confidence is right there in the response.
Where MiniCPM5-2B wins outright is everything downstream of the decision. It holds 131K tokens, so the state can be a whole repository file tree or a long document rather than the first page of one. It writes tool calls natively. It is a model you can build a small agent on, and it was post-trained explicitly for agentic work. And its Apache-2.0 license means the usual commercial-lawyer conversation does not happen.
Where Liquid AI d1-3B wins outright is the latency floor and the modality. A device running a decision in 50 ms with no GPU is not a device running MiniCPM5-2B; the two live in different hardware tiers despite the similar parameter counts. And the 3B sees — the vendor reports 74.1 across eleven public image benchmarks read as decisions, against 73.9 for the vision-language model it was built from, and reports that removing the images collapses the same questions to 45.1, which is their way of showing the answers come from the pixels. Visual inspection of a production line is not a task the text-only 2B can be handed at all.

Running them, and the layer around them
Both are self-host stories, and both have had the deployment work done. Liquid AI d1-3B ships with day-one llama.cpp support, the full NVIDIA stack from DGX to Jetson, NVFP4 quantisation, an 8-bit weight-and-activation build, and GGUF conversions already published. MiniCPM5-2B is a plain LlamaForCausalLM architecture with no custom kernels required, one BF16 safetensors shard, GGUF and MLX builds in the wider family, and a documented preference for SGLang when you need the tool-call parser. Neither is on our catalogue, and this article is not an availability claim for either one.
The hosted alternatives are the vendor's own API and third-party platforms. Liquid serves the hosted d1 priced on input tokens only at $0.04 per million, with images billed as input at 1.5 tokens per 32×32-pixel patch and no output line at all; MiniCPM5-2B has no first-party API, which is normal for an Apache-2.0 open-weights release.
What both of them feed into is the same architecture problem. A local 2B or 3B handling classification, routing or guardrail checks sits in front of the generative calls you do not host, and that mix — owned inference next to rented inference — is the thing a routing layer exists to absorb. One endpoint across 200+ models, provider list prices passed through at 0% markup so a vendor price change is live the same day, automatic failover when one upstream degrades. Owning the small model makes the routing question harder, not easier, because now the boundary between your hardware and someone else's is a decision you have to make per request.
Pick by the answer type
If the thing you need is a label, a probability or a rating, and the option set is known in advance, Liquid AI d1-3B is the more honest instrument — it was built for that request and the answer is not going to arrive malformed. Accept the vendor license, the 32K context and the two-day-old evidence trail as the price.
If the thing you need is a small model that can reason, call tools, hold a long document and answer in words, MiniCPM5-2B is the more capable machine by a wide margin, and it comes with Apache-2.0 and a million downloads' worth of other people's testing behind it.
The teams that will get the most from either release are the ones running both: a decision model in front, where the answer is a number and milliseconds matter, and a small generative model behind it, for the parts of the job that need language. They are not competitors. They are a filter and a speaker.

