A hero title card headed 'Liquid AI d1-3B vs LFM2.5-2.6B-Base' with the subheading 'the decision model and its own ancestor', drawn as a left-to-right family tree: a card labelled LFM2.5-2.6B-Base raw weights, an arrow labelled post-training, and a card labelled d1-3B answers questions, with a keep-training chip under the left card and yes/no, label and score chips under the right one.
Guides & Insights

Liquid AI d1-3B vs LFM2.5-2.6B-Base: The Decision Model and Its Own Ancestor

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

This is not a matchup between two rivals. Liquid AI d1-3B, the open-weight decision model Liquid AI published on October 7, 2026, was built on a base the vendor constructed by averaging the weights of LFM2.5-2.6B with the text backbone of LFM2.5-VL-3B. LFM2.5-2.6B-Base is that first parent — the raw pre-trained checkpoint Liquid uploaded on August 1, 2026, 2.69B parameters, 34 trillion training tokens, 131,072-token context, no instruction tuning and no chat template. So the honest framing of this comparison is descendant against ancestor: a model that has been post-trained into a very specific job, standing next to the unshaped substrate it was shaped out of. One of them answers questions tonight. The other one is where you start if the question you need answered is not one anybody has asked yet.

What each one actually is

The two model cards read as documents from different stages of a production line, and the differences are not stylistic.

Liquid AI d1-3B carries 3.12B parameters, a SigLIP2 NaFlex vision encoder at 400M, a 32,768-token context window, a 128,000-token vocabulary and sixteen documented languages. It is a decision model: you give it a state and named questions, and it returns a yes/no probability, a chosen label with confidence and a full option distribution, or an ordered score — in a single forward pass, with no output tokens generated. It ships a system_one call for one state with several questions, a batched variant that packs many requests with no padding, and a raw probabilities accessor for the underlying distributions. It is not a chat model and does not write text, and the card says so in those words.

LFM2.5-2.6B-Base carries 2.69B parameters in 30 layers — 22 double-gated short-convolution blocks and 8 grouped-query attention blocks — trained on 34 trillion tokens and extended to a 131,072-token context during mid-training, with the same 128,000-token vocabulary and the same sixteen languages. It is a pre-trained text-only checkpoint with no instruction tuning, no benchmarks on its card, and a recommendation that reads as a warning: use it only for tasks that require heavy fine-tuning. It completes text. It does not follow instructions.

Both are ungated downloads under Liquid's own open license. Both are text-first. Neither is on our catalogue, and nothing here is an availability claim for either.

A two-column scoreboard for Liquid AI d1-3B and LFM2.5-2.6B-Base. The d1-3B column reads: role post-trained decision model; 3.12B parameters; 32,768-token context; text and image input; Decision Index 48.57 (vendor); answers a question in 8 ms. The base column reads: role raw pre-trained base; 2.69B parameters; 131,072-token context; text-only input; Decision Index not scoreable; answers a question only after you fine-tune. The footer reads 'd1-3B figures vendor-reported; the base checkpoint publishes no benchmark table.'

The lineage is the interesting part

Liquid did not fine-tune a fresh model to make the d1 release. It averaged two checkpoints — LFM2.5-2.6B and the text tower of LFM2.5-VL-3B — to get a better starting point, then fine-tuned several runs with different random seeds and data mixtures and merged those back together. The vendor's account of what improved the result is conspicuously free of exotic technique: training on long inputs, shuffling the order in which answer options appear, and removing shortcuts from the training data did more than any advanced method they tried.

That shortcut-removal detail is worth pausing on, because it names the failure this category exists to avoid. A model trained to answer multiple-choice questions will happily learn that option B is usually right, or that the longest option wins. Shuffling option order during training is what stops it. A decision model that has learned a position prior instead of reading the state is worse than useless — it is confidently wrong at scale — and it is the specific thing that separates a real decision model from a classifier with a prompt.

There is also a regression worth naming plainly, because the vendor does not: the base checkpoint has a 131,072-token context window and Liquid AI d1-3B has 32,768. Post-training into a decision model cost three-quarters of the usable context. If you were imagining that the descendant strictly dominates its ancestor on every axis, it does not, and long-document decisions are the axis where the older checkpoint has more room — in principle, though it has no decision head to use it with.

The scoreboard, with the asymmetry attached

Comparing these two on benchmarks is a category error, but it is a category error with an informative shape, so here it is with the caveats in place. Every d1-3B figure is Liquid's own, scored by the vendor with the official scorer rather than submitted to a public leaderboard, and unreproduced by any third party. Every LFM2.5-2.6B-Base figure is absent because a base model has nothing to measure.

• Decision Index 0.2.1 — Liquid AI d1-3B 48.57, first among everything under 10B in the vendor's table and ahead of a decision model twelve times its size. LFM2.5-2.6B-Base — not scored, not scoreable without a decision head.

• Text benchmarks — Liquid AI d1-3B averages 82.9 across seven public benchmarks spanning comprehension, toxicity, intent, medical QA and cross-lingual understanding. LFM2.5-2.6B-Base publishes no benchmark table at all.

• Vision — Liquid AI d1-3B averages 74.1 over eleven image benchmarks read as decisions; removing the images drops the same questions to 45.1. LFM2.5-2.6B-Base is text-only, with no vision tower.

• Parameters — 3.12B against 2.69B. The decision model is the larger of the two, which is the opposite of the usual post-training story.

• Context — 32,768 tokens against 131,072. The ancestor wins this row and it is the only row it wins.

• Latency — Liquid AI d1-3B answers a single question in 8 ms on an RTX 4090 and 50 ms on a Jetson Orin Nano, and runs 64 packed states at 475 per second on the 4090. LFM2.5-2.6B-Base has no latency to quote until you fine-tune it into something that answers a question, at which point the number is your fine-tune's.

A capture of Liquid AI's documentation page for its decision models, showing the left-hand navigation including Decision Models and Liquid Nanos, the 'Open d1' banner announcing that visitors can now run d1-3B and d1-omni-600M locally, and the page's opening description of purpose-built models for structured decisions such as classification, routing and scoring in a single call with zero generation.

What you are choosing between, in practice

The decision rule here is unusually clean, because the two checkpoints are not competing for the same budget line.

Take Liquid AI d1-3B when the question already exists and is one of the three shapes a decision model handles: a yes or no, a pick from a named set, or a rating on an ordered scale. The published applications are all recognisable production chores — triage and routing, moderation, intent and topic classification, extraction checks, reranking, agent guardrails, and visual inspection. The demo numbers the vendor ran on October 5 put d1 against GPT-6.1 Sol and Claude Opus 5.5 on six such tasks and claim it matches or beats GPT-6.1 Sol on four while costing 19x to 200x less; the methodology page states plainly that each application was run once per model, so treat the shape of the claim as the finding, not the decimal places. The genuinely striking one is the inspection demo: good-versus-defective sorting across four production lines at 85 to 97% accuracy on a task the model was never trained for, understood from a short description.

Take LFM2.5-2.6B-Base when the question does not exist yet. It is the cheaper starting point for a fine-tune you intend to own — a domain decision head, a bespoke classifier, a task nobody has a checkpoint for. You inherit the architecture, the tokenizer and the long context, and you pay in compute, data and evaluation. Two of the four applications the vendor built around d1 started from open-source projects rather than from scratch, which is a fair picture of what a base checkpoint plus discipline actually produces.

The practical trap is treating the base checkpoint as a drop-in. It is not an assistant, it will not follow a prompt, and evaluated as a chat model it scores near zero — which is not a fact about its quality, it is a fact about what "base" means.

Self-hosting either one, and the layer above

Both arrive as weights, and the vendor did the deployment work for the d1 side: day-one llama.cpp support, the full NVIDIA stack from DGX down to Jetson, NVFP4 quantization, an 8-bit weight-and-activation variant, and GGUF conversions on Hugging Face. The base checkpoint has GGUF, ONNX and MLX variants of its own through the wider LFM2.5 family. If you would rather not host anything, Liquid serves the hosted d1 from its own API priced on input tokens only, with images billed at 1.5 tokens per 32×32-pixel patch; text-only access is also available through third-party platforms.

What neither path changes is the rest of the stack. A model you host is a model you have to keep alive, version and fail over — and the decisions it feeds still land next to the generative calls you do not host. That mix is the layer a router collapses: one endpoint across 200+ models, provider list prices passed through at 0% markup so a vendor price change is live on your side the same day, and automatic failover so one upstream's bad afternoon does not become your outage. Self-hosting a 3B decision model next to it makes the routing question sharper rather than softer, because now you own half the pipeline and rent the other half.

Which one, for whom

If the question you need answered this week is one of the three shapes a decision model takes, and it is a question somebody else has already imagined, the descendant is finished and free and runs in 50 ms on a device with no GPU — take Liquid AI d1-3B and be done with it.

If you are building the thing that answers a question nobody has asked yet, and you have the data and the compute to own that, LFM2.5-2.6B-Base is the honest starting point, and it is worth knowing that the descendant in this article was made from it by exactly the route you are about to take.

And if you are not sure which of those two situations you are in, the answer is almost always the first one. Post-training into a decision head is the expensive move; running somebody else's is free.

A capture of our own models catalogue page, showing the filter rail for input modalities, context length, input price, status, series and supported parameters, a header reading '207 models, 16 providers, one API key, one bill', sort control set to Newest, a model listing beginning with openai/gpt-5-search-api-2025-10-14, and a 'How to call any model' panel with the OpenAI-compatible endpoint https://api.orcarouter.ai/v1/chat/completions.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily