Generated hero card for Kolibri vs LFM2.5 2.6B Base, headed 'Kolibri vs LFM2.5 2.6B Base' with the subline 'server-grade reasoning against an on-device checkpoint', a hummingbird icon on the left and a small chip outline on the right either side of a thin divider. The OrcaRouter logo sits in the strip below the artwork.
Guides & Insights

Kolibri vs LFM2.5 2.6B Base: a 78B Sovereign Deployment Against a Phone-Sized Checkpoint

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put Kolibri next to LFM2.5 2.6B Base and the obvious reading is David and Goliath: 78.1 billion parameters against 2.69 billion, a two-GPU deployment floor against a model Liquid AI says runs on a phone. The obvious reading is wrong, and not in the direction you would expect. Neither of these is a model you can point at a task today. Kolibri was released by Aleph Alpha on 3 October 2026 as a complete, post-trained, tool-calling German-and-English assistant with a serving plugin and a recommended sampling configuration. LFM2.5 2.6B Base, published by Liquid AI in early August 2026, is a pre-trained checkpoint with no instruction tuning at all, released specifically to be fine-tuned. One is a product you deploy. The other is raw material. Comparing their quality is a category error, and the interesting question is which kind of raw material your project actually needs.

The gap that decides most real procurements between these two is not parameters. It is the licence. Kolibri ships under Apache 2.0. The LFM2.5 2.6B Base ships under the LFM Open License v1.0, a bespoke licence (the model card files it as "license: other"), and that difference is the one that goes to a legal team rather than an engineering team.

Two different meanings of "open weights"

Both models let you download the weights and run them. What you are allowed to do afterwards is where they part.

• Kolibri — Apache 2.0, no acceptable-use rider, no user-count threshold, no separate commercial terms. Aleph Alpha's card notes only that the lab is a signatory of the EU GPAI Code of Practice.

• LFM2.5 2.6B Base — the LFM Open License v1.0, which is Liquid AI's own licence rather than an OSI-standard one. It is the same licence that governs the post-trained LFM2.5 2.6B and the rest of the family.

For a research team either is fine. For an enterprise that has to state its licence position in a procurement document, one is a known quantity and the other is a document somebody has to read. That is a real cost, it is paid before a single token is generated, and it is invisible in any benchmark table.

There is a second distinction in the same category, and it is the one Liquid AI's own card is at pains to make. The Base is not an assistant. It carries no instruction tuning, which means it will not follow a prompt, will not abstain, will not emit a tool call, and will not stay on topic — not because it is weak, but because it was never trained to. The card says plainly that it is recommended only for tasks requiring heavy fine-tuning: language-specific or domain-specific assistants, training on proprietary data, or experimenting with novel post-training approaches. Everything downstream of the base checkpoint — supervised fine-tuning, teacher specialisation, on-policy distillation, agentic reinforcement learning — is the buyer's problem. Kolibri arrives with all of that already done, at four reasoning effort levels, with a Hermes-style tool-call parser shipped next to the weights.

What the sizes actually buy

Strip out the framing and the two models are optimised for opposite constraints.

• Parameters — 78,103,074,560 total with 3,457,573,120 active per token on Kolibri, against 2.69 billion total on the LFM2.5 Base, which uses a hybrid stack of 22 double-gated short-convolution blocks and 8 grouped-query attention layers rather than a transformer of uniform blocks.

• Context — 131,072 tokens on the LFM2.5 Base, against a 262,144-token native window on Kolibri and a validated 1,048,576 beyond it.

• Training data — 34 trillion tokens on the LFM2.5 Base, against 20 trillion on Kolibri drawn from a raw pool of over 200 trillion after filtering and deduplication, of which 13.6 per cent was code.

• Languages — sixteen on the LFM2.5 Base, including Arabic, Chinese, Japanese, Korean, Thai and Vietnamese, against exactly two on Kolibri, German and English, by explicit design.

• Memory — the LFM2.5 Base has GGUF, ONNX and MLX builds published alongside it and is built for edge deployment; Kolibri's FP8 weights are a ~78 GB footprint with a minimum of two A100 80 GB cards, two H100 SXM5s, one H200, one B200 or one B300.

Read those five lines together and the picture stops being lopsided. Liquid AI's model covers ten times the languages on hardware three orders of magnitude smaller. Aleph Alpha's model covers two languages with a context window eight times longer and a depth of specialisation that only shows up inside its own vertical evaluations. Neither is a worse version of the other; they are answers to different questions, and the questions are "can this run without a server" and "can this reason through a German regulatory filing".

Generated two-column scoreboard for Kolibri and LFM2.5 2.6B Base. Left column Kolibri: role 'post-trained assistant', parameters '78.1B MoE, 3.46B active', context '262,144 native, 1M validated', languages 'German and English', licence Apache 2.0, deployment 'two 80 GB accelerators'. Right column LFM2.5 2.6B Base: role 'pre-trained checkpoint', parameters 2.69B, context 131,072, languages 'sixteen', licence 'LFM Open License v1.0', deployment 'phone, GGUF and MLX'. The footer reads 'Kolibri figures vendor-reported; LFM2.5 2.6B Base has no published evaluation.'

Only one of them has a measured score, and it is not the Base

Here is the part the pairing hides. Artificial Analysis does have a model page for LFM2.5-2.6B — but for the post-trained model, not the Base. That page reports an Intelligence Index of 8, ranked #9 of 49 models in its class, with a 128,000-token context, 2.7 billion parameters and the LFM Open License v1.0. It is a modest score and an honest one: the model is measured, priced at effectively zero, and evaluated as what it is — a small reasoning model.

The Base checkpoint has no score, and it cannot have one. An Intelligence Index is computed by running a model against reasoning and knowledge tasks; a checkpoint that has not been instruction-tuned has no meaningful behaviour on those tasks. Liquid AI publishes no evaluation table for it, and that absence is a statement about the artifact, not a gap in the release.

Kolibri's position is different again, and worth stating precisely because it is easy to misread. There is no Artificial Analysis page for Kolibri either — we checked, and the model is absent from that comparison entirely. What exists is Aleph Alpha's own post-training table, where Kolibri scores 75.5 on the English average and 70.8 on the German average across its harness, against 54.1 and 46.4 for its own predecessor Kolibri Origin. Those are vendor-run numbers on a vendor-built evaluation suite, and they are not convertible into an Intelligence Index. So of the two models in this headline, one has an independent score for a sibling version, one has a vendor table, and neither of the exact artifacts being compared has an independent measurement at all.

Screenshot of the Artificial Analysis model page for LFM2.5-2.6B, showing an Intelligence Index of 8 against a median of 5, a #9/49 intelligence rank, 128k-token context window, the summary line that the model is amongst the leading models in intelligence and well priced compared with other open-weight models of similar size, $0.00 input and output pricing against $0.00 medians, and the 49 models in this class count.

The sibling that changes the recommendation

Judging LFM2.5 2.6B Base on its own is the wrong way to evaluate Liquid AI's release, because the Base exists to serve the post-trained LFM2.5 2.6B, and the two ship together. If you have no intention of writing a fine-tuning pipeline, the Base is not your model — the post-trained sibling is, and it is the one with a public score and a published price of essentially nothing per token when you host it yourself.

That is the same relationship Kolibri has with Kolibri Origin, and the comparison is worth making because Aleph Alpha published the timeline. Origin finished pre-training at 30.6 billion parameters and 3.27 billion active on 11 June 2026 and was never released; Kolibri finished on 11 September 2026 at 78.1 billion and was released on 3 October. Two models, three months, one of them internal only. Liquid AI's equivalent split is public both ways: Base and post-trained, side by side, with quantised builds for llama.cpp, ONNX Runtime and Apple's MLX published alongside the native checkpoint. If your plan is to fine-tune, that surrounding tooling is worth more than a benchmark row, because it determines how much of the work you have to do yourself.

What to deploy, by requirement

This one reduces to a short decision list rather than a winner.

• If the model must run without a server — on a phone, a laptop, an appliance in a factory — LFM2.5 2.6B Base is the only one of the two that is even a candidate, and the post-trained LFM2.5 2.6B is the one you would actually start from. Kolibri cannot run there at any quantisation.

• If the workload is German or English document reasoning inside your own perimeter, with regulated-data constraints and a need for abstention behaviour, the LFM2.5 Base is not the right shape of artifact and Kolibri is aimed exactly at it — provided you can supply two 80 GB accelerators and accept a licence position you can state in one line.

• If you need more than German and English, Liquid AI's family covers sixteen languages at 2.6B, and the language coverage argument reverses at that size.

• If you need a deployed assistant this quarter and do not want to run a post-training pipeline, Kolibri is post-trained; the LFM2.5 Base is not, and the LFM2.5 post-trained model is a much smaller assistant than Kolibri rather than a cheaper version of it.

For teams whose workload genuinely sits in the middle — a few billion tokens a month, moderate context, no regulatory requirement to own the hardware — neither of these is the first thing to reach for. Sub-4B models are cheap to self-host and awkward to route at scale because the operational cost of a deployment can exceed the token bill; a 78B model is the opposite problem. The tier that actually gets routed is the one in between, and that tier is on OrcaRouter today: Qwen3.5 35B-A3B at $0.057 per million input and $0.459 output, Qwen3.8-Flash at $0.15 and $0.47 with a 1M-token context, or the mixture-of-experts Gemma 4 26B-A4B IT at $0.06 and $0.33. Those are provider list prices passed through with nothing added per token, on one OpenAI-compatible key with failover, and they are the honest answer for a team that wants to measure its real token volume before committing to either a fine-tuning project or a rack.

To be explicit about what we do not offer: neither Kolibri nor LFM2.5 2.6B Base is routable on OrcaRouter today. We probed the catalogue for both under every vendor and model spelling and neither is there. Kolibri is self-host-only, and the LFM2.5 Base is a fine-tuning artifact that was never meant to sit behind an API.

Screenshot of the OrcaRouter model page for qwen/qwen3.5-35b-a3b, showing the 32K-token context badge, 65K maximum output, text, image and video input, the Vision, Tools, JSON and Reasoning capability tags, the attribution 'Public benchmarks by Qwen - 2026-02-25', the description as an open-weight mixture-of-experts multimodal model with 35B total and 3B active parameters, the pricing tiles $0.06 and $0.46, and a pricing table with two tiers: under 128K input tokens at $0.057 input and $0.459 output, and above that at $0.229 and $1.835, selected by each request's input token count.

The choice in one line

Kolibri is a finished, licensed, documented 78-billion-parameter German-English reasoner that costs two accelerators to run. LFM2.5 2.6B Base is an unfinished 2.69-billion-parameter multilingual starting point that costs a fine-tuning pipeline and fits in your pocket. The parameter ratio between them is 29 to 1, and it is the least informative number in this article.