
TwIL-LM3-Pro vs LFM2.5 2.6B Base: One Is Finished, the Other Is a Starting Point
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 118 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1064 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 104 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 213 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The most revealing thing about the published comparison between TwIL-LM3-Pro and LFM2.5 2.6B Base is that webAI did not publish one against the 2.6B at all. webAI's Track A and Track B tables pit its 3.66B formal-logic specialist against a range of arms — its own Granite base, VibeThinker-3B, Qwen3-8B, Qwen3.5-4B, Llama-3.2-3B, gpt-oss-120b — and the Liquid AI entry they chose is LFM2.5-8B-A1B, not the 2.6B. The 2.6B that does appear in the table is `LFM2-2.6B`, the previous generation, which is a different model with a different pretraining run. That omission is not an oversight. LFM2.5 2.6B Base is a pre-trained checkpoint with no instruction tuning, and instruction-following benchmarks are not a test it was built to take.
So this page compares a finished artifact against a raw material. The Aion-style framing — one model has a score and the other does not — fits, but the reason is more specific and more interesting: one is post-trained and merged, the other deliberately stops before post-training, and the two documents that describe them are answering different questions on purpose.
Two parameters counts that are not the same measurement
TwIL-LM3-Pro is 3.66B parameters, dense, all of them active on every token. It descends from `ibm-granite/granite-4.2-3b` through LoRA fine-tuning, multipath distillation, checkpoint fusion, a WiSE-FT merge back toward the base, and an entropy-weighted GRPO stage — and the artifact webAI shipped is the merged policy, not a LoRA adapter, so it runs standalone.
LFM2.5 2.6B Base is 2.69B parameters across 30 layers: 22 double-gated short-convolution blocks interleaved with 8 grouped-query-attention layers. That hybrid conv/attention design is Liquid's on-device architecture, and it is the reason the family's throughput numbers are what they are rather than a function of parameter count alone. It was trained on 34 trillion tokens with a 128,000-token vocabulary and carries a 131,072-token context window.
The two counts sit within about 36% of each other and are arrived at through entirely different architectures, so neither "3.66B beats 2.69B" nor the reverse means anything as a quality claim. webAI's own model card makes this point sideways when it declines to compare corpus perplexity across tokenizers — TwIL-LM3-Pro's vocabulary is 100,352, LFM2.5-2.6B's is 128,000, and perplexity is per token.
What "Base" removes, and why it changes the scoreboard
Liquid AI publishes LFM2.5 2.6B in two forms: the post-trained checkpoint, tuned for agentic workloads in popular agent harnesses, and the Base, described in its own card as "the pre-trained text-only checkpoint, used to create all the LFM2.5-2.6B variants." The Base is the input to fine-tuning, not a product. It has no instruction tuning, no chat template worth the name, and no published evaluation — because evaluating a base model on instruction-following tasks measures the wrong thing.
TwIL-LM3-Pro's position is the inverse. Its entire value is the post-training stack: webAI reports that on its own harness the pipeline lifted the in-domain macro gate from its base's 0.4313 to 0.5539, with the held-out 10-dataset macro staying level (0.7942 to 0.7901) rather than collapsing. The held-out retention is the hard part of specialist post-training, and it is the part the Base cannot answer.
This is also where the size comparison in webAI's tables pays off. TwIL-LM3-Pro is reported ahead of LFM2.5-8B-A1B on both held-out macros — 0.7901 versus 0.7884 on the ten-dataset set, 0.7425 versus 0.7378 on the fourteen — at less than half that MoE model's parameter count. That is a genuine like-for-like result, and it is available precisely because LFM2.5-8B-A1B is a post-trained model that can be run through an instruction harness. The Base cannot, which is why the 2.6B never got a column.
The dimensions worth lining up
• Parameters — TwIL-LM3-Pro 3.66B dense vs LFM2.5 2.6B Base 2.69B (30 layers: 22 conv + 8 GQA).
• Post-training — LoRA SFT, distillation, WiSE-FT merge and GRPO at step 2580 vs none; the Base is the pre-training output by design.
• Context window — 131,072 tokens both, inherited from their respective bases.
• Vocabulary — 100,352 tokens vs 128,000, which is why their perplexities are not comparable.
• Languages — English only vs seventeen, including Arabic, Chinese, Japanese, Korean, Spanish and Thai.
• Thinking format — TwIL-LM3-Pro emits a `<think>` block by default and reasons at length (1,902 tokens on average in-domain, 24.2% of generations hitting the length cap) vs a base checkpoint with no instruction format at all.
• Licence — webAI Non-Commercial License ver. 1.0 vs Liquid's LFM Open License v1.0.
• What it is for — a formal-logic inference endpoint vs a fine-tuning starting point.
The context lines deserve a footnote both models supply: each carries 131,072 tokens through from pretraining, and neither vendor validated long-context behaviour — TwIL-LM3-Pro's card says every published score was measured inside an 8,192-token window. Matching numbers here are inherited, not demonstrated.
Licence, and what each one is legally for
TwIL-LM3-Pro's webAI Non-Commercial License permits research and personal use and requires a separate agreement with webAI for revenue-generating deployment. LFM2.5 2.6B Base ships under Liquid's own LFM Open License v1.0, which the Hugging Face API reports as `license: other` rather than a standard SPDX identifier — read the licence file before planning around it, because the terms are Liquid's own rather than a recognised template.
Neither licence is Apache 2.0, and it is a mistake to treat "open weights" as a synonym for "commercially permissive." The practical difference between the two is in what the licences are protecting: a non-commercial researcher can use both, a company building a product needs to check both, and the Base's whole purpose — being fine-tuned into somebody else's product — makes its exact terms more consequential than they look.
Where each one belongs in a stack
The Base is the right choice when you intend to train. If your problem is a narrow labelling task, a domain-specific classification head, or a fine-tune on a few thousand labelled examples, starting from a 2.69B pre-trained checkpoint with a wide multilingual vocabulary and a hybrid conv architecture designed for throughput is a sound engineering decision, and the cost of the post-training you would skip is the cost you are deliberately paying instead.
TwIL-LM3-Pro is the right choice when you do not intend to train and the task is already formal logic. Entailment labels, Lean statement formalisation, semantic parses and proof critique are the tasks its entire pipeline was aimed at, and starting from a checkpoint that has already been through the merge and the reinforcement stage saves you the one part of this that is genuinely hard to reproduce — keeping the held-out suite level while gaining in-domain.
The mistake is to read "3.66B post-trained specialist" as strictly better than "2.69B base checkpoint." They are at different stages of the same production line.
Serving them, or serving around them
Neither model is a route in the OrcaRouter catalogue today, and the reason differs for each. TwIL-LM3-Pro is non-commercially licensed and has no hosted endpoint; LFM2.5 2.6B Base is a fine-tuning artifact that nobody would serve as an inference endpoint on purpose. Where a router earns its place is one layer up, in the work that surrounds a specialist: generating and filtering the training data you would fine-tune the Base on, expanding the prompts you would feed a formal-logic model, and grading the outputs. Those are text calls, and they run over one OpenAI-compatible endpoint against 200-plus models at provider list price with 0% markup added, so a provider price cut lands the same day and a swap from one data-generation model to another is a config edit rather than a second vendor relationship.
Automatic failover matters more than usual in a fine-tuning pipeline, because a batch job that dies at 80% wastes the whole run. One key across the generation models, with a fallback that reissues on a provider error, is a smaller piece of infrastructure than three API integrations.

Which one, decided by your intent
If you are training, take the Base. If you are inferring on formal logic and the licence permits your use, take TwIL-LM3-Pro. If you need an instruction-following small model today and neither restriction fits, use LFM2.5 2.6B's post-trained sibling — the model Liquid actually publishes for that purpose — and leave the Base for the fine-tuning you were planning anyway.


