A generated title card reading 'FrogNano-4B-2609 vs LFM2.5 2.6B Base' with the subtitle 'a finished 4B agent against a 2.69B pretrained checkpoint', three chips reading 'RL post-training finished', 'pretraining only, 2.69B dense in 30 layers' and 'no instruction tuning', a footer reading 'Both columns vendor-reported; no third party has scored either model', and the OrcaRouter logo composited in the bottom-right corner.
Engineering & Research

FrogNano-4B-2609 vs LFM2.5 2.6B Base: a Finished Specialist and a Bag of Pretrained Weights

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Type a question into LFM2.5 2.6B Base and it will continue your sentence instead of answering it. That is not a defect — Liquid AI published it as the pre-trained checkpoint underneath the LFM2.5 family, with instruction tuning deliberately left to its non-Base sibling, and the card says plainly that it is meant for heavy fine-tuning. Microsoft's FrogNano-4B-2609 is what the other end of that pipeline looks like: the same kind of small language model, put through five iterations of reinforcement learning on synthetic repository tasks until it emits structured tool calls and candidate patches through a five-tool harness. Neither one has a published independent benchmark. The difference is that one of them has had its post-training finished, and knowing which is the entire decision.

The FrogNano figures below come from Microsoft's model card and technical report. The LFM figures come from Liquid AI's card, its architecture config, and its licence file. Every number attributed to either vendor is labelled as vendor-reported, and neither of these models has been through a third party's evaluation — for LFM2.5 2.6B Base that is largely by design, since a checkpoint with no instruction tuning is not a fair target for a chat benchmark.

The comparison only works if you know what you are comparing

Put the two spec sheets side by side and the first thing that goes wrong is the parameter count. LFM2.5 2.6B Base is 2.69 billion dense parameters in 30 layers — 22 double-gated short-convolution blocks and 8 grouped-query-attention layers, trained on roughly 34 trillion tokens across 16 languages, with a 128,000-token vocabulary and a 131,072-token context. Its quantised build lands around 1.67 GB in Q4_K_M.

FrogNano-4B-2609's card gives its parameter field as the band "500M-5B", and the model summary describes it as approximately 4.66 billion. The download settles the argument: two safetensors shards totalling 9.32 GB of BF16 weights, 738 tensors, on an inherited dense 32-layer hybrid Gated DeltaNet and gated-attention stack. A 24-layer vision tower sits in the checkpoint and was never post-trained.

So the size ratio is roughly 1.7 to one, and the memory ratio is closer to 5.6 to one once quantisation is on the table. That gap matters more than the parameter line, because both of these models are self-hosted prospects rather than endpoints — which is the only reason they belong in the same article.

• Intended use — LFM2.5 2.6B Base is raw material for a fine-tuning run. FrogNano-4B-2609 is a finished policy for repository-level software engineering inside the Leaf harness.

• Instruction following — LFM2.5 2.6B Base has none out of the box; Liquid AI's card recommends it for tasks requiring heavy fine-tuning. FrogNano-4B-2609 follows a task description and emits structured tool calls, because that is exactly what its RL objective trained.

• Context — 131,072 tokens for LFM2.5 2.6B Base, against roughly 131K combined tokens for FrogNano's evaluated configuration. Nominally identical, but FrogNano's window has to hold tool results, shell output and test logs, not just a prompt.

• Output ceiling — FrogNano's validated configuration allows 8,192 generated tokens per assistant turn. LFM2.5 2.6B Base publishes no equivalent, since it is not built to end an answer.

• Modality — both are text only. FrogNano's vision components are inherited and explicitly unsupported; LFM2.5 2.6B Base is text-in, text-out by construction.

• Licence — LFM2.5 2.6B Base sits under the LFM Open License v1.0, which is free below $10M in annual revenue and needs a separate licence at or above that line, with derivative works inheriting the cap. FrogNano's card says MIT in its front matter and Apache 2.0 in its body.

The thing Microsoft paid for, and the thing you would have to pay for yourself

This is the load-bearing section, because it is the one where a spec comparison misleads.

FrogNano-4B-2609's entire value-add is post-training, and Microsoft has published the method. Start from Qwen3.5-4B, which scored 39.4% on SWE-bench Verified through the Leaf harness as a general model. Generate candidate repository tasks from real snapshots, run rollouts from the current checkpoint to find tasks it sometimes solves, keep those, train, and repeat — five iterations, roughly 1,500 validated synthetic environments in total, and no distillation from a larger model at any point. The result on Microsoft's own card is 61.5% on SWE-bench Verified, 37.6% on SWE-bench Pro, 31.1% on Terminal-Bench 2.0 and 47.3% on PatchEval-Verified. The paper's efficiency appendix reports the same climb as 48.2%, 53.4%, 58.3%, 58.6% and 61.6% across iterations, which is a useful reminder that even the lab's own two tables of the same experiment do not agree on the rungs.

Now read that against LFM2.5 2.6B Base. It is the checkpoint before that work begins. Its card's own recommendation — fine-tune it — is precisely the job FrogNano documents: a task environment, an executable reward, a longer-serving harness, and several training iterations against a moving policy. The paper does not claim that is easy. It reports that intermediate checkpoints became progressively more expensive to train as reasoning traces lengthened, that a log-length penalty had to be added to stop the drift, and that later iterations lost the parallel tool-call ability the base model had, ending at a 1.71% concurrent-call rate. Anyone planning to run the FrogNano recipe on a different 2.6B base should budget for those three problems specifically.

What LFM2.5 2.6B Base buys is a different kind of head start: a 34-trillion-token pretraining budget, a documented hybrid architecture, an existing quantised build, and a documented context of 131,072 tokens, all at a size that runs on hardware a 9.32 GB BF16 checkpoint would not fit on. If your task is narrow enough that a small fine-tune beats a general model, that is a real position — and if it is not, you have saved yourself a training run.

A generated two-column scoreboard titled 'FrogNano-4B-2609 vs LFM2.5 2.6B Base'. The left column reads State RL post-training complete, Params approx 4.66B with the card saying 500M-5B, Weights 9.32 GB BF16, Context about 131K evaluated, Licence MIT in front matter and Apache 2.0 in body, Independent scores none. The right column reads State pretrained checkpoint only, Params 2.69B dense, Weights 1.67 GB in Q4_K_M, Context 131,072 tokens, Licence LFM Open License v1.0, Independent scores none. A footer reads 'Both columns vendor-reported; no third party has scored either model.'

What you have to build either way

Neither of these is a network call, and that shapes the practical comparison more than any spec row.

LFM2.5 2.6B Base wants an inference runtime and a training run. Liquid AI's card lists native-format, GGUF, ONNX and MLX builds, so the serving half is well covered — llama.cpp on a laptop is a real option for the quantised checkpoint. The training half is yours: data, objective, evaluation, and a way to tell whether the result got better or just different.

FrogNano-4B-2609 wants an OpenAI-compatible endpoint with the right parsers and a Kubernetes cluster with permission to manage pods and network policies. Microsoft's README is unusually direct that the harness is a dependency rather than a convenience, and the card states that identical scores require a matching checkpoint, tokenizer, serving configuration, task images and evaluation protocol. Configuration is not a detail here — serve it without a Qwen3 reasoning parser and a Qwen3 coder tool-call parser and the model writes prose where the harness expects a tool call.

That is the shape of this matchup: on one side, a training project with a small serving problem; on the other, a serving project with a large evaluation problem and no training left to do. Both are evenings of work. Only one of them is evenings of work that can end in "the model is not good enough" through no fault of yours.

A screenshot of the Hugging Face model card for microsoft/FrogNano-4B-2609 showing the model summary table with the Parameters row reading 500M-5B, the Context length row reading approximately 131K tokens in the evaluated coding-agent configuration, the Training Dates row reading Jun 2026 to Aug 2026, the Release date row reading 22-SEP-2026, the License row reading Apache License 2.0, the Qwen/Qwen3.5-4B model dependency, and the model overview naming the dense 32-layer hybrid Gated DeltaNet and gated-attention architecture.

Where a router earns its place in a build like this

Both of these models end up inside a pipeline that also calls something hosted. FrogNano's own evaluation rig is the proof: the Leaf harness talks to an OpenAI-compatible endpoint, which means the model under test and any model you compare it against are reached through the same interface. That is a pattern worth copying even when the reason is only hygiene, and it is where a single endpoint does something concrete.

OrcaRouter carries 200-plus models behind one OpenAI-compatible key at 0% markup, so provider list prices pass through untouched and a vendor price change is live on our side the same day — which is the figure that actually moves when you are deciding whether a self-hosted specialist beats a hosted general model on cost per task. Automatic failover covers the hosted half of that comparison, not the checkpoint you are serving yourself. We do not host FrogNano-4B-2609 or LFM2.5 2.6B Base; both are downloads. What a routing layer removes is the second integration every time the hosted side of your benchmark changes.

Picking one, honestly

Choose LFM2.5 2.6B Base if your task is narrow, your hardware is small, and you are prepared to own the training run — the licence is free below $10M in revenue, the quantised build is 1.67 GB, and Liquid AI's own card recommends it for exactly this. The trade is that you get a starting point, not a capability: nothing in the release tells you what a FrogNano-shaped RL run on top of it would score, and nobody has published one.

Choose FrogNano-4B-2609 if the task is repository-level software engineering and you want the post-training already done. The 61.5%, 37.6%, 31.1% and 47.3% are genuine published results from a lab that described its method in detail — and they are also, at this moment, a closed loop: same lab, same harness, same base-model baseline. The 9.32 GB download, the harness dependency and the compound licensing are the price of skipping a training project you would otherwise have to run.

A generated pipeline card headed 'The same pipeline, two positions' showing three stages connected by arrows: 'Pretrained checkpoint' captioned LFM2.5 2.6B Base, 'RL on synthetic tasks, 5 iterations' captioned Microsoft's loop, and 'Finished agent' captioned FrogNano-4B-2609, with a footer reading 'Neither model has been independently evaluated; both appear as vendor-reported figures only.'

The useful reframe is that these are not competitors. LFM2.5 2.6B Base is upstream of a process that produced something like FrogNano-4B-2609. If you find yourself choosing between them, the real question is whether your problem is worth owning a training run for — and if the answer is no, the 4B checkpoint with the harness, the published method and the vendor-reported SWE-bench ladder is the one that arrives finished.