A generated title card reading 'Clef vs LFM2.5 2.6B Base', subtitled '55 GB on an H200 against 5.4 GB on a laptop', with chips reading '27B decision model' and '2.69B pre-trained base'. The OrcaRouter logo is composited in the bottom-right corner.
Engineering & Research

Clef vs LFM2.5 2.6B Base: 55 GB on an H200, or 5.4 GB on a Laptop

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Here is a comparison where the hardware decides the argument before the models do. Cloudflare/clef is a 27-billion-parameter multimodal decision model, post-trained from Qwen3.8-27B, whose repository measures a little over 55 GB across twelve backbone shards plus a small joint schema head. LiquidAI/LFM2.5-2.6B-Base is a 2.69-billion-parameter text-only checkpoint whose weights are a single 5.4 GB file, published by Liquid AI on 1 August 2026 under the LFM Open License. Cloudflare's own published latency for Clef — 209.3 ms median, 238.6 ms p95 — was measured on a single H200. Liquid AI's intended deployment for its LFM2.5 family is a laptop, a phone, or a Ryzen handheld. Both vendors describe their model as small, fast and efficient, and both are telling the truth about a different machine.

There is a second gap behind the first, and it is the one that actually changes plans. Neither Clef nor LFM2.5 2.6B Base answers a question out of the box in the way you probably expect. Clef arrives fully trained for a task it cannot deviate from: it answers typed questions, and it will not do anything else. The Liquid checkpoint arrives untrained for anything at all — it is a base model, deliberately without instruction tuning, and Liquid AI's own card says it is recommended only for heavy fine-tuning. So the real question is not which one is cheaper to run. It is whether you are buying a finished component or a starting material, and the answer is different for each of them.

Where each one runs, and why that is the whole comparison

Deployment shape is not a side note for these two models. For a decision model it is the architecture, because a 209 ms forward pass and a "runs on the machine in front of you" story are the same claim told from different ends.

• Parameters — 27B for Clef, with the Qwen3.8-27B vision encoder kept; 2.69B text-only for LFM2.5 2.6B Base.

• Disk — a 55 GB repository for Clef; one 5.4 GB safetensors file for the Liquid checkpoint, with quantised GGUF, ONNX and MLX builds published alongside it.

• Hardware — Cloudflare tested Clef with torch 2.11 and transformers 5.10.2 on a single H200, and its latency figures come from that run. Liquid AI's post-trained sibling of this checkpoint runs at 220 tokens per second on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU, in under 2.5 GB of memory, and the vendor notes 30 tok/s is achievable on a phone.

• Context — 65,536 tokens for Clef, per the Workers AI model page and the launch post; 131,072 tokens for LFM2.5 2.6B Base.

• Input — Clef reads text, JSON, images and video frames and scores them jointly. The Liquid checkpoint is text-only.

• Licence — Apache-2.0 for Clef. LFM Open License v1.0 for LFM2.5, which is source-available rather than OSI open source: commercial use is conditioned on your organisation staying under $10 million in annual revenue, and above that line the licence simply does not grant it. That threshold is a procurement fact, not a footnote.

A two-column comparison scoreboard titled 'Clef vs LFM2.5 2.6B Base — the scoreboard'. The left column 'Clef' reads: Parameters: 27B multimodal; On disk: 55 GB repository; Context: 65,536 tokens; Output: probability per option, no text; Hardware: H200 class, 209.3 ms median; Licence: Apache 2.0. The right column 'LFM2.5 2.6B Base' reads: Parameters: 2.69B text-only; On disk: 5.4 GB single file; Context: 131,072 tokens; Output: untuned text continuation; Hardware: laptop, 220 tok/s on M5 Max; Licence: LFM 1.0, $10M revenue threshold. A footer reads 'Clef figures per Cloudflare; LFM figures per Liquid AI's cards. Neither independently reproduced.' The OrcaRouter logo is composited in the bottom-right corner.

Cost per decision is not the same question as cost per token

The tempting move here is to divide two prices and declare a winner. That does not survive contact with what the two models do.

Clef is $0.24 per million input tokens on Cloudflare's Workers AI. Its output is not prose, so the usual output-token line does not exist; you pay for state, once, and get a distribution back. For a routing layer making millions of small decisions a day, that number is the whole operating cost, and it is a number you can only get from a vendor rather than from your own electricity bill.

LFM2.5 2.6B Base costs nothing per call and cannot make a decision. To get a cost-per-decision out of it you first have to fine-tune it on your task, which means assembling data, running a training job, evaluating the result, and only then discovering what it costs you in kVA and engineering time. The attractive economics of a 2.6B checkpoint are real, but they are downstream of work that has not happened yet, and the person comparing per-token prices has skipped straight past it.

The honest framing is that these are two different purchases. Clef is a component with a list price and a hosting bill. The Liquid checkpoint is a raw material whose cost is the training run you are going to pay for either way — and whose payoff is that afterwards you own the artifact outright, at which point the marginal cost of a decision really is electricity. For a narrow, high-volume, stable task, that trade is often correct. For a task you have not yet specified, it is backwards, because you cannot fine-tune toward a target you cannot describe.

An untrained base is not a worse model, it is a different starting line

It is tempting to read the Liquid checkpoint's absent benchmark table as a hidden weakness. It is not. Scoring a pretrained base model on instruction-following benchmarks would measure the absence of fine-tuning, not the quality of the substrate, which is precisely why Liquid AI benchmarks the post-trained LFM2.5-2.6B and leaves the base unevaluated. The blank is verifiable from your side, and that is the point: download 5.4 GB, run your own evaluation on your own task, and know the answer before you commit to serving anything.

Clef's table is the opposite shape of evidence and has the opposite problem. Cloudflare publishes 40-odd benchmark rows, a full workflow evaluation set and a latency distribution, and every one of them was produced by Cloudflare on Cloudflare's own Decision Index, with no third party having reproduced any of it. The Clef column is far more informative than the Liquid column, and not one figure in it has been checked by anyone else. That is a stronger claim than a blank, and a weaker one than it looks.

What you can measure yourself splits the same way. With the Liquid checkpoint, you can measure everything — it is 5.4 GB on your own disk. With Clef, the Apache-2.0 weights are equally yours to download and run, but matching Cloudflare's 209 ms median requires the class of GPU Cloudflare used, so in practice most teams will measure it through the hosted endpoint and inherit the vendor's availability along with its latency.

Where the routing layer fits, for each of them

Neither Clef nor any LFM2.5 variant is a route on OrcaRouter, and this article is not an availability claim — the catalogue returns a 404 for both, so Clef comes from Cloudflare's Workers AI or your own GPU and the Liquid checkpoint comes from its own distribution.

What a routing layer is genuinely for here is the pattern both models imply. A small bounded scorer that decides, and a larger generalist that acts on the decision, is a two-model architecture by construction — and Clef's own backbone makes the second half of that concrete, since the model Cloudflare post-trained from, Qwen3.8 27B, is a listed route at $0.33 per million input and $2.40 per million output tokens over a 262,144-token context. One endpoint in front of 200-plus models means the cheap decider and the capable actor sit behind the same key, composed into a single call through the routing DSL rather than wired together by hand, with automatic failover so a provider's bad afternoon does not take the decision path down with it. If instead you are self-hosting the Liquid checkpoint and comparing your fine-tune against the models it has to beat, the same endpoint is what makes that comparison a config change instead of a procurement cycle.

TypeSafe's Jev 1.13 is worth naming for the self-hosting case specifically: it is the decision model Cloudflare benchmarked against and deliberately speaks the same POST /v1/systemone request shape as Clef, it is a listed route at $0.042 per million input tokens over a 65,536-token context, and it is the low-effort way to find out whether a bounded decision model helps your pipeline before you spend a training run on a substrate that may not need to exist.

A headless capture of the OrcaRouter model page for qwen/qwen3.8-27b, showing the Qwen breadcrumb, the Qwen3.8 27B title with a 262K context marker, the text, image and video input modalities, and the site's Code samples, Pricing, Performance, Public benchmarks and FAQ navigation tabs.

Which one, for what

Take the Liquid checkpoint when the task is narrow, high-volume, specified well enough to write training data for, and bound by one of the two hard constraints a routed API cannot fix: the data cannot leave your infrastructure, or the machine it has to run on is the one on your desk. The LFM 1.0 revenue threshold is the thing to check first, because it decides whether the licence is even available to you.

Take Clef when the decision is already enumerable, the state fits in 64K, the answer needs a calibrated probability rather than a sentence, and 209 ms in a hot path is the property you are buying. Its fixed schema is the feature — an auditable, reproducible, parser-free output shape — and its 55 GB footprint is the price of the backbone that makes it accurate on the first try instead of after a training run.

What you should not do is pick by parameter count. A 2.69B model that has to be trained before it can answer is not a smaller version of a 27B model that already can, and the two numbers in the title — 55 GB and 5.4 GB — describe two different budgets, not two points on one scale.

A headless capture of the LiquidAI/LFM2.5-2.6B-Base model card on Hugging Face, showing the model title, the TextGeneration and Transformers and Safetensors tags, a 16-languages marker, the Liquid AI organisation, and the opening line describing the LFM2.5 family as hybrid models designed for on-device deployment.

The open question, and it is the same one for both

Neither vendor has published evidence that survives independence. Cloudflare's Decision Index run is self-administered and unreproduced; Liquid AI's throughput figures are the vendor's own measurements on hardware it chose. The LFM2.5 2.6B Base has no benchmark at all by design, and the Clef table has a great many by choice.

For the Liquid checkpoint the missing evidence is cheap to fix and entirely in your hands — the file is small enough to evaluate on a laptop in an afternoon. For Clef it is not: reproducing the 209 ms median takes the hardware Cloudflare used and the state nobody outside Cloudflare has assembled. That asymmetry, more than anything in either specification, is what should shape which of these two you plan around this month.