A generated title card reading 'AREX-2 vs LFM2.5 2.6B Base - Two Kinds of Unfinished', with two panels. The left panel, headed LFM2.5 2.6B Base, lists: shipped August 2026; 2.69B params, dense; 131,072-token context; downloadable, no instruction tuning. The right panel, headed AREX-2, lists: repo created 29 Sep 2026; zero bytes stored; no model card at all; nothing to download. A footer reads 'One is a checkpoint without instructions. The other is a name without a checkpoint.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

AREX-2 vs LFM2.5 2.6B Base: An Empty Repo and a Checkpoint Without Instructions

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Set AREX-2 next to LFM2.5 2.6B Base and the first thing they share is that neither one is a product you can use today — for reasons that have nothing to do with each other. AREX-2 is a Hugging Face repository BAAI created on September 29, 2026 at 17:56 UTC; it contains a .gitattributes file, stores zero bytes of model data, and carries no model card, no licence tag and no announcement of any kind. LFM2.5 2.6B Base, published by Liquid AI in August 2026, is the opposite kind of unfinished: a complete, downloadable 2.69-billion-parameter text-only checkpoint under the LFM Open License (lfm1.0) that ships deliberately without instruction tuning, because it is meant to be fine-tuned rather than asked questions.

One is an absence of an artifact. The other is an artifact that is missing a step it was never supposed to have. Those are the same category only if you are skimming, and treating them as the same category is exactly how a team ends up waiting on the wrong thing. The useful comparison between these two is not capability — there is no AREX-2 to measure — but what each one is a raw material for, and what it costs to find out whether it is any good.

The difference in kind, stated first

LFM2.5 2.6B Base is a substance. It is the pretrained checkpoint of Liquid AI's LFM2.5 hybrid family: 30 layers, of which 22 are dual-gated short-convolution blocks and 8 are grouped-query attention, trained on roughly 34 trillion tokens across 16 languages, with a 131,072-token context window and a quantised build that lands near 1.67 GB at Q4_K_M. It has no instruction tuning. Feed it a question and it continues text; it does not answer. Liquid AI's own card points at heavy fine-tuning as the intended use, and there is a post-trained sibling for anyone who wants a model that responds to prompts. It is a substrate, and the fact that it scores badly on prompt-following benchmarks is not a defect — it is the definition of the thing.

AREX-2 is not a substance or a product; it is a namespace. The only facts available are the org (BAAI), the creation timestamp (29 September 2026) and the name. Nothing has been uploaded and nothing has been said. The name itself belongs to a family that does exist: on 23 July 2026 BAAI released AREX-Base, a 122-billion-parameter mixture-of-experts with 10 billion active parameters built on Qwen3.5-122B-A10B, and AREX-Turbo, a dense 4B built on Qwen3.5-4B — both Apache 2.0, both deep-research agents with a two-loop design that gathers evidence, produces a candidate answer with a confidence figure, then verifies that answer against the original constraints and refines or restarts. Its published numbers, vendor-reported and not independently reproduced, run to 82.5 on BrowseComp and 85.4 on GAIA for the Base, and 70.7 / 81.6 for the Turbo.

• Availability — LFM2.5 2.6B Base is downloadable now. AREX-2 has no files in its repository.

• Size — 2.69B dense parameters for the Liquid model, all visible in the checkpoint. Unknown for AREX-2; the family spans a 122B MoE and a dense 4B, so the "2" predicts nothing.

• Context — 131,072 tokens for LFM2.5 2.6B Base. Nothing published for AREX-2.

• Licence — LFM Open License v1.0, free below $10M annual revenue and requiring a separate licence at or above that line. No licence tag yet on AREX-2; the first AREX generation was Apache 2.0.

• What it is — a fine-tuning substrate versus, if the family is anything to go by, a hosted agent whose value is in a search-and-verify loop rather than in its weights.

A screenshot of the Hugging Face model card for LiquidAI/LFM2.5-2.6B-Base, showing the Liquid AI author line, tags for Text Generation, Transformers, Safetensors and 16 languages, a Model Details table listing LFM2.5-2.6B-Base at 2.6B parameters as a pre-trained base model for fine-tuning alongside the post-trained LFM2.5-2.6B, and the feature list beginning with total parameters 2.69B, 30 layers, a 34-trillion-token training budget, a 128,000 vocabulary and a 131,072-token context length.

Blank scorecards that mean opposite things

Neither model has a benchmark you should plan around, and the reasons diverge completely.

Liquid AI is not hiding anything by leaving its Base unscored. Scoring a pretrained checkpoint on instruction-following benchmarks would measure the absence of fine-tuning, not the quality of the substrate — which is why the company benchmarks the post-trained sibling and leaves the Base unevaluated. That blank is verifiable from your side, and that is the whole point: you can download 1.67 GB, run your own evaluation on your own task, and know the answer before you spend anything on serving.

AREX-2's blank is not a design decision, it is a not-yet. There is nothing to download, so there is nothing to test, and no evaluation you could run this week would tell you anything about it. Anyone who publishes AREX-2 benchmark numbers in the next few days is publishing something they cannot have measured.

A generated two-column card headed 'Two blank scorecards, for opposite reasons'. The left column, LFM2.5 2.6B Base, reads: no published evaluation - deliberate; scoring a pretrained checkpoint measures the missing fine-tuning; downloadable, verify it on your own task today; 1.67 GB at Q4_K_M, runs on a single card. The right column, AREX-2, reads: no published evaluation - nothing to evaluate; no weights, no card, no licence tag; not downloadable, nothing to verify; any figure quoted this week is not a measurement. A footer reads 'One blank is a design decision. The other is a not-yet.' The OrcaRouter logo is composited in the bottom-right corner.

Two different fine-tuning questions

Because both of these are starting points rather than finished services, it is worth being precise about what each would be a starting point for, since that is where they genuinely stop resembling each other.

A 2.69B hybrid checkpoint is a tool for making a small, cheap, specialised model. The interesting uses are the unglamorous ones: a classifier that reads a document type in a regulated workflow, an extraction model that turns a form into structured JSON, a domain-specific rewriter that runs on a single card close to the data. The value comes from the fact that you own the artifact afterwards and the marginal cost of a call is electricity. The training loop is the work, and it is work you can start today.

AREX-2 points the other direction. The AREX family is not designed to be fine-tuned into a product; it is designed to be called as an agent that runs searches, integrates evidence and verifies its own answers against constraints. If a second generation continues that, the thing you would be evaluating is a trajectory — how many steps it takes, whether it notices a contradiction, whether the confidence figure on its candidate answer means anything. That is not a fine-tuning question and it is not a weight question. It is a serving question, and it starts with somebody publishing weights and a card.

These are not substitutes at any size. A team that needs a model it can own and retrain is not waiting for AREX-2. A team that needs a research agent is not going to fine-tune a 2.6B hybrid into one.

Where the routing layer fits, for each of them

The practical difference between these two shows up the moment a model goes into an application, because the two have opposite relationships with hosting.

LFM2.5 2.6B Base is not on OrcaRouter's catalogue and no LFM2.5 variant is, so it comes from Liquid AI's own distribution and the usual third-party hosts — a local artifact rather than a routed endpoint. AREX-Base and AREX-Turbo are not on our catalogue either. What a routing layer is genuinely for in this picture is the other side of the architecture: the day you compare your fine-tune against the models it has to beat, you want that comparison to cost a config change rather than a procurement cycle. One API in front of 200-plus models, provider list price passed through with nothing added per token, one key for the whole panel, and failover so that a provider's bad afternoon does not become your evaluation's bad afternoon. Those are the conditions under which an A/B test on an unproven model is cheap enough to actually run.

That is also the honest framing for AREX-2 itself. When it ships — assuming it ships open weights, as its two predecessors did — the sensible way to try a brand-new agent on a real workload is on a side path, behind failover, not as the single provider your production loop depends on.

What a reasonable person does about each this week

If you have a small, narrow task and data that cannot leave your infrastructure, the Liquid checkpoint is available, small, permissively licensed below a revenue threshold, and testable this afternoon. Download it, run it on your own evaluation set, and let the result decide. Nothing about AREX-2 changes that plan.

If you need long-horizon research behaviour today, the model to reach for is the one that already exists: AREX-Base, shipped in July, with published agent benchmarks you can at least check against a paper. AREX-2 is a name with a timestamp, and the two things that would change its status are visible from outside — whether files appear in the tree, and whether a card with a parameter count and a licence appears with them.

The failure mode this article exists to prevent is the one where a reserved name gets treated as a roadmap item. It is not. It is a repository that BAAI made yesterday, and the correct amount of planning to do around it right now is none.