A hero title card for Liquid AI's open-weight decision models, headed 'Two models that never write a word' with the subheading 'Liquid AI d1-3B and d1-omni-600M, open weights, October 7 2026', three chips labelled yes/no, pick a label and rate a scale, and a diagram of one state entering a single forward pass that returns a probability, a label and a score.
Guides & Insights

Liquid AI d1-3B and d1-omni-600M: Open Weights for Models That Never Write a Word

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Liquid AI d1-3B and Liquid AI d1-omni-600M went up on Hugging Face on October 7, 2026, and the two checkpoints have a property that makes every comparison table you have read this year the wrong shape. Neither one generates text. You hand Liquid AI d1-3B a state — a support ticket, a JSON payload, a photograph, or all three at once — and a set of named questions, and it returns answers plucked straight off the model's distribution over your options. Yes/no as a probability. A choice among labels with a confidence and a full probability vector. A position on an ordered scale. There is no sampling step, no JSON to parse, no runaway generation, and the vendor's own usage counter reports the thing plainly: output_tokens: 0. The 3B model is the headline — it posts 48.57 on the vendor-scored Decision Index 0.2.1, ahead of everything under 10B and ahead of a decision model twelve times its size. The 600M sibling is the one worth watching: it reads images and up to thirty seconds of speech through the same trunk weights, and it is labelled an early research release.

What a decision model is, in one paragraph

The category has been building for a year under names like system-one models and classifiers-that-are-not-classifiers, and the d1 pair is the cleanest statement of it so far. A generative model is asked a question and writes an answer; the answer has to be parsed, the parse can fail, and every token it writes costs you money and time. A decision model is asked a question and reports what it already believes. Liquid's questions come in three types. noul is a yes/no question, answered with P(yes) between 0 and 1. choice picks one label from a named set and returns the label, a confidence, and the probability of each option. score places the state on an ordered scale of two to ten levels and returns the expected level with its distribution and legend. One state can carry several questions at once, and the state and its images are read once for all of them.

That last property is the whole business case. Three questions about one ticket cost 1.3x the time of one question, not 3x, because the model does one forward pass over the input and reads several answers out of it. There is nothing to write, so there is nothing to write badly.

The 3B and the 600M are built from opposite directions

This is where the release gets technically interesting, and it is the part the launch coverage has mostly skipped. The two models do not share a lineage.

Liquid AI d1-3B descends from LFM2.5-VL-3B, the vendor's decoder-only vision-language model. It carries 3.12B parameters, a SigLIP2 NaFlex shape-optimized 400M vision encoder, a 32,768-token context window, a 128,000-token vocabulary, and sixteen documented languages. To build it, Liquid averaged the weights of LFM2.5-2.6B and the text backbone of LFM2.5-VL-3B, fine-tuned checkpoints from several random seeds and data mixtures, then merged the results again. The vendor's own summary of what mattered is refreshingly unglamorous: training on long inputs, shuffling the order of answer options, and hunting down shortcuts in the training data did more for the final number than any advanced technique.

Liquid AI d1-omni-600M descends from LFM2.5-Encoder-350M, a bidirectional encoder — a different species of network entirely. It totals 587M parameters split into a 381M shared trunk and decision head, a 94M vision encoder borrowed from LFM2.5-VL-450M, and a 112M audio encoder built from a 17-layer FastConformer. Every modality runs through the same trunk weights. Its context is 16,384 tokens covering text, image and audio positions together, and with images attached the state and question text is cut to 896 tokens because that is what it was trained on. Audio gets one warning the card states without hedging: the capability was trained on requests between an English speaker and an assistant, and clips are cut at thirty seconds.

So the family is not one model at two sizes. It is a decoder and an encoder, post-trained to do the same job with two completely different mechanisms, one of which sees and one of which hears.

A two-column scoreboard for Liquid AI d1-3B and d1-omni-600M showing d1-3B at Decision Index 48.57 with 3.12B parameters, text and image input and 8 ms per question on an RTX 4090, against d1-omni-600M at 15.95 with 587M parameters, text, image and audio input and no reported latency, footed 'All figures vendor-reported by Liquid AI, Oct 7 2026; no independent reproduction.'

The numbers, and whose they are

Every figure in this section is Liquid's own. The Decision Index rows for the d1 models were scored by the vendor using the official scorer — not submitted to the leaderboard — while every competing row comes from the public leaderboard at v0.2.1. No independent lab has reproduced any of it yet, and the models are two days old as this is written.

• Decision Index 0.2.1 — Liquid AI d1-3B scores 48.57 with sub-scores of Knowledge 23.8, Language 56.4, Retrieval 52.8, Tools 74.5 and Arts 36.3. Liquid AI d1-omni-600M scores 15.95, with Tools at 15.1.

• Where 48.57 sits — first among everything under 10B in the table the vendor published, and ahead of Decider 35B-A3B at 47.11. It is second overall, behind Winnow-12B at 50.02, which is four times the size.

• Text benchmarks, vendor-reported — d1-3B averages 82.9 across seven public benchmarks, ahead of Decider 4B's 81.1; d1-omni-600M averages 78.4, which beats Decider 2B's 77.1 at roughly a quarter of the parameter count. The 600M posts the table's best toxicity score (Civil Comments 95.8) and its best paraphrase score (PAWS-X 79.5).

• Vision, vendor-reported — d1-3B averages 74.1 across eleven public image benchmarks read as decisions over each benchmark's options, against 73.9 for its LFM2.5-VL-3B backbone. Removing the images drops the same questions to 45.1, which is how the vendor shows the answers are coming from the pixels.

• Latency, vendor-measured — one question takes 8 ms on an RTX 4090 and 9 ms on an AMD MI325X; 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin 64 GB, 50 ms on the smallest Jetson Orin Nano, and 30 ms on an Apple M5 Pro. Sixty-four states packed into a single pass run at 475 per second on the 4090.

• What is missing — d1-omni-600M has no inference table at all. Liquid says it is under active development and simply does not report latency for it. There is also no audio decision benchmark anywhere in the release, because, in the vendor's words, dedicated audio decision benchmarks are currently an open problem.

The claim the release is really making

Two days before the weights went up, Liquid published the hosted version of this model and ran it against GPT-6.1 Sol and Claude Opus 5.5 on six applications. Its summary: d1 matches or beats GPT-6.1 Sol on four of the six, costs 19x to 200x less, and answers faster on every task. The published methodology is worth reading before repeating any of that — each application was run once per model on October 5, 2026, the chat models answered in JSON at their default reasoning setting, and d1's cost was computed at $0.04 per million input tokens.

The demos are the more persuasive half. Sorting good and defective parts from four production lines — circuit boards, candles, cashews, gum — at 85 to 97% accuracy, on a model that was never trained for the task and understood it from a short description. A SQL predicate that answers yes or no for each of 150 support tickets. A context-compaction loop that reads each tool output in a coding agent's session and keeps, trims or drops it, removing 52% of the tokens while keeping every output the task needs. In Tetris, adding the screen to a fully text-describable game raised the score from 70 cleared lines to 81 — a vision result you cannot get from a text-only classifier no matter how good it is.

Those are vendor-run demonstrations with vendor-chosen tasks, and the methodology section says so. They are still the clearest picture anyone has published of what a decision model is for.

A capture of Liquid AI's own blog post 'Open d1: Edge decision models for text, vision, and audio' dated Oct 7, 2026, showing the opening paragraph that announces d1-3B and d1-omni-600M as open-weight models, d1-3B's Decision Index score of 48.57, and its latency figures of 8 ms on an RTX 4090, 16 ms on a Jetson AGX Thor and 26 ms on a Jetson AGX Orin.

Where it runs, and what it costs to call

The open weights are built for local execution and the vendor did the integration work up front: day-one support for llama.cpp, the full NVIDIA stack from DGX down to Jetson, NVFP4 quantization, and an 8-bit weight/activation variant published alongside the base checkpoint. GGUF conversions are already on Hugging Face. If you would rather not host anything, Liquid serves the hosted d1 from its own API — input tokens only, no output tokens, with images billed as input at 1.5 tokens per 32×32-pixel patch, which makes a 1024×1024 image 1,536 tokens. Text-only access is also available through third-party platforms.

The honest routing note: neither Liquid AI d1-3B nor Liquid AI d1-omni-600M is on our catalogue, and this article is not an availability claim for either one. What the d1 pair changes for a router is the shape of the pipeline around it. A decision model that answers in milliseconds and costs nothing in output tokens is a filter, not a replacement — the good-and-defective sort and the ticket triage happen in front of the generative call, and the generative call is where a single endpoint across 200+ models, list prices passed through at 0% markup, and automatic failover actually earn their keep. Different layer, same pipeline.

What would settle the argument

Three things are unresolved, and all three are the kind that only time closes.

The first is independent reproduction. The Decision Index numbers were produced by the vendor with the official scorer and every competitor row was lifted from a public leaderboard, which is a defensible method and not the same thing as a third party running your model. The second is audio. A 600M checkpoint that routes a voice command in one forward pass is a genuinely new capability, and there is no benchmark in existence that measures whether it is doing that well. The third is the category itself: when the answer is read off a distribution rather than written out, the usual failure modes of a small model — looping, drifting, refusing, hallucinating a format — simply cannot occur, and nobody has published a good account of what replaces them. Calibration error is the obvious candidate, and it is exactly what the confidence field on every answer invites you to measure yourself.

Ten demos run Liquid AI d1-3B in a loop over live camera input in a public Hugging Face space, which is the cheapest way to form your own opinion. The weights are free and the questions are specific. That is a better starting position than most releases this year offer.

A capture of Liquid AI's documentation page for its decision models, showing the left-hand navigation including Decision Models and Liquid Nanos, the 'Open d1' banner announcing that visitors can now run d1-3B and d1-omni-600M locally, and the page's opening description of purpose-built models for structured decisions such as classification, routing and scoring in a single call with zero generation.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily