
Liquid AI d1-omni-600M vs Liquid AI d1-3B: Which Half of the d1 Family Do You Actually Need?
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 56 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 56 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 348 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Liquid AI d1-omni-600M and Liquid AI d1-3B were uploaded to Hugging Face within eight hours of each other on October 5, 2026 and released together in the same October 7 announcement, which makes the usual question — which one is newer, which one is better — the wrong one. They are the two ends of a deliberate trade. Liquid AI d1-3B is the finished article: 3.12B parameters, a 48.57 on the vendor-scored Decision Index 0.2.1, benchmark tables, latency measured down to a Jetson Orin Nano, and a place described in the release post as the highest decision quality at its size. Liquid AI d1-omni-600M is the experiment: 587M parameters, a 15.95 on the same index, audio input that the 3B does not have, and a model card that says outright it is an early research release with no inference numbers because it is still under active development. Picking between them is not a quality decision. It is a decision about whether you need the extra modalities at the bottom of the family or the extra accuracy at the top, and the numbers line up behind that split rather than blurring it.
Everything below comes from the two model cards and the October 7 release post, with the release post's own labelling respected: the d1 rows on the Decision Index were scored by Liquid AI with the official scorer rather than submitted to the public leaderboard, and nothing here has been independently reproduced.
Two backbones that were never going to converge
The d1 family did not scale a single recipe down. The two checkpoints start from opposite ends of Liquid's model catalogue and meet in the middle.
Liquid AI d1-3B is built on top of LFM2.5-VL-3B, the vendor's decoder-only vision-language model from August 2026. Its base was made by averaging the weights of LFM2.5-2.6B with the text backbone of LFM2.5-VL-3B, then fine-tuning checkpoints under different random seeds and data mixtures before merging them again. It carries a SigLIP2 NaFlex shape-optimized 400M vision encoder, a 32,768-token context, a 128,000-token vocabulary, and sixteen documented languages.
Liquid AI d1-omni-600M comes from the other direction. Its trunk is LFM2.5-Encoder-350M, a bidirectional encoder, fine-tuned on decision tasks first and then extended in stages — a 17-layer FastConformer encoder plus adapter for audio, with the audio encoder later fine-tuned against a frozen text backbone, then a SigLIP2 tower lifted from LFM2.5-VL-450M with an adapter and LoRA updates to the backbone for vision. The final model was merged from the LoRA updates and averaged with the previous checkpoint. It ends at 587M parameters total: a 381M shared trunk and decision head, a 94M vision encoder and a 112M audio encoder.
Decoder-only versus bidirectional is the part to hold on to. The 3B reads a state and produces a decision the way a language model produces a token sequence, one direction at a time. The 600M reads the whole state at once and decides, which is what you would expect from an encoder that was never built to generate. Both are trained to report answers off the model's distribution with zero output tokens, but the machinery underneath is not the same class of model, and the accuracy gap below is the visible cost of the smaller, encoder-shaped design.
The Decision Index spread is large, and the sub-scores are more interesting than the total
On Decision Index 0.2.1, Liquid reports 48.57 for Liquid AI d1-3B and 15.95 for Liquid AI d1-omni-600M, against 50.02 for Winnow-12B. That is a 32-point gap between two checkpoints released the same day by the same lab, and reading the five sub-scores explains where it comes from.
• Knowledge — 23.8 for Liquid AI d1-3B against 8.3 for Liquid AI d1-omni-600M
• Language — 56.4 against 12.9
• Retrieval — 52.8 against 35.0
• Tools — 74.5 against 15.1
• Arts — 36.3 against 6.8
Retrieval is the one place the small model holds its ground, losing under a third of the 3B's score where the other four categories it loses 60 to 80 percent. That pattern is consistent with what the 600M is: a trained encoder with real representational capacity for matching a state against content, and much less of the layered capability that the 3B inherits from a decoder that was pretrained on far more language. If your workload is a retrieval-shaped decision — does this passage answer this question, which of these documents is relevant — the 600M's profile is less bad than its total suggests. If your workload is a tool-routing decision, the 51-point gap in that column is the number to stare at.
The text benchmark table tells a milder story than the index, which is worth knowing before either number is used to make a case. On seven public benchmarks the 3B leads with a mean of 82.9 and the 600M reaches 78.4. The 600M actually loses narrowly on SQuAD 2.0 (74.0 to 85.3), PubMedQA (61.3 to 66.0), BoolQ (77.7 to 86.7) and XNLI (74.7 to 85.0), but it wins on Civil Comments toxicity detection (95.8 to 93.0) and on PAWS-X paraphrase identification (79.5 to 76.9). Liquid's own framing is that the 600M beats Decider 2B's 77.1 mean at a quarter of the parameters. Two benchmark suites, two different apparent verdicts, both vendor-reported — that is what the evidence supports and no more.

What the 600M has that the 3B does not
The reason to tolerate a 32-point index gap is that Liquid AI d1-omni-600M does one thing Liquid AI d1-3B cannot, and it is not a fidelity difference.
• Audio — Liquid AI d1-omni-600M takes up to 30 seconds of speech per request through its FastConformer encoder; Liquid AI d1-3B takes none
• Modality mixing — the 600M accepts text with images or text with audio, and raises a ValueError if both arrive together; the 3B takes text and images
• Context window — 16,384 tokens across text, image and audio positions for the 600M, with text trimmed to 896 tokens when images are present; 32,768 tokens for the 3B
• Vocabulary — 65,536 for the 600M, 128,000 for the 3B
• Precision — the 600M card recommends float16 on GPU and warns that bfloat16 changed the top answer on some rows; the 3B ships 15 quantizations including an w8a8 build
• Languages — the 600M lists 16 languages in a different set from the 3B's 16, and its audio is described as trained on exchanges between an English speaker and an assistant, which is a narrow slice of what a production audio feed contains
The audio training note is easy to skim past and should not be. A model trained on English speaker-to-assistant exchanges has seen one speaker geometry, one turn structure, and one accent distribution. Deploying it on call-centre audio or field recordings is asking for behaviour the card does not claim, and the release post is candid that no audio decision benchmark exists to check it against — Liquid calls that "currently an open problem" and invites the community to build one.

Latency: one sibling has the tables, the other has a footnote
For decision models the interesting number is end-to-end latency, because there is no decoding to time. Liquid publishes a full set for the 3B and none for the 600M.
• One question — 8 ms on an RTX 4090, 9 ms on an MI325X, 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin 64 GB, 50 ms on an Orin Nano, 30 ms on an Apple M5 Pro
• Three questions over one state — 21 ms on the RTX 4090 and 20 ms on the AGX Thor, roughly 1.3x the cost of a single question rather than 3x
• A 3.4K-token state — 102 ms on the 4090, 220 ms on the Thor, 1,640 ms on the Orin Nano
• Packed throughput — 475 decisions per second on the RTX 4090, 1,106 per second on the MI325X
• A 384px image — 17 ms on the 4090, 18 ms on the MI325X
Those figures describe Liquid AI d1-3B only. For Liquid AI d1-omni-600M the model card states that inference numbers are not reported because the model is an early research release under active development. It is not that the small model is slower — the opposite is nearly certain, since a fifth of the parameters does not get slower at the same precision — it is that no figure exists, and quoting the 3B's milliseconds for the 600M would be a fabrication with a plausible shape. What can be said without inventing anything is that at the float16 precision the card recommends, 587M parameters is on the order of 1.2 GB of weights before activations, which is arithmetic on a published parameter count rather than a measurement.
The cascade is the real answer for most workloads
Because both checkpoints were released together and return the same kind of object — a probability, a label with a confidence, or an ordered score — they compose in a way that two arbitrary models do not. The 600M can screen and the 3B can adjudicate. Score incoming items with Liquid AI d1-omni-600M, and escalate the ones it puts near the middle of its scale to Liquid AI d1-3B for a sharper call. The escalation rules are the confidence and the distribution the 600M already returns, so the routing logic needs no extra model. On a workload with a strong majority of easy items, most of the traffic never reaches the 3B and most of the money is never spent.
That pattern is also the reason the two models are worth running behind a router. Through OrcaRouter both would sit behind one API key at each provider's list price passed through with 0% markup, so the cascade is a routing rule rather than a second integration, and an escalation that fails at the provider layer is retried on a fallback instead of failing the request. Automatic failover matters more here than it does for a settled model, because one half of this pair is a checkpoint whose behaviour the vendor itself describes as under active development.
None of that is an availability claim, and the distinction is worth making plainly: the open d1 checkpoints are not in our catalogue. The vendor's route is to download the weights and run them locally — llama.cpp support landed day one across Apple, AMD, Qualcomm and NVIDIA hardware — or to reach them through the vendor's own API and third-party platforms.
Choosing, in one pass
If you need text and images, and the answer has to be right, take Liquid AI d1-3B. It has the benchmarks, the latency tables, the wider context, the larger vocabulary and the quantizations, and it is the member of the pair Liquid positions as the quality leader at its size.
If you need speech in the decision path, take Liquid AI d1-omni-600M, because it is the only open-weight option in this family that accepts audio at all, and accept that you are adopting it on vibes and a demo until someone publishes an audio decision benchmark or the withheld vision split.
If you do not yet know which of those describes your workload, start with the 3B and instrument the confidence it returns. The sub-scores are the tell: a task that lives in the Tools or Language columns is going to be badly served by the 600M, while something retrieval-shaped is the one place the small checkpoint is closer than its total implies. The family exists so you can trade accuracy for footprint, and the trade is only safe if you know which column your task occupies.

Through OrcaRouter both models sit behind one API key at a routing rule rather than a second integration and an escalation that fails at the provider layer is retried on a fallback instead of failing the request.
