
Liquid AI d1-omni-600M vs LFM2.5-VL-3B: One Decides, One Describes
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 56 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 56 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 348 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Liquid AI d1-omni-600M and LFM2.5-VL-3B are both small, both multimodal, both published by the same lab on Hugging Face under the same lfm1.0 licence — and pairing them is a category error that turns out to be a useful one. LFM2.5-VL-3B landed on August 12, 2026 as Liquid AI's edge vision-language model: 3.1B parameters, a 32,768-token window, and an output that is prose. Hand it a scanned page and it returns the transcription, with layout, as text. Liquid AI d1-omni-600M was uploaded on October 7, 2026 with the rest of the open d1 family, and it does the opposite thing with the same inputs. It is a 587M-parameter decision model: you attach named questions to a state and it reports what it already believes — P(yes), a chosen label with a confidence, a position on an ordered two-to-ten scale — in a single forward pass, with zero output tokens and nothing downstream to parse. If you searched for "the small multimodal Liquid model," there are two shapes in front of you now, and the shape is the decision.
This page is what the two model cards, the October 7 release post, and the repositories themselves actually establish. It is also explicit about where they stop. Nothing in this comparison has been reproduced outside Liquid AI, and for one of the two models the vendor has published no accuracy benchmark specific to its own headline capability at all.
The two checkpoints, side by side
The spec sheets diverge on almost every axis, which is what you would expect from two models that were never meant to compete for the same slot.
• What it is — LFM2.5-VL-3B is a generative vision-language model that writes text; Liquid AI d1-omni-600M is a decision model that returns structured answers and emits no tokens
• Parameters — 3.1B in BF16 for LFM2.5-VL-3B versus 587M for Liquid AI d1-omni-600M, itself a 381M shared trunk and decision head plus a 94M vision encoder and a 112M audio encoder
• Backbone — LFM2.5-VL-3B pairs an LFM2.5-2.6B language model with a SigLIP2 NaFlex shape-optimized 400M vision encoder; Liquid AI d1-omni-600M is trained from LFM2.5-Encoder-350M, a bidirectional encoder, with a SigLIP2 tower taken from LFM2.5-VL-450M and a 17-layer FastConformer for audio
• Inputs — text and images for LFM2.5-VL-3B; text plus images or text plus up to 30 seconds of speech for Liquid AI d1-omni-600M, which raises a ValueError if images and audio arrive in the same request
• Context — 32,768 tokens for LFM2.5-VL-3B versus 16,384 combined text, image and audio positions for Liquid AI d1-omni-600M, whose card adds that with images present the state and question text are trimmed to 896 tokens to match training
• Vocabulary — 128,000 tokens for LFM2.5-VL-3B, 65,536 for Liquid AI d1-omni-600M
• Runtime — LFM2.5-VL-3B runs in transformers and vLLM and has a separate DSpark draft model for speculative decoding; Liquid AI d1-omni-600M ships custom code, needs trust_remote_code=True, and its card recommends float16 on GPU while warning that bfloat16 changed the top answer on some rows
• Licence — lfm1.0 for both, which permits fine-tuning and deployment
One line in that list is doing more work than the rest. Liquid AI d1-omni-600M is built on a bidirectional encoder rather than a decoder. That is not a size reduction of LFM2.5-VL-3B; it is a different ancestry, and it is the architectural reason the two models cannot be swapped for one another no matter how the benchmark tables are read.
The output contract decides more than the benchmarks do
Ask LFM2.5-VL-3B a question about an image and you get language back. That is worth money when the answer has to be read by a person, logged as a string, or fed to a step that expects prose. It is also a liability you have to engineer around: generated output can wander, a JSON-shaped request can come back shaped differently than you asked, and every token is billed and waited on. The model is explicitly scoped for single-turn, high-throughput, low-latency vision work — batched OCR of scanned documents, real-time detection in a vehicle, on-device translation of signage — and explicitly steered away from long-context, reasoning-heavy questions such as "what is wrong with this blueprint."
Liquid AI d1-omni-600M answers a strictly narrower question in a strictly more reliable envelope. Its interface is a state — text, JSON, one or more images, or up to 30 seconds of speech — plus a set of named questions, each declaring its own type. A noul question is a yes/no and comes back as P(yes) between 0 and 1. A choice question picks one label from a set you name and returns the label, a confidence, and the probability of each option. A score question places the state on an ordered scale of two to ten levels and returns the expected level with its distribution. Several questions can hang off one state and are read from a single pass, so the model card's usage counter reports output_tokens: 0. There is no generation step, which means there is no such thing as output that fails to parse.
What you give up is everything a generated answer carries that a label does not: the reasoning, the hedge, the phrase you can paste into a ticket, the ability to ask a follow-up in the same conversation. Liquid states it plainly — the model is not a chat model and does not write text. If your pipeline needs an explanation next to the verdict, Liquid AI d1-omni-600M is the wrong half of the pair, and the honest answer is to run both: LFM2.5-VL-3B to produce the reading of the image, Liquid AI d1-omni-600M to produce the decision about it.

There is no head-to-head vision number, and that is the finding
The instinctive next move is to line up the vision benchmarks. You cannot. On LFM2.5-VL-3B the vendor publishes a full set — RefCOCO Macro Precision@1 at 87.9, ScreenSpot-v2 at 80.7, ChartQA at 81.3, POPE at 88.7, MMStar at 63.3, BLINK at 61.5, MuirBench at 58.3, ToolSandBox at 59.5, OCRBench v2 English at 47.5, and a RealWorldQA of 73.1 that edges the previous LFM2-VL-3B's 71.1. All of those are vendor-reported, none has been re-run independently, and they are nonetheless specific claims about specific capabilities that a reader can hold the lab to.
For Liquid AI d1-omni-600M the release post says the Decision Index v0.3 includes a private vision split "which we do not report on in this release," and that instead Liquid validated that the 600M checkpoint "handles all three modalities," with the evidence placed in playground demos. That is the whole of the published vision record. There is no image benchmark number for this checkpoint, in either its model card or the launch post. The same post confirms what is missing on the audio side even more directly: dedicated audio decision benchmarks are, in the vendor's own words, "currently an open problem."
So the correct reading of the matchup is asymmetric and should be stated as such. LFM2.5-VL-3B has demonstrated vision ability with numbers attached. Liquid AI d1-omni-600M has a vision encoder borrowed from a sibling model, an adapter trained against a frozen backbone, a demo, and a promise. If your deployment needs the model to be right about images this month, the 3B is the only one of the two with anything to stand on.
Speed: only one of these has been measured
Because a decision model generates nothing, latency is the number that matters, and again only one side is documented. LFM2.5-VL-3B is quoted at 228 tokens per second on an Apple M5 Max and 116 tokens per second on an AMD Ryzen AI Max+ 395, both inside 3.3 GB of memory, with about 20 tokens per second on a Galaxy S26 Ultra and roughly 11,000 tokens per second on a single H100 under vLLM. Those figures describe decode throughput, which is the right measurement for a model that writes.
For Liquid AI d1-omni-600M the card states that inference numbers are not reported because the model is an early research release and under active development. The d1-3B sibling in the same family does have latency tables — 8 ms for one question on an RTX 4090, 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin 64 GB, 50 ms on an Orin Nano, with three questions over one state costing about 1.3x the time of one. Those numbers belong to the 3B decision model and must not be read across to the 600M. For Liquid AI d1-omni-600M, the honest position is that the architecture implies a small, fast model and nobody outside the lab has published a millisecond figure.
Weight footprint can at least be estimated. At the float16 precision the card recommends, 587M parameters is roughly 1.2 GB of weights before activations — arithmetic on the published parameter count, not a vendor measurement — against the under 3.3 GB Liquid reports for LFM2.5-VL-3B. On a device where megabytes are the constraint, that difference is the argument for the 600M.

How to pick between them
The choice resolves cleanly once the output contract is treated as fixed rather than negotiable.
Reach for LFM2.5-VL-3B when the deliverable is text: full-page OCR with layout annotation, grounding and detection from natural-language queries, signage translation on a device, any step where a human or a downstream language model consumes the answer. It also has the wider context at 32,768 tokens and the deeper language coverage in practice, because its answers are language rather than labels.
Reach for Liquid AI d1-omni-600M when the deliverable is a decision: routing a ticket, triaging a frame, moderating a clip, gating a guardrail, scoring a response on a two-to-ten scale — and, uniquely between these two, when the input is speech. It is the only one of the pair that takes audio at all, up to 30 seconds per request, though its card notes the audio was trained on exchanges between an English speaker and an assistant, which is a narrow description of the world your calls will come from.
Reach for neither, and stay with a general vision-language model, if the task needs reasoning about the image rather than a reading or a verdict about it. Both model cards point away from that use case, and Liquid AI d1-omni-600M has no mechanism to attempt it.
Trying an experimental checkpoint without betting the pipeline on it
Liquid AI d1-omni-600M is, by the vendor's own label, an early research release, and its vision and audio behaviour are unbenchmarked. That is exactly the profile of a model you want to evaluate behind a fallback rather than in front of customers. OrcaRouter routes 200+ models through one API key with automatic failover across providers, which is the cheap way to put a checkpoint like this on a live path: if the experimental route degrades or a provider goes down, the request lands on the fallback without a code change. The same key carries every model the pipeline hands off to, at each provider's list price passed through with 0% markup, so a vendor price cut is your price the same day.
One caveat, stated because it is load-bearing: the open d1 models and the LFM2.5-VL family are not in our catalogue today. Self-hosting the weights is the route the vendor recommends for both, with day-one llama.cpp support across Apple, AMD, Qualcomm and NVIDIA hardware, and running them on a hosted endpoint means the vendor's own API or third-party platforms. Reading this page as an availability claim would be a mistake.

What to watch
The interesting question is not whether the 600M eventually beats the 3B at anything — it will not, and it was not built to. It is whether Liquid AI publishes the private vision split it withheld, and whether anyone builds the audio decision benchmark the lab says does not exist yet. Until one of those lands, a comparison between Liquid AI d1-omni-600M and LFM2.5-VL-3B is a comparison between a measured model and a plausible one. For unmeasured work — speech in, verdict out, small enough to sit on the device — the 600M is the only open-weight checkpoint in this family trying it, and being first with no scoreboard is still being first.
OrcaRouter routes 200+ models through one API key, with each provider's list price passed through with 0% markup so a vendor price cut is your price the same day.
