A generated title card for Microsoft-Decision-1 vs Liquid AI d1-omni-600M subtitled 'the widest sensor surface in the category against the narrowest input list', with chips reading 587M parameters with vision and audio, a text-only Foundry API, 30 seconds of speech against 32,768 tokens of text, a bidirectional encoder against a post-trained decoder, and no published calibration on either side.
Engineering & Research

Microsoft-Decision-1 vs Liquid AI d1-omni-600M: A Scorer That Listens Against a Scorer That Only Reads

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

You are routing claims at an insurance desk, and the file arrives as three things: a PDF, a photograph of a dented door, and a ninety-second phone call. Liquid AI d1-omni-600M can read the document, look at the photo and listen to the call, then answer a set of questions you wrote in advance — is this covered, which category, how severe, on a scale of two to ten — in one pass, with no text generated and nothing to parse. Microsoft-Decision-1 can read the document. It went generally available on Microsoft Foundry on October 8, 2026 as a hosted, text-only scorer post-trained on Qwen3.5-9B with a 32,768-token window, and its input surface ends there: no image, no audio, no video. Both return calibrated probabilities over options you supply, both emit zero output tokens, and neither has published a calibration figure.

That last shared sentence is the uncomfortable part of this comparison. The two are the widest-sensor and narrowest-sensor decision models to ship this month, and for all the architectural distance between them they have landed in exactly the same evidentiary position: real weights or a real endpoint, documented contracts, and no number anyone outside the vendor can point at when someone asks whether the probabilities are trustworthy.

Two ancestries, not two sizes

The 587M figure invites you to read Liquid AI d1-omni-600M as a shrunken relative of the models this blog has covered before. It is not. The model is trained from LFM2.5-Encoder-350M, a bidirectional encoder, with a 381M shared trunk and decision head, a 94M vision encoder taken from the LFM2.5-VL-450M line and a 17-layer FastConformer for audio. Microsoft-Decision-1 is the other lineage entirely: a Qwen3.5-9B decoder, post-trained by Microsoft, hosted on Foundry, weights not distributed and no fine-tuning path offered.

Bidirectional versus decoder is the architectural fact that decides how each one behaves, and it is also why the parameter counts are a distraction. A bidirectional encoder reads the whole state at once, which is the right shape for a judgment over a fixed input; a post-trained decoder gives up the generation path but keeps the tokenizer, the window and the enterprise packaging that come with a foundry deployment. Neither is a scaled version of the other, and they cannot be swapped for one another no matter how a benchmark table is read.

What the input list actually buys you

This is where the d1-omni-600M diverges hardest from everything else in the category, and where a claim needs to be read carefully.

• Speech — Liquid AI d1-omni-600M accepts text plus up to 30 seconds of speech and reports a decision over it, which no text-only scorer can do at any accuracy. Microsoft-Decision-1's non-text modalities do not exist.

• Mutual exclusion — the d1-omni-600M raises a ValueError if images and audio arrive in the same request. The widest sensor surface in the category is still one non-text modality at a time, which is a real constraint on that claims desk.

• Trimming — its card states 16,384 combined text, image and audio positions, and adds that with images present the state and question text are trimmed to 896 tokens to match training. Whatever document you attach a photo to is competing with that trim. Microsoft-Decision-1's 32,768 tokens are all text and are not trimmed.

• Formats — the d1-omni-600M has three question types: noul for a binary decision returning P(yes) between 0 and 1, choice for one label from a set you name with a confidence and a probability per option, and score for a position on an ordered scale of two to ten levels with its distribution. Several questions hang off one state and are read from a single pass, and the usage counter reports output_tokens: 0. Microsoft documents yes/no, multiple-choice, rating, classification and rubric shapes, plus an explicitly supported abstention option such as "cannot tell" when the evidence is insufficient — the one design detail on the Microsoft page with no direct counterpart here.

• Runtime — the d1-omni-600M ships custom code, needs trust_remote_code=True, and its card recommends float16 on GPU while warning that bfloat16 changed the top answer on some rows. That is a serving-time sensitivity documented by the vendor, not a defect. Microsoft-Decision-1 is a managed endpoint where the deployment shape is your only runtime variable, and it is worth noting that batch inference is disabled: there is no offline channel to amortize a bulk scoring run through.

A two-column generated scoreboard titled Microsoft-Decision-1 vs Liquid AI d1-omni-600M. Left column Microsoft-Decision-1 rows read: availability Foundry GA, October 8, 2026; base post-trained Qwen3.5-9B decoder; inputs text only; context 32,768 tokens of text, untrimmed; weights not distributed; calibration not published. Right column Liquid AI d1-omni-600M rows read: availability uploaded October 7, 2026; base LFM2.5-Encoder-350M bidirectional encoder with 587M total; inputs text, images, or 30 seconds of speech, one non-text modality at a time; context 16,384 positions with text trimmed to 896 tokens when images are present; weights open under lfm1.0; calibration not published. A footer line reads that neither vendor has published an accuracy or calibration benchmark for these models.

The evidence gap, in both directions

Liquid AI published the open d1 family, including d1-omni-600M, on 7 October 2026 with a release post and a documentation set. What is not published is the number that would make the multimodal claim concrete. There is no accuracy benchmark specific to its own headline capability — the vision-and-audio decision quality that justifies the 94M and 112M encoders — and the vision split is withheld from the released evaluation material. No latency figure is published either, so the throughput of the model is something you discover on your own hardware.

Microsoft's position is structurally identical and one step further along. Its Benchmarks tab names the metrics — accuracy, calibration error, safety recall, false-positive rates, fairness consistency — states that evaluation ran on public and community decision benchmarks plus held-out internal test sets not used in training, that option order was varied, that paired statistical tests were applied, and claims the model "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology." It prints none of the results. Twenty-five languages are listed as supported, including Japanese, Korean, Arabic, Vietnamese, Thai, Turkish, Hindi, Bengali, Swahili, Hebrew, Persian and Ukrainian, with the candid warning that coverage, quality and calibration "may vary by language" and that non-English, especially lower-resource languages, is a weak spot. Pricing is not on the model page either; it links out to Microsoft's pricing surface.

So the comparison between these two is not one of evidence against evidence. It is a choice between two kinds of missing number. Liquid AI has shipped the sensor surface and left the accuracy of that surface unmeasured in public. Microsoft has shipped the procurement path — Azure authentication, governance, a Responsible AI assessment, a documented methodology — and left the calibration of a probability product unquantified.

A capture of Liquid AI's own blog post 'Open d1: Edge decision models for text, vision, and audio' dated Oct 7, 2026, showing the opening paragraph that announces d1-3B and d1-omni-600M as released that day, d1-3B's 48.57 Decision Index v0.2.1 score, and its 8 ms, 16 ms and 26 ms latencies on an RTX 4090, a Jetson AGX Thor and a Jetson AGX Orin.

The calibration question neither of them answers

For a decision model, the probability is the product. Everything downstream is a threshold: 0.7 escalates, 0.95 auto-accepts, and the cost of drawing that line wrong is paid in bad automated decisions rather than in tokens. In this category there is at least one family that publishes the figures — InternLM's Intern-Decision checkpoints print Brier and expected-calibration-error numbers on their model cards — and neither model in this article does. That asymmetry is the practical reason to keep the field wide rather than committing on architecture alone.

The good news is that the test is cheap and does not require the vendor's cooperation. Assemble a couple of hundred labelled cases that look like your real traffic, write the question schema your application would actually send, run both models over them, and compute expected calibration error on the output. You will then know two things nobody outside the two vendors currently knows: how well Microsoft-Decision-1's probabilities track its accuracy on your distribution, and whether the d1-omni-600M's audio and image paths are worth the encoders they carry. Microsoft's own documented guidance points the same way — validate on representative data, set thresholds from the cost of your errors, always include an abstention option, randomize option order, keep a human in the loop for consequential decisions.

Where each belongs, and where OrcaRouter does

Microsoft-Decision-1 belongs wherever the deciding material is text and the blocker is procurement: a long document, a retrieved passage, a proposed tool call, a generated answer being graded against a rubric, with Azure authentication, unified billing and a vendor assessment required before anything reaches production. It is the narrower instrument and the easier one to get approved.

Liquid AI d1-omni-600M belongs wherever the deciding material is not text — a photo attached to a form, a short call recording, a screenshot — and wherever you want lfm1.0 weights you can run and fine-tune inside a boundary you control. It is the wider instrument and the harder one to validate, because the multimodal claim is exactly the part with no published benchmark behind it.

A capture of Liquid AI's documentation site showing the left navigation (Liquid Foundation Models with Text, Vision, Audio, Decision Models and Liquid Nanos entries, plus Fine-tuning, Edge Inference and llama.cpp), the 'New: Open d1' banner, and the opening description of Liquid Foundation Models as a class of multimodal architectures built for fast inference and on-device deployment.

Neither is in our catalogue, and nothing here is an availability claim for either. A model that returns probabilities instead of text is not something you route chat completions to, and staying out of the scoring seat is the whole design of the instrument. What OrcaRouter carries is the generative side of the same loop: the more than 200 models behind one OpenAI-compatible key that write the rubric your scorer grades against, transcribe or summarise the material before it becomes a state, and emit the tool call your scorer approves before it runs. Provider list price is passed through at 0% markup, so a vendor price cut on the generating side is live on ours the same day. Automatic failover keeps that side alive when a single provider degrades — which a pipeline that scores every retrieved document notices immediately — and if you want several writers judged rather than one, the routing DSL composes them into a single call and model fusion reports their agreement as a field your scorer can read like any other.

Bottom line

Microsoft-Decision-1 reached general availability on Microsoft Foundry on October 8, 2026: text-only, 32,768 tokens, built on a post-trained Qwen3.5-9B, hosted with no weights, no price on the model page, no latency figure and a benchmarks section that describes its methodology without printing a result. Liquid AI d1-omni-600M was uploaded on 7 October 2026 as a 587M bidirectional-encoder decision model that reads text, images or 30 seconds of speech — one non-text modality at a time, with text trimmed to 896 tokens when images are present — under the lfm1.0 licence, with no accuracy benchmark for its headline capability, no vision split and no published latency. Choose on modality and control rather than on scores, because the scores do not exist: if your material is a photograph or a phone call, only one of these can answer, and if it is a forty-page document with a procurement review attached, only the other one can be deployed.

What OrcaRouter carries is the generative side of the same loop: the more than 200 models behind one OpenAI-compatible key that write the rubric your scorer grades against, transcribe or summarise the material before it becomes a state, and emit the tool call your scorer approves before it runs.