
Microsoft-Decision-1 vs Liquid AI d1-3B: Same 32K Window, Two Different Senses
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 113 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 52 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 62 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 399 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
For once the spec lines up. Microsoft-Decision-1, generally available on Microsoft Foundry since October 8, 2026, takes 32,768 tokens of context. So does Liquid AI d1-3B, uploaded to Hugging Face on October 5, 2026 and announced by Liquid AI two days later. Both accept a state plus a set of typed questions — yes/no, choice among named options, a position on an ordered scale — and both answer in a single forward pass with a probability distribution instead of generated text; Laya, Intern-Decision-4B, Kev and the rest of this niche disagree with each other about context length, so this is the one pairing where that dimension drops out and the comparison has to be argued on something else.
The something else turns out to be the input channel. Microsoft's model is text-only, explicitly: it "does not accept or produce image, audio, or video content." Liquid AI's carries a 400-million-parameter SigLIP2 NaFlex vision encoder on a 2.6B-parameter language backbone and reaches 3.12 billion parameters total, and it was built for exactly the thing Microsoft ruled out — scoring a screenshot, a scanned form or a camera frame against a fixed question. That split, plus a licence difference that has nothing to do with capability, is the whole matchup.
Where the two agree, which is more than you would expect
• Context window — 32,768 tokens for Microsoft-Decision-1 and 32,768 tokens for Liquid AI d1-3B.
• Output shape — calibrated probabilities over supplied options; neither generates text, and Liquid AI's usage counter reports the fact as output_tokens: 0.
• Question types — both document choice, ordered score and binary yes/no, and both let you define the option set per request rather than retraining a vocabulary head.
• Abstention — Microsoft documents explicit abstention options such as "cannot tell"; Liquid AI's Decision Index schema supports the same through its option set.
• Single-pass inference — one call per decision on both sides, which is what makes either one viable at pipeline volume.
Two models agreeing on the two hardest constraints in this niche is worth pausing on. A 32K window is what makes a decision scorer usable on a full contract, a multi-page policy or a long support thread; the 512-token English Laya checkpoint and the per-call question limits on Intern-Decision-4B do not. If your input is short, most of this field is interchangeable and the decision is about hosting. If your input is long, the field narrows to these two.
The sense that only one of them has
Liquid AI d1-3B is multimodal, and the model tree makes the lineage plain: LFM2.5-2.6B-Base, then LFM2.5-VL-3B, then d1-3B post-trained for single-pass calibrated decisions. Images enter through the SigLIP2 NaFlex encoder, and the card reports a mean of 74.1 across eleven public image benchmarks against 73.9 for the base vision model — so the decision-tuning did not cost it vision quality. It also reports the reverse experiment: strip the images out and the same questions drop to 45.1, which is the cleanest evidence you will get that the perception channel is load-bearing rather than decorative.
Microsoft-Decision-1 has no equivalent, and the omission is deliberate rather than unfinished. Microsoft's card lists the out-of-scope uses as text generation, open-ended question answering, conversation, translation and summarization, and separately states the model is text-only. For a pipeline that needs to score a screenshot of a form against a rubric, or decide whether a camera frame contains a safety event, that is a hard stop — no prompt engineering recovers a modality the model was never given.
The language counts run the other way, and this is the second real divide. Microsoft-Decision-1 lists 25 supported languages, and the company is candid that coverage, quality and calibration "may vary by language," naming non-English and especially lower-resource languages as an underperformance area. Liquid AI d1-3B states 16 languages. On breadth Microsoft is ahead on paper; on honesty about what breadth means, both vendors say something similar, and neither has published per-language calibration.

The benchmark tab, again
Microsoft-Decision-1's Foundry page carries a Benchmarks tab with a methodology section and no figures: public and community decision benchmarks plus held-out internal sets, metrics covering accuracy and calibration error, option order varied, paired statistical tests, and a claim that the model "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology." Nothing to check, nothing to reproduce.
Liquid AI d1-3B publishes 48.57 on Decision Index 0.2.1 — and the qualifier matters more than the number, because Decision Index is Liquid AI's own index and d1-3B is Liquid AI's own model. It is a vendor-scored result on a vendor-authored benchmark, positioned by the company as best among decision models under 10B and ahead of a 35B model on the same scale. Treat 48.57 as a directional claim, not an audited figure. Alongside it the card lists internal decision-format evals at a mean of 77.1, SQuAD 2.0 at 85.3, PAWS-X at 76.9 and DecisionBench 71.8 — all self-reported, and the "under 10B" framing is a genuine qualifier rather than a dodge, since the model is 3.12B.
So the calibration question resolves the same way it does across this whole field: Liquid AI gives you a number produced on its own scale, Microsoft gives you a methodology and no number, and neither gives you an independent third-party measurement. The only calibration figure that will matter to your deployment is the one you compute on your own labeled data.

Serving, licences, and what you are actually buying
d1-3B's latency is published and it is the reason the model exists. Eight milliseconds per decision on an RTX 4090, nine on an AMD MI325X, with packed throughput of 475 and 1,106 per second at 64 states. On edge hardware: 30 ms on an Apple M5 Pro, 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin 64 GB, 50 ms on an Orin Nano. Microsoft-Decision-1 publishes no latency figure at all, and its designed-in options — a hosted Foundry API on pay-as-you-go or reserved provisioned throughput, with batch inference disabled — means you measure it through your own deployment. That is not a small gap for a model whose entire purpose is to be called once per retrieved document or proposed action.
The licence is where the two diverge in a way that surprises people. Liquid AI d1-3B is tagged lfm1.0, not Apache 2.0 — the same family licence as the rest of Liquid's LFM line, which carries usage conditions that a permissive licence does not. If your legal review reads OpenAI-compatible-with-no-strings, this is not that, and it is worth reading before the model ships rather than after. Microsoft-Decision-1 has no weights at all: it is a hosted endpoint under the Direct from Azure portfolio, with Azure authentication, unified billing, no download, and the licence question replaced by a service agreement. Neither model can be fine-tuned through the paths the vendors document, which is the sharpest limitation for a decision model — a scorer you cannot adapt to your own labels is a scorer whose error modes you cannot fix, only route around.
Routing around them is where OrcaRouter belongs in this pattern, and only on one side of it. We host neither Microsoft-Decision-1 nor Liquid AI d1-3B: neither returns text, neither is in our catalogue, and a probability-scoring endpoint is not a chat-completions target. What we carry is the generating half of the loop — the model that drafts the answer your scorer will grade, extracts the fields your scorer will judge, or proposes the action your scorer will approve. That is more than 200 models behind one OpenAI-compatible key at provider list price passed through with 0% markup, so a vendor price change is live on our side the same day, and automatic failover covers the generation call in a loop that has no spare latency to retry one.

Two clean picks, and one that is not close
Take Liquid AI d1-3B if anything in your pipeline is a picture. Screenshot triage, form extraction, visual QA routing, a moderator scoring camera frames — Microsoft-Decision-1 cannot be made to do this, and d1-3B was built for it with the vision benchmarks published to back the claim. Take it too if you need edge inference measured in tens of milliseconds on a Jetson or a laptop, if you want the model on your own hardware, or if a licence you can read is enough for your counsel.
Take Microsoft-Decision-1 if your input is text, your documents are long, your gate is procurement — Azure authentication, unified billing, a Responsible AI assessment, a vendor to escalate to — or your traffic covers more than the 16 languages d1-3B lists. Accept that its missing latency figure and missing benchmark table mean your first two weeks with it are measurement, not integration.
The one case that is not close is multimodal work: there, the choice was made when Microsoft wrote "text-only" into the model card.
