
Liquid AI d1-3B vs Intern-Decision-4B: Which Calibrated Decider Do You Actually Need?
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 56 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 56 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 348 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Two open-weight models now answer questions by not writing anything, and until you put their I/O schemas side by side they look like the same idea twice. Liquid AI d1-3B, uploaded to Hugging Face on October 5, 2026 and announced by Liquid two days later, is a 3.12B multimodal decision model whose card says flatly that it "is not a chat model and does not write text." InternLM's Intern-Decision-4B, uploaded quietly to the InternLM organisation on September 26, 2026 with no blog post, no tweet and no launch page behind it, is a 4B decision model fine-tuned from Qwen3.5-4B that likewise returns an answer distribution for every question in one forward pass. Same contract on the surface. The differences underneath — a licence that will complicate you, a contract that can write text by accident, and machines nobody has compared — are what decide which one belongs in your stack.
How close the two actually are
The first thing to get straight is how much of this comparison is a tie. Both models take a state and a schema of named questions, both read the signal from the logits at a placeholder position rather than sampling, both are multimodal, and both report that they wrote nothing. On the shape of the call they are near-identical, and the interesting differences are in the details around the edges.
• Size — Liquid AI d1-3B 3.12B total, built on Liquid's LFM2.5-VL-3B; Intern-Decision-4B 4B (about 4.54B by the tensor count the card prints), fine-tuned from Qwen3.5-4B
• Question types — d1-3B and Intern-Decision-4B use the same three: noul for yes/no, choice for one of a named set, score for a point on an ordered scale
• Answer fields — d1-3B returns the answer plus confidence and probabilities; the InternLM schema returns distributions plus a confidence value that the model card warns is not a calibrated probability
• Input cap — d1-3B 32,768 tokens of context; Intern-Decision-4B an 8,192-token ceiling that refuses longer requests outright rather than truncating, with images counted against it
• Licence — d1-3B under Liquid's LFM Open License v1.0, Apache 2.0 on the InternLM side
• Languages — d1-3B lists 16; Intern-Decision-4B's card states none
• Interface — d1-3B ships as a model you load and call with trust_remote_code=True; Intern-Decision-4B ships as a Python library you pip install, with one implemented backend and the request's own model field explicitly not able to switch checkpoints
• Vendors' own score — d1-3B 48.57 on Decision Index 0.2.1 public split; Intern-Decision-4B 90.02 averaged across its own seven-set table with a Brier score of 0.347 and an expected calibration error of 0.065
Those last two numbers are not comparable, and treating them as a scoreboard would be the single easiest mistake to make on this page. They come from different harnesses, different question sets and different scorers.

Calibration is the whole product, and only one of them stands behind it in print
A decision model is only worth its latency if the number it returns next to the answer means something. That is what calibration is: a stated confidence of 0.9 should be right about nine times in ten, and a model that is confident and wrong is worse than one that says it does not know.
InternLM is the one that engages with this. Its card documents a fitted temperature of 1.99241824, derived by negative log-likelihood minimisation over 1,728 designated calibration cases with 1,693 held out for validation, and printed to eight decimal places so it can be reproduced. It also publishes a before-and-after pilot over 96 cases that shows the mechanism working rather than asserting it: Brier 0.628 and ECE 0.213 before calibration against 0.550 and 0.089 after. Brier and ECE are the two standard measures of whether stated probabilities are honest, and the direction of travel is large.
Liquid publishes no equivalent. d1-3B's own card carries accuracy-style benchmarks — SQuAD 2.0 85.3, Civil Comments 93.0, MASSIVE intent 87.3, PubMedQA 66.0, BoolQ 86.7, XNLI 85.0, PAWS-X 76.9, a 77.1 mean — alongside its 48.57 on the Decision Index, and the model's stated selling point is calibrated typed answers. But there is no ECE row, no Brier row, and no published fitted temperature.
That asymmetry does not mean Qwen3.5-4B's descendant is better calibrated than a 3B built on LFM2.5-VL-3B; it means that between these two cards, one of them gives you the arithmetic to reproduce the claim and the other gives you the claim. If your pipeline puts a threshold on a confidence — route above 0.8, escalate below 0.4 — you can validate the InternLM side tonight and you can only spot-check the Liquid side.
The caveat cuts both ways. Both sets of numbers are first-party. Neither model has been scored by an independent party, and the InternLM calibration claims rest on InternLM's own 96-case pilot against InternLM's own fit.
The licence that asks you for a number
Intern-Decision-4B is Apache 2.0, inherited from Qwen3.5-4B, with the upstream notices retained — commercial use, redistribution, fine-tuning, no revenue question anywhere in it.
d1-3B is under Liquid's LFM Open License v1.0, and Section 5 is the part nobody reads until procurement does. Commercial use is granted only on condition that you or your legal entity stay below a Threshold, defined in that same document as annual revenue of ten million US dollars or more; use above it is simply not licensed. Non-profits and research users are carved out.
For an individual or a small team both are free. For anything with a finance department the two repos are different products, and the difference is a number you either are or are not above.
The tail of d1-3B's benchmark table is its real claim
The headline figure on Liquid's card is 48.57 on Decision Index 0.2.1, ahead of the other sub-10B rows in a table that runs from Winnow-12B at 50.02 down to Decider 2B at 28.97. What makes that number worth quoting is the discrimination index that sits underneath it. On the Decision Index's own sub-scores d1-3B scores 74.5 on Tools and 36.3 on Arts while managing only 23.8 on Knowledge — against Decider 35B-A3B, a 36B mixture-of-experts model 12x its size, which scores 56.5 on Tools, 32.6 on Arts and 31.8 on Knowledge.

In other words the 3B is not a small model that is uniformly a bit worse. It is a small model that is unusually strong at structured tool-style decisions and unusually weak at questions that need world knowledge, and it beats the 36B on the two categories that most taste like the job it was built for. If your decisions are "which of these five named actions" and "is this grounded in the retrieved passage", the size gap is doing less work than you would expect. If your decisions lean on facts the model has to already hold, expect the 23.8 to show up.
There is a second version of that argument on the vision side. d1-3B averages 74.1 across eleven public image benchmarks against 73.9 for its own base model, and with the images stripped out the same questions score 45.1 — so on image questions the answers are genuinely coming from the picture. It is level with its base model rather than better than it, and the per-benchmark rows cut both ways: CV-Bench 82.1 against 87.6, POPE 88.5 against 90.1, BLINK 59.2 against 58.7.
Worth noting the pause in the InternLM card too: its high aggregate comes from question-driven tasks such as Typed Decision 80.55 and ToolACE 96.45, while Jevbench-Hard sits at 73.87. Where the schema gets hard, the score drops.
Speed: the gap is not where you expect it
d1-3B advertises 8 ms for a single question on an RTX 4090 and 9 ms on an MI325X, 21 ms for three questions sharing one state, 16 ms on a Jetson AGX Thor, and packed throughput of 475 decisions per second on the 4090 and 1,106 on the MI325X.
Intern-Decision-4B's card measures 44.16 ms mean end-to-end on the same RTX 4090, with a 44.03 ms median and 44.60 ms at P95.
That looks like a 5x win and it is not one, because the two are measuring different things. Liquid's numbers are warm model calls, one request at a time, with kernel selection and compilation paid up front — the card says explicitly that a single question costs 16 ms on the 4090 without CUDA graph compilation, and that the first call with a new shape pays for compilation. InternLM's 44 ms is per-query end-to-end through a Python library whose only implemented backend is local Hugging Face inference, so it carries the surrounding machinery. Neither number is wrong. Put both under one harness, on one machine, on the same shape of request, and the gap shrinks to something you would want to measure rather than assume.
What is not in question is the ceiling. 8,192 tokens is a hard refusal on the InternLM side, and image tokens spend from the same budget, so a long document plus a screenshot is a request that comes back rejected rather than truncated. d1-3B gives you 32,768 tokens to work with. If your states are tickets and short exchanges, that gap never bites. If they are page-sized, it decides the question before any benchmark does.
Run them as a panel, not a swap
Answering a specific yes/no and answering "which of six named things is this, and how sure are you" are different asks, and the two cards are shaped differently enough — one with a hard input ceiling and a published calibration fit, the other with 32K of context and a better scoring tail — that the useful deployment for a team evaluating both is often not a replacement but a panel: the same state posed to both models, with disagreements routed to a human or to a larger model.
Neither model is on our catalogue, and this is not an availability claim; both are self-hosted artefacts. But a panel is exactly the shape where a routing layer does the work rather than the paperwork. OrcaRouter exposes more than 200 models behind one API key with provider list prices passed through at 0% markup — so when a vendor you route behind us cuts its price, your bill moves the same day — and automatic failover across providers stops a two-model panel from becoming two single points of failure. The models you route today are not these two; they are the hosted generalists you would be replacing, such as Qwen3.5-27B at $0.086 per million input tokens and $0.688 per million output, which is the number a self-hosted 3B has to beat once you have counted the GPU.
What to do this week
If your decisions are short states, the licence has to survive your company's review, and you want the widest range of build targets — llama.cpp, Ollama, LM Studio, a packed batch path that runs 64 states through one pass — d1-3B is the more productised of the two and its tool-and-arts discrimination is the strongest thing either card demonstrates. If your decisions run over long states, or you need a licence with no revenue clause, or you want a calibration fit you can reproduce and a harness where the surrounding cost is already counted, Intern-Decision-4B is the one whose paperwork you can actually check.

The uncomfortable part is the same for both: no independent party has run either model, and no calibration comparison exists that was not commissioned by the vendor who wrote the card. For a model class whose entire value is the meaning of a number next to an answer, that is the gap worth closing before you put a threshold on anything.
