A hero title card headed 'Liquid AI d1-3B vs Granite 4.2 3B' with the subheading 'one decides, one writes', showing a left card labelled d1-3B with a checklist icon, the caption '3.12B - 32,768 tokens' and chips reading probability, label and score, and a right card labelled Granite 4.2 3B with a speech-bubble icon, the caption '3B - 128K tokens' and a chip reading 'a written answer'.
Guides & Insights

Liquid AI d1-3B vs Granite 4.2 3B: Two 3B Models That Never Do the Same Job

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put Liquid AI d1-3B and Granite 4.2 3B on the same single GPU and they will both fit comfortably — 3.12B parameters on one side, 3B on the other. They will also behave as though they came from different disciplines. Granite 4.2 3B, which IBM's own model card dates to August 25, 2026, is a reasoning model: three thinking modes, a <think> block, and an answer you read. Liquid AI d1-3B, uploaded to Hugging Face on October 5, 2026 and announced by Liquid two days later, is a decision model: it takes a state plus an explicit set of typed questions and returns a probability, a label or a position on a scale in one forward pass, reporting output_tokens: 0 because it never writes anything. The useful question is not which one is better. It is which of your calls is actually a decision, because the two repos are built for opposite halves of that split.

What is actually in each repository

Everything below comes from the two model cards and Liquid's announcement post. Nothing here has been independently reproduced, no third party has scored either model on the same harness, and the numbers in the cards are the vendors' own.

• Size — Liquid AI d1-3B 3.12B total, built on Liquid's LFM2.5-VL-3B; Granite 4.2 3B a 3B decoder-only dense transformer built on granite-4.1-3b-base

• Output — d1-3B returns probabilities, labels and confidences and writes no text at all; Granite 4.2 3B generates tokens, including an optional chain of thought inside <think> tags

• Context — d1-3B 32,768 tokens; Granite 4.2 3B 128K natively, with a long-context extension to 512K on the card

• Input — d1-3B text, JSON, images or a mix, through a SigLIP2 NaFlex 400M vision encoder; Granite 4.2 3B text only

• Licence — d1-3B under Liquid's LFM Open License v1.0; Granite 4.2 3B under Apache 2.0

• Languages — d1-3B lists 16; Granite 4.2 3B lists twelve tested, with the card noting others are untested

• Serving — d1-3B ships custom code (trust_remote_code=True, transformers>=5.14), with GGUF and w8a8 repos in the same family and cards documented for llama.cpp, Ollama and LM Studio; Granite 4.2 3B runs under vLLM 0.20+ or SGLang 0.5.18+ with a custom thinking parser

Two of those lines do more work than the rest. The output row is the whole argument of this article, and the licence row is the one nobody puts in the comparison table.

A two-column scoreboard for Liquid AI d1-3B and Granite 4.2 3B. The d1-3B column reads: job answers a typed question; parameters 3.12B; output tokens billed 0; context 32,768 tokens; image input yes; one question 8 ms on an RTX 4090. The Granite 4.2 3B column reads: job writes a reasoned answer; parameters 3B; output tokens billed every token it writes; context 128K, to 512K extended; image input no; one question a full thinking pass. The footer reads 'Both columns vendor-reported, neither independently scored.'

One writes a chain of thought, the other refuses to write

Granite 4.2 3B earns its keep by thinking out loud. Its card documents three modes: thinking on by default, non-thinking, and a low-effort mode that keeps the reasoning trace but shortens it. The serving stacks split the trace from the final answer so your application receives clean content plus a separate reasoning field. That is a good design for a question that needs to be worked out — a maths problem, a branch of logic, a piece of code that has to be written before it can be judged.

d1-3B has no such mode, and its card says so bluntly: "It is not a chat model and does not write text." A call declares questions with a type. noul is a yes/no returning P(yes) between 0 and 1. choice picks one label from a named option set and returns the label, a confidence and a probability per option. score places the state on an ordered two-to-ten scale and returns the expected level with its distribution and legend. The signal is read from the logits at a placeholder position in one pass, not sampled token by token, which is why the usage block can report zero output tokens without lying.

That difference changes your bill before it changes your accuracy. A Granite classification call that thinks for 400 tokens is a call you pay 400 output tokens for, and the label you wanted is buried in prose that a parser then has to retrieve. A d1-3B call bills input only. If most of your traffic is a label with a rule attached, you have been paying for the reasoning whether or not it changed the label.

Where Granite 4.2 3B is plainly the better call

Take the reasoning benchmarks at face value — they are IBM's own, run through an internal NeMo Evaluator pipeline, and no external party has reproduced them — and the 3B is unusually strong for its size: AIME25 78.33, HMMT February 2025 66.67, GPQA 54.80, LiveCodeBench v6 69.71, MMLU-Pro 67.84, IFBench 74.33, BFCL v4 52.41, tau3-bench 45.78, and RULER 64K 67.52 falling to 55.30 at 128K.

Four jobs follow from that list, and d1-3B is not in any of them.

• A 128K-native context, extensible to 512K, against 32,768 tokens on d1-3B — if your state is a long document, the choice is made for you

• Agentic tool calling with a documented parser and harness integrations for OpenCode, Pi and OpenHands, which a decision model has no equivalent of

• Any task whose answer is a paragraph: summarisation, drafting, extraction into free-form structure, code generation

• Fine-tuning and redistribution without a revenue ceiling, because Apache 2.0 has no threshold clause

Note what is not on the Granite card: SWE-Bench and Terminal-Bench rows are marked NA for the 3B. IBM's agentic reinforcement-learning stage was applied to the 8B and 30B members of the family only, so the smallest model does not carry the agentic scores its larger siblings do. A team that reads "Granite 4.2" and assumes the 3B inherits the family's agentic profile will be surprised.

A capture of Liquid AI's own blog post 'Open d1: Edge decision models for text, vision, and audio' dated Oct 7, 2026, showing the opening paragraph that releases d1-3B and d1-omni-600M as open-weight models, d1-3B's Decision Index score of 48.57 and the claim that it is ahead of every model under 10B, and its latency of 8 ms on an NVIDIA GeForce RTX 4090, 16 ms on a Jetson AGX Thor and 26 ms on a Jetson AGX Orin.

Where d1-3B wins before Granite finishes thinking

Liquid's own latency table, measured warm, one request at a time, is the strongest part of its case: a single question in 8 ms on an RTX 4090 and 9 ms on an AMD MI325X, 16 ms on a Jetson AGX Thor and 26 ms on a Jetson AGX Orin 64GB, with 64 states packed into one pass at 475/s on the 4090 and 1,106/s on the MI325X. For scale, that is roughly the time Granite 4.2 3B needs to emit a few tokens of its reasoning trace.

Three question types over one state cost d1-3B 21 ms on the 4090 rather than three separate calls, because the state and its images are read once for all questions. In a moderation or routing stack that used to fire one HTTP request per rule, that is the difference between a queue of calls and one.

It also sees. d1-3B scores a mean of 74.1 across eleven public image benchmarks against 73.9 for its base model, and the card notes that with the images removed the same questions score 45.1 — the answers are coming from the picture, not from the text prompt. Individual rows cut both ways (CV-Bench 82.1 against the base model's 87.6, POPE 88.5 against 90.1), so the vision claim is "level with the base model", not "better than it".

The honest counterweight: d1-3B's 48.57 on Decision Index 0.2.1 was produced by Liquid running the official scorer itself rather than by a leaderboard submission, and the competing rows come from the public leaderboard. Its 77.1 mean across eight benchmark-as-decision tasks and its 71.8 on DecisionBench are likewise first-party. Everything on the d1 side of this comparison is unverified.

A capture of the Hugging Face model card for ibm-granite/granite-4.2-3b, showing the Apache 2.0 licence, the 3B parameter count, the decoder-only dense architecture built on Granite-4.1-3B-Base, the 128K native context with long-context extension to 512K, bfloat16 precision, the tested-language list and the built-in chain-of-thought reasoning mode.

Why a 3B reasoner is not a 3B decision model

The tempting move is to treat Granite 4.2 3B as a cheaper substitute for d1-3B, or the reverse. Both fail, and the failure is structural rather than a matter of quality.

Ask Granite to decide and you get a generation you must parse, a bill that scales with how hard the model found the question, and a latency that depends on the thinking mode you selected. Ask d1-3B to reason and you get nothing at all — it has no token stream to reason in. The two models are on opposite sides of a line that most pipelines cross several times per request: something has to decide, and something has to speak. d1-3B belongs at the front, where a triage rule picks a route and returns a calibrated confidence. Granite 4.2 3B belongs behind it, on the branch that needs an answer written.

There is a second, quieter trap. A decision model's value is calibration — a confidence of 0.9 that is right nine times in ten. Granite 4.2 3B's card publishes no calibration metrics and no probabilities; it publishes accuracy and reasoning scores. Comparing the two on a benchmark leaderboard tells you nothing about which of them you should trust with a threshold.

The licence line nobody puts in the table

Granite 4.2 3B ships under Apache 2.0. d1-3B ships under the LFM Open License v1.0, which reads as permissive until Section 5: commercial use is granted on condition that you or your legal entity stay under a threshold defined in the same document as annual revenue of ten million US dollars or more, and any commercial use above that line is simply not licensed. Non-profits and research users are carved out.

For a hobbyist or a startup this is a non-issue. For a company that crosses that line mid-year — or that might be acquired by one — the two repos are not equivalent artefacts no matter how similar their parameter counts look, and the licence review happens long after someone has already shipped the prototype. It is the cheapest thing on this page to check early.

Running the pair as one pipeline

Neither model is on our catalogue today, so this is not an availability claim: both are self-hosted artefacts, and Granite 4.2 3B in particular has no inference providers listed on its card at all. But the pattern a team actually ends up with is one call that decides and another that speaks, and that pattern is where a routing layer earns its place. OrcaRouter puts more than 200 models behind one API key with provider list prices passed through at 0% markup, so a vendor price cut reaches your bill the same day rather than at the next contract renewal, and automatic failover across providers means the shared component in the middle of a pipeline does not become a single point of failure while you are still evaluating it. If you want to A/B d1-3B against a hosted generalist on your own traffic before committing to a self-hosted deployment, the nearest routed neighbour to Granite's lineage is Gemma 4 31B at $0.13 per million input tokens and $0.38 per million output — a much larger model, but the same "generalist that writes" role in the pipeline.

The verdict, stated as a rule

If the thing you need is a judgement — spam or not, which queue, how urgent, does this image match that description — d1-3B is the only one of the two that does the job in one pass, and it does it in single-digit milliseconds. If the thing you need is a written answer, a long document read, a tool call or a piece of code, Granite 4.2 3B is the one with a token stream at all, and the 3B's reasoning scores are genuinely impressive for its size — on IBM's own evals, with nobody having checked them yet.

What would change the picture is not a new benchmark row. It is a third party scoring both models on the same calibration harness, because the property that decides whether you can put a threshold on a model is the one neither card lets you compare today.