A hero title card headed 'Liquid AI d1-3B vs Gemma 4 12B' with the subheading 'a decision model against a generalist', showing a small 3.12B checklist card with a probability chip beside a larger 11.96B card carrying document, waveform and film-frame icons and a written-answer chip.
Guides & Insights

Liquid AI d1-3B vs Gemma 4 12B: A Decision Model Against a Generalist

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The fastest way to get this comparison wrong is to put Liquid AI d1-3B and Ge​mma 4 12B on the same benchmark and wait for a winner. They do not answer the same question. Liquid AI d1-3B is a 3.12B decision model published as open weights on October 7, 2026: you give it a state and a set of named questions and it returns probabilities, a chosen label and a confidence, with zero output tokens and no generation step. Ge​mma 4 12B is Goo​gle's encoder-free any-to-any generalist, an 11.96B checkpoint that takes text, image, audio and video and writes an answer over a 256,000-token window in more than 140 languages. One of them tells you which of four things a circuit board is. The other tells you why. Compare them on quality alone and you learn nothing, because the outputs are not commensurable — and the vendors have never benchmarked them against each other. Compare them on what a call actually returns, what it costs, and where each one runs out of road, and the pairing becomes a genuinely useful boundary test.

Two different output contracts

The distinction that does all the work here is what comes back on the wire.

Liquid AI d1-3B returns a structured answer object. Every question in a request declares a type — noul for yes/no with P(yes), choice for one of a named option set with a confidence and a probability per option, or score for a position on an ordered two-to-ten-level scale with the expected level and its distribution. The model reports output_tokens: 0 because nothing is written. Several questions over one state are read from a single forward pass, and the state and its images are encoded once for all of them.

Ge​mma 4 12B returns prose. Sometimes it returns JSON you asked for, sometimes it returns JSON you did not ask for exactly, and it can reason before answering. It is a language model first: anything it knows how to classify, it expresses as text. That is a strength when the answer must be explained to a human or fed to a downstream step that expects language, and a cost when the only thing you needed was a probability.

Everything downstream follows from that. A d1-3B response cannot fail to parse. It also cannot explain itself, and the model card says so directly: it is not a chat model and it does not write text.

Where the evidence trails stand

The provenance of the two number sets is not equivalent, and that matters more here than the numbers do.

• Decision Index 0.2.1 — Liquid AI d1-3B scores 48.57, scored by the vendor with the official scorer, first among models under 10B and ahead of Decider 35B-A3B at 47.11. Ge​mma 4 12B is not scored on the Decision Index at all, so this row has no counterpart.

• Reasoning benchmarks — Ge​mma 4 12B posts 77.5 on AIME 2026, 78.8 on GPQA Diamond and 1,659 Codeforces ELO, and carries an Artificial Analysis Intelligence Index of 22 in reasoning mode and 13 with reasoning off. Liquid AI d1-3B has no reasoning benchmark, because a decision model does not produce a reasoning trace to score.

• Long context — Ge​mma 4 12B accepts 256,000 tokens and scores 43.4% on MRCR v2 eight-needle at 128k on Goo​gle's own card. Liquid AI d1-3B accepts 32,768 tokens. If your "state" is a forty-page document plus scans, that gap is the whole decision and it is not close.

• Vision — Liquid AI d1-3B averages 74.1 across eleven public image benchmarks read as decisions over each benchmark's option set, against 73.9 for the LFM2.5-VL-3B backbone it was built from, and the vendor reports that removing the images drops the same questions to 45.1. Ge​mma 4 12B's vision is stronger and broader, but it is reported as generation quality, not as decision accuracy over named options.

• Languages — sixteen documented for Liquid AI d1-3B; more than 140 for Ge​mma 4 12B.

• Speed — Liquid AI d1-3B answers one question in 8 ms on an RTX 4090, 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin 64 GB and 50 ms on a Jetson Orin Nano, and packs 64 states into a single pass at 475 per second on the 4090. Ge​mma 4 12B generates, so its latency is a function of how many tokens it writes.

• Evidence quality — Ge​mma 4 12B's results are Goo​gle's, published in June 2026 and since poked at, quantized and re-run by a large community. Liquid AI d1-3B's are Liquid's, published two days ago, reproduced by nobody outside the lab. Read the second column accordingly.

A two-column scoreboard for Liquid AI d1-3B and Gemma 4 12B. The d1-3B column reads: output a label, a confidence, zero tokens; 3.12B parameters; 32,768-token context; text and image input; 8 ms per question; Decision Index 48.57 (vendor). The Gemma 4 12B column reads: output generated text; 11.96B parameters; 256,000-token context; text, image, audio and video input; latency that scales with output length; Decision Index not scored. The footer reads 'd1-3B figures vendor-reported and unreproduced; Gemma 4 12B figures per Google, June 2026.'

The one place they genuinely compete

There is a real overlap, and it is not on a leaderboard. It is the LLM-as-a-judge call.

Plenty of production pipelines already use a mid-size generalist as an assessment layer: score this ticket's urgency, decide whether this output passes a rubric, pick which tool the agent should call next, check whether a retrieved passage answers the question. Today most of those calls go to a generative model that is asked to emit a number or a label as text, and the number it emits is whatever the sampler produced rather than a softmax over your option set. That is the workflow Liquid AI d1-3B was post-trained to replace, and the published demos are all versions of it — a SQL predicate that answers yes or no for 150 support tickets, an agent context-compaction loop that keeps, trims or drops each tool output and removes 52% of the tokens, a visual-inspection app that sorted good and defective parts from four production lines at 85 to 97% accuracy on a task it was never trained for.

Ge​mma 4 12B does all of those jobs, and does more than those jobs. It can also write the summary, hold a 256K document in view, take a recorded phone call as input, and answer in one of 140-plus languages. If your pipeline needs an assessment and a narration, the 12B is doing two jobs with one call and the 3B decision model is doing one of them.

The boundary is context length and output shape. Long state, spoken input, many languages, or an explanation the reader expects — Ge​mma 4 12B, no contest. Short state, a fixed option set, a per-request latency budget in single-digit milliseconds, or a hard guarantee that the answer will parse — that is the d1-3B column, and the generalist is paying for capability you will not use.

What each one costs to call

Neither model is on our catalogue, so there is no routing price to quote for either — Ge​mma 4 12B is not among the Ge​mma sizes we serve, and Liquid AI d1-3B is open weights you host yourself or reach through the vendor's own API and third-party platforms. For context on the Gemini-adjacent side of that family, the two Ge​mma 4 sizes we do route are listed at $0.06 per million input and $0.33 per million output for Ge​mma 4 26B A4B, and $0.13 and $0.38 for Ge​mma 4 31B, both with 262,144-token contexts and text, image and video input. Median provider pricing quoted for Ge​mma 4 12B elsewhere sits near $0.10 and $0.30.

Liquid's hosted d1 bills input tokens only, at $0.04 per million, with images counted as input at 1.5 tokens per 32×32-pixel patch. A decision request therefore has no output-token line at all, which is the structural reason the vendor's own cost comparison against much larger models lands where it does. The open weights have no per-call price; you pay in hardware. Both have to sit in the same pipeline as everything else you call, which is the layer a router exists for — one API across 200+ models, provider list prices passed through at 0% markup, and automatic failover so a single upstream blip does not become your outage. That is the plumbing around the comparison, not the comparison.

A capture of Liquid AI's blog post 'Open d1: Edge decision models for text, vision, and audio' dated Oct 7, 2026, showing the opening paragraph announcing d1-3B and d1-omni-600M as open-weight models, the d1-3B Decision Index score of 48.57, and the latency figures of 8 ms on an RTX 4090, 16 ms on a Jetson AGX Thor and 26 ms on a Jetson AGX Orin.

How to decide, honestly

Start from the answer format rather than the model. If the downstream system can consume a probability, a label and a confidence, and the answer set is small and known, Liquid AI d1-3B is not a downgrade from a 12B — it is a different instrument, and the latency numbers make it a categorically different deployment: 50 ms on a Jetson Orin Nano is a device that has no business running a 12B at all.

If the answer has to be a sentence, or the state is longer than thirty-two thousand tokens, or someone is going to speak it, or the output language is Bengali, then Ge​mma 4 12B is the only one of the two that can do the job and the comparison is over before it starts.

The genuinely useful position for most teams is both, in different places. A decision model in front of the expensive call, doing the triage, the routing and the guardrail checks that do not need language; a generalist behind it for the work that does. Those are the same pipeline, one is a filter and one is the payload — and the teams that will get the most out of the d1 release are the ones already paying a 12B to answer yes or no.

A capture of our own catalogue page for Google Gemma 4 31B, showing the model id google/gemma-4-31b-it, a release date of 2026-04-02, a 30.7B dense multimodal description with a 256K token context window, and headline pricing of $0.13 per million input tokens and $0.38 per million output tokens.