
Liquid AI d1-3B vs Gemma 4 12B: A Decision Model Against a Generalist
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 128 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 62 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 358 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The fastest way to get this comparison wrong is to put Liquid AI d1-3B and Gemma 4 12B on the same benchmark and wait for a winner. They do not answer the same question. Liquid AI d1-3B is a 3.12B decision model published as open weights on October 7, 2026: you give it a state and a set of named questions and it returns probabilities, a chosen label and a confidence, with zero output tokens and no generation step. Gemma 4 12B is Google's encoder-free any-to-any generalist, an 11.96B checkpoint that takes text, image, audio and video and writes an answer over a 256,000-token window in more than 140 languages. One of them tells you which of four things a circuit board is. The other tells you why. Compare them on quality alone and you learn nothing, because the outputs are not commensurable — and the vendors have never benchmarked them against each other. Compare them on what a call actually returns, what it costs, and where each one runs out of road, and the pairing becomes a genuinely useful boundary test.
Two different output contracts
The distinction that does all the work here is what comes back on the wire.
Liquid AI d1-3B returns a structured answer object. Every question in a request declares a type — noul for yes/no with P(yes), choice for one of a named option set with a confidence and a probability per option, or score for a position on an ordered two-to-ten-level scale with the expected level and its distribution. The model reports output_tokens: 0 because nothing is written. Several questions over one state are read from a single forward pass, and the state and its images are encoded once for all of them.
Gemma 4 12B returns prose. Sometimes it returns JSON you asked for, sometimes it returns JSON you did not ask for exactly, and it can reason before answering. It is a language model first: anything it knows how to classify, it expresses as text. That is a strength when the answer must be explained to a human or fed to a downstream step that expects language, and a cost when the only thing you needed was a probability.
Everything downstream follows from that. A d1-3B response cannot fail to parse. It also cannot explain itself, and the model card says so directly: it is not a chat model and it does not write text.
Where the evidence trails stand
The provenance of the two number sets is not equivalent, and that matters more here than the numbers do.
• Decision Index 0.2.1 — Liquid AI d1-3B scores 48.57, scored by the vendor with the official scorer, first among models under 10B and ahead of Decider 35B-A3B at 47.11. Gemma 4 12B is not scored on the Decision Index at all, so this row has no counterpart.
• Reasoning benchmarks — Gemma 4 12B posts 77.5 on AIME 2026, 78.8 on GPQA Diamond and 1,659 Codeforces ELO, and carries an Artificial Analysis Intelligence Index of 22 in reasoning mode and 13 with reasoning off. Liquid AI d1-3B has no reasoning benchmark, because a decision model does not produce a reasoning trace to score.
• Long context — Gemma 4 12B accepts 256,000 tokens and scores 43.4% on MRCR v2 eight-needle at 128k on Google's own card. Liquid AI d1-3B accepts 32,768 tokens. If your "state" is a forty-page document plus scans, that gap is the whole decision and it is not close.
• Vision — Liquid AI d1-3B averages 74.1 across eleven public image benchmarks read as decisions over each benchmark's option set, against 73.9 for the LFM2.5-VL-3B backbone it was built from, and the vendor reports that removing the images drops the same questions to 45.1. Gemma 4 12B's vision is stronger and broader, but it is reported as generation quality, not as decision accuracy over named options.
• Languages — sixteen documented for Liquid AI d1-3B; more than 140 for Gemma 4 12B.
• Speed — Liquid AI d1-3B answers one question in 8 ms on an RTX 4090, 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin 64 GB and 50 ms on a Jetson Orin Nano, and packs 64 states into a single pass at 475 per second on the 4090. Gemma 4 12B generates, so its latency is a function of how many tokens it writes.
• Evidence quality — Gemma 4 12B's results are Google's, published in June 2026 and since poked at, quantized and re-run by a large community. Liquid AI d1-3B's are Liquid's, published two days ago, reproduced by nobody outside the lab. Read the second column accordingly.

The one place they genuinely compete
There is a real overlap, and it is not on a leaderboard. It is the LLM-as-a-judge call.
Plenty of production pipelines already use a mid-size generalist as an assessment layer: score this ticket's urgency, decide whether this output passes a rubric, pick which tool the agent should call next, check whether a retrieved passage answers the question. Today most of those calls go to a generative model that is asked to emit a number or a label as text, and the number it emits is whatever the sampler produced rather than a softmax over your option set. That is the workflow Liquid AI d1-3B was post-trained to replace, and the published demos are all versions of it — a SQL predicate that answers yes or no for 150 support tickets, an agent context-compaction loop that keeps, trims or drops each tool output and removes 52% of the tokens, a visual-inspection app that sorted good and defective parts from four production lines at 85 to 97% accuracy on a task it was never trained for.
Gemma 4 12B does all of those jobs, and does more than those jobs. It can also write the summary, hold a 256K document in view, take a recorded phone call as input, and answer in one of 140-plus languages. If your pipeline needs an assessment and a narration, the 12B is doing two jobs with one call and the 3B decision model is doing one of them.
The boundary is context length and output shape. Long state, spoken input, many languages, or an explanation the reader expects — Gemma 4 12B, no contest. Short state, a fixed option set, a per-request latency budget in single-digit milliseconds, or a hard guarantee that the answer will parse — that is the d1-3B column, and the generalist is paying for capability you will not use.
What each one costs to call
Neither model is on our catalogue, so there is no routing price to quote for either — Gemma 4 12B is not among the Gemma sizes we serve, and Liquid AI d1-3B is open weights you host yourself or reach through the vendor's own API and third-party platforms. For context on the Gemini-adjacent side of that family, the two Gemma 4 sizes we do route are listed at $0.06 per million input and $0.33 per million output for Gemma 4 26B A4B, and $0.13 and $0.38 for Gemma 4 31B, both with 262,144-token contexts and text, image and video input. Median provider pricing quoted for Gemma 4 12B elsewhere sits near $0.10 and $0.30.
Liquid's hosted d1 bills input tokens only, at $0.04 per million, with images counted as input at 1.5 tokens per 32×32-pixel patch. A decision request therefore has no output-token line at all, which is the structural reason the vendor's own cost comparison against much larger models lands where it does. The open weights have no per-call price; you pay in hardware. Both have to sit in the same pipeline as everything else you call, which is the layer a router exists for — one API across 200+ models, provider list prices passed through at 0% markup, and automatic failover so a single upstream blip does not become your outage. That is the plumbing around the comparison, not the comparison.

How to decide, honestly
Start from the answer format rather than the model. If the downstream system can consume a probability, a label and a confidence, and the answer set is small and known, Liquid AI d1-3B is not a downgrade from a 12B — it is a different instrument, and the latency numbers make it a categorically different deployment: 50 ms on a Jetson Orin Nano is a device that has no business running a 12B at all.
If the answer has to be a sentence, or the state is longer than thirty-two thousand tokens, or someone is going to speak it, or the output language is Bengali, then Gemma 4 12B is the only one of the two that can do the job and the comparison is over before it starts.
The genuinely useful position for most teams is both, in different places. A decision model in front of the expensive call, doing the triage, the routing and the guardrail checks that do not need language; a generalist behind it for the work that does. Those are the same pipeline, one is a filter and one is the payload — and the teams that will get the most out of the d1 release are the ones already paying a 12B to answer yes or no.

