A hero title card headed 'Liquid AI d1-3B vs Qwen 3 8B' with the subheading 'should an 8B classification workload move down?', showing a pipeline of an incoming request branching into a small d1-3B box labelled 8 ms and 0 output tokens and a larger Qwen 3 8B box labelled writes an answer, with label and score chips leaving the small box and a prose chip leaving the large one.
Guides & Insights

Liquid AI d1-3B vs Qwen 3 8B: Should an 8B Classification Workload Move to a Decision Model?

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Somewhere in most production stacks there is a Qwen 3 8B being paid to answer yes or no. It moderates a comment, tags a support ticket, decides whether a retrieved passage is relevant, or scores another model's output against a rubric — and it does that by generating text which something downstream then parses. Liquid AI d1-3B, the 3.12B decision model Liquid AI open-weighted on October 7, 2026, exists to do exactly that job without generating anything: a state in, a set of named questions in, and probabilities, labels and confidences out in a single forward pass with zero output tokens. Qwen 3 8B is the vendor's April 2025 open-weights workhorse — roughly 8.2B dense parameters, 119 languages, thinking on or off per request, and sixteen months of community deployment behind it. The question this pairing actually asks is narrow and practical: for the slice of your traffic that is a classification, is an 8B generalist the cheapest way to get a label in 2026, or has the shape of that job changed underneath it?

What you are actually paying an 8B to do

Put the two output contracts side by side and the inefficiency is structural rather than incremental.

A Qwen 3 8B classification call is a generation. You write a prompt that describes the labels, the model optionally reasons about the input — and on this model you can turn that reasoning on or off per request, which is the single most useful knob it has for this kind of work — and then it writes back an answer. That answer is text. If you asked for JSON you will usually get JSON, and occasionally you will get JSON inside a sentence, or a preamble, or a trailing explanation you did not want. The label you care about is somewhere in the token stream, and a parser retrieves it. You are billed for every token the model writes, including the ones you throw away.

Liquid AI d1-3B has no token stream. Questions are declared with a type: noul for yes/no, answered as P(yes) between 0 and 1; choice for one label from a named option set, answered as the label plus a confidence plus a probability per option; score for a position on an ordered two-to-ten scale, answered as the expected level with its distribution and legend. The response carries a running total that reads output_tokens: 0. There is nothing to parse because nothing was written, and there is a raw probabilities accessor when you want the unrounded distribution rather than the argmax.

The part that changes your bill is what happens with several questions. One state can carry many: is this a refund request, which team should own it, how urgent is it. The state and its images are read once, and Liquid reports three questions over one state costing 1.3x the time of one rather than 3x. What it does not collapse is the input charge — the vendor's billing note is explicit that each question is billed as its own prompt, including its text and all of its images. Fewer latency, same input tokens. That is an honest asterisk on the headline, and it is the kind of thing you only find by reading the pricing paragraph rather than the benchmark.

A two-column scoreboard for Liquid AI d1-3B and Qwen 3 8B. The d1-3B column reads: job a closed question answered; 3.12B parameters; output tokens billed 0; 32,768-token context; image input yes; one question in 8 ms on an RTX 4090. The Qwen 3 8B column reads: job text in, text out; about 8.2B dense parameters; output tokens billed every one it writes; 32,768 native context extending to 131,072 via YaRN; image input no; one answer as long as the answer is. The footer reads 'd1-3B figures vendor-reported; Qwen 3 8B figures per Alibaba, April 2025.'

What sixteen months of Qwen 3 8B buys you

Before treating the smaller model as an upgrade, be precise about what the incumbent brings, because it is more than parameter count.

• Context — Qwen 3 8B runs 32,768 tokens natively and extends to 131,072 validated through YaRN. Liquid AI d1-3B runs 32,768 tokens flat, with no documented extension path. For a state that is a long document, the 8B is the only one of the two that can hold it.

• Languages — 119 languages and dialects for Qwen 3 8B against sixteen documented for Liquid AI d1-3B.

• Reasoning control — Qwen 3 8B lets you turn thinking on or off per request. Liquid AI d1-3B has no reasoning mode to control, which is either a simplification or a limitation depending on what you were using the thinking for.

• Breadth — the 8B writes, explains, summarises, translates and codes. Liquid AI d1-3B does none of those things. Its card states it plainly: it is not a chat model and it does not write text.

• Evidence — Qwen 3 8B has sixteen months of public benchmarking, quantisation, fine-tuning and production use. Liquid AI d1-3B's Decision Index score of 48.57 was produced by Liquid using the official scorer rather than submitted to the public leaderboard, and it is days old with no third-party reproduction.

• Vision — this row runs the other way. Liquid AI d1-3B takes images, with the vendor reporting 74.1 across eleven public image benchmarks read as decisions over each benchmark's option set, and reporting that stripping the images collapses the same questions to 45.1. The text-only Qwen 3 8B needs a separate vision model in front of it for any of that.

• Speed — one question in 8 ms on an RTX 4090, 16 ms on a Jetson AGX Thor, 50 ms on a Jetson Orin Nano, and 64 packed states per pass at 475 per second on the 4090. A generation-based classifier has no comparable number, because its latency depends on how long the answer is.

The honest summary is that the 8B is the better general instrument and the 3B decision model is the better specialised one, and the only question that matters is how much of your traffic is the specialised case.

The cost arithmetic, done carefully

This is where the comparison either pays for itself or does not, so it is worth doing with real structure rather than a slogan.

The claim Liquid published two days before the open-weights release is that its hosted decision model matches or beats GPT-6.1 Sol on four of six real applications, at 19x to 200x lower cost, and answers faster on every task. Read the methodology before repeating it: each application was run once per model on October 5, 2026, the chat models answered in JSON at their default reasoning setting, and d1's cost was computed at $0.04 per million input tokens. That $0.04 figure is the hosted d1 input rate, and it is the only rate that exists — there is no output line to add.

Now build the same shape of estimate for a Qwen 3 8B classifier and the asymmetry is visible without quoting anybody's provider price. The 8B has two billable lines where the decision model has one, and the second line is the one that grows: a classification that reasons first pays for the reasoning, and a classification that pads its JSON pays for the padding. The decision model's input charge is fixed by the state and the questions. The 8B's is not fixed by anything you can predict from the state alone.

Two counterweights keep this from being a rout. First, the decision model bills each question as its own prompt, so asking five questions about one long document quintuples the input — the 8B asks all five in one call and pays for the document once. Second, and larger, a decision model can only answer questions it was post-trained to answer in that shape. Anything that requires an explanation, a summary or a judgement expressed in words puts you back on the generalist and, in practice, back on paying for both.

A capture of Liquid AI's blog post 'Open d1: Edge decision models for text, vision, and audio' dated Oct 7, 2026, showing the opening paragraph announcing d1-3B and d1-omni-600M as open-weight models, the d1-3B Decision Index score of 48.57, and its latency of 8 ms on an RTX 4090, 16 ms on a Jetson AGX Thor and 26 ms on a Jetson AGX Orin.

What the two are genuinely good at

Strip the benchmarking out and the split is cleaner than the parameter counts suggest.

Liquid AI d1-3B is for the yes/no, the pick-from-a-list and the rating, at a latency floor nothing generative reaches, on hardware that includes a Jetson Orin Nano at 50 ms and a laptop. The applications Liquid demonstrated in October were all versions of that shape — a SQL predicate answering yes or no for 150 support tickets, an agent context-compaction loop that keeps, trims or drops each tool output and removes 52% of the tokens, an inspection app sorting good and defective parts from four production lines at 85 to 97% accuracy on a task it had never been trained for and understood from a short description. That last demo is the clearest statement of what the category is for: a job where the answer is one of a few things, the state is an image, and no explanation was ever requested.

Qwen 3 8B is for the work that has to come back as language, or span 119 languages, or hold a document longer than thirty-two thousand tokens, or reason out loud before committing. It is also the right answer when the classification is not one of the three shapes above — when the label set is open-ended, when the criteria are nuanced enough that you wanted the thinking, when the same call has to handle the easy and the hard case without you routing between two models. And it is the safe pick when a reviewer asks what independent benchmarks exist, because the answer is: a lot, going back a year and a half.

The teams that get this right do not pick. They put the decision model in front of the generalist, on the portion of traffic where the answer is a number, and leave the 8B handling everything that needs a sentence. The filter is 3B and 8 ms; the 8B stays for the payload.

Deploying the pair

Liquid AI d1-3B arrives as weights with the deployment work already done: day-one llama.cpp support, the full NVIDIA stack from DGX down to Jetson, NVFP4 quantisation, an 8-bit weight-and-activation variant, and GGUF conversions on Hugging Face. Qwen 3 8B has the deepest self-hosting ecosystem of any model in this weight class. Neither is on our catalogue, and nothing here is an availability claim for either — the hosted route for Liquid AI d1-3B is the vendor's own API and third-party platforms, where the hosted d1 is priced on input tokens only at $0.04 per million with images billed at 1.5 tokens per 32×32-pixel patch.

Once both are in the same pipeline the routing problem stops being theoretical. You have a small model you own, an 8B you may host or rent, and the generative calls you certainly rent — and deciding per request which one a given input should hit is exactly the decision a routing layer exists to make. One endpoint across 200+ models, provider list prices passed through at 0% markup so a vendor's price change is live on your side the same day, automatic failover so an upstream's bad afternoon does not become your outage. When your classification traffic is measured in billions of calls a year, that layer is where the savings from a 3B decision model actually get collected.

The verdict

If your 8B is being asked closed questions — is this toxic, is this relevant, which queue does this belong in, how urgent is it — you are paying a generalist's two-line bill and a generation's latency for a label, and Liquid AI d1-3B is a straight replacement with a weaker evidence trail and a real license difference to weigh. Accept the 32K context ceiling, the sixteen languages, and the fact that nobody outside Liquid has scored it yet.

If your 8B is being asked to explain, summarise, translate, hold a long document or reason before deciding, this comparison does not apply to you. Keep the 8B, and consider the decision model only as a pre-filter for the traffic you can identify in advance as the easy kind.

The interesting number to watch is not either model's benchmark score. It is how quickly the first independent reproduction of the d1 Decision Index lands, because that is the moment the comparison stops being a vendor's claim and starts being a fact you can plan against.

A capture of our own catalogue page for Qwen3 VL 8B Instruct, showing the model id qwen/qwen3-vl-8b-instruct, a release date of 2025-10-14, text, image and video input, a 131,072-token context window with 32K max output, a P50 time-to-first-token of 3.59 seconds, and pricing of $0.18 per million input tokens against $0.70 per million output tokens.