A generated title card for Decision 3.0 subtitled 'six open multimodal decision models, 0.6B to 27B', with chips reading 'weights Apache-2.0', 'images and video', 'no launch post yet', and a footer reading 'All figures in this article are vLLM-SR's own; nothing here is independently reproduced.' The OrcaRouter logo is composited in the bottom-right corner.
Engineering & Research

Decision 3.0 Is on Hugging Face: Six Open Multimodal Decision Models, Announced Only on X

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Decision 3.0 exists in a form you can download right now, and almost nobody has written it down. Six checkpoints — d3, d3-flash, d3-mini, d3-nano, d3-lite and d3-edge — landed in the vllm-sr organisation on Hugging Face on 10 October 2026, newest-first starting at 09:38 UTC and finishing with d3-edge at 23:49 UTC. They are Apache-2.0, they are built on Qwen3.5 and Qwen3.8 backbones, they answer choice, yes/no and score questions with a probability per option instead of generating a token, and five of the six now read images and video. Every model card describes itself as "the 27B multimodal foundation decision model of Decision 3.0, the decision models of vLLM Semantic Router," which is the vLLM project's open decision layer.

What is missing is the announcement. The vLLM project's X account congratulated the vLLM-SR team on the release on 11 October, quoting the maintainer's own introductory thread — that is the entire public record. The project's blog index still ends on 10 October with a post about routing decision models, not about these ones. The GitHub releases page for vllm-project/semantic-router still tops out at v0.4.0 from 27 September. Weights shipped; the launch note has not.

That distinction matters for how you read everything below. Every performance figure in this article comes from vLLM-SR's own model cards, measured with their own harness on their own hardware. Nothing here has been reproduced by an outside party, and the board they publish on is one they maintain. This is an early, honest read of an artifact that is real and a claim that is not yet audited.

What is actually on Hugging Face

The six repositories are simple, complete, and consistent with each other. Each carries safetensors weights, a chat template, a decision_config.json that pins the base-model revision, custom modelling code (modeling_d3.py, d3_engine.py, d3_server.py), a readout head in its own readout.safetensors, and a MODEL_MANIFEST.json with a SHA-256 per file. A seventh repository, the collection page, lists all six.

The family, in the order the sizes fall:

• d3 — 26.09B parameters including a 0.46B vision encoder, built on Qwen/Qwen3.8-27B. Uploaded 10 October, 09:38 UTC.

• d3-flash — 8.39B including a 0.46B vision encoder, built on Qwen/Qwen3.5-9B.

• d3-mini — 4.54B including a 0.33B vision encoder, built on Qwen/Qwen3.5-4B.

• d3-nano — 2.21B including a 0.33B vision encoder, built on Qwen/Qwen3.5-2B.

• d3-lite — 0.85B including a 0.10B vision encoder, built on Qwen/Qwen3.5-0.8B.

• d3-edge — 0.59B including a 0.10B vision encoder, also built on Qwen/Qwen3.5-0.8B. Uploaded last, 23:49 UTC.

All six carry the same request contract: a state (text or JSON, optionally with images and videos) plus a dictionary of named questions, and a response of typed probability distributions. Three question kinds exist — Choice, Noul (yes/no), Score — and the model answers all of them in one call by running one forward pass per question over the same input, with no decoding loop anywhere.

A screenshot of the vLLM-SR collection page on Hugging Face, headed Decision 3.0 with the subtitle 'Towards Open Multimodal Foundation Decision Models' and marked as updated about 15 hours ago with 25 upvotes. It lists six repositories, each tagged Zero-Shot Classification: vllm-sr/d3 at 26B with 130 downloads and 18 likes, vllm-sr/d3-flash at 8B with 69 downloads, vllm-sr/d3-mini at 5B with 62, vllm-sr/d3-nano at 2B with 82, vllm-sr/d3-lite at 0.9B with 83, and vllm-sr/d3-edge at 0.6B with 33. The left sidebar lists the organisation's other collections: Vela 2.0, Decision 2.0, Decision 1.0, Vela 1.0, MoM 1.0 and MoM Nano.

The numbers, and whose they are

vLLM-SR publishes a leaderboard it calls the Jev Decision Index 0.3.1 and puts d3 at the top of it. The headline rows, all vendor-reported:

• d3 — Index 64.7, public suite 65.5, same-skill tests 62.8, new-domain tasks 56.6.

• Perplexity Decider v1.1 (27B) — 62.8, 62.3, 61.1, 55.6.

• Fastino GLiDE no-thinking (28B) — 60.2, 59.1, 59.5, 52.9.

• Jev — 60.1, 58.0, 58.0, 55.0.

• Torchcast Decision 27B — 59.9, 65.1, 58.1, 50.8.

• Decision 2.0 (27B), the previous generation — 55.9, 57.0, 55.7, 47.9.

A footnote on d3's card states plainly that d3's own row is an internal evaluation while the comparison rows are live board data read on 10 October. That is the right disclosure to make and it is worth reading twice: the competitor numbers were pulled from a board, and the subject number was produced by the team that built the model. The card also claims all 140,178 public requests were answered with none unsupported — a coverage claim, not an accuracy claim, and an unusual one to publish.

A screenshot of the vllm-sr/d3 model card on Hugging Face. The header shows 18 likes, the vLLM Semantic Router organisation, and tags for Transformers, Safetensors, qwen3_5, feature-extraction, decision-model, classification, system-one, multimodal, vision, video, custom_code and an Apache-2.0 licence, with the model size listed as 26B parameters in BF16. The card graphic reads 'Decision 3.0 / d3 27B'. A specification panel gives Parameters as 26.09B including the 0.46B vision encoder, Inputs as 'Text or JSON, plus images and videos (several per request)', Decision types as 'Choice - Yes / No - Score' and License as Apache-2.0. The Highlights section states a Jev Decision Index 0.3 public suite score of 65.45 measured with the official 0.3 kit on the released weights, all 140,178 public requests answered and none unsupported, and +8.5 on the public suite over Decision 2.0. The sidebar shows 130 downloads last month, a model tree pinned to Qwen/Qwen3.8-27B, and two Spaces using the model.

The smaller checkpoints report themselves as marching in step. Each card gives a public-suite figure and a delta against the equivalent size in Decision 2.0: d3-flash 57.79, up 11.0 over the previous 9B; d3-mini 54.90, up 10.7; d3-nano 42.22, up 12.2; d3-lite 36.93, up 16.4; d3-edge 28.82, up 12.1. The relative gains are largest at the bottom of the range, which is what you would expect if the multimodal pretraining recipe is doing more work for small models than large ones — and also what you would produce by tuning the small models hardest. There is no way to tell which from the outside.

The multimodal part is the real change

Decision 2.0's models read text. Decision 3.0's read text, images and video, and that is the substantive difference between the generations — not the index movement. Each card lists image input as PNG, JPEG or WebP, supplied as paths, URLs, PIL images or base64 data URLs, with several per request at up to 1.6 megapixels each; and video as MP4, WebM, MOV or MKV, or as frame arrays, read at 2 frames per second with at most 32 frames spread across the whole clip and each frame at up to 0.2 megapixels. Every question in a request sees every image and every video attached to it.

The supporting evidence for video is a Perception Test result on all 19,140 validation questions: d3 at 77.4, d3-flash 73.3, d3-mini 70.3, d3-nano 63.5, d3-lite 59.3, d3-edge 46.6, against a documented chance floor of 33.3 on three-option questions. vLLM-SR labels this an internal evaluation. d3 also carries a vision board placing it at 71.6 overall, ahead of Perplexity Decider v1.1 at 70.6 and JEV-27B-VL at 69.6, with the comparison rows again drawn from board data rather than run by the authors.

The throughput figures are single-request, single-GPU, and stated as such: d3 answers a text request in a median 55 ms, a request carrying an image in 273 ms, and a request carrying a ten-second video in 553 ms, on one AMD Instinct MI325X. d3-edge does the same in 6.7 ms, 41.9 ms and 308.3 ms. Those are per-question medians on a model designed to be called many times inside one workflow, which is the number that matters for the architecture these models are built for — twenty sequential decisions at an extra 100 ms each costs two seconds, as Microsoft made the same point about its own decision model this month.

What the repositories do not say

Three gaps are worth naming before anyone plans around this.

The first is the input budget. decision_config.json for the largest model sets max_length to null, and no card states a token ceiling for state plus questions plus option descriptions. Decision 1.0's cards published explicit budgets — 1,024 tokens for the encoder branch, 16,384 for the decoder branch — and those numbers are a real planning constraint. Their absence here is conspicuous, and it is the single most useful thing a follow-up commit could add.

The second is calibration. A decision model's most important number is not accuracy, it is whether a returned 0.9 means nine times out of ten. The vision and text indices say nothing about that. There is no Brier score, no expected-calibration-error figure, and no temperature in any of the six cards. InternLM published both on its own decision model in September; vLLM-SR has not, for any member of this family.

The third is independence. Decision 2.0's models have had two weeks of outside use; Decision 3.0's have hours. The download counts on 11 October — 130 for d3, 83 for d3-lite, 82 for d3-nano — are the footprint of an upload, not of adoption. Treat every index figure above as a hypothesis with a repository attached.

Where this sits in a stack you can call today

The practical question about a decision model is whether it drops into a pipeline that already exists, and here the ecosystem is further along than the models are. The System One request format these models use is not proprietary to vLLM-SR: it is the same state-and-named-questions contract the hosted Jev models are called with, which is why "Jev" appears both as the name of the index in d3's model card and as a row on its leaderboard.

OrcaRouter carries that side of the family. typesafe/jev-1.13 is in our catalogue and is served over POST /v1/systemone — the same contract, at $0.042 per million input tokens with no completion charge because there is no completion, up to roughly 64K input tokens, text in and structured JSON out, non-streaming. It is the kind of model a decision layer calls today. Two things we do not claim: d3 and the rest of Decision 3.0 are not in our catalogue, and Decision 3.0's local inference path is a Python class in a Hugging Face repository, not an endpoint we could route even if we wanted to. What a router does buy you here is on the other side of the loop — the generative model that acts on a decision, routed through one key across 200-plus models with automatic failover when a provider wobbles, so that adding a scoring step in front of a workflow does not also mean adding a second vendor relationship.

What would change this picture

Four things, in rough order of how much they would tell you. A blog post from the project stating the input budget and the intended deployment shape, which would confirm this is a release rather than a push. A Brier score or an ECE figure on any one checkpoint. A third party running the public suite on the released weights rather than reading the index. And a runtime — the semantic-router repository landed two commits on 11 October that wire d3 and d3-edge into its model runtime, including video input, which suggests the serving story is being built now and will be the thing that makes these models cheap to try.

Until at least the first of those lands, the honest summary is this: there is a complete, Apache-2.0, genuinely multimodal decision-model family sitting on Hugging Face from 0.6B to 27B, published with unusually good provenance — file hashes, pinned base revisions, a stated hardware target — and with unusually little of the surrounding documentation that would let you decide whether to use it. Downloading it costs nothing. Believing the index should wait.

A generated six-row scoreboard titled 'Decision 3.0 - the scoreboard', listing the family from largest to smallest with each checkpoint's parameter count and its vLLM-SR-reported Jev Decision Index 0.3 public-suite figure: d3 26.09B and 65.5, d3-flash 8.39B and 57.79, d3-mini 4.54B and 54.90, d3-nano 2.21B and 42.22, d3-lite 0.85B and 36.93, d3-edge 0.59B and 28.82, with a footer reading 'All figures are vLLM-SR's own; nothing here is independently reproduced.'

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily