
Decision 3.0 Is on Hugging Face: Six Open Multimodal Decision Models, Announced Only on X
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 70 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 116 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 48 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 478 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 404 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
Decision 3.0 exists in a form you can download right now, and almost nobody has written it down. Six checkpoints — d3, d3-flash, d3-mini, d3-nano, d3-lite and d3-edge — landed in the vllm-sr organisation on Hugging Face on 10 October 2026, newest-first starting at 09:38 UTC and finishing with d3-edge at 23:49 UTC. They are Apache-2.0, they are built on Qwen3.5 and Qwen3.8 backbones, they answer choice, yes/no and score questions with a probability per option instead of generating a token, and five of the six now read images and video. Every model card describes itself as "the 27B multimodal foundation decision model of Decision 3.0, the decision models of vLLM Semantic Router," which is the vLLM project's open decision layer.
What is missing is the announcement. The vLLM project's X account congratulated the vLLM-SR team on the release on 11 October, quoting the maintainer's own introductory thread — that is the entire public record. The project's blog index still ends on 10 October with a post about routing decision models, not about these ones. The GitHub releases page for vllm-project/semantic-router still tops out at v0.4.0 from 27 September. Weights shipped; the launch note has not.
That distinction matters for how you read everything below. Every performance figure in this article comes from vLLM-SR's own model cards, measured with their own harness on their own hardware. Nothing here has been reproduced by an outside party, and the board they publish on is one they maintain. This is an early, honest read of an artifact that is real and a claim that is not yet audited.
What is actually on Hugging Face
The six repositories are simple, complete, and consistent with each other. Each carries safetensors weights, a chat template, a decision_config.json that pins the base-model revision, custom modelling code (modeling_d3.py, d3_engine.py, d3_server.py), a readout head in its own readout.safetensors, and a MODEL_MANIFEST.json with a SHA-256 per file. A seventh repository, the collection page, lists all six.
The family, in the order the sizes fall:
• d3 — 26.09B parameters including a 0.46B vision encoder, built on Qwen/Qwen3.8-27B. Uploaded 10 October, 09:38 UTC.
• d3-flash — 8.39B including a 0.46B vision encoder, built on Qwen/Qwen3.5-9B.
• d3-mini — 4.54B including a 0.33B vision encoder, built on Qwen/Qwen3.5-4B.
• d3-nano — 2.21B including a 0.33B vision encoder, built on Qwen/Qwen3.5-2B.
• d3-lite — 0.85B including a 0.10B vision encoder, built on Qwen/Qwen3.5-0.8B.
• d3-edge — 0.59B including a 0.10B vision encoder, also built on Qwen/Qwen3.5-0.8B. Uploaded last, 23:49 UTC.
All six carry the same request contract: a state (text or JSON, optionally with images and videos) plus a dictionary of named questions, and a response of typed probability distributions. Three question kinds exist — Choice, Noul (yes/no), Score — and the model answers all of them in one call by running one forward pass per question over the same input, with no decoding loop anywhere.

The numbers, and whose they are
vLLM-SR publishes a leaderboard it calls the Jev Decision Index 0.3.1 and puts d3 at the top of it. The headline rows, all vendor-reported:
• d3 — Index 64.7, public suite 65.5, same-skill tests 62.8, new-domain tasks 56.6.
• Perplexity Decider v1.1 (27B) — 62.8, 62.3, 61.1, 55.6.
• Fastino GLiDE no-thinking (28B) — 60.2, 59.1, 59.5, 52.9.
• Jev — 60.1, 58.0, 58.0, 55.0.
• Torchcast Decision 27B — 59.9, 65.1, 58.1, 50.8.
• Decision 2.0 (27B), the previous generation — 55.9, 57.0, 55.7, 47.9.
A footnote on d3's card states plainly that d3's own row is an internal evaluation while the comparison rows are live board data read on 10 October. That is the right disclosure to make and it is worth reading twice: the competitor numbers were pulled from a board, and the subject number was produced by the team that built the model. The card also claims all 140,178 public requests were answered with none unsupported — a coverage claim, not an accuracy claim, and an unusual one to publish.

The smaller checkpoints report themselves as marching in step. Each card gives a public-suite figure and a delta against the equivalent size in Decision 2.0: d3-flash 57.79, up 11.0 over the previous 9B; d3-mini 54.90, up 10.7; d3-nano 42.22, up 12.2; d3-lite 36.93, up 16.4; d3-edge 28.82, up 12.1. The relative gains are largest at the bottom of the range, which is what you would expect if the multimodal pretraining recipe is doing more work for small models than large ones — and also what you would produce by tuning the small models hardest. There is no way to tell which from the outside.
The multimodal part is the real change
Decision 2.0's models read text. Decision 3.0's read text, images and video, and that is the substantive difference between the generations — not the index movement. Each card lists image input as PNG, JPEG or WebP, supplied as paths, URLs, PIL images or base64 data URLs, with several per request at up to 1.6 megapixels each; and video as MP4, WebM, MOV or MKV, or as frame arrays, read at 2 frames per second with at most 32 frames spread across the whole clip and each frame at up to 0.2 megapixels. Every question in a request sees every image and every video attached to it.
The supporting evidence for video is a Perception Test result on all 19,140 validation questions: d3 at 77.4, d3-flash 73.3, d3-mini 70.3, d3-nano 63.5, d3-lite 59.3, d3-edge 46.6, against a documented chance floor of 33.3 on three-option questions. vLLM-SR labels this an internal evaluation. d3 also carries a vision board placing it at 71.6 overall, ahead of Perplexity Decider v1.1 at 70.6 and JEV-27B-VL at 69.6, with the comparison rows again drawn from board data rather than run by the authors.
The throughput figures are single-request, single-GPU, and stated as such: d3 answers a text request in a median 55 ms, a request carrying an image in 273 ms, and a request carrying a ten-second video in 553 ms, on one AMD Instinct MI325X. d3-edge does the same in 6.7 ms, 41.9 ms and 308.3 ms. Those are per-question medians on a model designed to be called many times inside one workflow, which is the number that matters for the architecture these models are built for — twenty sequential decisions at an extra 100 ms each costs two seconds, as Microsoft made the same point about its own decision model this month.
What the repositories do not say
Three gaps are worth naming before anyone plans around this.
The first is the input budget. decision_config.json for the largest model sets max_length to null, and no card states a token ceiling for state plus questions plus option descriptions. Decision 1.0's cards published explicit budgets — 1,024 tokens for the encoder branch, 16,384 for the decoder branch — and those numbers are a real planning constraint. Their absence here is conspicuous, and it is the single most useful thing a follow-up commit could add.
The second is calibration. A decision model's most important number is not accuracy, it is whether a returned 0.9 means nine times out of ten. The vision and text indices say nothing about that. There is no Brier score, no expected-calibration-error figure, and no temperature in any of the six cards. InternLM published both on its own decision model in September; vLLM-SR has not, for any member of this family.
The third is independence. Decision 2.0's models have had two weeks of outside use; Decision 3.0's have hours. The download counts on 11 October — 130 for d3, 83 for d3-lite, 82 for d3-nano — are the footprint of an upload, not of adoption. Treat every index figure above as a hypothesis with a repository attached.
Where this sits in a stack you can call today
The practical question about a decision model is whether it drops into a pipeline that already exists, and here the ecosystem is further along than the models are. The System One request format these models use is not proprietary to vLLM-SR: it is the same state-and-named-questions contract the hosted Jev models are called with, which is why "Jev" appears both as the name of the index in d3's model card and as a row on its leaderboard.
OrcaRouter carries that side of the family. typesafe/jev-1.13 is in our catalogue and is served over POST /v1/systemone — the same contract, at $0.042 per million input tokens with no completion charge because there is no completion, up to roughly 64K input tokens, text in and structured JSON out, non-streaming. It is the kind of model a decision layer calls today. Two things we do not claim: d3 and the rest of Decision 3.0 are not in our catalogue, and Decision 3.0's local inference path is a Python class in a Hugging Face repository, not an endpoint we could route even if we wanted to. What a router does buy you here is on the other side of the loop — the generative model that acts on a decision, routed through one key across 200-plus models with automatic failover when a provider wobbles, so that adding a scoring step in front of a workflow does not also mean adding a second vendor relationship.
What would change this picture
Four things, in rough order of how much they would tell you. A blog post from the project stating the input budget and the intended deployment shape, which would confirm this is a release rather than a push. A Brier score or an ECE figure on any one checkpoint. A third party running the public suite on the released weights rather than reading the index. And a runtime — the semantic-router repository landed two commits on 11 October that wire d3 and d3-edge into its model runtime, including video input, which suggests the serving story is being built now and will be the thing that makes these models cheap to try.
Until at least the first of those lands, the honest summary is this: there is a complete, Apache-2.0, genuinely multimodal decision-model family sitting on Hugging Face from 0.6B to 27B, published with unusually good provenance — file hashes, pinned base revisions, a stated hardware target — and with unusually little of the surrounding documentation that would let you decide whether to use it. Downloading it costs nothing. Believing the index should wait.

Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
