A generated title card for the comparison of Decision 3.0 and Microsoft-Decision-1, subtitled 'same Qwen3.5-9B backbone, opposite ways to buy it', with chips reading 'weights Apache-2.0 vs hosted only', 'images and video vs text only', 'published 10 Oct vs GA 8 Oct', and a footer reading 'Decision 3.0 figures are vLLM-SR's own; Microsoft-Decision-1 figures are Microsoft's own; neither is independently reproduced.' The OrcaRouter logo is composited in the bottom-right corner.
Engineering & Research

Decision 3.0 vs Microsoft-Decision-1: One Is a Download, the Other Is a SKU

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The 9B tier of Decision 3.0 and Microsoft-Decision-1 are, on paper, the same product. Both post-train Qwen3.5-9B into a scorer that answers a fixed set of candidate answers with a calibrated probability each, without ever generating a token. Both are aimed at the same jobs — routing a request, grading an output against a rubric, gating an agent action, triaging a queue. Both came out within a week of each other: Microsoft-Decision-1 went generally available on Microsoft Foundry on 8 October 2026 and was announced the next day, and d3-flash was pushed to Hugging Face at 17:23 UTC on 10 October 2026.

The difference is not the model. It is what you are allowed to do with it. d3-flash is an 8.39-billion-parameter Apache-2.0 checkpoint you download and run, sitting third in a family that starts at 0.59B and tops out at 26.09B. Microsoft-Decision-1 is a hosted endpoint on Foundry, with weights that are not distributed at all, in a line Microsoft says it will rebase onto its own MAI models later. One of these is an artifact; the other is a service. Almost every practical question about the pair follows from that single fact, including the ones that look like they are about benchmarks.

Two decisions about what a decision model is for

Microsoft's post is explicit about scope in a way model cards rarely are. The supported question shapes are yes/no, multiple-choice, rating, classification and rubric-based grading. The exclusions are listed just as plainly: not designed for text generation, open-ended question answering, conversation, translation or summarization, and not intended for tasks requiring knowledge absent from the input. It runs one pass over up to 32,768 tokens and emits zero output tokens because there is no decoding loop. Microsoft also supports an abstention option such as "cannot tell" when the supplied evidence is insufficient, which is the detail that makes thresholding on the output meaningful at all.

d3-flash sits inside the same conceptual box — Choice, Noul (yes/no) and Score questions, one forward pass per question, a probability per option, no generation — and then widens the input side. Where Microsoft-Decision-1 is text-only by design, d3-flash accepts text or JSON, multiple images per request at up to 1.6 megapixels each, and multiple videos read at 2 frames per second. Every question in a request sees every image and video attached to it. It is a smaller model that does more with the input it is given.

That is the first real split, and it is not a matter of taste. If the decision you need to make is about a screenshot, a receipt, a chart or a clip from a camera, Microsoft-Decision-1 cannot make it. That is a capability boundary, not a quality gap, and no amount of accuracy work closes it.

What each one publishes, and how

Here the two releases are doing genuinely different kinds of proof, and it is worth separating them rather than lining up numbers.

Microsoft ran a 36-benchmark comparison spanning nearly 150,000 questions held blind from training, covering routing, ranking, long context, multilingual and out-of-distribution tasks, reasoning and safety, and says its model came out best across those sets. It reports latency as a ratio rather than a number — p50 around 35 times faster than GPT-6 Sol, and 2.5 times quicker than H2O-Lightning-4B v1.1, which it names as the runner-up. It documents robustness as a procedure: the same request perturbed eight ways, a 1.3% average decision-flip rate, and zero flips when option descriptions are paraphrased or options are reversed or shuffled. It tested safety on 5,250 requests across 11 benchmarks covering harmful content, jailbreaking and prompt injection. It states the calibration standard it holds itself to — a 90% prediction should be right about nine times out of ten on representative cases.

vLLM-SR published an index. The Jev Decision Index 0.3.1 places d3 at 64.7 with d3-flash at public-suite 57.79, a claimed gain of 11.0 over Decision 2.0's 9B model at 46.76. It also claims all 140,178 public requests answered with none unsupported, which is a coverage statement rather than an accuracy one. The card discloses that d3's own row is an internal evaluation while the comparison rows alongside it — Perplexity Decider v1.1, Fastino GLiDE, Jev, Torchcast Decision 27B — are live board data. And on the multimodal side, d3-family models report Perception Test scores on all 19,140 validation questions, with d3-flash at 73.3 against a three-option chance floor of 33.3.

So: Microsoft publishes a breadth claim with a documented method and a robustness number, and no per-checkpoint calibration figure. vLLM-SR publishes a leaderboard position with a stated provenance gap between its own row and the others, plus a video result, and no calibration figure either. Neither card carries a Brier score or an expected-calibration-error number for the model it describes. For two products whose entire value proposition is that the returned number means something, that is the shared hole, and it is the single most useful thing either vendor could ship next.

A generated two-column scoreboard headed 'Decision 3.0 d3-flash vs Microsoft-Decision-1 - the scoreboard'. The left column for d3-flash reads: base model Qwen3.5-9B fine-tuned, 8.39B parameters; distribution Apache-2.0 weights on Hugging Face from 10 October 2026; input text, images and video; context budget not stated on the card; reported accuracy Jev Decision Index 0.3 public suite 57.79, an internal evaluation; latency median 19.1 ms text, 149.4 ms image and 407.6 ms video. The right column for Microsoft-Decision-1 reads: base model Qwen3.5-9B post-trained, weights not distributed; distribution hosted on Microsoft Foundry, generally available 8 October 2026; input text only; context budget 32,768 tokens; reported accuracy best of 36 benchmarks covering nearly 150,000 blind questions; latency p50 around 35 times faster than GPT-6 Sol. A footer reads 'd3-flash figures are vLLM-SR's own internal evaluation; Microsoft-Decision-1 figures are Microsoft's own; neither is independently reproduced.'

Distribution decides more than the spec sheet does

This is where the comparison stops being about models.

With d3-flash you get weights, a decision_config.json pinning the base revision, a per-file SHA-256 manifest, custom modelling code, a readout head, and a documented dependency set that includes transformers==5.17.0 and an optional CUDA kernel package for the linear-attention layers. You can run it on your own hardware, fine-tune it, quantise it, air-gap it. You are also taking on the operational work: a Python class is not an endpoint, and the card does not state a token budget for state plus questions plus candidate descriptions — max_length is null in the shipped config, so that ceiling is yours to discover.

With Microsoft-Decision-1 you get an endpoint inside an existing cloud account, with the authentication, regional availability and provisioning controls that come with it, and you get none of the operational work. You also get none of the control. There is no repository to download, no fine-tuning path, and no self-hosting option; the model's lifecycle is Microsoft's, not yours, and a deprecation notice or a rebase onto a different backbone is a change you absorb rather than one you schedule. Microsoft has said a rebase onto its own MAI models is coming, which is a reasonable roadmap for a model whose whole point is fast single-pass scoring, and also means the specific checkpoint you benchmarked is not necessarily the one you will be calling in a year.

The latency story also needs reading carefully. Microsoft's headline is a ratio against a much larger general-purpose model, which is the correct comparison for its argument — a decision model's job is to replace a costly generative call with a cheap scoring one, and 35x against GPT-6 Sol at p50 is what that argument looks like when it lands. vLLM-SR's numbers are absolute, single-request and single-GPU, and they show the modality tax plainly: on one AMD Instinct MI325X, a text request to d3-flash takes a median 19.1 ms, the same request with an image takes 149.4 ms, and with a ten-second video 407.6 ms. Both sets are vendor-reported, on different hardware, and neither has been reproduced outside the lab that produced it.

A screenshot of the vllm-sr/d3 model card on Hugging Face, the flagship checkpoint of the Decision 3.0 family that d3-flash belongs to. The header shows 18 likes, the vLLM Semantic Router organisation, and tags for Transformers, Safetensors, qwen3_5, feature-extraction, decision-model, classification, system-one, multimodal, vision, video, custom_code and an Apache-2.0 licence, with the model size listed as 26B parameters in BF16. The card graphic reads 'Decision 3.0 / d3 27B'. A specification panel gives Parameters as 26.09B including the 0.46B vision encoder, Inputs as 'Text or JSON, plus images and videos (several per request)', Decision types as 'Choice - Yes / No - Score' and License as Apache-2.0. The sidebar shows 130 downloads last month, the model tree pinned to Qwen/Qwen3.8-27B, and a collection entry listing 6 items.

How to choose

The decision rule is short, which is a good sign that the two products are not actually competing for the same slot.

Take Microsoft-Decision-1 if the decision is text-only, if you already run on Foundry, and if being a call rather than a deployment is the whole point — no GPU to buy, no container to operate, no model lifecycle to own. Microsoft's supporting material is also the more usable of the two for a platform team: the perturbations, the safety test counts and the statement that an abstention option is supported are the things you need to write a threshold policy, and the 32,768-token ceiling is stated rather than left blank.

Take Decision 3.0 — realistically d3-flash or d3-mini — if any part of the decision depends on an image or a clip, if you need to run it somewhere a hosted API cannot reach, or if you want to fine-tune the scorer on your own labelled cases. The Apache-2.0 licence and the file-level hash manifest make that a genuine option rather than a theoretical one. What you accept in exchange is an unstated input budget, no published calibration, and the fact that the family is days old with essentially no outside use behind it.

If you need both — a hosted scorer for the routine text path and a self-hosted one for the multimodal or air-gapped cases, called through the same interface — then the honest position is that no vendor currently gives you that, and the two request formats are close enough to unify but not identical.

Where OrcaRouter sits, and where it does not

typesafe/jev-1.13 is the decision model in our catalogue, served over POST /v1/systemone with the same state-and-named-questions contract both models above implement: text in, structured JSON out, up to roughly 64K input tokens, non-streaming, $0.042 per million input tokens with no completion charge because it never produces one. Notably, the Android-style token budget question that neither card above answers in full is answered here.

We do not carry either model in this comparison. Decision 3.0's checkpoints ship as a local inference path in a Hugging Face repository rather than a routable endpoint, and Microsoft-Decision-1 is distributed through Microsoft's own platform and several third-party platforms — not through us. Nothing here should be read as an availability claim for either.

What the routing layer is actually worth in this decision is on the other half of the loop. A scorer is cheap by construction; the model that acts on its output is not, and pairing a decision model with a generative one normally means two integrations, two billing relationships and two failure modes. Calling both through one key across 200-plus models — with automatic failover when a provider wobbles, and provider list price passed through at 0% markup so a vendor price change is live the same day — means you can swap the scorer without touching the call site. That matters more than usual here, because both of these models are early enough that the decision you make this week is one you should be able to reverse next month.

A screenshot of the OrcaRouter models catalogue, headed 'Models - 207 models - 16 providers - one API key, one bill'. Filter tabs read All 207, Text 172, Image 10, Embeddings 5, Video 10 and TTS 10. A panel titled 'How to call any model' lists three steps - copy the model ID from any card, paste it into the model field, POST to our OpenAI-compatible endpoint - beside a code block showing POST https://api.orcarouter.ai/v1/chat/completions. Model cards below include OpenAI: GPT-6.1 Sol and Anthropic: Claude Sonnet 5.5 at $2.00 per million input tokens and $10.00 output, and TypeSafe: Jev 1.13 at 65K context with $0.042 per million input tokens and $0.042 per million output tokens. A credit panel above shows monthly plans from $50 to $1,000.

The question that decides this in six months

Two teams post-trained the same base checkpoint into the same shape of product in the same week and drew opposite conclusions about how it should reach a customer. That is not a coincidence of timing so much as a fork in how the category is being sold: as weights or as a service, as a thing you own or a thing you call.

Watch three signals. Whether Microsoft's rebase onto MAI or its own models changes the accuracy and latency behaviour it is currently selling — and how much notice customers get. Whether vLLM-SR publishes an input budget and a calibration figure for Decision 3.0, closing the gap that keeps its index from being actionable. And whether anyone outside either lab runs the public suite on the released d3 weights, because at the moment the only number that would settle this comparison is one neither vendor has produced.