
Microsoft-Decision-1: The Microsoft Model That Answers With a Number Instead of a Sentence
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 82 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 122 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 358 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
Microsoft-Decision-1 has a line in its model card that no other Microsoft model has ever carried: not designed for text generation. The model went generally available on Microsoft Foundry on October 8, 2026, and it is a deliberate narrowing of what a language model is for. You hand it a situation and a question with a fixed list of answers — yes/no, a multiple-choice set, a rating scale, a rubric — and it returns a calibrated probability for each option. No prose. No explanation. No rationale field. The output is JSON numbers and nothing else. It is built on the open-weight Qwen3.5-9B, post-trained by Microsoft, and Microsoft says it will rebase the model on other backbones later, naming MAI and other partner models.
That is a stranger product than it first sounds. Microsoft's competitive position in 2026 rests on frontier chat models and on Copilot, and Microsoft-Decision-1 is the opposite of both: a mid-sized, single-purpose scorer whose entire job is to tell an application which of your own options is most likely correct, and how confident it is. The interesting question is not whether it is good at writing — it is explicitly bad at that and does not try — but whether a probability distribution over options is a more useful primitive for the workloads enterprises actually run than another general-purpose model with a JSON mode.
What it does, precisely
The contract is one call in, one distribution out. Microsoft lists the supported question formats as yes/no, multiple-choice, rating, classification and rubric-based. Every option gets a score. The model runs in a single invocation over inputs of up to 32K tokens — 32,768 is the stated context window — and the practical ceiling is the input plus the option set, not a generation budget, because there is no generation.
The published use cases are the ones a platform team would recognise immediately:
• AI-output evaluation — grade a generated response against a supplied rubric, or decide whether it is grounded in the evidence it was given.
• Classification and routing — classify a request, judge relevance, triage a queue, pick a workflow branch.
• Agent guardrails — score a proposed tool call or agent action before the integrating application lets it execute.
• Content-safety screening — flag content against thresholds the application defines, rather than a fixed vendor policy.
• Search and document relevance — judge whether a retrieved document answers a supplied question.
• Confidence-based automation — auto-accept high-confidence outcomes and escalate the rest to a human.
The abstention option is the detail worth noticing. Microsoft explicitly supports options such as "cannot tell" when the supplied evidence is insufficient, which is the difference between a scorer that is calibrated and one that is merely confident. And because scores come back as numbers, the escalation threshold is a decision you own — you set where 0.7 sends something to a person and 0.95 does not.
What it will not do
Microsoft is unusually explicit about the exclusions, and they matter more than the feature list for anyone sizing this for a real pipeline. Microsoft-Decision-1 is not designed for text generation, open-ended question answering, conversation, translation or summarization. It is not intended for tasks without a closed question and a defined set of answer options, or for tasks that require knowledge absent from the input. It is text-only: no images, audio or video in, none out. It provides no explanations or rationales.
Read those together and a boundary appears that is easy to trip over. This is not a chatbot you can ask to also classify things, and it is not a summarizer you can bolt a score onto. It is a scoring function with a token budget. The team's own framing — that it should not be the sole automated decision-maker in consequential decisions about people, and should not be the sole basis for decisions involving credit, employment, housing, insurance, education, healthcare, legal rights "or similarly consequential domains" — points the same way. It is designed to sit next to a decision, not to be it.

The benchmark situation is the story nobody wants to print
There are no numbers. The Microsoft Foundry catalogue page for Microsoft-Decision-1 has a Benchmarks tab, and it is empty of figures. What it contains instead is a methodology paragraph and a qualitative claim: the model was "evaluated on public and community decision benchmarks and on held-out internal test sets not used in training," Microsoft reports that it "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology," and the metrics used were accuracy, calibration error, safety recall, false-positive rates and fairness consistency, with option order varied and paired statistical tests applied.

That is a serious methodology description attached to zero published results. It means every performance claim about Microsoft-Decision-1 today is vendor-reported and unreproduced, and the honest position for anyone evaluating it is that the calibration — the one property that makes a probability useful at all — is unverified outside Microsoft. The company does name where it believes the model is strongest and weakest, which is more useful than a headline score: strongest on reasoning, rule application and robustness to prompt formatting; competitive on classification, retrieval, fairness, tool use and most multilingual tasks; weaker on specialized domain-knowledge tasks.
The self-reported limitations are worth reading before the feature list. Scores can shift with phrasing and option ordering, and a poorly framed question still returns a score. Calibration is strongest on familiar task types. It may rely on outdated knowledge, and it gives no explanations. On multilingual coverage: 25 languages are listed as supported including Japanese, Korean, Arabic, Vietnamese, Thai, Turkish, Hindi, Bengali, Swahili, Hebrew, Persian and Ukrainian, but Microsoft states that coverage, quality and calibration "may vary by language," and lists non-English — especially lower-resource languages — as an area of underperformance. The underlying Qwen3.5-9B supports more than 200 languages; the post-trained model supports a quarter of that.
How you get it, and what it costs
Microsoft-Decision-1 is distributed as a hosted API in Microsoft Foundry under the "Direct from Azure" portfolio. Model weights are not distributed — this is not an open-weights release, and there is no Hugging Face repository for it to download. Any application that can issue HTTPS requests can integrate using the Foundry endpoints and standard Azure authentication. The deployment listing shows serverless and unified-endpoint options on pay-as-you-go or reserved provisioned throughput, standard SKU, with batch inference disabled, and the training disclosure reports the training dataset was first used in September 2026 with collection ongoing.
Pricing is not published on the model page. The catalogue's pricing field links out to Microsoft's model pricing page rather than printing an input and output rate, so the per-token cost of a Decision-1 call is something you have to look up in the Azure pricing surface or read off a bill. That is a real gap for anyone trying to model cost per decision at volume, and it is worth saying plainly rather than estimating. Two things are worth knowing when you do price it: 0% of the cost is output tokens, because there are none, and batch inference is switched off, so you cannot amortize a bulk scoring run through the batch channel the way you would with a generative model.
Where OrcaRouter fits into this is on the other side of the call. We do not host Microsoft-Decision-1, and it is not in our catalogue — a model that returns probabilities rather than text is not a model you route chat completions to. What we do carry is the half of the pattern that does generate: the models that write the rubric, draft the candidate responses, or produce the tool call that Decision-1 then scores. Those live behind one OpenAI-compatible key with more than 200 models on it, at provider list price passed through with 0% markup, so a vendor price cut on a judge model is live on our side the same day. If you are building an evaluation loop where one model writes and another scores, the scoring call goes to Microsoft and the generation call can go anywhere — including through the routing DSL, which composes several models into a single call when you want a panel rather than a single judge.

Why a scorer is a different bet from a better chatbot
The pattern Microsoft is selling here already exists in the open. InternLM's Intern-Decision-4B, Liquid AI's d1-3B, Convai Innovations' Laya and the Kev family from Jared Palmer all return calibrated distributions over supplied options without generating text, and most of them are Apache-2.0 weights you can run on your own hardware for nothing. Microsoft's entry differs in three ways that are not benchmark-dependent: it is a managed API with Azure authentication, billing and governance attached, so it fits an enterprise procurement path that a Hugging Face download does not; its base is a 9B model, larger than most of that field; and it comes with a Responsible AI assessment and a documented evaluation methodology, which is often the actual gating requirement for a regulated deployment.
What it does not come with is a number. Against open competitors that publish Brier scores and expected calibration error — the two figures that tell you whether a 0.8 means 0.8 — Microsoft has published a methodology and no results. Until independent calibration testing exists, the defensible way to use Microsoft-Decision-1 is the way its own documentation recommends: validate on data representative of your use case, set thresholds from the cost of your errors, always include an abstention option, randomize option order where ordering could bias the answer, and keep a human in the loop for anything consequential. That is good advice for any scorer. It is especially good advice for one whose calibration nobody outside the company has measured.
