
Microsoft-Decision-1 vs Kev: Buying a Score, or Owning the Recipe That Produces One
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 113 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 52 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 62 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 399 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
Jared Palmer put a receipt next to his model. The Kev family — Kev-0.8B, Kev-4B and Kev-9B, built on frozen open-weight base checkpoints with a rank-16 LoRA adapter and a small pointer head — shipped on September 24, 2026 with its training recipe, its stage-by-stage data counts and its evaluation numbers all published, and the honest headline for anyone comparing it to Microsoft-Decision-1 is that Kev tells you far more about how to reproduce itself. Microsoft's scorer went generally available on Microsoft Foundry on October 8, 2026: a Qwen3.5-9B base post-trained by Microsoft, 32,768-token context, text-only, weights not distributed, a published evaluation methodology and no published results. Both take a state and a set of typed questions and return a calibrated distribution instead of generated text. One is a service, the other is a method you can run in a notebook tonight.
The reason that distinction decides this matchup rather than capability: Kev-4B reports ECE 0.017 on its out-of-domain locked test and Microsoft-Decision-1 reports nothing, so the calibration argument that usually settles a scorer-versus-scorer comparison has only one side stating its position.
What Kev actually publishes
Kev-4B's card is unusually specific, and the specificity is the point. The backbone is Qwen3.5-4B-Base, frozen at a named revision, with 24 Gated DeltaNet linear-attention layers and 8 full-attention layers at hidden size 2,560. On top of it sits a rank-16 LoRA adapter — 33.8 million trainable parameters, applied to attention, MLP and DeltaNet projections — and a pointer head that scores each option's closing token against the question's final token and softmaxes. One temperature, T = 2.41, is stored in the head file and applied at load, and the card gives the fitted value and the file hash. Training ran in stages with published record counts: a base recipe of 12,576 records, 1,425 for dates and missing evidence, 5,219 real CFPB complaint narratives, 6,000 skill records and 5,320 developer-tooling records, with later stages replaying earlier ones. The card also states what was not used: "No output of Jev (TypeSafe's hosted decision model) was used."
The benchmark table is stated on the same page, and it comes with the failure cases attached, which is rarer than the successes. Transfer-v4 locked out-of-domain test: accuracy 0.838, Brier 0.224, ECE 0.017. Breadth-v1 over 14 held-out datasets and 3,089 questions: 0.690 accuracy, ECE 0.029. Hard-v1 test: 0.803. Documents-v1 test: 0.903. MMLU-Pro with ten options: 0.565. Date arithmetic: 0.65, named by the author as the weakest area, with a preprocessor flag offered as a partial fix and an explicit note that it was not re-measured. The limitations section adds that the calibration is a single in-distribution temperature so coverage degrades out of domain, that option order can flip an answer, and that lengths beyond 8,192 tokens are served but unvalidated. Serving figures are published too: 18.1 ms for six questions over a short state on an H100, 12.9 ms cached, 14.3 GB resident GPU memory.

What Microsoft publishes instead
Microsoft-Decision-1 covers most of the same ground with a different posture. It scores yes/no, multiple-choice, rating, classification and rubric questions, supports explicit abstention options such as "cannot tell," returns JSON probabilities, and runs in one invocation over up to 32K tokens — four times the context Kev validates. It is distributed as a hosted API in Microsoft Foundry under the Direct from Azure portfolio, with serverless and unified-endpoint deployment, standard SKU, Azure authentication and a unified billing path. Batch inference is disabled, and there is no weight download and no fine-tuning path.
The evaluation section describes public and community decision benchmarks plus held-out internal sets, metrics including accuracy and calibration error, option order varied, paired statistical tests, and a claim that the model "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology." There is no table. The model card names its own limits honestly — scores shift with phrasing and option ordering, calibration is strongest on familiar task types, no explanations are produced, and non-English coverage is weakest — but a limitation list is not a calibration measurement. Anyone setting a 0.9 auto-accept threshold on Microsoft-Decision-1 is setting it on faith and adjusting it in production.

The part of the comparison that costs money
Kev's economics are the reason this matchup is genuinely close. The recipe is a frozen base plus a 33.8M adapter, which means the marginal model you train costs adapter-sized compute, not foundation-model-sized compute. You can hold the base in memory once and load several adapters — one per decision domain — which is a deployment shape Microsoft's hosted API cannot express at all. If your routing rules differ from a generic judge's, Kev lets you train the difference; Microsoft-Decision-1 lets you write a better prompt and hope.
Against that, the hosted API carries no operational surface. There is no adapter to track, no serving stack to keep warm, no validated-versus-served context gap to reason about, and no risk that a temperature fitted on one dataset behaves differently on yours. For a team whose decision volume is modest, the service is the cheaper artefact in engineering hours even if the tokens cost money; for a team scoring millions of items a day, the self-hosted adapter wins on unit cost the moment the GPU is paid for.
There is a third path that uses neither model for the part where it is weakest, and it is where OrcaRouter sits. We do not host Microsoft-Decision-1 or Kev — a scorer that returns probabilities is not a chat-completions target and neither model is in our catalogue. The generative half of a decision pipeline is what we carry: the model that drafts the candidate answer, writes the rubric, or emits the tool call that gets scored. All of that is behind one OpenAI-compatible key with more than 200 models on it at provider list price passed through with 0% markup, so a vendor price change is live on our side the same day. If you are building a grading or triage loop, the writing call is routeable and the scoring call is not, and keeping that boundary clear is what stops a scoring pipeline from quietly becoming a chat pipeline.

How to decide
Choose Kev when you need numbers before you ship, when you want to fine-tune on your own decision labels, when you want to run scoring inside your own boundary at zero marginal cost, or when you want several decision domains sharing one resident base model. Its published ECE figures are still the author's own measurements on the author's own harness — treat them as credible self-assessment, not as an audited third-party result — but they are at least a stated scale you can try to reproduce, and the recipe is right there to reproduce them with.
Choose Microsoft-Decision-1 when the gate is procurement or context length: a managed Azure endpoint with unified billing and a Responsible AI package, 32K of input for long documents, and a vendor to escalate to. Its missing benchmark table is not a scandal — plenty of hosted models launch without one — but it does mean the first calibration curve for this model will be drawn by its users, and you should plan to be one of them, on your own labeled data, before you point a production threshold at it.
