
Microsoft-Decision-1 vs Intern-Decision-2B: Nine Milliseconds Faster and Measurably Worse Calibrated
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 44 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 120 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 52 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 61 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 369 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
There is a number in InternLM's own table that should not exist, and it is the reason this comparison is worth more than the spec sheet. Intern-Decision-2B — 2,213,241,664 parameters, fine-tuned from Qwen/Qwen3.5-2B, uploaded to Hugging Face on 26 September 2026 at 05:36:19 UTC without an announcement — posts a 33.28 ms mean latency per query on a single RTX 4090, which is marginally faster than its own 852-million-parameter sibling at 33.98 ms. It also posts the worst calibration in its family: an expected calibration error of 0.100 against the 0.8B's 0.066, with a fitted temperature of 2.100509348278 against the 0.8B's 2.747760550703. Meanwhile Microsoft-Decision-1, generally available on Microsoft Foundry since October 8, 2026, publishes neither a latency figure nor a calibration error at all — only a methodology paragraph describing how both were measured. So the question this matchup really asks is not which of the two is better. It is what you are buying when you buy the middle size of anything.
Both models are decision scorers: state in, bounded question set in, calibrated probabilities out, no generated text and no tokens to parse. That shared contract is what makes the differences legible. Microsoft-Decision-1 is a hosted, text-only Foundry API on a Qwen3.5-9B base with a 32,768-token window, no distributed weights and no fine-tuning path. Intern-Decision-2B is an Apache-2.0 checkpoint with the upstream Qwen licence preserved as LICENSE-QWEN, roughly 4.46 GB of repository, custom inference code, and no hosted endpoint anywhere.
What the middle size is actually for
InternLM pushed three checkpoints in forty seconds: the 0.8B at 05:35:57, this one at 05:36:19, the 4B at 05:36:37. On the vendor's own seven-suite average the family climbs in the order you would hope — 79.38, 84.68, 90.02 — which is the one place the 2B looks like a sensible purchase. Everywhere else it looks like the size nobody would have chosen on purpose.
Prompt processing dominates a decision call, and that is the mechanical explanation for the latency inversion rather than a mystery. A scorer does exactly one forward pass over a prompt whose length is set by the state, the schema and the option descriptions — never by anything the model writes, because it writes nothing. In that regime parameter count is a second-order cost, so the usual reason to reach for the smallest checkpoint does not apply: you are not saving time, you are saving memory. The 0.8B's whole repository is about 1.73 GB against this model's 4.46 GB, and that is the honest reason to prefer it. The 2B's only real claim is that it happens to be the fastest fast one, by a margin small enough to be noise.
Set against Microsoft-Decision-1, that claim is close to irrelevant, because the two models are not in the same latency regime. A hosted endpoint's speed is a function of your deployment shape — serverless versus provisioned throughput on the same standard SKU — before it is a function of the model, and Microsoft publishes no per-call figure to compare against. What Microsoft does publish is a hard operational constraint that cuts the other way: batch inference is disabled. There is no offline channel to amortize a bulk scoring run through, so a Microsoft-Decision-1 pipeline pays the interactive cost of every decision, while a self-hosted checkpoint pays in GPU hours whether or not it is scoring.
The calibration column, where the middle loses
Read the three Intern-Decision cards together and the family stops behaving predictably. Accuracy is monotone in size; calibration is not. The 2B carries an ECE of 0.100 — the weakest of the three — and a fitted temperature of 2.100509348278, well below the 0.8B's 2.747760550703. The card instructs you to use the inference module shipped with the size you downloaded, because the default calibrations are per-checkpoint; anyone who copies a wrapper from one sibling to another silently applies the wrong temperature.
The transform itself is worth understanding before you treat that 0.100 as a verdict on the weights. It is a softmax over the field's candidate logits, followed by a second softmax over the log of that distribution divided by the temperature. Because it runs after the first softmax and preserves ordering, it cannot change the argmax at all. It moves confidence, the yes-probability and the expected value of a score question, and leaves the label identical. If your pipeline reads labels, the temperature is a no-op and the ECE is a curiosity. If your pipeline reads probabilities — thresholds them, ranks by them, feeds them into an expected-value calculation — then a 0.100 ECE is the difference between a threshold that means what you wrote and one that does not. The right response is to fit your own temperature on your own labelled cases, not to conclude the weights are bad.
Microsoft-Decision-1 asks for exactly the same work, with less to start from. Its Benchmarks tab states that accuracy, calibration error, safety recall, false-positive rates and fairness consistency were measured on public and community decision benchmarks plus held-out internal test sets, that option order was varied, that paired statistical tests were applied, and that the model "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology." No ECE. No Brier. No temperature. No accuracy table. The model whose entire value proposition is a trustworthy probability is the one in this comparison with no published calibration number at all.

Contract edges: four things the hosted model will not do
The two models take what looks like the same call, and the divergences live at the edges of it.
• Modality — Microsoft-Decision-1 is explicitly text-only and accepts no image, audio or video. Intern-Decision-2B accepts up to eight images alongside the state, which makes it a candidate for screenshot triage and layout checks that the hosted API cannot serve at any accuracy.
• Input ceiling and failure mode — Microsoft-Decision-1 runs a single invocation over up to 32,768 tokens. Intern-Decision-2B declares DecisionEngine(max_length=8192) and rejects oversize input rather than truncating it, which is correct behavior for a scorer and also a hard wall, because the state, the schema and the skeleton must all be present in one pass; no chunking strategy preserves the contract.
• Question shape — InternLM documents one to sixteen questions per call with up to 62 options each across three field types (choice, score, noul), where noul is a binary yes/no returning a probability and score returns a probability-weighted expected value on a scale you name. Microsoft documents the formats — yes/no, multiple-choice, rating, classification, rubric — plus an explicitly supported abstention option such as "cannot tell" when the evidence is insufficient, which is the single most useful line on the page for anyone writing escalation logic.
• Price — the Microsoft-Decision-1 model page does not print a rate; pricing links out to Microsoft's pricing surface, so cost per decision is something you read off Azure or off a bill, with 0% of it attributable to output tokens because there are none. Intern-Decision-2B costs nothing per call and everything in GPU time, and it has no hosted provider. Its storage is roughly 4.46 GB across a 3.76 GB language shard, a 612.5 MB vision tower and a 50.3 MB projector.
What is confirmed, and what is only the vendor talking
Keeping those two categories apart is the whole discipline with this family. Confirmed by a file listing or an HTTP response: the parameter count, the shard map, the licence pair, the base model, the architecture underneath — a Qwen3_5ForConditionalGeneration with 24 layers, a 2,048 hidden size, 8 query heads against 2 key-value heads, a head dimension of 256, a repeating pattern of three linear-attention layers to one full-attention layer, one retained multi-token-prediction layer and a 262,144-position embedding ceiling that the 8,192-token engine ceiling makes practically irrelevant.

Vendor-reported and unreproduced: every accuracy number, every latency figure, and the temperature. There is no paper, no arXiv entry, no launch post, no changelog and no independent evaluation. The demo Space answers 401, which means it is not public rather than broken. The GitHub repository that appeared after the model did — three commits, training code, two inference backends, an evaluation bundle with 10,751 test rows, a 96-case calibration benchmark and a reproduction guide — is more documentation than most quiet releases get, and it ships no weights, no training data and no code licence. Microsoft is in a different but adjacent position: its methodology is real and its claim is qualitative, and neither company is currently in a position to have its central number checked by anyone but you.
Where OrcaRouter sits in a pipeline like this
Neither model is in our catalogue, and nothing here should be read as an availability claim. A model that returns probabilities instead of text is not something you route chat completions to, and that is true of both of these. What we do carry is the generative half of the loop those scorers exist to serve: the more than 200 models behind one OpenAI-compatible key that write the rubric, draft the candidate answers and emit the tool call that a scorer then grades before it executes. Provider list price is passed through at 0% markup, so a vendor price cut on the generating side is live on ours the same day, and if you would rather not bet your threshold on one judge, the routing DSL composes several models into a single call and model fusion reports their agreement as a field you can score. Automatic failover keeps that side alive when a single provider degrades, which matters more in a loop that scores everything it sees than in one that answers a user now and then.

Bottom line
Microsoft-Decision-1 has been generally available on Microsoft Foundry since October 8, 2026: hosted, text-only, 32,768 tokens, a Qwen3.5-9B base post-trained by Microsoft, no distributed weights, and a benchmark section that documents its methodology without printing a figure — including no latency, with batch inference switched off. Intern-Decision-2B is a 2,213,241,664-parameter Apache-2.0 checkpoint from 26 September 2026 that is the quickest of its three siblings at 33.28 ms on a 4090 and the worst calibrated at an ECE of 0.100, with a fitted temperature of 2.100509348278 that cannot change a label but will change every confidence you threshold on. If you want the middle size, the honest case for it is thin: pay the extra 2.7 GB for the 4B's accuracy or accept the 0.8B's footprint, and in either world fit your own calibration before a threshold goes anywhere near production.
What we do carry is the generative half of the loop those scorers exist to serve: the more than 200 models behind one OpenAI-compatible key that write the rubric, draft the candidate answers and emit the tool call that a scorer then grades before it executes.
