
Microsoft-Decision-1 vs Intern-Decision-0.8B: Half an Ounce of Jailbreak Trouble Against a Hosted Blank
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 44 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 120 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 52 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 61 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 369 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
Start with the column that a launch post would have left out. Intern-Decision-0.8B — InternLM's 852,985,920-parameter decision model, fine-tuned from Qwen3.5-0.8B and uploaded to Hugging Face on 26 September 2026 with no announcement, no paper and a GitHub link that still 404s — is measured at 64.48 on WildJailBreak against a reference of 96.29 for TypeSafe's Jev 1.13. That is not a rounding difference. It is a refusal-robustness collapse in a model whose entire job is to judge inputs, published on the vendor's own card, in a table whose benchmark belongs to a competitor. Set that next to Microsoft-Decision-1, which went generally available on Microsoft Foundry on October 8, 2026 as a text-only scorer post-trained on Qwen3.5-9B with a documented evaluation methodology and not one result printed under it, and you have the real shape of this matchup. One of these models will tell you a number that makes it look bad. The other tells you which metrics it would have used.
That inversion is worth more than the spec sheet. Both models take a state and a bounded question set and return probabilities over your options — neither generates text, and on that axis they are the same kind of instrument. The differences that matter are who runs it, how big it is, how much you can verify, and one safety column that the smaller, quieter, fully-downloadable model chose to publish anyway.
What each one actually hands back
Microsoft-Decision-1 is a hosted API, and that is the first structural difference. Weights are not distributed — no download, no repository, no fine-tuning path — and integration means a Foundry deployment with Azure authentication, billing and governance attached. It runs one pass over up to 32,768 tokens, supports yes/no, multiple-choice, rating, classification and rubric-shaped questions, and is text-only in and numeric out. It explicitly supports an abstention option such as "cannot tell" when the evidence is insufficient, which is what makes a threshold meaningful.
Intern-Decision-0.8B is the opposite arrangement. It is Apache 2.0 weights with the upstream Qwen licence preserved beside them as LICENSE-QWEN, about 1.73 GB of repository across three shards — a 1.50 GB language shard, a 176 MB vision shard and a 25 MB projector — which fits on a single consumer GPU or a well-provisioned laptop. It accepts a state, a schema of one to sixteen named questions with up to 62 options each, and optionally up to eight images, and its bundled inference.py documents the contract in more detail than most launched models bother with: options are mapped onto single-token symbols A–Z, then a–z, then 0–9; logits are read at the position immediately before each <decision> placeholder in a pre-rendered skeleton; a softmax is taken over only that field's legal candidates; a fitted temperature is applied.
The two licences are not the same purchase. Microsoft-Decision-1 gives you a managed endpoint and no artifact you own. Intern-Decision-0.8B gives you an artifact you own, with no endpoint, no support contract and no documentation beyond the repository itself.

The ceiling, and the calibration you can audit
Two numbers decide most real integrations, and both of them belong to the small model.
The first is 8,192 tokens. Intern-Decision-0.8B's inference engine refuses oversize input rather than truncating it, which is the correct behavior for a scorer and also a hard wall: there is no chunking strategy that preserves the contract, because the state, the schema and the skeleton all have to be in one pass. Microsoft-Decision-1's ceiling is four times higher at 32,768, and it is a hosted limit you can raise by provisioning differently rather than a constant in a Python file.
The second is calibration, and here the direction reverses. InternLM publishes a fitted temperature of 2.747760550703 for the 0.8B, selected by NLL minimisation over 1,728 designated calibration cases with 1,693 held for validation, with the card explicit that test-suite labels were not used to select it. The applied transform is a softmax over the field's candidate logits followed by a second softmax over the log of that distribution divided by the temperature. Because the transform runs after the first softmax and preserves ordering, it cannot change the argmax at all — it moves confidence, the yes-probability and the expected value of a score question, and leaves the label identical. If your pipeline reads the label, the temperature is a no-op. If it reads the probability, thresholds it, or feeds it into an expected-value calculation, the temperature is the difference between a number that means something and one that does not. Passing temperature=1 returns the uncalibrated distribution if you would rather fit your own.
Microsoft-Decision-1 has no equivalent artefact. Its Benchmarks tab states that accuracy, calibration error, safety recall, false-positive rates and fairness consistency were measured on public and community decision benchmarks plus held-out internal test sets, that option order was varied, that paired statistical tests were applied, and that the model "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology." There is no calibration error printed, no temperature to inspect, and no validation split described. The one property that makes a probability useful — whether 0.8 means 0.8 — is asserted and unquantified.
Read those two paragraphs against each other and the comparison stops being about size. A 0.8-billion-parameter checkpoint whose calibration you can inspect and whose software you can run is, for an evaluation team, a more tractable object than a hosted API whose calibration is a sentence. The small model also has a latency figure on the record — 33.98 ms mean, 33.44 ms median, 37.50 ms p95 per query on a single RTX 4090 through the local Hugging Face path, in InternLM's own measurement — while Microsoft's is something you measure through your own deployment, where the choice between serverless and provisioned throughput will move the number more than the model will.

Where the small model loses
None of the above makes Intern-Decision-0.8B the safe choice, and the same published table says why. WildJailBreak at 64.48 against Jev's 96.29 is a refusal-robustness gap of thirty points in the exact category where a scorer meets untrusted input. A decision model is often the component that decides whether a proposed action is allowed; one that is easy to talk past is a liability in that seat. Whether the deficit comes from the Qwen3.5-0.8B base, from the decision-tuning objective or from the small parameter count is not something the repository tells you, and there is no paper, no training recipe and no data disclosure that would.
The rest of InternLM's own table is more flattering than that column and still modest. Average across seven suites is 79.38 for the 0.8B against 84.68 for its own 2B sibling and 90.02 for the 4B — the family climbs steeply with size, which is a mild signal that the table was not reverse-engineered to flatter the flagship. AG News lands at 88.61 against Jev's 89.57, effectively level on four-way topic classification. On calibration the 0.8B reports a Brier of 0.530 and an expected calibration error of 0.066, against Jev's 0.358 and 0.095 — better ECE, meaningfully worse accuracy. That is a plausible profile for a small model with a deliberately fitted temperature, and it is not a combination a marketing team would choose to publish by accident.
There is also no release date, no announcement, no collection gathering the three checkpoints, no hosted endpoint anywhere, and no independent evaluation of any kind. The demo Space answers 401. The correct description of Intern-Decision-0.8B is a released checkpoint with unaudited claims — which is a better situation than a waitlisted API, because you can measure it in an afternoon, but it means the validation burden is entirely yours.
Choosing, and where OrcaRouter sits
Take Microsoft-Decision-1 if the blocker is procurement or scale: you need Azure authentication, unified billing and a Responsible AI assessment before anything reaches production; you need a 32K context; you need an endpoint you do not operate; or you need the abstention path designed in rather than hand-rolled. Accept in exchange that you cannot run it, cannot fine-tune it, cannot inspect its calibration, and will be measuring its core claim yourself.
Take Intern-Decision-0.8B if the blocker is cost, memory or control: you want an Apache-2.0 scorer that fits in 1.73 GB, runs on hardware you already own, produces the same label every time, and can be swapped for its 2B or 4B sibling by changing a checkpoint path. Accept in exchange the 8,192-token wall, the WildJailBreak column, and the fact that nobody outside InternLM has reproduced anything on the card.
The honest recommendation is to do both, because they are cheap enough that the argument is barely worth having. The scoring call itself is not something we route — a model that returns probabilities instead of text is not a chat completion, and neither of these is in our catalogue. What OrcaRouter carries is the other half of the loop: the more than 200 models behind one OpenAI-compatible key that draft the rubric, generate the candidate answers, or emit the tool call your scorer is about to grade, at provider list price passed through with 0% markup, so a vendor price cut on a generator is live on our side the same day. Automatic failover matters more than usual here, because a pipeline that scores everything it sees has no headroom to retry a stalled call, and the routing DSL lets you send one judge prompt to several writers and fuse their agreement rather than trusting a single one.

Bottom line
Microsoft-Decision-1 reached general availability on Microsoft Foundry on October 8, 2026: hosted, text-only, 32,768 tokens, built on Qwen3.5-9B, with a methodology paragraph in place of a benchmark table. Intern-Decision-0.8B is a 1.73 GB Apache-2.0 checkpoint uploaded on 26 September 2026 that publishes a Brier of 0.530, an ECE of 0.066, a fitted temperature of 2.747760550703, 33.98 ms on a 4090 — and a WildJailBreak score of 64.48 that no one asked it to print. If you can only validate one thing before adopting either, validate calibration on your own labelled cases; that is the number both vendors have left for you to find.
