
Microsoft-Decision-1 vs Intern-Decision-4B: One Publishes Its Calibration, the Other Publishes Its Methodology
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 113 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 52 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 62 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 399 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
Put Microsoft-Decision-1 next to Intern-Decision-4B and you get a comparison that is decided less by capability than by what each vendor was willing to publish. Microsoft's model reached general availability on Microsoft Foundry on October 8, 2026: a Qwen3.5-9B base, post-trained by Microsoft, text-only, 32,768-token context, weights not distributed, an evaluation methodology describing accuracy, calibration error and paired statistical tests — and not one result. InternLM's Intern-Decision-4B went up on Hugging Face on September 26, 2026 at 05:36:37 UTC with no announcement behind it, is fine-tuned from Qwen3.5-4B, accepts images as well as text, runs in 44 milliseconds on a single RTX 4090, and prints a full table of Brier and expected-calibration-error figures. Both return a probability distribution over options you supply. Only one of them tells you how well it does it.
That inversion — the larger, better-financed model being the opaque one — is the whole shape of this matchup, and it decides the choice for most teams faster than any spec line.
The one number that separates them
InternLM reports Brier 0.347 and ECE 0.065 for Intern-Decision-4B across its benchmark suite, with an average score of 90.02 over Jevbench-Easy (100.00), Jevbench-Original (98.61), Jevbench-Hard (73.87), Typed Decision (80.55), ToolACE (96.45), AG News (90.82) and WildJailBreak (89.86). Those are vendor-reported figures on the vendor's own harness, and the Jevbench rows in particular are the project's own benchmark family — read them as self-assessment, not independent verification. But they are numbers with a stated scale, and ECE in particular is the figure that tells you whether a returned 0.8 is worth acting on.
Microsoft publishes no equivalent. The Benchmarks tab on the Microsoft-Decision-1 catalogue page describes what was measured and claims the model "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology," with the strongest areas named as reasoning, rule application and prompt-format robustness, and the weakest as specialized domain knowledge. No Brier, no ECE, no accuracy table. If your adoption decision needs a calibration figure before you can size an escalation threshold, one of these two models hands it to you and the other asks you to measure it yourself.

Where the two contracts actually differ
On the surface these models take the same call. You supply a state, a schema of named questions, and a fixed set of options; you get back a probability per option without a token of generated text. The differences show up at the edges of that contract.
• Modality — Microsoft-Decision-1 is explicitly text-only and does not accept or produce image, audio or video content. Intern-Decision-4B accepts optional images alongside the state, up to eight of them per call, which is what makes it a candidate for screenshot triage, document-layout checks and visual QA routing.
• Question cardinality — InternLM documents 1 to 16 questions per call, up to 62 options each, across three field types (choice, score, noul). Microsoft documents the formats — yes/no, multiple-choice, rating, classification, rubric — and a single invocation over up to 32K tokens, but not a per-call question ceiling.
• Context — Microsoft states 32,768 tokens. The Intern-Decision-4B card states its limits in questions, options and images rather than as one context number, so you cannot line the two up on a single figure.
• Distribution — Microsoft-Decision-1 is a hosted Foundry API; model weights are not distributed and there is no download. Intern-Decision-4B is Apache-2.0 on Hugging Face with the upstream Qwen licence preserved as LICENSE-QWEN, which means retaining that licence and upstream notices if you redistribute.
• Cost shape — InternLM's weights cost nothing to run and are quoted at 44.16 ms mean, 44.60 ms P95, on one RTX 4090. Microsoft's per-token price is not printed on the model page at all.
• Governance — Microsoft ships a Responsible AI assessment, documented deployment practices, and explicit exclusions (not for text generation, open-ended QA, conversation, translation or summarization; not to be the sole automated decision-maker in consequential decisions about people). InternLM ships a model card that is candid about provenance — weights "modified by decision tuning" — and nothing resembling a compliance package.

Two very different reasons the latency numbers are hard to compare
Microsoft-Decision-1 does not publish a latency figure, so there is nothing to put next to InternLM's 44.16 ms mean. That is not a small omission for this workload. These models exist to be called many times inside a pipeline — once per retrieved document, once per proposed tool call, once per response being graded — and the difference between four decisions a second and forty is the difference between scoring a sample and scoring everything. InternLM's number is achievable today on hardware you can rent by the hour; Microsoft's has to be measured through your own Foundry deployment, and the deployment shape you pick (pay-as-you-go serverless versus provisioned throughput) will change it more than the model will.
This is also where OrcaRouter's half of the pattern becomes relevant, and it is the generative half only. We do not host Microsoft-Decision-1 or Intern-Decision-4B — a model that returns probabilities instead of text is not something you route chat completions to, and neither appears in our catalogue. What sits behind our one OpenAI-compatible key is the more-than-200-model pool that does the writing: the judge prompt drafted by one model, the candidate answers produced by two others, the tool call that gets scored before it runs. Provider list price is passed through at 0% markup, so a price cut on a generator is live on our side the same day, and automatic failover keeps the generation side of a scoring loop alive when a single provider degrades — which matters more than usual here, because a pipeline that scores everything it sees has no headroom to retry a stalled call.

Who should take which
Take Intern-Decision-4B if any of these are true: you need to score images as well as text, you need a calibration number before you can ship, you want the option of running the model on your own hardware inside a boundary you control, or you want to fine-tune it. The last point deserves emphasis — a 4B open-weights scorer with a documented wrapper is a model you can adapt to your own labels, and Microsoft-Decision-1 is a hosted API with no fine-tuning path and no weights to adapt.
Take Microsoft-Decision-1 if the blocker is procurement rather than performance: you need Azure authentication, unified billing, governance and a vendor Responsible AI assessment before anything reaches production, and you need a 32K context for long documents that the 4B's per-call limits complicate. Its lack of published benchmarks is a real cost, but it is a cost you can retire in an afternoon by running your own labeled set through both models and comparing ECE — which is what you would do anyway before trusting either vendor's table.
The uncomfortable answer is that both of these are cheap enough to test that the argument is barely worth having. Intern-Decision-4B needs a GPU you probably already have. Microsoft-Decision-1 needs a Foundry deployment and a sample of your own labeled decisions. Run your data through both, compute calibration on the subset that matters to you, and let the numbers decide — which is, fittingly, exactly what these models are for.
