
Microsoft-Decision-1 vs Laya: 25 Languages Behind an API, 100+ Behind a 322MB Download
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 113 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 52 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 62 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 399 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
Count the languages and the comparison looks decided. Microsoft-Decision-1, generally available on Microsoft Foundry since October 8, 2026, lists 25 supported languages and states plainly that coverage, quality and calibration "may vary by language," naming non-English and especially lower-resource languages as an area of underperformance. Laya, published by Convai Innovations on September 18, 2026 under Apache 2.0, describes itself as multilingual across 100+ languages and reports that 45 of 51 tested languages stay usable with routing, against 23 of 51 for its English-only sibling. One is a 9B model behind an Azure endpoint with a 32,768-token context; the other is a 421-million-parameter classifier running at 32.8 milliseconds per decision on a single Tesla T4 for free. Both refuse to write a token and both return probabilities over options you define.
But read the two model cards in full and the language count stops being the interesting number. It is the calibration story behind it, and the two projects tell that story in opposite directions.
Laya publishes the number that makes its own marketing awkward
The Laya model card states, in its own limits section, that the base checkpoints score 0.362 zero-shot on the typed-decisions benchmark — below the 0.461 majority-class baseline. That is a card telling you its model is worse than guessing "the most common answer" on that test. It also states that the base checkpoints are near chance zero-shot, that the headline 0.766 typed-decisions figure requires fine-tuning on that benchmark's own split, that ordinal score questions are the weakest primitive (SST-5 0.372), that high-cardinality choice questions suffer because 77 options leaves roughly three or four tokens each, that noul can follow its own labels, and that the action.act_probability field "carries no usable signal yet" at AUROC 0.30. Convai's own summary sentence is the one to carry into any evaluation: Laya is a fast base to specialise, not a zero-shot decision engine.
That candour is more valuable than it sounds, because it tells you exactly what Laya is for. It is an encoder you fine-tune on your own labels — the whole architecture is built so that new schemas need no retraining, but performing well on a task does. Used that way, the published figures are strong: 0.766 against Jev's 0.727 on typed decisions after fine-tuning, AG News 0.950 against 0.910, DAIR Emotion 0.595 against 0.480, and ECE 0.081 against Jev's 0.246 — with the honest note that it ships over-confident and needs temperature refitting, which moves mean ECE from 0.466 to 0.081.
Microsoft publishes the methodology and withholds the result
Microsoft-Decision-1's Foundry page runs the same exercise in the other direction. It describes the evaluation in detail — public and community decision benchmarks plus held-out internal sets, metrics including accuracy and calibration error, option order varied, paired statistical tests, identical harness for compared models — and reports that the model "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology." It names its strong areas (reasoning, rule application, robustness to prompt formatting) and its weak ones (specialized domain knowledge, non-English).
And then it prints nothing. No accuracy figure, no calibration error, no per-benchmark table. The self-reported limitations are real and useful — scores shift with phrasing and option ordering, a poorly framed question still returns a score, calibration is strongest on familiar task types, no explanations are produced — but a limitations list is not a measurement. So the trade is unusually clean: Laya gives you numbers you can attempt to falsify on your own data, and Microsoft gives you a compliance posture and a 32K context you can put a long document into.

What the size difference actually buys
Laya's smallness is not a compromise here, it is the design. Because every option is scored at its own [MASK] token and softmaxed over that question's options, the answer space is assembled at request time rather than baked into a vocabulary head — which is why a schema you invent this afternoon needs no retraining, and why the model fits in 322 to 421 million parameters. The multilingual checkpoint runs about 2.2 times faster than the English one, the router detects script in under half a millisecond, and batched throughput reaches 103 to 332 questions per second on one T4. For a workload that scores every ticket, every retrieved document or every proposed action, that throughput is the feature, not the parameter count.
The costs of that design are equally specific and worth stating against Microsoft's alternative. The English checkpoint takes a 512-token window; the multilingual one 1,024, extensible toward 8K. Microsoft-Decision-1 takes 32,768. A long support thread, a full contract or a multi-page policy document does not fit in Laya's window at all, and no amount of speed compensates for truncation. Laya answers one question set per forward pass with a documented token budget per option; Microsoft documents a single invocation over up to 32K tokens without a per-question ceiling. And Laya's competitive weak spot as a classifier is worth knowing: it scores 0.425 on Banking77 against Jev's 0.870, which is the kind of fine-grained, high-cardinality intent problem where a specialisation pass matters most.

The cost comparison that is not per token
Laya is free at the point of use — self-hosted, Apache 2.0, listed at a self-hosted cost of "$0" — and its real price is the fine-tuning run you owe it before it is good at your task. Microsoft-Decision-1 is a hosted Foundry API with a per-token rate that is not printed on the model page, no weights to download, no fine-tuning path, and batch inference disabled. One is an upfront data-labeling project, the other is a metered call. Which is cheaper depends entirely on volume, and the crossover is high: a 421M encoder scoring 300 items a second on a rented T4 is very hard to beat per decision once it has been trained.
Here is where OrcaRouter earns its place in a pipeline built on either model, and it is the generative side only. We host neither Microsoft-Decision-1 nor Laya: both return probabilities rather than text, neither is in our catalogue, and a scoring endpoint is not a chat-completions target. What we do carry is the model that produces the thing being scored — the draft reply, the proposed tool call, the candidate answer that Laya's fine-tuned head then judges. That is more than 200 models behind a single OpenAI-compatible key at provider list price passed through with 0% markup, with automatic failover when one serving path degrades. In a loop that scores everything it generates, the generation call is the part that can fall over, and failover is what keeps the scoring queue fed.

Which one to pick
Take Laya if your decisions arrive in many languages, if your volume is high enough that per-token pricing matters, if you have labels and a week to fine-tune, or if you need to run scoring on hardware inside your own boundary. Accept the 512-to-1,024-token window and the fact that the base checkpoint will be near chance until you specialise it, and you get a decision model that is honest about being a starting point.
Take Microsoft-Decision-1 if your inputs are long documents rather than tickets, if your deployment gate is Azure authentication, unified billing and a Responsible AI package rather than a download, if 25 languages covers your traffic, or if you would rather not own the model at all. Accept that you will be drawing its first calibration curve yourself, and budget an afternoon on a labeled sample to do it — because unlike Laya, this one will not hand you an ECE number to argue with.
