A generated title card for Microsoft-Decision-1 vs Laya subtitled 'language coverage against context length', with chips reading 25 supported languages vs 100+ languages, a 32,768-token context vs a 1,024-token window, hosted API only vs Apache-2.0 weights, and a footer noting that both vendors' figures are self-reported.
Engineering & Research

Microsoft-Decision-1 vs Laya: 25 Languages Behind an API, 100+ Behind a 322MB Download

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Count the languages and the comparison looks decided. Microsoft-Decision-1, generally available on Microsoft Foundry since October 8, 2026, lists 25 supported languages and states plainly that coverage, quality and calibration "may vary by language," naming non-English and especially lower-resource languages as an area of underperformance. Laya, published by Convai Innovations on September 18, 2026 under Apache 2.0, describes itself as multilingual across 100+ languages and reports that 45 of 51 tested languages stay usable with routing, against 23 of 51 for its English-only sibling. One is a 9B model behind an Azure endpoint with a 32,768-token context; the other is a 421-million-parameter classifier running at 32.8 milliseconds per decision on a single Tesla T4 for free. Both refuse to write a token and both return probabilities over options you define.

But read the two model cards in full and the language count stops being the interesting number. It is the calibration story behind it, and the two projects tell that story in opposite directions.

Laya publishes the number that makes its own marketing awkward

The Laya model card states, in its own limits section, that the base checkpoints score 0.362 zero-shot on the typed-decisions benchmark — below the 0.461 majority-class baseline. That is a card telling you its model is worse than guessing "the most common answer" on that test. It also states that the base checkpoints are near chance zero-shot, that the headline 0.766 typed-decisions figure requires fine-tuning on that benchmark's own split, that ordinal score questions are the weakest primitive (SST-5 0.372), that high-cardinality choice questions suffer because 77 options leaves roughly three or four tokens each, that noul can follow its own labels, and that the action.act_probability field "carries no usable signal yet" at AUROC 0.30. Convai's own summary sentence is the one to carry into any evaluation: Laya is a fast base to specialise, not a zero-shot decision engine.

That candour is more valuable than it sounds, because it tells you exactly what Laya is for. It is an encoder you fine-tune on your own labels — the whole architecture is built so that new schemas need no retraining, but performing well on a task does. Used that way, the published figures are strong: 0.766 against Jev's 0.727 on typed decisions after fine-tuning, AG News 0.950 against 0.910, DAIR Emotion 0.595 against 0.480, and ECE 0.081 against Jev's 0.246 — with the honest note that it ships over-confident and needs temperature refitting, which moves mean ECE from 0.466 to 0.081.

Microsoft publishes the methodology and withholds the result

Microsoft-Decision-1's Foundry page runs the same exercise in the other direction. It describes the evaluation in detail — public and community decision benchmarks plus held-out internal sets, metrics including accuracy and calibration error, option order varied, paired statistical tests, identical harness for compared models — and reports that the model "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology." It names its strong areas (reasoning, rule application, robustness to prompt formatting) and its weak ones (specialized domain knowledge, non-English).

And then it prints nothing. No accuracy figure, no calibration error, no per-benchmark table. The self-reported limitations are real and useful — scores shift with phrasing and option ordering, a poorly framed question still returns a score, calibration is strongest on familiar task types, no explanations are produced — but a limitations list is not a measurement. So the trade is unusually clean: Laya gives you numbers you can attempt to falsify on your own data, and Microsoft gives you a compliance posture and a 32K context you can put a long document into.

A generated two-column scoreboard for Microsoft-Decision-1 and Laya across six shared dimensions: architecture a 9B dense decoder post-trained by Microsoft versus a 421M ModernBERT-large with a from-scratch decision head; languages 25 supported with coverage caveats versus 100+ with 45 of 51 usable; context 32,768 tokens versus a 512-token English checkpoint and a 1,024-token multilingual one; published calibration none versus vendor-reported ECE 0.081 after temperature refitting; fine-tuning not available versus a documented specialisation path; and cost a Foundry per-token rate versus zero on your own hardware. A footer notes both sides are vendor-reported and neither is independently audited.

What the size difference actually buys

Laya's smallness is not a compromise here, it is the design. Because every option is scored at its own [MASK] token and softmaxed over that question's options, the answer space is assembled at request time rather than baked into a vocabulary head — which is why a schema you invent this afternoon needs no retraining, and why the model fits in 322 to 421 million parameters. The multilingual checkpoint runs about 2.2 times faster than the English one, the router detects script in under half a millisecond, and batched throughput reaches 103 to 332 questions per second on one T4. For a workload that scores every ticket, every retrieved document or every proposed action, that throughput is the feature, not the parameter count.

The costs of that design are equally specific and worth stating against Microsoft's alternative. The English checkpoint takes a 512-token window; the multilingual one 1,024, extensible toward 8K. Microsoft-Decision-1 takes 32,768. A long support thread, a full contract or a multi-page policy document does not fit in Laya's window at all, and no amount of speed compensates for truncation. Laya answers one question set per forward pass with a documented token budget per option; Microsoft documents a single invocation over up to 32K tokens without a per-question ceiling. And Laya's competitive weak spot as a classifier is worth knowing: it scores 0.425 on Banking77 against Jev's 0.870, which is the kind of fine-grained, high-cardinality intent problem where a specialisation pass matters most.

A screenshot of the Laya model card on Hugging Face, read 10 October 2026, showing the convaiinnovations organisation, a TextClassification pipeline tag, Transformers and Safetensors badges, and the Laya and system-one tags on the model page.

The cost comparison that is not per token

Laya is free at the point of use — self-hosted, Apache 2.0, listed at a self-hosted cost of "$0" — and its real price is the fine-tuning run you owe it before it is good at your task. Microsoft-Decision-1 is a hosted Foundry API with a per-token rate that is not printed on the model page, no weights to download, no fine-tuning path, and batch inference disabled. One is an upfront data-labeling project, the other is a metered call. Which is cheaper depends entirely on volume, and the crossover is high: a 421M encoder scoring 300 items a second on a rented T4 is very hard to beat per decision once it has been trained.

Here is where OrcaRouter earns its place in a pipeline built on either model, and it is the generative side only. We host neither Microsoft-Decision-1 nor Laya: both return probabilities rather than text, neither is in our catalogue, and a scoring endpoint is not a chat-completions target. What we do carry is the model that produces the thing being scored — the draft reply, the proposed tool call, the candidate answer that Laya's fine-tuned head then judges. That is more than 200 models behind a single OpenAI-compatible key at provider list price passed through with 0% markup, with automatic failover when one serving path degrades. In a loop that scores everything it generates, the generation call is the part that can fall over, and failover is what keeps the scoring queue fed.

A screenshot of the OrcaRouter models catalogue page headed 207 models from 16 providers behind one API key and one bill, with filters for input modalities, context length, input price, status, series and supported parameters. Neither Laya nor Microsoft-Decision-1 appears.

Which one to pick

Take Laya if your decisions arrive in many languages, if your volume is high enough that per-token pricing matters, if you have labels and a week to fine-tune, or if you need to run scoring on hardware inside your own boundary. Accept the 512-to-1,024-token window and the fact that the base checkpoint will be near chance until you specialise it, and you get a decision model that is honest about being a starting point.

Take Microsoft-Decision-1 if your inputs are long documents rather than tickets, if your deployment gate is Azure authentication, unified billing and a Responsible AI package rather than a download, if 25 languages covers your traffic, or if you would rather not own the model at all. Accept that you will be drawing its first calibration curve yourself, and budget an afternoon on a labeled sample to do it — because unlike Laya, this one will not hand you an ECE number to argue with.