Generated title card for RSI-Jev vs Laya reading 'Two open decision models. Opposite bets.', with a left card labelled 'RSI-Jev v6.1-VL' carrying a brain-cog icon and the line '4.69B params, zero-shot', a right card labelled 'Laya' carrying a feather icon and the line '421M params, fine-tune it', a thin vertical divider between the two cards, and the caption 'Both Apache-2.0. Neither one is hosted for you.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

RSI-Jev vs Laya: Two Open Decision Models, Opposite Bets

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

As of October 2026 there are two open-weight models worth considering if you want typed decisions — a yes/no, a pick-one-of-k, a rate-on-a-rubric — without a hosted endpoint in the path. They are RSI-Jev v6.1-VL 4B, published 2026-10-07 by the third-party Shanghua-Gao/RSI-Jev research loop, and Laya, released 2026-09-18 by Convai Innovations. Both answer in a single forward pass with no generated text; both return a calibrated probability for every option; both ship under Apache-2.0 weights with a server that speaks TypeSafe's decision API, so a client written for the commercial model runs against either by changing a base URL. They are also built on opposite bets, and the two numbers that separate them are 3 to 4 tokens per label and 421 million parameters.

The first bet is about scale. Laya's English checkpoint is ModernBERT-large at 421M parameters with a 512-token context, and its multilingual checkpoint is mmBERT-base at 322M with a 1,024-token window that reaches 8,192 with the encoder set wide. RSI-Jev v6.1-VL runs an entire Qwen3.5-4B-Base tower — 4.69B parameters, 3.57B of them in the 32 decoder layers, plus a vision tower and three decision heads at layers 16, 20 and 32. Laya is a model you can preload three of in a few gigabytes; RSI-Jev is a 9.7 GB bf16 checkpoint. The second bet follows from the first: Laya spends almost nothing per call and expects you to teach it your domain, while RSI-Jev spends four and a half billion parameters trying to answer your question without any training at all.

What each one is actually for

Laya's own model card contains the sentence that frames this comparison better than any reviewer could: "Laya is a fast base to specialise, not a zero-shot decision engine." On the typed-decisions benchmark — 2,000 decisions across four workflows — the base English checkpoint scores 0.362, against 0.318 for random guessing and 0.461 for always picking the majority class. On the same decisions, the checkpoint Convai fine-tuned on that benchmark's own training split scores 0.766, clearing the 0.735 teacher-agreement ceiling. That gap is the product: Laya is a 421M encoder you fine-tune on a narrow taxonomy, and Convai ships a Kaggle notebook that runs the whole loop — build the dataset, train, fit calibration temperatures, evaluate — on two free T4s.

RSI-Jev is the other trade. Its v6.1-VL release scores 50.98 on its own public Decision Index 0.3, a 140,178-request run of the default configuration, which on the project's board dated 2026-10-06 ties the best 4B model there and sits 27th of 113 overall. Its fifteen-benchmark suite is 0.793 and its held-out set is 0.729. Those are zero-shot figures — though the project is careful to say that ten of the fifteen benchmarks contribute training data in some form, so "zero-shot" applies to the release's search, not to every number on the page. What the figures do buy you is a model that answers a question about a document you have never shown it and does not need a training run first.

• Size — Laya 421M English / 322M multilingual vs RSI-Jev 4.69B, running the whole Qwen3.5-4B-Base tower.

• Zero-shot accuracy — Laya 0.362 on typed-decisions, below the 0.461 majority baseline vs RSI-Jev 50.98 on its own Decision Index 0.3, tying the best 4B entry there.

• Fine-tuned accuracy — Laya 0.766 with a checkpoint trained on the benchmark's own split vs RSI-Jev's figures being release-level, not per-domain.

• Languages — Laya 45 of 51 languages usable above three times random server-side routing vs RSI-Jev English-centric text.

• Modality — Laya text only vs RSI-Jev text plus up to four images per request.

• Context — Laya 512 tokens English, 1,024 multilingual and up to 8,192 for long documents vs RSI-Jev 32,768 tokens, refused rather than truncated.

• Latency — Laya 32.8 ms p50 on a T4, 7.2 ms per question at batch ten vs RSI-Jev 22.5 ms at effort low and about 40 ms at full depth, on an H200.

The option-count cliff is the sharpest difference

Both models define the answer space at request time, so a new schema needs no retraining — that is the shared structural win of this whole family. But they allocate the answer space differently, and the difference shows up on exactly the tasks enterprise routing tends to be. Laya scores every option at its own masked token, and the options share a fixed head budget: 192 tokens on the English checkpoint, 256 on multilingual. On Banking77, with 77 intents, that works out to roughly three or four tokens per label, and accuracy falls to 0.425. Convai's own card documents the cliff and offers the fix — raise head_max_len to 512 and the context to 1,024 or more so each label has room, or split a large option set into a coarse-to-fine two-step choice.

RSI-Jev's serving path admits up to 5,120 options per question. That is not a like-for-like benchmark against Laya's 0.425 — the two were measured on different harnesses, and the option ceiling is a configuration limit, not a score. It is a statement about which model will not fall over when your taxonomy has a hundred entries. If your choice questions are "billing / technical / sales / other," either works. If they are a hundred-odd intent labels, one of these two needs tuning before it is usable and the other does not.

Headless Chromium capture of the GitHub repository page for Shanghua-Gao/RSI-Jev: the repository header with the Public badge and the counters Fork 5 and Star 80, the repository description about typed-decision models (noul / choice / score) trained by a self-improving loop of AI agents with the checkpoints, the code that produced them and every version that failed, a commit list headed by the merge commit 'Merge pull request #35 from Shanghua-Gao/copy-no-ranking', the file rows for the v6.1-VL weight-averaging work and the v5.0-VL 3B quickstart, the counters 202 commits, 8 tags and 8 releases, the MIT license line, and the topic tags decision-model, jev, lm, system-one and typed-decisions.

Latency, the language question, and the images

Laya's latency claim is its loudest, and it is a real one: 32.8 ms p50 for a single question on a Tesla T4, 72.3 ms for a batch of ten, and 337 ms for fifty — 103 to 332 questions a second on one modest GPU. RSI-Jev's own numbers, clocked on an H200, are 22.5 ms at effort low, 26.8 ms at medium, 39.9 ms by default and 40.4 ms at full depth. Those are the same order of magnitude, and both are local forward passes rather than network calls, which is the comparison that actually matters once a hosted endpoint is out of the picture. Note what the two are doing with the time: RSI-Jev can deliberately stop at layer 16 to buy that 22.5 ms and run all 32 when the question is hard, and Laya has no such dial — it always runs its whole (small) encoder.

The language story runs the other way and is decisive for anyone outside English. Laya ships a router that detects the script in well under a millisecond and dispatches to the multilingual checkpoint, and its published table shows 45 of 51 languages usable above three times random, against 23 for the English checkpoint alone. Its card is also honest about why that matters: the English checkpoint collapses on non-Latin scripts — Khmer scores 0.000 accuracy at 0.952 confidence — so confidence gating cannot rescue a wrong route. RSI-Jev has no language story of this kind; it is an English-centric text model that happens to read images.

Images are the mirror image. RSI-Jev takes one to four per request as base64 data URLs and scored 0.834 on its held-out image set in this release; Laya's shipped checkpoints are text classifiers, and while community ports like laya-vision exist on the Hub, they are not the vendor's product. If your decision is made about a photo, a chart or a screenshot, that is one model's column and not the other's.

Neither one has been benchmarked against the other

This is the part a spec-by-spec comparison quietly hides: there is no head-to-head between these two models. What exists is a head-to-head between each of them and the same closed model — TypeSafe's Jev 1.13 — and neither can be put beside the other.

Convai published one: on typed-decisions, Jev 1.13.0 at 0.727 against Laya's routed 0.766; on Banking77, Jev 0.870 against Laya's 0.425; on calibration, Jev 0.246 against Laya's post-temperature-fix 0.081. Convai flags its own limits in the same table — the Jev figures are third-party published, they never had API access to measure it, and the sample sizes and prompts differ. RSI-Jev's comparison is the Decision Index, which is RSI-Jev's own public board and carries no Jev entry at all. So the only external numbers that touch both of these models come from harnesses built by the parties selling them, and the sensible reading of any single figure above is "this is what the maker measured, on the maker's task."

What both projects do well, and rarely, is publish their own weaknesses. RSI-Jev names the five non-commercial image sources behind its vision releases and says plainly that whether weights trained on them inherit those terms is not settled; it reports a calibration regression in its newest release and calls its default exit threshold unconfirmed. Laya documents that its base checkpoints sit below the majority baseline, that its ordinal score questions are its weakest primitive, that its noul can follow its own option labels rather than the state, and that one of its response fields carries no usable signal. That honesty is the most useful thing to inherit from either project: check the confidence values on your own labelled cases before you automate against them.

Headless Chromium capture of OrcaRouter's own model page for typesafe/jev-1.13: the breadcrumb 'Home / Models / TypeSafe', the page title Jev 1.13 above the slug typesafe/jev-1.13, the line 'by TypeSafe - 2026-09-24', the description that it is TypeSafe's structured decision and evaluation model taking noul / choice / score questions and returning a structured answer for each, the note 'POST /v1/systemone; non-streaming; up to ~64K input tokens; text in, structured JSON out.', the endpoint panel reading /v1/systemone with the price $0.04, our p50 TTFT of 161 ms, 363 ms and 58.9M, and the buttons 'Get the Jev 1.13 API', 'Try in playground' and 'Use via API'.

Where the contract they copy actually lives

Both of these models exist because one wire format was worth copying. TypeSafe's Jev defines the request shape — state, questions, three typed primitives — and the answer shape, and both Laya and RSI-Jev implement it so that an existing client works by changing a base URL. That reference model, typesafe/jev-1.13, is the one of the three we serve: it is on our catalogue on the dedicated systemone endpoint, a POST to /v1/systemone, non-streaming, against a 65,536-token context, at $0.042 per million input tokens with output billed at zero. Neither Laya nor RSI-Jev is in our catalogue — both are downloads, and that is the point of them.

The practical version of that matters more than the comparison. Decision layers are almost never alone in a stack; they sit next to a generative model that writes the reply, the summary or the code. Having the reference contract on the same key as 200+ other models, at provider list price passed through with 0% markup, means a vendor rate change reaches you the same day, and automatic failover means the generative half of that pair is not a single point of failure while you work out whether the cheap decision half is good enough. If you decide a self-hosted 421M encoder or a 4.69B checkpoint is the right call, you still want the contract it speaks to be reachable from the same place — and if you would rather not run either, the model both of them copy is one request away.

Which one to download

Choose Laya if you have labelled data, a taxonomy that does not move much, and a language mix that is not English-only. It is small enough to run many of, fast enough to put in front of every request, and designed from the ground up to be fine-tuned — the 0.766 fine-tuned score against 0.362 zero-shot is the whole argument. Budget for the training run, the labelling and the per-question-type temperature refit that takes its calibration error from 0.466 to 0.081, and keep the option sets under roughly twenty labels or raise the head budget before you trust a large classification.

Choose RSI-Jev v6.1-VL if you want a decision to work without any training first, if your questions are sometimes about an image, if your option sets are large, or if you want to trade latency for depth per request. Expect to run a 9.7 GB checkpoint rather than a 400M one, expect a project that has published eight releases in thirteen days and may publish another while you are evaluating this one, and expect to check its calibration yourself — this release's own card says it got worse, not better.

Whichever you pick, the same two things are true. Neither model generates text, so neither can fail by emitting a malformed field; both return probabilities, and the probability is the part that has to be validated per deployment rather than taken from a card. And both have moved the interesting question of a decision layer from "whose API" to "whose weights" — which is a better question to be asking, and one this pair answers very differently.

A generated two-column scoreboard titled 'RSI-Jev v6.1-VL vs Laya - the scoreboard', six rows across both columns: size, '4.69B parameters' against '421M English / 322M multilingual'; zero-shot, '50.98 Decision Index' against '0.362 typed-decisions'; fine-tuned, 'Not per-domain' against '0.766 on its own split'; languages, 'English-centric' against '45 of 51 usable'; modality, 'Text + up to 4 images' against 'Text only'; and latency, '~23-40 ms on an H200' against '32.8 ms p50 on a T4'. A footer reads 'Both vendor-reported; neither has been benchmarked against the other.'