Generated title card for RSI-Jev vs Jev 1.13 reading 'One you download. One you call.', with a left card labelled 'RSI-Jev v6.1-VL 4B' carrying a download-tray icon and the line 'Apache-2.0 weights, 4.69B', a right card labelled 'Jev 1.13' carrying a cloud-endpoint icon and the line 'Hosted API, $0.042 / 1M in', a connecting line between the two cards, and the caption 'Same wire format, different contract'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

RSI-Jev vs Jev 1.13: One You Download, One You Call

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Put RSI-Jev v6.1-VL 4B and Jev 1.13 side by side and the first thing a caller notices is that they are the same request. Hand either one a state and a set of typed questions — a yes/no, a pick-one-of-k, a rate-on-a-rubric — and both return a calibrated probability for every option, in one forward pass, with no generated text to parse. That is not a coincidence: RSI-Jev is built to speak Jev's wire format on purpose, so a client written against TypeSafe's API runs against it by changing a base URL. What is not the same is everything around the call. Jev 1.13 is TypeSafe's closed commercial model, served from an endpoint you meter; RSI-Jev v6.1-VL is a 4.69B-parameter checkpoint under Apache-2.0 weights that you download and serve on your own hardware. Neither one is a re-brand of the other, neither vendor endorses the other, and almost every figure in this comparison comes from the party that produced it.

The date on the subject matters, because this project ships a release roughly every day. RSI-Jev v6.1-VL 4B was published on 2026-10-07 by the third-party Shanghua-Gao/RSI-Jev project — a self-improving research loop that trains Jev-style decision models and publishes every arm that failed along with the ones that won. It is the eighth release in thirteen days on that line, and it is a weight average of the previous release with a second fine-tune of the same Qwen3.5-4B-Base. Jev 1.13 is TypeSafe AI's model, launched 2026-09-15 and carried on our own catalogue since 2026-09-24. Both dates matter below, because a comparison against a project that moves daily has a shelf life measured in days.

What the two actually are, in one line each

Jev 1.13 is a hosted decision model behind a dedicated endpoint — POST /v1/systemone, non-streaming, roughly 64,000 tokens of input budget across the state and all your questions together, priced at $0.042 per million input tokens with output billed at zero because there are no output tokens. Its architecture, parameter count and training compute are undisclosed; TypeSafe has said the details are being kept close and a paper may follow. You do not run it. You call it, and every call is a metered network request.

RSI-Jev v6.1-VL is the other arrangement in full. It is a Qwen3.5-4B-Base tower fine-tuned end to end, with decision heads at layers 16, 20 and 32, served from a checkpoint that is self-contained and 9.7 GB in bf16. The parameter count is 4.69B and it is worth knowing where it goes: 3.57B in the 32 decoder layers, 0.64B in the token embeddings, 0.33B in the vision tower, 0.05B in the main decision head and 0.10B in the two early-exit heads. There is no mixture of experts and no second model. You install it with a pip command from the repository, run its server, and from that point the decision never leaves your infrastructure.

• Who runs it — a metered hosted endpoint you do not control vs a 9.7 GB checkpoint on your own GPU, Apple Silicon or CPU.

• Price shape — $0.042 per million input tokens, output free, paid per call vs zero at the margin plus the cost of the machine and the operations.

• Input budget — about 64,000 tokens per request on the hosted model vs 32,768 text tokens plus an image budget on the checkpoint, with anything longer refused rather than truncated.

• Weights and licence — closed, undisclosed size vs Apache-2.0 weights, MIT code, 4.69B parameters.

• Modality — text for Jev's contract vs text plus up to four images per request on the RSI-Jev vision releases.

• Ownership — TypeSafe AI's commercial model vs a third-party research project that states in its own licence line it is "Not affiliated with TypeSafe AI."

The score on RSI-Jev's own board, and why it is only half a comparison

The number the project leads with is its Decision Index 0.3 score: 50.98 for v6.1-VL 4B, up from 46.23 for the release before it. That is a full run of the default configuration — 140,178 requests, coverage 1.0 — and on the project's own public board, dated 2026-10-06, it ties the best 4B model on that board (50.98 against ezjev 4B s2's 50.82, which the kit treats as a tie at 0.25) and sits 27th of 113 overall. On the older Decision Index 0.2.1 it reads 50.74 against v6.0-VL's 46.24. Its fifteen-benchmark suite, reported without the open_jev_ood task that was found to overlap training rows, is 0.793, and its held-out set is 0.729.

Every one of those figures is RSI-Jev's own, measured on RSI-Jev's harness. The Decision Index is a public benchmark board, but there is no reading of Jev 1.13 on it, because the project's suite was built to score open decision checkpoints and Jev is a closed endpoint. So the temptation to set 50.98 against Jev's 0.727 on the typed-decisions benchmark and declare a winner is exactly the mistake to avoid: those two numbers come from different harnesses, different sample sizes and different data, and nobody has run one harness over both models.

Headless Chromium capture of the GitHub repository page for Shanghua-Gao/RSI-Jev: the repository header with the Public badge and the counters Fork 5 and Star 80, the repository description about typed-decision models (noul / choice / score) trained by a self-improving loop of AI agents with the checkpoints, the code that produced them and every version that failed, a commit list headed by the merge commit 'Merge pull request #35 from Shanghua-Gao/copy-no-ranking', the file rows for the v6.1-VL weight-averaging work and the v5.0-VL 3B quickstart, the counters 202 commits, 8 tags and 8 releases, the MIT license line, and the topic tags decision-model, jev, lm, system-one and typed-decisions.

The one head-to-head that exists is Laya's, not RSI-Jev's

There is one published comparison that does put a Jev number next to an open checkpoint, and it was not run by either party here. Convai Innovations, the makers of the Laya decision model, tabulated TypeSafe's published Jev 1.13.0 figures against their own and flagged the limits themselves: the Jev numbers are third-party published and were never measured by Convai, the sample sizes and prompts differ, and the vendor does not list its own benchmarks for the model. That table is worth reading for calibration, not for a verdict — and it does not include RSI-Jev at all, because RSI-Jev did not exist when it was published.

What it does show is the shape of the hosted-versus-open question a reader is actually weighing. The hosted model leads where the option space is large and the model has to hold a wide answer set steady; the open models win on raw latency per call because there is no network in the path. Nothing in that pattern tells you which of these two specific models is better at your task, and the honest position is that the answer does not exist yet in public.

What the checkpoint buys that the endpoint cannot

The strongest argument for RSI-Jev is not a score. It is that the weights sit on your disk. For a routing decision made on a medical record, a legal document or a customer's account history, "the data never leaves the building" is not a preference you trade against a benchmark point — it is a hard requirement, and no hosted endpoint at any price answers it. The same property removes the rate limit: the vendor's own documentation for the hosted model notes that its limits are adjusting dynamically and can change without notice, and a self-hosted checkpoint has no such ceiling beyond your hardware.

The second thing the checkpoint buys is depth control, and it is unusual. Because the decision heads sit at three depths, an effort setting picks how many layers a request may use: low stops at layer 16 for about 23 ms median, medium at 20 for 27 ms, high at 32 for about 40 ms, and auto answers at the first exit confident enough, averaging 19.5 of 32 layers on the project's suite. Those latencies are the project's own numbers clocked on one H200 in bf16 and should not be blended with any hosted figure — a local forward pass and a metered API call are not the same measurement, and RSI-Jev's own documentation is explicit that its earlier comparison against Jev's published latency pitted local GPU work against a network round-trip.

The third thing is the images. Jev's contract is text in, structured JSON out. The RSI-Jev vision releases take one to four images per request as base64 data URLs, with the state referring to each by a marker, and v6.1-VL scores 0.834 on the project's held-out image set. If your decision is "does the photo show visible damage," that is a capability the hosted contract does not offer at all.

What you give up is real too, and the project publishes it. Calibration got worse in this release, not better: final expected calibration error is 0.048 at layer 32 and 0.055 with auto, against 0.036 and 0.024 for the previous release. The default single threshold of 0.95 ships explicitly unconfirmed — it is the fallback of a selection rule whose own pick, 0.85, missed the project's depth cap on half its development data. The early exits read text only, so any question with an image runs all 32 layers regardless of effort. And five of the image training sources are non-commercial or research-only, with the project stating plainly that whether weights trained on non-commercial data inherit those terms is not settled.

Where to call the hosted one, and where not to

This is the part of the comparison we have a stake in, so it is worth being precise. We serve TypeSafe's commercial model as typesafe/jev-1.13 on the dedicated systemone endpoint — a POST to /v1/systemone rather than the OpenAI chat-completions shape, non-streaming, against the 65,536-token context our catalogue lists. It is the same request and answer shape RSI-Jev implements, from the model whose contract the project copies. RSI-Jev itself we do not host; there is no rsi-jev id and no shgao id in our catalogue, and a reader who wants that model downloads it.

The reason that distinction matters here is narrow and concrete. A decision layer is rarely the whole of a workflow — it usually sits beside a generative model that writes the reply, the summary or the code. That has historically meant two contracts. It no longer has to for the hosted half: Jev 1.13 sits on the same key as 200+ other models at provider list price passed through with 0% markup, so if TypeSafe changes a rate, the change is live on our side the same day rather than on the next billing cycle. The self-hosted half never had that problem, because you are the provider. The clean way to decide between them is to try the commercial contract on a handful of your own labelled cases first, see whether the out-of-the-box behaviour is good enough to automate against, and only then work out whether running a 4.69B checkpoint yourself is worth the operations.

Headless Chromium capture of OrcaRouter's own model page for typesafe/jev-1.13: the breadcrumb 'Home / Models / TypeSafe', the page title Jev 1.13 above the slug typesafe/jev-1.13, the line 'by TypeSafe - 2026-09-24', the description that it is TypeSafe's structured decision and evaluation model taking noul / choice / score questions and returning a structured answer for each, the note 'POST /v1/systemone; non-streaming; up to ~64K input tokens; text in, structured JSON out.', the endpoint panel reading /v1/systemone with the price $0.04, our p50 TTFT of 161 ms, 363 ms and 58.9M, and the buttons 'Get the Jev 1.13 API', 'Try in playground' and 'Use via API'.

Which one you should actually pick

Pick RSI-Jev v6.1-VL 4B if the decision has to stay inside your perimeter, if you need a decision on an image as well as text, if your option sets run into the hundreds (the checkpoint admits up to 5,120 options per question), or if you want to tune depth and latency per request. Go in knowing that you are adopting a project that moved eight times in thirteen days, that its latest release traded calibration for accuracy, and that its own card names the parts of the exit policy it could not confirm.

Pick Jev 1.13 if you want the decision to work without a serving stack, if you value an endpoint that someone else keeps up, and if the $0.042-per-million-input price — with no output tokens to meter — is cheap against your call volume. Go in knowing that you are calling a closed model whose size is undisclosed, whose rate limits can shift without notice, and whose published benchmarks are not something you can re-run.

The thing both share is more useful than what separates them, and it is the reason a comparison like this is worth writing at all. Neither model generates text, so neither introduces the class of failure that comes from a model that forgets to close a brace or invents a field. Both return probabilities, and in both cases the probability is the part you have to validate on your own labelled data before you automate against it — the latency is already commoditised, and the confidence value is what has to be earned per deployment. Whichever side of the download-versus-call line you land on, test the calibration first.

A generated two-column scoreboard titled 'RSI-Jev v6.1-VL 4B vs Jev 1.13 - the scoreboard', six rows across both columns: who runs it, 'You, on your own GPU' against "TypeSafe's hosted endpoint"; weights, 'Apache-2.0, 4.69B' against 'Closed, undisclosed'; input budget, '32,768 tokens' against 'About 64,000 tokens'; price, 'Free at the margin' against '$0.042 per 1M input'; modality, 'Text + up to 4 images' against 'Text only'; and latency, 'Local pass, ~23-40 ms' against 'Metered network call'. A footer reads 'RSI-Jev figures vendor-reported; Jev 1.13 pricing per our catalogue.'