
What Is RSI-Jev? A Self-Improving Loop That Builds Jev-Style Decision Models
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 147 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 80 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 53 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 296 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
RSI-Jev is a third-party, open research project that builds Jev-style System One decision models, and the model this page is about is its 4B release, RSI-Jev v6.0-VL, dated 2026-10-06. It is written by Shanghua Gao (@gasvn), with Sufian (@SufianTA) credited in the repository's acknowledgements, and it is not TypeSafe's Jev and not affiliated with TypeSafe AI — the project's own licence line says exactly that. The idea underneath it is narrow enough to state in a sentence: ask a typed question about a document, a chat or an image — yes/no, pick-one-of-k, rate-on-a-rubric — and one forward pass returns a calibrated probability for every option. Nothing is generated, so there are no reasoning tokens to spend, and none are spent. v6.0-VL is not the project's first release; it is its seventh in twelve days, which is the single most important thing to understand about it, because the useful information here is in the shape of the line rather than in any one checkpoint.
One thing has to be said before any of the numbers, because it dates all of them. v6.0-VL was the head of the line for exactly one day. On 2026-10-07 at 07:56 UTC — this morning — the project published RSI-Jev v6.1-VL, which is v6.0-VL averaged, weight 0.5 each, with a second fine-tune of the same Qwen3.5-4B-Base trained on other data, and which scores 50.98 on the project's Decision Index 0.3 kit against v6.0-VL's 46.23 on that same kit. Nothing was trained after the average. Its calibration is worse than v6.0-VL's, and its own card says so. That release is real and current; this page is not about it. Every figure below is read from v6.0-VL's release record, dated 2026-10-06, and where a count has moved since — the release tally, the experiment tally — this page gives both the figure as it stood for v6.0-VL and the figure as it reads today.
What RSI-Jev is not is also worth putting early, because two of the three obvious assumptions are wrong. It is not a hosted product you can call today through a general API, and it is not served by OrcaRouter — our catalogue carries no rsi-jev id, no shgao id and no model card for it. The one thing we do have is the model whose HTTP contract this project copies: TypeSafe's commercial Jev, which we serve as typesafe/jev-1.13 on the systemone endpoint. One of those two you call, and the other you download and serve yourself. Everything below comes from the project's own repository, release cards and serving documentation, read on 2026-10-07, and where a number is the project's own rather than an outside measurement, this page says whose it is.

What the thing actually does
The project describes itself in one line as "a recursively self-improving research system that builds Jev-style System One models," and the artifacts it produces are deciders rather than generators. You hand it a state — a document, a chat transcript, a transaction, and in the vision releases up to four images — and one or more typed questions with named criteria. It returns, for each question, a probability for every option. Three question types cover the space, and they are the ones TypeSafe's API defines:
• noul — a true/false judgement, returned as a single probability, with no distribution and no confidence value, matching the reference's answer shape exactly.
• choice — pick one of a set of labelled options, returned with the full probability distribution and a confidence statistic.
• score — rate on an ordered rubric, returned as a probability-weighted zero-based index into the levels, plus the legend and the distribution.
Because there is no generation step, there is no second model call and no sampling. A decision about a document the model has already read is a cheap operation by construction, and the project's own line for that cost — "about 10 ms" — belongs to its 2B era, not to the current release; the measured figures for v6.0-VL are given further down.
Two facts about ownership matter more than anything else on this page. RSI-Jev is not TypeSafe's work and TypeSafe has not endorsed it. The licence line, quoted in full: "Code: MIT. Weights: Apache-2.0, following the base model; some image training sources are non-commercial, listed on each model card. Not affiliated with TypeSafe AI." And the relationship runs one way: the project copies Jev's wire format deliberately, and says so, because a compatible server is the point. "Jev-style" is the project's own phrase for the kind of model it builds. TypeSafe's Jev is a different, closed, commercial model, and the two are not the same thing under a shorter name.
Inside the current model: a 4B Qwen tower with three exits
RSI-Jev v6.0-VL is a Qwen3.5-4B-Base tower with the tower fine-tuned and a trained decision head on top. That is the whole architecture — there is no mixture of experts, no router, and no second model. It runs the entire base model, which is why its parameter count is 4.69B and not something smaller: 3.57B sits in the 32 decoder layers, 0.64B in the token embeddings, 0.33B in the vision tower, 0.05B in the main decision head and 0.10B in the two early-exit heads. The released checkpoint is self-contained and 9.7 GB in bf16.
Three decision heads are attached, at layers 16, 20 and 32 of the base, and they are the mechanism behind everything the current release is known for. A fourth exit at layer 12 was built, measured and dropped — "The layer-12 exit lost to the cascade from 16 in every comparison and is not in the package" — so three ship and four do not. The exits read a detached copy of their layer, a detail the project found the hard way: retuning heads on a trunk whose exits had been attached during training had not recovered the deep layers' accuracy, so detaching them is what restored depth.
One number here is the easiest way to misread the project. Everything up to and including v3.0 was a 2B model on Qwen3.5-2B-Base, and that is the lineage, not the current model. v4.0-VL was 2B, v5.0-VL cut a model to 3B, and v6.0-VL is 4B. A page that calls the current RSI-Jev model 2B is three releases out of date.
The loop is the actual project
The models are the output; the thing being built is the process. The project states that "the loop running the research is the next version of AutoScientists," the self-organising agent-team system published by the Zitnik lab at Harvard, and it works the way that sentence implies. AI agents propose hypotheses, register their predictions before spending GPU time, run the experiments, and retire their own champions when the evidence says to. Two counts make that concrete. Read on 2026-10-07, the repository's headline is eight releases in thirteen days, from v1.0 to v6.1-VL; for v6.0-VL's own release on 2026-10-06 it read seven releases in twelve days, each one trained, evaluated and documented by the loop. And the experiment count, which stood at 471 when this page's release shipped, reads 496 today — every one written up, failures included. Both counts are the project's own, and both move.
The discipline is what makes those numbers mean something, and the project lists it plainly. Null floors are measured rather than assumed — arms that are provably identical to the control, verified by object identity before any GPU time, so that the spread between them is the noise floor and a difference smaller than that spread is not a result. Predictions are registered before the run, so a version that misses its own bar ships as a failure rather than being quietly re-cut. Artifacts are verified: a checkpoint is reloaded from disk and re-scored, and published only if it reproduces its training run's per-question predictions, which both v1.0 checkpoints do at 1.0000. Contamination is "checked rather than asserted." And failures ship, including the ones that killed the project's own champion.
What a contribution is, in the project's own words from its contributing guide: "A contribution here is usually a measurement, not a patch." The released record is kept as a chain rather than as a snapshot — "versions/ keeps one card per release, all of them, on main forever… That chain IS the project" — which is why an old release's numbers can be checked against what the project says about them later, and why the one correction discussed below is visible instead of silent.
What v6.0-VL changed: it spends depth instead of tokens
The current release's mechanism is a setting called effort, and it controls something unusual: how many layers of the model a request is allowed to use. Because the heads sit at three depths, a question that is easy can be answered at layer 16 and a question that is hard can run all 32. low stops at layer 16, medium at 20, high at 32, and auto answers at the first exit whose calibrated probability clears that exit's threshold. Median latency per request on the Decision Index sample, measured on one H200 in bf16: 23 ms at low, 27 ms at medium, 40 ms at high, and 40 ms for the unset default. These are the project's own measurements on its own hardware and they should not be blended with the serving documentation's GB10 numbers, which are a different machine.
The measured behaviour of auto is the interesting part: on the project's fifteen-benchmark suite, 20% of questions stop at layer 16, 46% at layer 20 and 34% run to 32, which averages 23.3 of 32 layers. A single fixed threshold averages 20.9, and over-stopping is what a single threshold buys you. auto is not a compromise on quality, which is worth stating because that is usually what an adaptive setting is: it posts the best suite row of any setting (0.771 against 0.770 for the default), the best MMLU-Pro row (0.444 against 0.440) and the best final calibration (ECE 0.024 against 0.036). The one place high wins is the held-out set, 0.702 against auto's 0.696. Some tasks get measurably worse with depth — BANKING77 by 0.035, New Yorker caption matching by 0.060 — which is why the effort level is a choice the caller makes rather than a rule the server imposes.
The result that made this a release rather than an experiment is on the project's public Decision Index 0.2.1, where the score went from 38.38 for v5.0-VL to 46.24 for v6.0-VL in one release. On the public board dated 2026-09-28 that is the highest score among 4B models and anything smaller, and 14th of 71 overall; the next entry at that size is JPT-4B at 43.04. The earlier releases on this line are lineage and not the current model: v5.0-VL (2026-10-02) cut the model to the first 20 of 32 layers and made it say "unknown" when a question has no answer; v4.0-VL (2026-10-01) was the first one that reads images; and v3.0 (2026-09-28) is the release where reinforcement learning first helped, via a listwise ranking reward — NDCG@5 over 16 candidates — that lifted reranking R@1 from 0.192 to 0.308 against its supervised parent at a cost of 0.0028 on the suite. That is where the project's RL story starts, and it is three releases back now.

The numbers, with the caveats that travel with them
The Decision Index 0.2.1 headline for v6.0-VL is 46.24, on a full run in which all 150,759 requests of the suite were answered: knowledge 28.8, language 46.2, retrieval 55.5, tools 65.8, arts 37.1. The fifteen-benchmark suite reads 0.770, the held-out set 0.698, MMLU-Pro 0.440, and final ECE 0.024 with auto. Two caveats have to sit in the same breath as those figures, because without them the numbers mislead.
The first is a cut in the suite. The internal suite task open_jev_ood overlapped 579 training rows, so its number was inflated by an unknown amount; from v6.0-VL onward the project reports the suite without it, at 0.770. v5.0-VL's 0.764 was with it, and v6.0-VL's card restates that release as 0.763 without it. The two figures are not comparable, and if you do compare them you have to use the restated 0.763 and say that is what you are doing. The held-out set, MMLU-Pro and BBH have no overlap and are unaffected. A related audit found about 1,000 items of the Decision Index kit's test rows in the training corpora — ANLI 274, RouterBench-GSM8K 90, ARC 5, and BRIGHT/ToolRet query text without labels, roughly 0.3% of the kit's rows — and re-scoring without them moves the index by at most 0.04 on the project's sample. That is a correction to v5.0-VL's record, published in v6.0-VL's card, and it is the project's own contamination rule costing it a number.
The second is whose benchmark this is. The Decision Index is RSI-Jev's own public board, not a third-party verdict, and 46.24 is a score on that board. It is not comparable to anything TypeSafe has published, because the two numbers do not come from the same harness, and nobody has run an independent head-to-head between RSI-Jev and Jev 1.13. What can be said is structural rather than numerical: one is a hosted commercial model on a vendor's endpoint, and the other is a checkpoint you download and serve yourself.
Two further pieces of context belong with the set. The project is explicit that "ten of the fifteen benchmarks contribute training data in some form, so none of these numbers is zero-shot"; the held-out set is the comparison that is held out, and even it "is held out from training, not sealed from the search." And v6.0-VL ran 97 arms on its own line — 93 if the four data audits are left out — which is the scale of search that produced a 7.86-point jump on that index.
It speaks Jev's wire format, with four differences a caller should know
The compatibility surface is the reason this project exists in the shape it does. The same request shape ({state, model, questions}), the same three question types with the same criteria shapes, the same answer shapes, the same 1-to-64 questions per request, the same error envelopes, and the same confidence statistic — the peak, (K · p_max − 1) / (K − 1), clamped to 0..1. The project's own claim about that surface is that "anything written against Jev works against this without changes," and the server documents what is copied and what is not, which is more useful than the claim.
• The prompt and the readout are its own. RSI-Jev is a base model with a trained readout head, served with the encoder it was trained with, because using the reference's prompt "would take the model off its training distribution." The wire contract is the compatibility surface; the prompt is not.
• Option keys are visible to the model. The reference hides them, so renaming a key provably cannot change an answer there. Here it can, and the server reports this honestly as option_keys_visible_to_model: true.
• Criteria must be strings or null. A structured criterion — an object — is rejected with a 422, because no release was trained on one. This is the one place where a request the reference accepts will not run here.
• Nothing is cut. Serving takes up to 32,768 text tokens plus the image budget, and a longer request is refused with a 422 that says so rather than silently truncated. The 2,048-token figure that appears in the older cards is the length the models were trained on, not a serving cap, and describing a silent truncation to 2,048 as current behaviour is wrong.
Option counts differ by a wide margin in serving's favour: up to 5,120 options per question (RSIJEV_MAX_ANSWERS), against 160 in training and 64 admitted by the reference. Do not quote 160 as the serving cap. One thing is new in v6.0-VL's responses rather than inherited: every answer reports which layer answered, in usage.depth, alongside the calibrated confidence, so an adaptive decision can be audited after the fact. Images are an extension the reference does not have — one to four per request, as base64 data URLs, with the state referring to each by a literal marker.
Where a reader can actually run it, and where they cannot
RSI-Jev is a download. The project ships its own server, which speaks that Jev-compatible API, and the documented route is a pip install from the repository followed by its serve command with the checkpoint alias and an effort setting. The weights live on Hugging Face under the shgao organisation, released under Apache-2.0 following the base model, with one open question the project states itself: five of the image training sources are non-commercial or research-only, and "whether weights trained on non-commercial data inherit those terms is not settled." The code is MIT.
Hardware is not the constraint. The project develops on an HP ZGX Nano, an NVIDIA GB10 machine it credits HP and NVIDIA for, and the server runs on any CUDA GPU, on Apple Silicon, or on plain CPU — the last of which the docs price at 733 ms for a single question on the GB10's own Arm CPU, so usable rather than fast.
What it does not do is appear in a general model catalogue, and this is where we have to be precise about our own position. RSI-Jev is not on OrcaRouter and there is no model card for it to route to. What we serve is TypeSafe's commercial Jev, typesafe/jev-1.13, on the dedicated systemone endpoint, reached with a POST to /v1/systemone rather than the OpenAI chat-completions shape — the same request and answer shapes this project implements, from the model whose contract it copies. That is the whole of the relationship: the two speak the same protocol, we serve one of them, and you run the other yourself. If you already hold a key with us, the Jev 1.13 call shape is a first-class route on one API for 200+ models with 0% markup (provider list price passed through, so vendor price cuts are live here the same day) — which matters for the comparison in one specific way. A page like this one is cheap to act on if you can try the commercial contract first and only then decide whether running an open 4B checkpoint yourself is worth the operational work.

How to read RSI-Jev on 2026-10-07
The weaknesses the project publishes are as specific as its results, and the dates matter. External testing of v2.1 found that the model leans toward the more severe or expensive option on ordered choices and rubric scores, rarely chooses "unknown" when the document cannot answer, and answers a question and its negation inconsistently. Later releases aimed data at the "unknown" case — KoBBQ unknown-when-ambiguous went 0.891 at v4.0-VL, 0.932 at v5.0-VL, and 0.918 / 0.939 at v6.0-VL — and the v6.0-VL card is candid that the gains came from data rather than from depth, and that the 10% of questions which would go to layer 32 and stop at 16 or 20 are where the depth policy is still guessing. Reranking has further to go: hippo-memory's own retrieval order scores 0.484 R@1 and remains ahead of the 0.308 the model reached at v3.0, a figure the project has not claimed to close since. The RL evidence is one seed, and v3.0 "itself has no matched SFT control." The training corpora and policy development sets are not public, so the stages cannot be rerun from the repository alone, and the reranking corpus builder — about 96 GB of memory — was not re-run end to end.
Traction, read the same day: 73 stars, 5 forks, 0 open issues, 7 GitHub releases. Those move daily, and a repository three weeks old is not an established project, whatever its release cadence. The honest summary is that RSI-Jev is one of the more legible research efforts in this corner of the field — a chain of dated, measured, sometimes-losing releases with the search published alongside the scores — and one of the least independently verified, since almost every number on this page is its own and no outside party has benchmarked it against the commercial model it is compatible with.
What to watch is not the next release, because there will be one within days at this cadence; it is whether anything outside the project starts measuring. The two things that would change the picture are an independent benchmark run on the published checkpoints, and a decision-model comparison that puts both the open 4B and TypeSafe's hosted model through one harness. Neither exists today. Until one does, the useful way to read a score like 46.24 is as a well-documented claim from a project that registers its predictions in advance and ships the arms that failed — which is a stronger evidence trail than most, and still not a third-party result.
