A generated title card headed 'Recursive Self-Improvement, Explained' with the subtitle 'A closed loop is a scored loop, and the scoring is the whole question', above a row of three rounded cards reading 'Predictions registered before the run', 'Null floors measured, not assumed' and 'Failures ship, including the champion's'; a footer reads 'Worked example: RSI-Jev v6.0-VL, published 2026-10-06; the project's own record, read 2026-10-08.', with minimal flat line icons and the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

Recursive Self-Improvement, Explained Through a Project That Actually Does It

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

"Recursive self-improvement" is one of those phrases that mostly gets used to mean a feeling. Defined without mystique, it is narrower than that and more interesting: a system that proposes its own next experiment, runs it, and keeps or retires the result by a rule fixed in advance, with the results feeding the next round. The recursion is not magic and it is not unbounded — it is a loop with a scoring function, and its quality is entirely determined by how honestly that scoring function is applied. RSI-Jev, the third-party open project that builds Jev-style System One decision models, is a rare case where you can read the loop instead of arguing about it: every hypothesis, every failed arm and every release is published with its numbers, and the repository is dated day by day. The model this page uses as its worked example is RSI-Jev v6.0-VL, a 4B typed-decision model published on 2026-10-06, which is inside the last week. It is not TypeSafe's Jev and not affiliated with TypeSafe AI — the project's own licence line says exactly that, and what we serve at OrcaRouter is the other one, typesafe/jev-1.13, on our systemone endpoint.

One clarification before the word "self-improving" does any work, followed by one date. This is not a model rewriting itself without limit, and nothing on this page should be read that way. The loop in question rewrites a recipe: it proposes a training change, registers what it expects that change to do, spends GPU time, and then either keeps the change or records why it lost. The model is a Qwen3.5-4B-Base tower with trained decision heads — a fixed architecture trained by a fixed training script, with the search happening over data, objectives and stages. The loop runs its own experiments and retires its own champions. People decide what is worth measuring. And the date: v6.0-VL was published on 2026-10-06, and the project has published a further release since — the line moves roughly daily, and this page's figures are the ones attached to the release named here, dated where they were read. That is worth saying up front on a page whose subject is a loop that keeps running.

What the term means, stated flatly

Strip the phrase down and there are three parts, all of them ordinary. First, a search space: the set of things that could be changed. Second, a scorer: something that says whether a change helped. Third, a record: what was tried, what happened, and what was thrown away. A system is doing recursive self-improvement when the output of round N is the input to round N+1 in all three of those, without a human re-deriving the search space or re-deciding the threshold each time.

What most writing on the topic skips is the second part, and that is where the whole question lives. A loop with a weak scorer optimises the scorer. It will produce a monotonically rising line and a system that has learned the shape of its own exam. That failure does not require malice or a bug — it is what happens by default when the same number both selects and reports. So when you read a claim that some system is self-improving, the useful question is never "how much better did it get." It is "who set the bar, when, and could it move after the result was seen."

The popular versions of this concept — the ones ranking for the term today — are mostly forward-looking: Wikipedia's entry describes a system rewriting its own code toward an "intelligence explosion," and the wider coverage tends to be about whether that timeline is closer or further than forecast. That is a legitimate argument, and it is unanswerable with the evidence anyone currently has. What is answerable is the smaller question, which is whether a closed loop of the kind described above can be built and held to its own rules right now. For that, one project with public artifacts beats a decade of speculation, and this is what the rest of the page is about.

Closed, not merely iterative: five mechanisms

A loop is not closed because it repeats. Plenty of automated pipelines repeat without ever being closed, because a pipeline that can adjust its threshold after seeing the result is doing something categorically different from one that cannot. RSI-Jev's own rules are unusually explicit about the difference, and they can be read as definitions rather than as a manifesto. Five of them do most of the work.

Predictions are registered before the run. A hypothesis is written down with the number it expects to move and by how much, before GPU time is spent. A version that misses its own bar ships as a failure rather than being quietly re-cut. This is the mechanism that stops the loop from becoming a machine for writing post-hoc explanations, and it costs something real: the record contains entries whose only content is that someone was wrong on the record.

Null floors are measured, not assumed. The team runs arms that are provably identical to the control — verified by object identity before any GPU time — and the spread between those arms is the noise floor. A difference smaller than that floor is not a result, whatever it looks like. The project's own bar reflects this: on its internal suite the threshold is +0.006, with per-seed standard deviation on a single benchmark of 0.011–0.016, so single-benchmark differences below that are not findings. Most of the noise in this business is measured, not statistical.

A held-out set is spent once it is read. Each release reads the held-out set once, so two releases can still be compared on equal terms — and because reading it spends it, each release freezes the next one's comparison. The project's own phrasing is worth keeping intact: the held-out set "is held out from training, not sealed from the search." It is a check on memorisation, not a guarantee of novelty. And the suite figure itself has a hole in it that the project published: an internal task overlapped several hundred training rows, so from v6.0-VL onward the suite is reported without it — 0.770 — and v5.0-VL's earlier 0.764 was restated as 0.763 without it. That restatement is the point. A number was inflating, the cause was found, and the older release's card was corrected rather than left alone.

Failures ship, including the ones that killed the project's own champion. The repository's headline counts, read on 2026-10-08, are of a piece: eight releases in thirteen days, and hundreds of experiments written up with the failures included. A negative result is treated as the product. The contributing guide is blunt about why — the expensive part of a search is not running the winner, it is running the losers, so a well-measured negative from outside removes a branch and is worth more than a small positive.

The release chain is the project. One card per release, all of them, on the main branch permanently. Each card carries its own release's numbers, so one version's speed and calibration never overwrite another's. The project states the reason directly: that chain is the only way to see whether a self-improving loop is improving. Everything else — the recipe, the scripts, the pinned stack — carries only the current release, because the opposite rule would make it impossible to tell what is current. A published checkpoint carries its own code so that old releases stay runnable without old branches.

A screenshot of the subject project's own GitHub repository page, Shanghua-Gao / RSI-Jev, showing the repository description 'Typed-decision models (noul / choice / score) trained by a self-improving loop of AI agents - checkpoints, the code that produced them, and every version that failed', the sidebar counters 80 stars, 5 forks and 1 watching, 198 commits and 8 releases, the newest release headed 'RSI-Jev v6.1-VL 4B', an MIT license line, the topic tags ai-agents, autonomous-research, decision-model, jev, recursive-self-improvement, system-one and typed-decisions, and the merge commit 'Merge pull request #33 from Shanghua-Gao/release-v6.1-vl' at the top of the commit list.

What the discipline costs, and what it buys

Two stories from the record are the honest counterweight to the word "self-improving," because both are cases where the loop's own rules made the work slower and better at the same time.

The first is the bug behind v1.0. Finding it took seven registered negatives. Each of the seven was an optimiser-side fix that reduced a training instability without removing it, because the cause was a precision mistake nowhere near the optimiser — and the eventual fix was one line. Not one of the seven was worth publishing on its own. Together they are what made the cause findable at all, and that is the whole argument for registering and keeping negatives: their value is joint, not individual.

The second is less flattering and the project publishes it anyway, in a footnote. An early arm failed exactly one guard — MMLU-Pro came in 0.026 below the bar against a 0.020 limit — and was logged as rejected. The run owner then widened the limit to 0.030, on the grounds that MMLU-Pro is a guard against forgetting rather than a target, and the arm was confirmed on four fresh seeds and became the next release. The original rejection was left in the log, with the project's own comment that a bar moved after seeing the result is the kind of thing a reader should be able to catch it doing.

That footnote is the most useful paragraph in the repository for anyone trying to assess this kind of claim in the wild. A bar that moved is not proof of bad faith — the reasoning given is defensible, and the widened bar then had to survive four fresh seeds. But a bar that moved silently is a loop that is no longer closed, and the difference between the two is entirely whether it was written down. The general rule the project settled on afterwards is the one worth carrying to other systems: a near-miss that fails exactly one guard gets a diagnosis and a targeted repair instead of being discarded — and if the repair fails, it becomes a dead end with a record.

How often the loop is wrong: fourteen setups, two kept

Here is the number to hold on to. Through the release that reads images, the project counts fourteen reinforcement-learning setups and 63 reward-trained arms. Two were kept.

Each setup was judged against a supervised control trained on the same items for the same number of steps, which is the comparison that makes the count meaningful — an RL arm beating a baseline it never had to match proves nothing. The instructive losses, in the project's own words:

• Binary correctness RL — the probability output collapsed to 0 and 1. Rewarding the act of picking right rather than the honesty of the reported odds pushes a distribution to the corners, and a decision model whose confidence is always total is useless for the thing decision models are for.

• Proper-score RL, the reconstruction of the Laya-style objective — fine at 300 steps, diverged at 1,500. Stable while it was doing nothing much, and unstable exactly when it started to matter.

• RLCR — tied on accuracy, worse raw calibration. It bought no capability and paid for it in the one property the model exists to provide.

• Bandit RLCD, where only the chosen option's outcome is revealed — no better than supervised training on the same feedback. The project killed its own pre-registered test here: the arm had to beat supervised training on at least two of four calibration metrics and won one.

The winner, and the reason it won, is the most transferable finding on the page. It is a listwise reward — ranking quality measured over the order of many separately scored candidates, rather than treating each one in isolation. Reranking recall at the first position moved from 0.192 for the supervised parent to 0.308, with wins and losses on a paired test. The project's explanation is not a tuning story: a per-item training target scores each candidate on its own, and no single label encodes the quality of an order across candidates. Where the reward said something the labels cannot express, RL beat supervised training on the same rows. Where it did not, supervised training matched it.

That generalises into a rule about when this kind of loop can find anything at all. With a gold label in hand, the expected policy-gradient update equals the gradient of a supervised loss — the project's own phrasing — so an "RL arm" scored against labelled decisions is really a loss-design arm. And loss-level changes moved the internal suite by at most 0.002, while new data moved it by 0.13. Read together: the loop's leverage was never in the objective. It was in what got measured and what got fed in.

Which is the honest answer to a question a reader is right to ask. RL setups are cited elsewhere as evidence about self-improvement in general; here they are evidence about one loop on decision tasks. The listwise result is one seed, and the project says so: the paired RL-versus-supervised comparison was run on a different parent from the one the release applied it to, and that release "has no matched supervised control." A finding with that caveat attached is worth more than one without it, and it is why the page you are reading does not generalise it.

Where the human sits

The project is explicit, and the explicitness is the interesting part: the loop runs its own experiments and retires its own champions, but it does not decide what is worth measuring, and it does not notice on its own when a number is technically true and practically misleading. People do that.

Two of the project's own rules exist because someone pushed back. The MMLU-Pro guard was widened rather than letting a near-miss be discarded. The seed requirement now scales with the effect size instead of spending four seeds on every difference — because the confirmation seed, the one number an arm was not selected on, is the load-bearing one, and spending compute on differences you can already see buys nothing. One of the releases in the chain started as a refusal to drop an arm that had failed one guard. The project credits the people who sent that feedback by name, in the acknowledgements.

So "self-improving" here describes the middle of the loop, not the whole of it. The loop is a search that runs without a human turning the crank. The human is still at the top of the loop choosing the objective and at the bottom reading the result critically — which is precisely the arrangement that keeps the loop from becoming a machine for confirming its own preferences. Any account of recursive self-improvement that omits that seat is describing a different system than this one, and probably a hypothetical one.

What the project's own numbers do not show

The discipline above is only worth describing if its limits are stated in the same breath, and RSI-Jev's record states them for itself.

• It is one project, one model size at a time, one seed. The published checkpoint is a single seed's, fixed in advance as primary rather than chosen for scoring best. The RL evidence is one seed.

• It is decision tasks, not generation: a yes/no, pick-one-of-k or rate-on-a-rubric question over a document, a chat or an image, answered in one forward pass with a calibrated probability per option. Nothing is generated, so there are no reasoning tokens to spend. What the loop demonstrates here is that it can improve a scorer of that shape.

• Some of it is not reproducible from the repository alone. The training corpora and the policy development sets are not public, and the project says so plainly in the release card: "the stages cannot be rerun from this repository alone." One corpus builder needs about 96 GB of memory and was not re-run end to end.

• Not everything is checkable at all. The internal suite was never fully held out. Of the fifteen benchmarks, ten contribute training data in some form, so none of those numbers is zero-shot — the held-out set is the comparison that is held out, and it is the only one.

And the external parts have their own scar tissue. An audit after one release found roughly a thousand items of the public benchmark kit's test rows in the training corpora — about 0.3% of the kit's rows, with the largest single contribution a few hundred rows from one source. Re-scored without them, the index moves by at most 0.04. That correction was published in the next release's card rather than quietly applied, under the rule the project states as "no contamination, checked rather than asserted." It is a rule costing them a number, which is the only kind of rule worth having.

Where you can actually run it — and where you cannot

Nothing here is served by OrcaRouter. Our catalogue carries no RSI-Jev id and no model card for it, and there will not be one until somebody serves it. RSI-Jev is installed from its own repository and runs its own server, which speaks the Jev API: you point an existing Jev client at it and change the base URL. It runs on Linux, Windows and macOS, on a CUDA GPU, Apple Silicon or plain CPU.

The one model we do host is the other side of that contract. TypeSafe's Jev 1.13, which we serve as typesafe/jev-1.13, is the commercial decision model whose request and answer shapes RSI-Jev reproduces, and it is reached on our dedicated systemone endpoint rather than through the chat-completions shape. That is the whole relationship, and it is a structural one rather than a quality claim: one is a hosted commercial model on a vendor's endpoint, the other is a checkpoint you download and serve yourself. Nobody has run an independent head-to-head between them, and this page is not going to declare a winner from two different harnesses.

A generated single-column scoreboard headed 'RSI-Jev - the loop's own ledger', with six rows reading 'Predictions: registered before the run', 'Null floors: measured from identical arms', 'Held-out set: read once, then spent', 'Failures: published, including the champion's', 'Release chain: one card per release, on main forever' and 'RL setups kept: 2 of 14, each judged against a matched control'; a footer reads 'All figures from the project's own record, read 2026-10-08.'

If what you want is to put a decision model like this behind an application, the routing question is separate from the model question, and it is the part that decides whether the experiment is worth running at all. OrcaRouter is one API for 200+ models with 0% markup — provider list price passed through, so a vendor's price change is your price the same day — plus automatic failover and a routing DSL for composing several models into one call. For a loop model you cannot yet call through a general API, that matters in a specific way: it means the alternative candidates you would benchmark it against are already behind one key, and the comparison is a config change rather than a second integration.

A screenshot of OrcaRouter's own model page for typesafe/jev-1.13 showing the left navigation, a PERFORMANCE panel with prefill and decode lines, the heading 'Jev 1.13' with the provider line 'typesafe', the price fields $0.11 prefill and $0.36 decode per million tokens, a benchmark box with MED 0.3, Response Trust 0.784, Structured Output 0.964, Refusal Correctness 1.0 and Consistency 0.59, an accuracy-versus-cost scatter with a 'Jev 1.13' marker, an 'Individual Runs' table, API and Agent curl snippets pointing at the systemone endpoint, the OrcaRouter logo and a sign-up button.

The part worth keeping

The reason to write about recursive self-improvement through a project like this, rather than through the argument about whether it leads somewhere dramatic, is that the loop's rules are the transferable part and the speculation is not. Registered predictions, measured null floors, spent held-out sets, published failures, an immutable release chain: none of those is a property of a superintelligence. They are properties of a lab that decided to be checkable, and they are available to any team running automated search today, at any scale, on any model.

Read that way, the interesting claim in RSI-Jev's record is not the score. It is that the loop's own ledger says it was wrong far more often than it was right — two kept recipes out of fourteen setups, seven negatives to find a one-line bug, a bar that moved and got written down — and that the project published the ledger. A rise from 38.38 to 46.24 on a public index in one release is a curiosity without the failures next to it. The failures are what make the number mean something.

Which is the honest position on the term itself. The question is not whether a system can improve itself. It is whether the improvement is measured by something that could have said no. Where that separation is real and written down, the loop is worth reading. Where it is not, a rising line is a description of a scorer, not of a system getting better — and no amount of recursion fixes that.