
RLCD Explained: Why TypeSafe Trains Jev to Be Honest About Confidence Instead of Liked
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 349 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 208 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 680 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 49 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 105 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Jev 1.13 (typesafe/jev-1.13) is trained with a method its maker calls Reinforcement Learning for Calibrated Decisions — RLCD — and that acronym is TypeSafe's own coinage rather than an industry term you are supposed to already know. The launch post says it in as many words: the company built "a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD)". It is a third answer to a question that used to have two, and the reason it exists is a mismatch most teams run into the first time they try to put a language model inside a decision. Before getting to it, two dates matter, because this page is not a launch piece. TypeSafe shipped the model itself on 2026-09-15, which sits outside the seven-day window this blog writes to, and nothing here should be read as framing Jev as new. The dated event is 2026-09-24, when OrcaRouter added typesafe/jev-1.13 to its own catalogue and opened the model card for it — the first time Jev has been callable through a third-party gateway rather than only through TypeSafe's own endpoint. That is the change this page runs on, and the practical consequence is that the technique below is now something you can try in code on a key you may already hold, rather than a research idea you read about.
What follows is the concept, not the model. The first third of this page is about the two training methods RLCD was designed against, because RLCD is only legible as a repair for what those two do when the task stops being a conversation and becomes a judgment. If you already know what RLHF and RLVR optimise for, the section you want is the third one, where TypeSafe's own three-way table does the work.
RLHF optimises for the answer a person prefers
Reinforcement learning from human feedback is the method that turned pretrained language models into assistants. TypeSafe's own primer states the objective flatly, in a card headed RLHF: it "turned pretrained models into chatbots. It trains models to produce responses people prefer." InstructGPT and ChatGPT were trained with it, and the primer adds a detail that is relevant here for a different reason — the approach was co-invented by Diogo Almeida, who is a cofounder of TypeSafe and the author of Jev's launch post. The company is not dismissing the method it was founded by someone who helped build. It is arguing that the objective is wrong for a particular job.
The clean way to see the mismatch is to ask what the reward signal actually measures. Under RLHF it measures a rater's preference between two candidate responses. That is an excellent proxy when the product is a conversation, because a conversation's success criterion really is whether a person finds the reply good. It is a broken proxy when the product is a decision, because the success criterion there is whether the stated confidence matches reality, and a rater comparing two plausible paragraphs has no way to see the difference between a well-calibrated 0.6 and a confident-sounding 0.95. Two answers can be equally preferred and differ enormously in how much a piece of software should trust them.
The failure modes TypeSafe names in the primer follow directly from that:
• Sycophancy — the model learns to produce what the rater wants to hear, which is a different target from what is true.
• Confident-sounding hallucination — fluency and certainty are rewarded by preference even when they are not backed by anything.
• Mode dropping — preference optimisation narrows the output distribution, "favoring a particular style, such as instruction following, while reducing the probability of other possible outputs." Mode dropping is the mild version of the classic mode-collapse failure that plagues generative adversarial networks, where a generator converges on one output that keeps fooling the discriminator.
The primer's own warning paragraph is the sentence worth keeping: "An output can be compelling to a person without being reliable enough for unattended automation. Human preference and machine trustworthiness are different optimization targets." That is not a criticism of RLHF as a method. It is the observation that a preference-trained model has never been asked the question automation needs answered — how often, exactly, is this thing right when it says it is sure.
RLVR optimises for outputs a program can check — and decisions rarely have one
Reinforcement learning with verifiable rewards is the second adaptation, and it is the one behind the reasoning models. TypeSafe's primer describes what it produced: models that "are strong at tasks such as mathematics, but slower and more expensive." The mechanism is a checker. If a task has an answer a program can test — a unit test, a proof checker, a numeric answer — then a reward can be computed without asking a human anything, and the model can be trained against that signal at scale. It works, and it is why reasoning models got good at exactly the domains where cheap automatic verification exists.
The limitation is the shape of that word "verifiable". A verifiable reward requires a verifier, and a verifier requires that the task has a right answer somebody can compute. Consider the questions a production system actually asks. Should this support ticket go to billing or to technical? Is this refund request inside policy? Does this transaction look like fraud? Each has a defensible answer most of the time, none has an answer a program can check, and the cases that matter most are precisely the ones where experienced humans disagree. There is no function to run. RLVR has nothing to reward, so it contributes nothing.
The tempting workaround is to manufacture a verifier by labelling a dataset and training against the labels. That gives the method something to chew on, but it changes the objective in a way that matters. Labels encode a decision, not the uncertainty around it. A model trained to reproduce one team's judgements on the hard cases learns to be as confident as those labels were — which is to say, exactly as overconfident as the humans who wrote them. And even where a genuine verifier does exist, there is a second gap. A verifier scores the answer. It does not score the stated confidence. A model that is right on 95% of cases and reports certainty on all of them takes a perfect reward and is, as a component in an automated pipeline, useless — because the 5% is the only part the pipeline needed to be told about. TypeSafe's launch materials make the same point from the other direction: "If a model can do a task 95% of the time but doesn't say when it's in the 5%, it can't automate that task."
What RLCD does, in TypeSafe's own framing
RLCD changes the output contract rather than the answer quality. The primer's card reads: "Reinforcement learning for calibrated decisions trains TypeSafe to return decisions and calibrated probabilities instead of generated text." The launch post's terse version is "calibrated decisions: answers with epistemically honest probabilities on System One tasks." Both are describing one move: train the model against whether its stated probability matched the frequency with which that answer turned out to be right, rather than against whether a person or a checker liked the answer.
The launch post sets the three methods side by side, and the contrast is the clearest statement of the idea that exists. Read as a set of contrasts rather than a table:
• What it optimises for — RLHF optimises human preference, "writeups and chat responses that human raters prefer"; RLVR optimises "outputs that can be programmatically verified"; RLCD optimises calibration, "answers with epistemically honest probabilities on System One tasks."
• What goes in — the older two take unstructured data "with an emphasis on sequential messages"; a calibrated decision model takes unstructured data "with an emphasis on structured program state."
• What comes out — generated strings that "need to be parsed + validated," with "always some risk that the AI goes off the rails," against type-safe structured values where "possible outputs and structure are defined in advance," the model "never makes type errors," and "all answers are accompanied with calibrated probabilities and confidence scores."
• How it is sampled — one token at a time, each conditioned on the last, against all outputs generated in a single query. This is the mechanical reason the third method is cheap: there is no decoding loop to pay for.
• What it costs — input tokens from $0.20 to $10 per million for the comparison models with output roughly five times the input price, against $0.042 per million input tokens with output billed at zero for Jev.
• How fast it answers — 3 to 329 seconds end to end for frontier models against 70ms to 500ms, which the vendor characterises as 40x to 200x faster on System One shaped queries.
• What it says about its own confidence — the older two "tend to be overconfident and inconsistent" even when prompted for a confidence estimate; RLCD "always communicates confidence and uncertainty with every output," where "higher confidence means higher accuracy."
The last line is the actual product claim, and it is falsifiable in a way the others are not. "Higher confidence means higher accuracy" is a statement about a curve: bucket a model's answers by the probability it attached, and the buckets should be right at roughly the rate the probabilities claim. TypeSafe's confidence documentation spells out the contract with unusually concrete numbers:
• Outcomes assigned a probability of 0.2 should occur about 20% of the time.
• Outcomes assigned a probability of 0.8 should occur about 80% of the time.
• Outcomes assigned a probability of 1.0 should occur 100% of the time.
And then the sentence that keeps the claim honest, in the vendor's own words: "These rates describe groups of predictions, not a guarantee about any single answer." That is not a hedge bolted on for legal reasons. It is the whole meaning of calibration. A well-calibrated model that says 0.8 is not promising to be right this time; it is promising that across every answer it labelled 0.8, about four in five were correct. One answer tells you nothing. A thousand answers across a week tell you whether the curve is real.

The same three-card contrast sits on TypeSafe's own documentation, which is the source for the comparison above and the clearest place to check the wording rather than take a summary's word for it. The capture below is that page as it stands today: three cards for the three post-training approaches, with the third one naming RLCD in full.

Two further details in the vendor's documentation show how far the method reaches into the product. The first is that confidence is derived rather than generated: the model returns a full probability distribution across the options or levels you supplied, and the confidence value is a statistic computed from the shape of that distribution. That is why the docs can tell you the definition is not load-bearing — you get the raw distribution either way and can compute your own statistic if yours fits better. The second is that RLCD is the only thing that shapes the weights. TypeSafe's models page states: "Jev is not fine-tuned or LoRA-adapted with customer data. It is trained with RLCD to return calibrated decisions, and the same weights serve every account." Domain adaptation happens in the request — your state, your criteria — not in a per-customer checkpoint. Whatever calibration the method produced is the calibration every customer gets.
Why calibration is the thing that makes a cheap decision model usable
A calibrated probability is not interesting on its own. It becomes the architecture at the moment your code branches on it, and TypeSafe's confidence documentation describes exactly that pattern as three ranges, each producing a different system behaviour.
• High confidence — act automatically. The model has a clear read and you can proceed without human involvement.
• Medium confidence — proceed with caution. The model has a reasonable answer but is not certain, so you confirm with the user, flag for review, or gather more information before acting.
• Low confidence — do not act. Route to a human, request clarification, or fall back to a different system, because the model is telling you it does not have enough to go on.
The docs are explicit that the boundaries are yours to draw and should differ by consequence: "A confidence threshold is not one number. Different actions within the same system should be gated at different levels depending on the consequences of getting it wrong." Their worked example puts a hard floor at 0.5 — anything the model reports below that is routed to a human without further inspection — and then applies a higher bar for a destructive action than for a read-only one. Your code encodes the risk tolerance; the model supplies the honest input to it.
That pattern is the entire argument for a two-model workflow, and it is worth stating as an argument rather than a feature list. Suppose you want an automated pipeline that handles the confident majority of cases and escalates the rest to a bigger model or a person. The escalation decision has to come from somewhere. If the cheap model reports 0.98 on everything, including the cases it is guessing at, then the branch has nothing to test and you either automate everything — including the calls it should have escalated — or you automate nothing. A model whose confidence is informative is the only kind that lets you automate a subset safely, because it is the only kind that can tell you which subset it is unsafe on. The docs put the same point in one line worth quoting for its bluntness: "If an intelligent system, whether human or machine, cannot express honest uncertainty, the system cannot be trusted."
There is a second reason this matters more for a cheap model than an expensive one, and it is the reason the routing story and the RLCD story are the same story. A model priced at $0.042 per million input tokens with no output charge is cheap enough to consult constantly — on every turn of an agent loop, on every record in a batch, on every ticket as it arrives. Being consulted constantly is exactly the situation in which a model's mistakes compound, because nobody is reading its output before it is acted on. Confidence is what makes that safe. The cheapness is what makes the escalation branch affordable, since the expensive path only runs on the fraction of cases the cheap model declined. Neither half works without the other, and the routing decision that joins them is a threshold on a number RLCD is the reason to believe.
The honest limit: calibrated is not correct
The most important thing to get right about RLCD is what it does not claim. Calibration is a property of the confidences, not a guarantee about the answers, and the vendor says so in its own documentation rather than leaving it to critics. The System One page: "System One models are trained for calibrated decisions: their probabilities are optimized against outcomes to reflect uncertainty. Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." A model can be perfectly calibrated and still make the wrong call on your ticket, because 0.9 means nine in ten, and this could be the tenth.
Our own serving figures are the useful counterweight here, precisely because they are measurements of the model in production rather than claims about what the method achieves. Over the seven days ending 2026-09-30, on traffic through OrcaRouter's playground since the model was added to the catalogue, the Jev 1.13 card reports a 0.49% error rate across 76.2 million tokens, alongside a p50 time to first token of 151 ms, a p95 of 247 ms, and about 349 output tokens per second. Two things about that number deserve to be said plainly. It is ours, not the vendor's, and it is a rolling window rather than a fixed test set — the same field read 0.57% earlier in the window, because it is recomputed over the trailing seven days of live traffic and yesterday's calls age out. It is also not a calibration measurement. An error rate tells you how often something went wrong on our traffic; it does not tell you whether the confidence values were honest, which is a different question and one that needs labelled data to answer.
Which is the practical instruction the vendor gives as well, in a note attached to its threshold guidance: "The correct threshold values depend on your domain and the performance of the model for your use case. Start with conservative thresholds, test with your own data, and adjust as you observe results." RLCD is a claim about how the model was trained. Whether the claim holds on your inputs is an empirical question, and it is one of the few model properties you can test without any machine-learning infrastructure — take a few hundred cases you already have labels for, bucket the answers by the confidence the model reported, and check whether the buckets are right at the rate they claim. If the 0.9 bucket is right about 90% of the time on your traffic, the threshold is real and you can automate above it. If everything clusters above 0.9 and the accuracy does not follow, you have learned something more useful than any headline number.
Two further limits belong in the same breath. The first is that there is no public benchmark card for this model to check any of it against — the vendor has not published one, and no third-party leaderboard carries the model; the Artificial Analysis model page for it returns a 404 as of 2026-09-30. So the calibration argument rests on the training description, the documented contract, and whatever you measure yourself, not on a published curve. The second is that the vendor's own performance claims are its own: the launch post openly notes that the workflow evaluations behind the speed and cost headline were built by its model capabilities team, that the reference answers they are measured against are the average of two external models, and that the numbers are "on the higher end of real world gains." It also says the pricing cannot be proven unsubsidised. None of that undermines the training method, which is a separate claim from the speed claim, but it does mean the case for RLCD is an argument about objective design rather than a settled empirical result. Treat it as a hypothesis you can test cheaply, which is a better position than most training-method claims leave you in.
What you can do with this today
The two terms of the argument meet in one place. RLCD is the reason a decision model's confidence is worth branching on; a threshold in your code is where that branch lives; and escalation is only affordable if the common path is cheap enough to run everywhere. Jev 1.13 is callable as typesafe/jev-1.13 on OrcaRouter — one API for 200+ models, 0% markup, provider list price passed through, so a vendor price cut is live here the same day — which means the confident-majority path and the generative escalation path bill on the same key instead of two vendor contracts. You still call it in its own shape, POST /v1/systemone, non-streaming, against a 65,536-token context, because that is not the OpenAI chat-completions route and it is not folded into the chat endpoint. Two dated notes from the vendor's SDK releases are worth knowing if you are wiring it up: version 0.7.1, released 2026-09-21, added examples for usage with AI gateways, and version 0.7.2, released 2026-09-26, added an http2 extra to the Python package. The second is the kind of detail that only shows up in release notes — an HTTP/2 client is worth having for a model whose whole value proposition is sub-200-millisecond round trips.
If you take one thing from the page, take the shape of the question RLCD answers. It is not "can a model be smarter." It is "can a model tell me when it is not smart enough, often enough and accurately enough that I can automate the rest." That is a different research target from the two the field spent the last few years on, and it is the only one that produces a number your code can act on. The confidence value is that number. Test it on your own labels before you trust it, and start with a threshold you would be embarrassed to be wrong about rather than one you would like to be right about.
One last piece of the picture is worth carrying alongside all of it, because it is the number the whole argument is aimed at and it is measured rather than claimed. The card below is our own seven-day serving record for typesafe/jev-1.13 — the model on the wire, not the training method, and not a benchmark. Read it as the second half of the calibration question: the confidences tell you which calls to act on, and this tells you how close the rest of the routing decision is to a system you would leave unattended.

