A generated title card headed 'System One, Explained' with the eyebrow 'TYPESAFE SYSTEM ONE' and the subtitle 'A model category that returns typed decisions instead of sentences'. Three cards on the right read 'Two questions, two models', 'No prose in, no prose out' and 'A value a program can branch on'. A footer line reads 'Jev 1.13 launched 2026-09-15; callable as typesafe/jev-1.13'. The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

"System One" as a Model Category: Where Jev 1.13 Sits in It

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

"System One" is the category term TypeSafe uses for its split between a model that decides and a model that writes, and Jev 1.13 (typesafe/jev-1.13) is its first member — a model that returns typed answers instead of sentences. It is not a new model. TypeSafe shipped Jev on 2026-09-15, and this page is not a launch piece: the model is fifteen days old and outside the seven-day window this blog writes to. What happened inside the window is that OrcaRouter added the model to its catalogue on 2026-09-24 and opened the model card for Jev 1.13 at https://www.orcarouter.ai/models/typesafe/jev-1.13 — the first time it has been callable through a third-party gateway rather than only through TypeSafe's own endpoint. The category idea is the reason this page exists; the serving change is the reason it is dated today.

The plain version of the category: an LLM is asked a question and writes an answer for a person to read. A System One model is asked a question and returns a value for a program to branch on. TypeSafe's own line is that "LLMs produce words for people" while "Jev produces typed decisions and is more like code: reliable, fast, self-consistent, and type-safe." That sentence is the whole category compressed into one clause, and it is worth unpacking slowly, because the four adjectives are doing different amounts of work and one of them is doing more than the others.

What "more like code" actually claims

Take the four claims in order, because they are not four restatements of "it's better."

• Reliable — the output shape is fixed in advance. You declare the question; the answer can only come back as one of the values you allowed. TypeSafe states plainly that "the model never makes type errors," and notes this is the one claim of theirs that is "mathematically impossible" to falsify with a counter-example, because a value that is not in your declared set is not a value the model can emit.

• Fast — all answers are produced in a single pass rather than one token after another. TypeSafe's launch post puts it as "Jev outputs all probabilities in parallel instead of autoregressively generating by token." On our own seven-day serving window ending 2026-09-30, the median time to first token on typesafe/jev-1.13 is 151 ms and the p95 is 247 ms.

• Self-consistent — the same state with the same questions tends to produce the same answers. A coding analogy is what makes this legible, but it is also where the analogy stops being a proof: a compiler's determinism is a property of its construction, whereas this is a claim about a behaviour. Our own measurements are the honest read on it — the error rate on our playground traffic across the same seven-day window is 0.49%, so it is self-consistent the way a good function is self-consistent, not the way arithmetic is.

• Type-safe — and this is the one carrying the most weight. Type-safe is not a quality adjective here; it is a statement about where the model sits relative to a type checker. In an ordinary generative pipeline the type system starts after the model finishes: the model writes text, a parser guesses at the shape, a validator checks it, and a failure path handles the cases where the guess was wrong. A System One model moves the type declaration to before the call. The three primitives our card documents are the type system: noul, a true/false judgment returned with a calibrated probability; choice, one label picked from up to 255 labelled options; and score, a rating on an ordered scale of 2 to 10 levels. You pick the primitive, you supply the labels or the criteria, and the value that comes back is drawn from that set.

TypeSafe does publish one difference between its own docs and ours worth stating rather than resolving: the vendor's documentation shows a zero-indexed Score example, while our card documents the scale as 2 to 10 levels. Both are describing the same primitive. If you are building a threshold, read the vendor's page for the exact indexing your SDK is on.

The two failure modes that stop existing

The interesting consequence of "no prose" is not aesthetic. It is that the two failures that dominate production generative pipelines are absent from this design rather than mitigated by it.

Format drift is the first. An LLM told to return JSON returns JSON most of the time, and something adjacent to JSON the rest of the time — a trailing comment, a markdown fence, a field renamed to a synonym, a nested object where the schema wanted a string. The prompt-level fixes (stronger instructions, few-shot examples, a schema in the system message) are all attempts to hold a shape that the model is free to abandon, because the shape is a request, not a constraint. TypeSafe's framing makes the contrast explicit: with strings, "possible outputs and structure" are requested and responses "need to be parsed + validated," with "always some risk that the AI goes off the rails." When the possible outputs are declared in advance, the drift has nowhere to go.

Unparseable output is the second, and it is really the same failure at a worse moment — not a field that came back slightly wrong, but a response the parser cannot read at all, arriving at the least convenient point in a workflow. A model that emits a typed value has no such state.

This is a structural argument, and it should be stated as one. It says nothing about whether an individual answer is correct — a choice question can pick the wrong label, and a noul can return true with high confidence when the honest answer is false. What disappears is the category of failure that a parser would have caught. That is a real and useful reduction, and it is not the same claim as "the answers are right."

Why the price is a shape, not a discount

The model is priced at $0.042 per million input tokens, with output billed at zero — and that zero is not a promotional rate, it is an artefact of the design. A model that emits three tokens of structured answer has no output volume to meter, so per-output-token pricing has nothing to attach to. The billing shape is a per-input-token charge and a decision. Our catalogue passes the provider's list price through at 0% markup, so the $0.042 is TypeSafe's number rather than a number we set, and a vendor price change would be live the same day.

Put the two shapes side by side and the difference is not a percentage. A generative pipeline's cost scales with how much the model says: a verbose answer costs more than a terse one for the same decision, and a chain-of-thought reasoning model bills for the tokens it spends thinking before it answers, whether or not the answer improves. A System One call's cost scales with how much you show it — the state and the questions. Ask one question against a long document and you pay for the document. Pack forty questions against the same state (the input budget on our card is 65,536 tokens across the combined state and questions, roughly 64K; if you have seen a "roughly 32,000 tokens" figure in earlier OrcaRouter articles, that is the state budget alone, not a competing total) and you pay for the document once and get forty decisions back.

That is why cost-per-decision, not cost-per-token, is the right unit for this class — and why the meter runs the opposite way from what most teams expect. The typical generative cost-reduction move is "make the model say less." Here there is nothing to say less of.

A headless-browser capture of the OrcaRouter model card for TypeSafe: Jev 1.13 at orcarouter.ai/models/typesafe/jev-1.13. The header reads 'Jev 1.13' with the badge '65K tokens', the slug typesafe/jev-1.13, the byline 'by TypeSafe - 2026-09-24', and the summary 'TypeSafe's structured decision and evaluation model... Served via POST /v1/systemone; non-streaming; up to ~64K input tokens; text in, structured JSON out.' A stats row reads '$0.04  151 ms  247 ms  76.2M', above a code sample pointing at https://api.orcarouter.ai/v1/systemone with "model": "typesafe/jev-1.13", and the buttons 'Get the Jev 1.13 API', 'Try in playground' and 'Use via API'.

TypeSafe's own numbers, which are vendor-reported and have not been independently replicated, are aimed squarely at that comparison: "193.6x Faster, 444.6x Cheaper," footnoted as "based on workflows for System One tasks (proof)," with a worked example reading "TypeSafe AI Cost $0.000081 Completed in 0.114s / LLMs Cost $0.013880 Completed in 8.566s." The homepage also lists "$42 Per Billion input tokens" against "238x Lower input price than Claude Fable 5.1." Treat all of it as the vendor's argument, not as a measured result: the launch post concedes that "our published evals are generally run from our laptops on the West Coast" and that "we can't prove it isn't subsidized; we'll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up)." Those two concessions are the vendor's own, and they are the right frame for every multiplier on the page.

Calibration is the second half of the idea

If the category were only "structured output," it would describe function calling with extra steps. The part that makes it its own thing is that every answer arrives with a probability, and the probabilities are the training target. TypeSafe calls the method Reinforcement Learning for Calibrated Decisions (RLCD) — their term, not a generic acronym — and the comparison table in the launch post sets it beside RLHF and RLVR: RLHF optimises for what human raters prefer, RLVR for outputs that can be checked programmatically, and RLCD for "answers with epistemically honest probabilities on System One tasks."

The practical difference is what the probability is for. In a generative pipeline the confidence estimate is a second generation: you ask the model how sure it is and it writes a number, which is itself prose with the same failure modes. Here the probability comes back with the decision, in the same pass, and it is the thing you branch on. TypeSafe's own framing of the payoff is that a model which can do a task 95% of the time but "doesn't say when it's in the 5%" cannot be used to automate the task; the confidence gives you a place to put the escalation, to a person or to a reasoning model.

TypeSafe's homepage states this as "Zero Hallucinations," explaining that every decision carries a confidence estimate so software can "act when confidence is high and escalate when it is not." Read that carefully: it is a claim about confidence estimates, not a claim that no answer is ever wrong. Our own card is the counterweight — a 0.49% error rate over the seven days ending 2026-09-30, on our traffic, measured by us. That figure is a rolling window, not a fixed test set: it was reading 0.57% a few days earlier on the same window, and it will move again.

Where System One sits next to System Two

The fast/slow vocabulary long predates TypeSafe. It comes from Kahneman's Thinking, Fast and Slow, and it has been borrowed by AI researchers for years before this — the "System 2" label was attached to chain-of-thought and deliberate reasoning models well before TypeSafe existed, and TypeSafe does not claim to have coined either term. What they have done is apply the distinction to a product boundary rather than to a prompting mode.

• A System Two reasoning model spends more compute before answering, and gets better at problems that need it. Its output is still prose, and the extra compute is billed as output tokens.

• A System One model in TypeSafe's sense does not think longer to answer better. It answers in one pass, and what it gives up for the speed is the ability to produce anything but a typed value.

• The two are complements in a workflow, not rivals in a comparison. A System One call handles the decisions that have to be fast, cheap and legible; the reasoning model gets the cases the confidence score flagged as uncertain. The typed output is what makes the handoff clean — you are passing a value and a probability to the next stage, not a sentence to re-parse.

Where the vocabulary gets slippery is in treating "System One model" as an established category that other vendors have adopted. There is no evidence of that, and this page should not be read as claiming it. TypeSafe uses the term for its own class of models; the disclaimer in our own card says the same thing by omission, listing a single endpoint type for a single model. If another lab starts using the phrase for the same architecture, that will be a fact worth reporting and it will need their own words to report it with.

A headless-browser capture of the TypeSafe blog post announcing System One models. The masthead reads 'TypeSafe AI | Manifesto | Our Team | Docs | Contact Sales', under the section heading 'Company News' with the date 'Sep 15, 2026' and the byline 'Diogo Almeida, founder, TypeSafe'. The opening paragraph asks 'Models have been superhuman at chat for years, so where is all the automation?', followed by 'After two years in stealth... I am beyond excited to announce that today, TypeSafe AI is releasing our first System One Model: a new class of frontier models built to make fast, structured decisions that software can use directly.' A later paragraph reads 'Our first public model is Jev, available today in early access.'

Our card also lists jaggedness as part of the honest boundary rather than as a surprise: nine named failure modes. The literal-reading and indirection entries are the ones that follow directly from the "more like code" analogy — a model that answers the question you wrote rather than the one you meant is behaving like a function that did exactly what the code said. The counting entry does not. A model that "recognizes the shape of an answer rather than tallying" is not code-like at all, which is why TypeSafe's own recommendation is to count in code and, where a judgment is genuinely needed, to ask one question per item and add up the answers yourself.

Two limits that shape the design, not the score

Both come from the same place: no strings means nothing to stream and nothing to send in pieces.

• Non-streaming — the first output is the finished answer, so a System One call is a single response, not a stream. The question is not whether it can stream but what would stream.

• One request shape — the model is served through POST /v1/systemone on our catalogue rather than the chat-completions shape, and that is the honest version of an older claim that it "speaks its own request shape." It is a real difference in how you call it: a state object and a named-questions map go in; a structured answer per question comes out. You will write a mapper for it, and because the output is typed, the mapper is the entire integration — there is no defensive parsing layer underneath it.

Worth knowing before you load-test: latency is not flat across question types. TypeSafe explains why, in their own words — "For higher cardinality choices, we do a 2 stage-system of scoring independently then making an explicit choice, hence the occassional slowdown." A 4-option routing decision and a 200-option classification are the same primitive on paper and different amounts of work in practice. Our daily medians across the seven days ending 2026-09-30 run 175, 170, 163, 161, 170, 147, 143 ms. One day in that series, 2026-09-28, had a p95 of 2,448 ms — a real single-day outlier that sits in the series honestly, but is not the shape of the service.

A generated single-column scoreboard titled 'Jev 1.13 - the scoreboard' with six rows: 'Median time to first token: 151 ms', 'p95 time to first token: 247 ms', 'Error rate, seven days: 0.49%', 'Input price: $0.042 / M tokens', 'Output billing: $0.000000 / M tokens' and 'Endpoint: POST /v1/systemone', plus a footer reading 'Serving figures: OrcaRouter Playground, seven days ending 2026-09-30. Price per the OrcaRouter catalogue; TypeSafe's own speed and cost multipliers are vendor-reported.' The OrcaRouter logo is composited in the bottom-right corner.

The other thing to know before a first integration is what you are connecting to. The tooling around Jev is open source under MIT and Apache-2.0 licences — the Python and JavaScript SDKs, an adapter that presents the same client backed by ordinary LLM APIs, the workflow-evals code, and a set of agent skills, all in TypeSafe's public repositories, with star counts and push dates that moved as recently as 2026-09-26 and 2026-09-29. The model is not. There is no weights repository: Jev's architecture, parameter count, training compute and weights are unpublished, and a reader checking that should not be misled by the three repositories in that organisation which are forks of unrelated projects — a vLLM fork, a diffusion-language-model release from 2025, and a Pulumi provider. None of them says anything about how Jev is built. The one-line answer is that the tooling is open and the model is not.

Running it today, and what changes for a reader

Jev 1.13 is on OrcaRouter as typesafe/jev-1.13, reachable on the same key as 200+ other models, with the provider's list price passed through at 0% markup. The practical value of that on a page about a category is narrow and worth stating exactly: trying a System One model no longer requires a separate account, a separate key and a separate invoice for a model you may not yet know you want. It sits beside the generative half of the same workflow — the classifier and the writer on one credential, in one place, with the counts of what you actually called.

Nothing here changes what the model is. It launched on 2026-09-15 and TypeSafe still describes it as early access; what it is has not moved since. What changed on 2026-09-24 is that a reader can now find out what it costs them in practice without committing to a second vendor relationship first. If you have been waiting to see whether the category is worth a prototype, that is the thing that moved.