A generated hero card for the article 'Jev 1.13 Explained', eyebrow 'TYPESAFE SYSTEM ONE', subtitle 'Why the model answers in labels instead of sentences', with three cards on the right reading 'Typed answers only - no generated text', 'No output tokens, so nothing to bill' and '$0.042 per million input tokens', and a footer line reading 'Callable as typesafe/jev-1.13'.
Guides & Insights

Jev 1.13 Explained: Why the Model Answers in Labels Instead of Sentences

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Jev 1.13 (typesafe/jev-1.13) is not a chat model, and the fastest way to understand it is to stop reading its spec sheet the way you read every other one. TypeSafe shipped it on 2026-09-15 as the first member of a class the company calls System One models: you hand it a piece of state and a set of named questions, and it hands back one typed answer per question — a label from a list you supplied, a level on a scale you defined, or a true/false with a probability attached. No prose, no code, no explanation. This is not a launch piece. The model itself is from 2026-09-15, fifteen days old and outside the seven-day window this blog writes to, so it does not earn a page on its own release. What happened inside the window is that OrcaRouter added typesafe/jev-1.13 to its own catalogue on 2026-09-24 and opened the Jev 1.13 model card at https://www.orcarouter.ai/models/typesafe/jev-1.13 — the first time Jev has been callable through a third-party gateway rather than only through TypeSafe's own endpoint, and the first live serving data anyone outside TypeSafe has published on it. That is the change worth reading about: the model became runnable where it was not.

The practical shape of that change is small and specific. Before 2026-09-24, adopting Jev meant a second vendor relationship — a TypeSafe account, a TypeSafe key, a TypeSafe invoice, and a bespoke request shape to write against. After it, Jev sits on the same key as the rest of a stack: one API for 200+ models, 0% markup (provider list price passed through, so vendor price cuts are live here the same day), and the model reachable at typesafe/jev-1.13 on POST /v1/systemone. You still call it in its own shape — the endpoint is not the OpenAI chat-completions route, and pretending otherwise would produce a 404 rather than a decision — but the contract you sign and the key you rotate are the same ones you already have.

What kind of model Jev is

TypeSafe's launch post states it in one sentence: "Our first public model is Jev, available today in early access." The post then frames the model as "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out." That is a fair description of the interface and an important one, because almost every wrong expectation about Jev comes from evaluating it as a small language model. It is not a small language model. It is a decision model with a fixed output grammar, and the grammar is the product.

The interface has exactly two inputs. The state is the material to be judged: an email, a support ticket, a log line, a JSON record, an array of game coordinates. The questions are a map of named items, each carrying a type, its own instructions, and — for the two structured types — its criteria. Every question is evaluated against that same state, and the answers come back as one structured JSON payload. TypeSafe's docs describe the questions as running concurrently and independently, and make two claims that follow from that design rather than from tuning: adding questions barely changes response time, and adding questions does not create context rot, because each question is judged in isolation rather than downstream of the previous ones.

TypeSafe's own design guidance is worth repeating because it is the clearest statement of what the model is for. Keep each question atomic and well-scoped — "the kind of judgment a highly knowledgeable person could make in a few seconds." If a decision needs extended reasoning or genuinely combines several independent factors, split it into separate questions and recombine them in your own code. Their example: instead of one "rate this startup pitch" prompt, ask separately about market size, technical feasibility and differentiation, then apply your own weighting. The reason that matters is that the weighting then lives in a coefficient you can change, instead of in a prompt you have to rewrite.

The model is closed in every sense that matters to an engineer reasoning about risk. Jev's architecture, parameter count, training compute and weights are unpublished. There is no weights repository on the TypeSafe GitHub organisation — the eleven public repositories there are tooling, SDKs, workflows and three unrelated forks, and none of them is the model.

"Typed" is the whole product

The OrcaRouter model card publishes the three primitives and, importantly, the limits on each. Every question you ask Jev is one of exactly three shapes:

• noul — a true/false judgment, returned with a calibrated probability rather than a bare boolean, so "probably true" and "certainly true" are distinguishable values.

• choice — pick one of up to 255 labelled options, each option carrying its own criteria text so the model knows what distinguishes your labels.

• score — rate on an ordered scale of 2–10 levels, with the level definitions supplied as criteria.

The consequence of that restriction is that there is nothing to parse out of prose and nothing to validate against a schema you hoped the model followed. TypeSafe's launch post is unusually direct about the guarantee: the "0%" schema-mismatch figure in its plots "is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots." When your answer space is a closed set you supplied, a returned value is either in the set or it did not come from the model — there is no third outcome where the model wrote something plausible in the wrong shape and your regex quietly accepted it.

The three types are also the reason a reviewer cannot evaluate Jev the way they evaluate a chat model. There is no MMLU-Pro score to compare, no writing sample to read, no reasoning trace to inspect. The only question that means anything is whether the typed answer is right, and whether the probability attached to it is honest. Both of those are measurable, but only against your data and your labels.

One documented difference is worth flagging rather than resolving: TypeSafe's own documentation shows a Score example that is indexed from zero, while the OrcaRouter card publishes the scale as 2–10 levels. The vendor documents levels; our card publishes 2–10. If you are building a rubric on top of Score, read the level definitions in your own response rather than assuming an index.

Why there are no output tokens to bill

The pricing is the cleanest expression of the architecture. Jev costs $0.042 per million input tokens on OrcaRouter, and the output rate is $0.000000 per million — not a discount, not a launch promotion, but the absence of a meterable quantity. A generative model is billed for the text it writes; Jev writes no text. It returns a label, a level and a probability. There is nothing to count on the output side, so nothing is charged there.

TypeSafe states the same number from the other direction on its homepage — "$42 Per Billion input tokens" — and attaches a comparison claim to it: "238x Lower input price than Claude Fable 5.1". That comparison, like everything else on the homepage, is the vendor's own, unreplicated by anyone. But the arithmetic it rests on is easy for a reader to check against their own bill, which is the useful part. The volume of a decision workload is set almost entirely by how much state you push in, and state is cheap in a way that generated tokens are not.

The vendor's headline figures are bigger than the price line and deserve the same labelling. TypeSafe advertises "193.6x Faster, 444.6x Cheaper" with a footnote restricting it to "workflows for System One tasks", and publishes a worked example beneath it: TypeSafe AI at $0.000081 completed in 0.114s against LLMs at $0.013880 completed in 8.566s. The launch post itself concedes the framing risk — the 193.6x and 444.6x are described as likely sitting "on the higher end of real world gains" — and notes that the side-by-side demo used a "highly simplified" query with human-readable keys chosen by the vendor to "paint our model in an advantageous light." None of these numbers has been independently replicated, and the vendor's own benchmark card is still marked pending.

What "calibrated" means, and what RLCD is

TypeSafe names its training method itself: "Reinforcement Learning for Calibrated Decisions (RLCD)". RLCD is TypeSafe's own term, not a general machine-learning acronym that predates the company, and its optimisation target is stated in the launch post's comparison table as "calibrated decisions: answers with epistemically honest probabilities on System One tasks." The contrast the same table draws is with RLHF, which optimises for human preference — writeups and chat responses raters like — and RLVR, which optimises for outputs that can be programmatically verified. RLCD optimises for a third thing: the probability attached to an answer being an accurate statement of the model's own uncertainty.

Practically, "calibrated" is a claim about the confidences, not a guarantee that the answers are right. A calibrated model that says 0.8 on a set of questions should be right about 80% of the time across that set; it can still be wrong on any individual one. That distinction is the honest way to read TypeSafe's homepage line "Zero Hallucinations — Every Jev decision comes with a confidence estimate, so your software can act when confidence is high and escalate when it is not." It is a confidence-estimate claim, not a proof of zero errors, and the counterweight is our own data: over the seven days ending 2026-09-30, our card measures a 0.49% error rate on Jev traffic through OrcaRouter — a figure that read 0.57% earlier in the same window, because it is computed over a rolling seven days of live playground traffic rather than a fixed test set. Both facts belong in the same paragraph: the confidences are the point of the model, and the model still fails roughly one call in two hundred on our traffic.

The calibration story also explains a latency behaviour that would otherwise look like a bug. TypeSafe's launch post says: "For higher cardinality choices, we do a 2 stage-system of scoring independently then making an explicit choice, hence the occasional slowdown." A 255-option choice is not one forward comparison; the vendor scores and then chooses. If you see a request against a large label set take noticeably longer than a noul, that is the documented mechanism, not congestion.

How you call it today

A screenshot of the OrcaRouter model card for TypeSafe: Jev 1.13 at orcarouter.ai/models/typesafe/jev-1.13, showing the slug typesafe/jev-1.13, the byline 'by TypeSafe · 2026-09-24', the list price of $0.042 per million input tokens with output at $0.000000, a 65,536-token context and the single supported endpoint type systemone.

On OrcaRouter the model is typesafe/jev-1.13, named "TypeSafe: Jev 1.13" in the catalogue, with a context_length of 65,536 tokens and exactly one supported endpoint type: systemone. You call it with POST /v1/systemone on your OrcaRouter key, sending a model field, a state field (string, object or array), and a questions map where each entry carries a type (noul, choice or score), its instructions, and its criteria. Responses are a single structured JSON payload and are not streamed — there is no streaming mode to opt into.

The card's own wording for the contract is "text in, structured JSON out", and the published limits are the ones to design against: non-streaming, up to roughly 64K input tokens across the combined state and questions, with requests over that limit rejected before reaching the model. The catalogue entry lists the list price as $0.042 per million input tokens and shows the completion rate as zero. Those are the vendor's numbers passed through unchanged — the proportional form of how we price: 0% markup on provider list rates.

Two token budgets circulate for this model and they are not in conflict, so keep them apart. The 65,536 figure is the card's context_length and is documented as roughly 64K of input across state plus questions combined. The "roughly 32,000 tokens" that earlier OrcaRouter articles quote is the state budget alone — the room your material gets before the questions take their share. If you are budgeting a request, the state budget is the number that constrains the payload you build; the combined figure is the ceiling on the whole call.

What Jev cannot do, stated plainly

It cannot write prose, summarise, translate or converse. That is the design, not a limitation to apologise for: the launch post says Jev "gives up string generation", and TypeSafe's own jaggedness page lists "Generation" as a named failure mode with the instruction "Use a generative model" beside it. Forced generation is slow and poor. If your pipeline needs a written summary, Jev is the wrong component and no amount of prompt craft changes that.

It is not a replacement for a generative model. The workflow Jev belongs to has two models in it: a generative one that reads, writes and reasons in text, and Jev sitting beside it making the typed calls in milliseconds. That is the honest framing of every cost comparison on the vendor's homepage — the "LLMs" column is not a competitor being displaced, it is the other half of the same system, and the reason the pairing is interesting is that the decision half is now on the same key as the generative half rather than behind its own contract.

And "calibrated" does not mean correct. It means the number attached to an answer is meant to be readable as a probability. A 0.62 on a noul is the model telling you it is not sure, which is useful information that a bare yes/no would have destroyed — and it is not a promise that the yes is right. Escalation logic built on the confidence is the intended pattern; treating the answer as ground truth is not.

Reading the numbers honestly

A generated figures card titled 'Jev 1.13 - the numbers we measured' with six rows: median time to first token 151 ms; p95 time to first token 247 ms; output throughput about 349 tokens/second; error rate over the window 0.49%; tokens served over the window 76.2 million; daily median 175, 170, 163, 161, 170, 147, 143 ms. The footer reads 'OrcaRouter Playground, seven days ending 2026-09-30. TypeSafe's own multipliers are vendor-reported and unreplicated.'

Every serving figure on our card comes from our own traffic through OrcaRouter's playground over a rolling seven-day window, not from the vendor's benchmark, and the window moved while this piece was being written — treat it as a reading, not a specification. For the seven days ending 2026-09-30: median time to first token 151 ms, p95 247 ms, output throughput around 349 tokens per second, error rate 0.49%, and 76.2 million tokens served. Daily p50 across the window runs 175, 170, 163, 161, 170, 147 and 143 ms — a gently improving line. The 09-28 p95 of 2,448 ms is a genuine single-day outlier sitting in that series, and quoting it as the norm would be wrong in the same way that dropping it entirely would be dishonest.

One further caveat on the traffic figure: 349 output tokens per second sounds like a generative model's throughput until you remember that Jev produces no generated text. The meter is measuring whatever our playground counts on the response side for a structured payload, and it is useful for spotting degradation between days rather than for comparing Jev to a chat model.

Those are our numbers. The vendor's numbers are the 193.6x, the 444.6x, the $0.000081 worked example and the 238x comparison against Claude Fable 5.1 — all TypeSafe's own, none independently replicated, all restricted by their own footnotes to System One task workflows specifically. The one performance claim TypeSafe makes that is not a benchmark at all, and is worth more than the multipliers, is a structural one: because the answers are typed, the integration has no parsing step and no schema-validation step, and that is a cost that does not show up in any latency table.

What is open, and what is not

Checked on 2026-09-30, the TypeSafe GitHub organisation published eleven repositories. None contains Jev. The ones that matter to a developer integrating the model are all MIT or Apache-2.0: skills (MIT), system-one-adapter-python (MIT, described as a "Drop-in TypeSafeClient replacement backed by LLM APIs"), typesafe-sdk-js (MIT), typesafe-sdk-python (MIT), daggerverse (Apache-2.0), WorkflowEvals (Apache-2.0, with workflow code published at evals.typesafe.ai), n8n-nodes-typesafe-ai (MIT), typesafe-ai.github.io and pulumi-clickhouse. Star counts and push dates move, so if you are reading this later, re-check rather than trusting the list.

Three of the eleven are forks of unrelated projects and prove nothing about how Jev works: a vLLM fork last pushed in May 2025, a LLaDA fork from June 2025 — LLaDA is an unrelated diffusion-language-model release — and a Pulumi provider for ClickHouse Cloud. It is tempting to read architecture out of a fork list. Do not: nothing about Jev's design follows from those three, and in particular Jev is not a diffusion model, however much the presence of a LLaDA fork might suggest it.

The honest one-line answer to "is Jev open source" is that the tooling is open and the model is not. That is a normal arrangement for a hosted frontier model, and it is the arrangement you should assume when you plan around Jev: an API with a published price, a documented contract and an unpublished parameter count.

Where the vendor says Jev is unreliable

TypeSafe publishes its own jaggedness page for jev-1.13, last reviewed 2026-09-17, naming where the model breaks. It is unusually candid and it is the right place to start a limitations section, because it is the vendor's own list rather than a competitor's:

• Literal reading — it takes wording at face value. Scoping words, negations and implied conditions are not inferred; it "answers the question you wrote, not the one you meant." The vendor's fix is to write the exact condition and the criteria for every option.

• Math and numbers — it is not a calculator, and it does not count reliably. Keep arithmetic in code.

• Date and time comparison — dates are read as text, not ordered quantities, so ordering, gaps and windows are unreliable, worse with mixed formats.

• Indirection — double negatives and multi-hop reasoning reduce accuracy. Reduce hops and point at the relevant state.

• Large state full of irrelevant detail — unrelated content acts as a distractor and accuracy drops as state grows. Filter first.

• Adversarial content — state is not treated as hostile, so injected instructions or misleading framing can shift answers.

• Contradictory instructions and criteria — when the two ask for different things, the model "might get confused."

• Common-sense structural invariants — P(noul) and 1 − P(not noul) are not guaranteed to be consistent. Ask each decision one way and enforce identities in code.

• Generation — already covered, and the vendor's own advice is to use a generative model.

Two of these deserve emphasis. The adversarial one matters because Jev's whole value proposition is judging untrusted material, and a state containing instructions can move an answer; if your state arrives from users, that is a prompt-injection surface with the same shape as any other. The structural-invariants one matters because a "calibrated" model invites you to do arithmetic on its probabilities, and the vendor is telling you not to assume the arithmetic closes.

What the launch post commits to, and what it does not

A screenshot of TypeSafe's own launch post, headed 'Introducing System One Models & Jev' and dated Sep 15, bylined 'Diogo Almeida, founder, TypeSafe', showing the opening paragraphs that frame the model as a frontier-intelligence function call taking unstructured state in and returning typed probabilistic decisions out.

Almost every vendor claim quoted in this piece traces back to one page: TypeSafe's own announcement, filed under Company News and dated 2026-09-15, signed by founder Diogo Almeida. Reading it directly is worth the two minutes, because the wording of one line sets the terms for everything since. "Our first public model is Jev, available today in early access." Early access is the vendor's own description of availability on the vendor's own platform, and it is a narrower statement than it looks — it commits TypeSafe to serving the model to approved users, and says nothing about who else may serve it. That is precisely the gap the 2026-09-24 catalogue addition closed, and the reason the model card matters more than the post for anyone evaluating Jev today.

The same page is equally clear about its own limits, which is why it is quoted rather than paraphrased above. It publishes no parameter count, no architecture description beyond "a new model architecture", no training compute and no weights repository, and it gives no date or conditions for general availability. Those absences are the planning constraints: an API with a published price on one side, and a stack whose internals you cannot inspect on the other. The post also frames the model honestly enough to be useful as a spec — "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out" — which is the one sentence in it that describes the interface rather than the ambition.

Questions this interface raises

What happens when a choice set is bigger than 255 options? It is capped — 255 labelled options is the ceiling on a choice question, and the vendor's two-stage scoring-then-choosing approach is what it does as cardinality climbs, which is also the documented source of occasional slowdowns on large label sets. If your taxonomy is bigger than that, the design answer is to decompose it into several questions and recombine in code, which is the same advice TypeSafe gives for compound judgments.

Does the 65,536-token context mean 65,536 tokens of state? No. The published budget is roughly 64K tokens across the combined state and all questions, and the "roughly 32,000 tokens" figure that appears in older coverage is the state budget alone. Budget your payload against the state figure, not the combined one, and remember that requests over the limit are rejected before the model is reached.

What to do with this

Jev 1.13 is worth a look for one specific reason and not for a general one. If you have a step in your pipeline that is currently a chat model being asked to return a label and being trusted to return it in the right shape — a router, a grader, a policy check, a rubric score applied to thousands of records — that step is what this model replaces, at $0.042 per million input tokens with nothing metered on the output side. If you have a step that needs a written answer, Jev is not the tool and its own vendor says so.

The thing that changed in the last week is not the model. It is that trying it stopped requiring a second vendor relationship. Eight days ago a Jev evaluation meant a separate account and a separate integration; today it is one model id on a key that already reaches 200+ models, with the vendor's list price passed through unchanged and the typed answers coming back from the same place as everything else. For a model this unusual, the ability to test it against your own labels without committing to a new contract is most of the decision.