
What Is TypeSafe AI? The Lab Behind Jev 1.13 Makes Decisions, Not Sentences
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 349 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 208 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 680 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 49 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 105 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 219 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
TypeSafe AI is a small San Francisco lab that sells a model which cannot write a sentence, and Jev 1.13 (typesafe/jev-1.13) is the model in question. It takes a piece of state — an email, a log line, a support ticket, a JSON blob — plus a set of named questions, and returns one typed answer per question, each with its own confidence value. No prose, no code, no explanation, and no chat box. The model itself launched on 2026-09-15, so it is not new and nothing here is a launch story: this page exists because of a smaller, datable change. On 2026-09-24, OrcaRouter added typesafe/jev-1.13 to its catalogue and opened the model card at its own address, which was the first time Jev could be called through a third-party gateway rather than only through TypeSafe's own endpoint. That is the event, it is six days old as of 2026-09-30, and it is the reason a reader who does not already have a TypeSafe contract can now put a System One model on the same key as the generative models it is designed to sit beside. Everything else on this page is background on the company that made it.
The honest framing of the timing, because it matters for how much of this is verified: Jev launched more than two weeks ago and has been generally available since 2026-09-21, when TypeSafe removed its waitlist. What is inside the seven-day window is the routing change, not the model. If you came here expecting a launch review, the launch already happened and a handful of other pieces covered it.
A company page with a manifesto and no architecture diagram
TypeSafe describes itself, in its own meta description, as "an AI lab building machine-native intelligence infrastructure for automation," with systems "designed to make decisions within software." Its homepage is stamped Version 0.01 and footered "Made in SF." There is a manifesto arguing that the shortest path to an AI-shaped economic shift runs through making intelligence composable, so that software can invoke semantic judgement the way it invokes a function; the company's own summary of its plan is to ship the form of machine-native composable AI, then make it dependable enough for real automation, then offer higher-level abstractions stable enough to layer on. Its tagline is "We're building prod, not God." The team page names three founders — Diogo Almeida as CEO, Sasha Sheng as COO and Erik Gafni as CTO. That is the whole of what the company says about itself: a position, a product, three names, and no numbers about the company itself.
What the site does not carry is the thing an engineer evaluating a new dependency reaches for first. There is no architecture page, no parameter count, no training-compute disclosure, and no model card in the academic sense. TypeSafe's own benchmark dashboard is still marked pending. For a company whose pitch rests on reliability, the disclosure surface is thin, and that is a fact about what has been published rather than an accusation about what is being hidden.

System One: the category, and why the name is borrowed
Jev is the first of what TypeSafe calls System One models. The name comes from Daniel Kahneman's split between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning, and the company's launch post says so directly — it also concedes that "System 1 thinking" has carried an overtone of error-prone, and argues that these models can be made more reliable than the alternatives. TypeSafe did not coin the term and does not own it.
The practical content of the category is a deliberate split of labour: a model that decides and a model that writes. You are meant to keep arithmetic, date ordering, counting and string comparison in ordinary code, where they are exact, and hand the semantic judgement to Jev. Three question types are what it can answer:
• noul — a true/false judgment, returned with a calibrated probability
• choice — pick one of up to 255 labelled options
• score — rate on an ordered scale; our card publishes 2-10 levels, and TypeSafe's own documentation shows a 0-indexed example, so treat the vendor's levels as the definition and the 2-10 as what the card publishes
Each question carries its own instructions, and for choice and score, its own criteria. Requests above roughly 64K input tokens are rejected before they reach the model, and responses are not streamed. The vendor separately documents the state budget — state plus the single longest question — at 32K tokens, which is a narrower figure than the 65,536-token context on the card and not a contradiction of it.
Their thesis, in their own words
TypeSafe states the thesis better than any summary would, so here it is verbatim: "LLMs produce words for people. Jev produces typed decisions and is more like code: reliable, fast, self-consistent, and type-safe." The training method behind that is theirs too, and it is a term worth attributing precisely because it has been picked up elsewhere as if it were generic. TypeSafe calls it Reinforcement Learning for Calibrated Decisions, or RLCD, and contrasts it with RLHF and RLVR: where those optimise human preference or programmatically verifiable rewards, RLCD targets, in the company's phrasing, "answers with epistemically honest probabilities on System One tasks." Treat RLCD as TypeSafe's own coinage, not as an established acronym in the literature.
The numbers they lead with, and who chose the workload
The homepage leads with "193.6x Faster, 444.6x Cheaper," footnoted to workflows for System One tasks. Under it sits a worked example: TypeSafe AI at $0.000081 and 0.114 seconds against LLMs at $0.013880 and 8.566 seconds. Further down, "$42 Per Billion input tokens. 238x Lower input price than Claude Fable 5.1." And a section headed "Zero Hallucinations," which on inspection is a claim about confidence estimates rather than a proof of zero errors: every Jev decision carries a confidence estimate, so software can act when confidence is high and escalate when it is not. All four are TypeSafe's figures, on workloads TypeSafe chose, and none of them has been independently replicated. The launch post itself says the team expects its 193.6x and 444.6x to sit at the high end of real-world gains, and notes the workflows were built by its own capabilities team, with reference answers produced by rival models. Our own seven-day serving data on the same model puts the error rate at 0.49%, which is a different measurement on a different workload and is the honest counterweight to reading "zero" as an absolute.
What they have not published
Jev's architecture, parameter count, training compute and weights are unpublished, and there is no weights repository in the company's GitHub organisation. Checked on 2026-09-30, that organisation has eleven public repositories; the substantial ones are tooling, all MIT or Apache-2.0 — an agent-skills repo, a Python SDK, a TypeScript SDK, a drop-in client adapter backed by other LLM APIs, a Dagger module collection, some published workflow-eval code, an n8n node, and the org's own site. Star counts and push dates move, so read them on the day rather than off this page.
Three of the eleven are forks of unrelated projects: vLLM, LLaDA, and a Pulumi provider for ClickHouse Cloud. They are forks of somebody else's work and say nothing about how Jev is built — in particular, nothing about Jev being diffusion-based. The one-line answer to whether Jev is open source is that the tooling is open and the model is not.
What they concede themselves
Two admissions from the launch post are worth more than most vendor caveats, because they name the exact places the evidence is weak. On price: "We can't prove it isn't subsidized; we'll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up)." On speed: "our published evals are generally run from our laptops on the West Coast (this is where our service is currently based)." There is a third, more useful one for anyone building on it: for high-cardinality choice sets, Jev runs a two-stage system that scores independently and then makes an explicit choice, "hence the occassional slowdown." That is the vendor explaining why latency is not flat across question types, and it is the sort of thing you normally learn from a support thread.
Where you can actually call it, as of today
Through TypeSafe's own API, where the model is generally available and the homepage still uses the vendor's own phrase "early access" — that wording is current and worth keeping, whereas "waitlisted" is retired: the docs publish concrete operating limits (100K tokens per second, 40 requests per second, 429 with SDK backoff past either) and carry no waitlist language at all. And, since 2026-09-24, through OrcaRouter.
What we serve is worth stating precisely, because the call shape is the part that differs. The catalogue entry is typesafe/jev-1.13, named TypeSafe: Jev 1.13, with supported endpoint type "systemone" — so it is reached via POST /v1/systemone rather than the OpenAI chat-completions shape, non-streaming, text in and structured JSON out, up to about 64K input tokens across state and questions combined. Input is $0.042 per million tokens and output is billed at zero, because there are no output tokens to meter: a typed decision is not prose. That is the provider's list price passed through, at 0% markup, on the same key as the other 200+ models in the catalogue — one API for 200+ models, 0% markup (provider list price passed through, so vendor price cuts are live here the same day). If the reason you are reading about TypeSafe AI is that you want to find out whether Jev's calibration holds up on your own data before you commit a production path to it, running it beside a generative model you already trust is the cheap way to find out; failing over to that model when Jev's confidence comes back low is the cheap way to ship it.
Our own seven-day serving figures for the window ending 2026-09-30, which are our numbers from our own traffic rather than the vendor's benchmark: p50 151 ms, p95 247 ms, about 349 output tokens per second, a 0.49% error rate, and 76.2M tokens served. The daily p50 across that window ran 175, 170, 163, 161, 170, 147, 143 ms, so the median has been drifting down modestly. One day in the series, 09-28, has a p95 of 2,448 ms — a real outlier in the data, not the norm, and not a number to plan a latency budget around. These roll daily; the card is the source.

Where the model is weak, according to TypeSafe
The vendor publishes a jaggedness page for Jev 1.13, last reviewed 2026-09-17, that names the failure modes more candidly than most launch material. It is the right page to read before building on the model, and its own list runs like this. Jev answers the question you wrote rather than the one you meant, so scoping words, negations and implied conditions need to be spelled out. It is not a calculator: it performs worse on mathematical than semantic questions. It reads dates as text rather than as ordered quantities, so ordering, gaps and window membership are unreliable, and worse with mixed formats. Double negatives and multi-hop indirection cost accuracy. Accuracy drops as state grows with irrelevant detail — the page says it "suffers from context rot" — so filtering in code first is the fix. State is treated as data, not as hostile by default, so injected instructions can move answers. Mismatched instructions and criteria confuse it. Structural invariants are not guaranteed: a noul and a choice output need not correspond, and a question plus its negation need not sum to one, so thresholds should not be ported between the two. And it is not trained to generate text — forcing it through chained choices "will not work well and will be very slow."

How to read TypeSafe AI six days into the routing change
The company is making a strong, specific, falsifiable claim: that a narrow decision model can be faster, cheaper and more trustworthy than a general one on the subset of work where you need a judgement rather than a paragraph. The claim is plausible and partly self-evidenced — the confidence estimates are a real design difference, the price is real, and the latency we measure on our own traffic is in the same order as the vendor's. What is missing is the part that would let an outsider check the rest: no architecture, no parameters, no weights, no independent benchmark, and a price the company itself says it cannot prove is sustainable. The failure modes are documented, which is more than most labs do and is the single best reason to take the model seriously.
If you are deciding whether to pay attention: run Jev on input you already have labelled, with the labels hidden, and look at whether its confidence numbers separate the cases it gets right from the ones it gets wrong. That test costs almost nothing at $0.042 per million input tokens and it is the only one that answers the question you actually have. If you are deciding whether to trust the company: the disclosures are what they are, and the honest answer is that the evidence is currently the vendor's word plus whatever you generate yourself.
