A generated title card reading "Where Jev 1.13 Breaks" under the eyebrow "TypeSafe System One" and the subtitle "The vendor's own list of what the model cannot do", with three stacked cards reading "No counting, no date maths, no generation", "Choice questions cap at 255 options" and "64K request budget - 32K of it for state", and a footer reading "Callable as typesafe/jev-1.13".
Engineering & Research

Where Jev 1.13 Breaks: TypeSafe's Own Limits List

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Jev 1.13 (typesafe/jev-1.13) was released on 2026-09-15, which puts it two weeks outside the last seven days, so its launch is not the story. The dated event is 2026-09-24: that is the day OrcaRouter added typesafe/jev-1.13 to its catalogue and opened a model card for it — the first serving support for Jev in a third-party gateway, after a fortnight in which the only way to call it was TypeSafe's own endpoint. That matters here for one specific reason. Jev is unusual in that its vendor publishes a list of the ways it fails, and a list you can only read is far easier to skip than a model you can actually call.

This page is that list, kept to what TypeSafe itself says, plus the operating limits and the bill.

TypeSafe publishes its own jaggedness list

A screenshot of the TypeSafe documentation index at docs.typesafe.ai showing the Reference section with the page "Model jaggedness" and the entry "Jev 1.13", beside the Models, API reference, Agent skill, Legal, Client SDKs and Cookbooks sections.

docs.typesafe.ai carries a page titled Jev 1.13 jaggedness. It applies explicitly to jev-1.13, carries a review date of 2026-09-17, and opens with the vendor's own framing: "Jev isn't perfect. Here are some jagged edges we are aware of with jev-1.13. Many of these will be fixed in later versions." Nine named modes follow, each with a concrete case and an "Instead:" remedy. Nothing below is inferred, and nothing is softened — the wording is TypeSafe's, and where the company gives its own example, the numbers in it are theirs.

Literal reading: it answers the question you wrote

Scoping words, negations, and implied conditions are read at face value. A question is answered on the words in the instruction, "whereas a person might have read the intent behind the instructions."

The vendor's diagnostic is the useful part: when you look at a wrong answer and find yourself explaining what you really meant, that explanation is the missing half of the instruction. The remedies are to state the exact condition in instructions, put boundary cases in the criteria, and where interpretation is genuinely unavoidable, split the question into two literal ones and combine them in code.

Math and numbers: it is not a calculator

TypeSafe says plainly to implement mathematical logic in code. Three specific failures sit under that:

• Counting is unreliable. This covers characters in a word, occurrences of a term in a passage, and items in a long list. "The model recognizes the shape of an answer rather than tallying, and the error grows with the size of the thing being counted." The vendor's own test for whether to ask at all: if a regular expression or a parser can find the unit, the count belongs in code and the model adds nothing.

• Numeric representations underperform semantic ones. Questions about colours using hex values are worse than the same questions using English colour names; given RGB triples or hex, Jev cannot reliably judge whether two values are near each other. The same gap shows on low-level code — assembly, or binary-encoded instructions — against high-level languages. Convert or bucket in code, and keep the model for the part that is genuinely a judgment.

• Score outputs do not carry exact magnitudes. The vendor states that Jev's score levels are weak in numerical calibration. An expectation can be used to test whether something passes a threshold; it cannot be used to reconstruct the number by interpolating between the two nearest levels. That is a hard no on a whole class of misuse — reading a score as a measurement.

Date and time: dates are read as text, not as quantities

Ordering two dates, measuring the distance between them, or deciding whether one falls inside a window is unreliable, and it degrades further with mixed formats, relative references, and domain boundaries such as quarters, settlement windows, and accrual periods.

The recommended split is clean. Extraction is a judgment, so give it to the model. Every component of a date is a small closed set — twelve months, thirty-one possible days, a bounded range of years — which turns extraction into a choice over enumerated options rather than free-form parsing, and gives you somewhere to put an explicit "not stated" so a missing part is reported instead of guessed. Code assembles the parts and owns everything after that, including ordering, duration, offset and weekday.

Indirection: double negatives and extra hops cost accuracy

Instructions carrying double negatives or layered indirection are answered less reliably. A question about a property of a property, or one that requires several hops of reasoning, costs accuracy. The remedy is to write instructions as directly as possible and to name the relevant parts of the state rather than describing them.

A large state full of irrelevant detail costs accuracy

Accuracy falls as the state grows with content unrelated to the decision. Unrelated detail acts as a distractor, and a large state makes it harder to tell which part of the input produced a wrong answer. TypeSafe's own reminder in the closing note is blunt: "Jev suffers from context rot, so unrelated material in the state costs you accuracy."

Retrieve and filter in code first, and send only the fields the question needs. Where filtering before the request is not possible, the vendor suggests using a noul to filter for relevance, then judging the survivors.

Adversarial content in the state moves the answer

State is data, and jev-1.13 does not treat it as hostile by default. An injected instruction, a deliberately misleading framing, or text that argues for its own classification can shift the result. This is the one mode where the vendor explicitly frames the fix as future work — "We expect to improve on this in the future" — and the interim advice is to be explicit in the criteria and to test the integration thoroughly before putting it in front of many users.

Contradictory instructions and criteria confuse it

When the instructions and the criteria ask for different things, the model can get confused. TypeSafe's example is a noul where true maps to no and false maps to yes, which performs worse than the same question phrased consistently. The instruction is to treat the criteria as an extension of the instruction and align the two in language an average person could read and understand.

Structural invariants it does not guarantee

This is the mode most likely to break a system that was built on an assumption nobody wrote down. Jev is extremely consistent in the ordinary sense — semantically similar inputs produce quantitatively similar outputs — but structural identities you might expect to hold are not guaranteed. The vendor publishes two worked cases.

• One question, two question types. "Is the customer asking for a refund?" asked as a noul and asked as a yes/no choice, on the ticket "I'm not happy with the fit. What are my options here?" returns a noul of 0.22, and a choice of yes 0.01, no 0.99, confidence 0.97. Those are answers to the same question.

• A question and its negation. "Is the customer asking for a refund?" and "Is the customer asking for something other than a refund?", asked as two nouls on the ticket "I was charged twice for the same order. Can someone look into this?", return 0.72 and 0.47. They sum to 1.19.

The remedies are operational, not rhetorical: do not rely on expected structural invariance, do not carry a threshold tuned on a noul over to a choice, and do not hold the model to arithmetic identities between separate questions. The reason is that a choice is relative — it settles which option — while each noul is absolute and can come back low for all of them.

Generation: it was not trained to write

jev-1.13 is not trained to generate text. You can force output by chaining choices, and TypeSafe says directly that this "will not work well and will be very slow." For extraction, the guidance is to pull candidate values out with a regular expression or a generative model and let Jev pick the correct one, or — when the answer space is bounded — to turn the extraction into a choice over the options rather than asking for the value itself.

The 255-option ceiling on choice questions

A generated scoreboard titled "Jev 1.13 - seven days on OrcaRouter" listing six cards: "Median latency: 151 ms", "p95 latency: 247 ms", "Output throughput: 348 tokens/second", "Error rate over the window: 0.49%", "Tokens served over the window: 76.2 million" and "Daily median, last seven days: 175, 170, 163, 161, 170, 147, 143 ms", with a footer reading "Serving figures measured by OrcaRouter, seven days ending 2026-09-30. Limits per docs.typesafe.ai/models.md."

A choice question a Jev integration is not a jagged edge, it is the shape of the product. These are worth separating out because no amount of prompt work changes them:

• No text generation. It returns a decision, not prose. That is the design, not a defect.

• No conversation. Jev is a structured decision model rather than a chat model. You send a state and a set of named questions; it returns one structured answer per question. There is no turn-taking to design around.

• No multimodal input. Input is text only — string, JSON object, or array of text values, with no image, audio or video. Non-text material has to be pre-processed into text or structured fields before it becomes part of the state.

• Non-streaming responses. There is a single structured response and no streaming mode. The reason this one does not matter is the same reason it is worth stating: there is nothing to stream. A typed decision — a boolean with a probability, one label out of a set, or a level on a scale — has no partial form worth revealing token by token.

• English is the primary language. Other languages, including CJK scripts, are handled but not equally well. TypeSafe's advice is to test on your own content before relying on Jev for a non-English workload, and to lean on confidence when routing.

The 255-option ceiling on choice questions

A choice question picks one of up to 255 labelled options, and that ceiling is a hard one. TypeSafe also explains why large choice sets run slower, in the vendor's own words: "For higher cardinality choices, we do a 2 stage-system of scoring independently then making an explicit choice, hence the occassional slowdown." So the latency cost of a big option set is structural rather than incidental, and it is the vendor telling you where it comes from.

Our own serving window for typesafe/jev-1.13, read from the model card on 2026-09-30, shows what that looks like in practice across seven days of our own traffic: a median of 151 ms and a p95 of 247 ms, 348 output tokens per second, and a 0.49% error rate across 76.2 million tokens served. The daily medians move in a narrow band — 175, 170, 163, 161, 170, 147 and 143 ms from 2026-09-24 through 2026-09-30 — but the daily p95 for 2026-09-28 is 2,448 ms, roughly ten times the days either side of it. We cannot attribute that single-day outlier to choice cardinality and are not going to; the honest reading is that the tail exists, and a latency-sensitive workflow should be designed against the tail rather than the median.

The input bill is the whole bill

Output is billed at zero on Jev, which is sometimes read as "Jev is free." It is not, because input is metered and a large state is not free simply because there is nothing on the output side. The vendor price is $0.042 per million input tokens — the same number TypeSafe states as $42 per billion — and OrcaRouter passes the provider list price through at 0% markup, so a vendor price cut lands here the same day.

Here is what that does to a realistic shape, using the vendor's own rate:

• A small request. A 1,200-token support ticket plus roughly 300 tokens of rubric and questions is 1,500 input tokens, which is $0.000063 a call.

• A large request. A 55,000-token contract plus questions that bring the request to 60,000 tokens is 40 times the tokens, so $0.0025 a call — still small per call, and 40 times larger than the first case for the same one answer.

• At volume. 60,000 tokens a call and 10,000 calls a day is 600 million input tokens a day, which is 0.6 billion, so $25.20 a day and about $756 across a 30-day month. The same call count against the 1,500-token request is 15 million tokens a day: $0.63 a day, about $18.90 a month.

The gap between those last two lines is not a pricing trick, it is the metered state. Which is why the filtering advice in the context-rot section is not only an accuracy measure — trimming the state is also the only lever that moves the bill.

The published operating limits, so nobody has to guess

A screenshot of the OrcaRouter model page for TypeSafe Jev 1.13 showing the title "Jev 1.13" with the 65k context badge, the id typesafe/jev-1.13, the release date 2026-09-24, input text, a p50 latency of 151 ms, and the description "Served via POST /v1/systemone; non-streaming; up to ~64K input tokens; text in, structured JSON out."

TypeSafe's models page publishes concrete numbers, so a planner does not have to infer them:

• Throughput and rate. 100K tokens per second and 40 requests per second, per docs.typesafe.ai/models.md. A request over either limit returns 429 Too Many Requests; the vendor's client SDKs retry with backoff by default and honour the retry-after header when the response carries one.

• The limits move. The vendor states that rate limits are adjusting dynamically and "can change without notice" as capacity comes online, with higher limits available on custom and enterprise plans. Treat 100K/40 as the number today rather than a contract.

• Context budget. The request budget is roughly 64,000 tokens across the combined state and all questions — the model card publishes 65,536 — and the vendor's models page separately caps the state plus the single longest question at 32,000 tokens. That second figure is the state budget, not a smaller version of the first; both are real and neither contradicts the other.

• Aliases move under you. jev-latest and jev-preview both resolve to jev-1.13.0 today, and the vendor notes there is no preview build available right now. An alias moves when a new release ships, so if you have tuned confidence thresholds against a specific version, pin the versioned ID and move on your own schedule.

What your use case needs to look like

Read end to end, the vendor's own list describes a narrow, useful tool. Jev is a fit when the judgment is bounded and the arithmetic is not the model's job: is this record in policy, which of these forty labels applies, how does this read on a five-level scale — asked over state you filtered yourself, with a literal instruction and criteria that agree with it, and with every count, comparison and date measurement done in code around it.

It is not a fit when the task needs counting, ordering or date arithmetic, when it needs several hops of reasoning, when the input material is not text, when the state is a haystack and the question is a needle, or when anything about the source is hostile. Those are not gaps in a prompt; they are places the model does not work, and TypeSafe is the party saying so.

One more thing worth knowing before you wire it up: the honest difference in how Jev is called. On OrcaRouter the catalogue reaches Jev through the dedicated systemone endpoint, POST /v1/systemone, rather than through the OpenAI chat-completions shape. That is a real difference in the request you write, and it is the correct version of the older claim that Jev "speaks its own request shape". Everything else is the same as any other model on the account — one key for 200+ models, no per-token fees from us, and automatic failover if a route goes bad. TypeSafe removed the waitlist on 2026-09-21; the vendor's own homepage still describes Jev as early access, and its own benchmark page is still marked pending, so the only performance figures on this page are the serving numbers we measured ourselves and the vendor's own claims, labelled as theirs.