
Where OpenAI's Decisions API Ends and Jev Begins: What the First Outside Client Shows
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 147 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 80 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 53 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 296 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
At 23:05 UTC on 2026-10-06, a plugin called llm-openai-decisions landed on PyPI, and its own README contains the sentence the launch coverage never printed: "Unlike Jev, the new gpt-6-luna decision model supports image input in addition to text." The author is Simon Willison, who also wrote the first client for TypeSafe's decision model, and the comparison he draws is between GPT-6 Luna — the model behind OpenAI's Decisions API — and Jev 1.13, TypeSafe's text-only System One model. The same blog post notes that the plugin itself was written by GPT-6 Astra reading OpenAI's documentation.
That is the useful thing about a client shipping a week after an API: it is written against the request shape, not the keynote, so it surfaces the differences that marketing copy smooths over. Nothing here is a leak and nothing here is unconfirmed — every figure below is from a page published by OpenAI, TypeSafe, or the plugin's own repository, all read on 2026-10-07. What is still missing is independent measurement of the endpoint itself, and that gap is stated at the end rather than papered over.
The three surfaces, kept apart
The first thing to get straight is that these are three different layers, and press coverage blurs them.
• The endpoints. OpenAI's is POST /v1/decisions, a dedicated route rather than a mode on the Responses API. TypeSafe's is POST /v1/systemone. Both take a body of evidence plus a list of typed questions and return one answer per question, keyed by the name you gave it.
• The models. The Decisions API accepts exactly one model today, gpt-6-luna, which is also the cheap tier of OpenAI's general GPT-6 line and shipped on 2026-09-22 as an ordinary text and image model. Jev 1.13 is a decision model end to end — it cannot emit free text at all. TypeSafe shipped it on 2026-09-15 and it has been generally available since 2026-09-21.
• The clients. OpenAI's own SDKs have supported the endpoint since it opened in public beta on 2026-10-06, so the plugin is not the first way to call it. It is the first client outside OpenAI's own SDK that we have been able to verify, and the first one written for a command-line tool rather than an application library.
The two metres, side by side
Here is where the client's README is worth more than the announcement, because the numbers sit next to the wording.
• Price — the Decisions API bills $0.10 per million input tokens with no output charge, no cache-read charge and no cache-write charge. Jev 1.13 bills $0.042 per million input tokens with output free. Both charge for what you send and nothing for what comes back, so a decision costs a fraction of a cent at either price and the difference between them is a factor of 2.4, not a different billing model.
• Input — the Decisions API takes text, or user messages mixing text with images, and the plugin's README states that PNG, JPEG, WebP and GIF attachments are supported with at most 128 images in a request. Images must go in as inline base64 data URLs; hosted HTTP or HTTPS image URLs and file_id inputs are not accepted by the endpoint, so the plugin converts a URL or a local path before sending it. Jev 1.13's documentation is blunt in the other direction: "No image, audio, or video input," and non-text inputs should be pre-processed into text or structured fields before they reach the state.
• Question types — OpenAI documents three: predicate (a probability from 0 to 1 that a stated condition holds), choice (one value from your list, plus a distribution and a separate confidence field), and score (a probability-weighted average across ordered levels). TypeSafe's three are the same three ideas under different names: noul for yes/no, choice, and score. "Noul" is short for Bernoulli, and the name is a fair marker of the cultural difference — one vendor ships a plain-English noun, the other ships a statistics joke.
• Score arithmetic — the two agree exactly, which is the strongest sign that this is one product category rather than two. OpenAI's documentation feeds three severity levels probabilities of 0.1, 0.7 and 0.2 and gets back a score of 1.1 — deliberately between two levels rather than snapping to the nearest. TypeSafe's levels are zero-indexed the same way, and the plugin for Jev takes between two and ten ordered levels. OpenAI's page states no ceiling; that is an absence in their documentation, not a limit we can assert.
• Budgets — Jev documents its window precisely: roughly 64,000 tokens per request, about 32,000 of them covering the state plus the longest single question, against the 1,050,000-token context our own catalogue carries for GPT-6 Luna as a general model. OpenAI's decisions page publishes no token budget at all, only the note that regional processing premiums and long-context input multipliers still apply to the rate.
What the client reveals that the announcement did not
Three details in the plugin and its README are worth pulling out, because each one changes how you would wire the endpoint up.
The first is that a decision can come back as a refusal. OpenAI's documentation never explains this in prose, but every SDK sample branches on it — answer.type === "refusal" in JavaScript, an OpenAI::Models::Decision::Answer::Refusal case in Ruby — and the plugin's README spells out that a declined question is preserved as {"name":"...","type":"refusal"}. So the answer type is effectively four values, not three, and any production loop has to handle a fourth branch that no announcement mentioned.
The second is what is missing from the plugin rather than present in it. Willison's own post frames the work as having GPT-6 Astra read the new documentation and build the client from it — documentation-to-client, one pass, no human tutorial. That is now a normal way for a client to get written, and it means the question of whether an endpoint's docs are complete enough to generate a working client from has become a practical one rather than an editorial one. For this endpoint the answer is mostly yes, with the refusal type as the visible seam.
The third is the shape of the questions array. OpenAI lets you put independent questions in one request against shared input — check a product photo for damage and classify its category in the same call — but requires separate requests when a later question depends on an earlier answer. Jev takes the same position for the same reason: its questions are evaluated in parallel against one state, so anything sequential has to become two calls. Both vendors have designed for fan-out, and both are telling you the same thing about where the latency budget goes.
The category now has three occupants, and two of them are not general models
It is worth naming the third, because the decision-endpoint frame only makes sense with all three in view. Perplexity ships a Decisions API of its own, served by pplx-decider-v1-27b — a 27-billion-parameter decision model released under Apache 2.0 with weights on Hugging Face on 2026-10-01. That gives the category a genuinely different shape from OpenAI's: Jev 1.13 is closed and text-only, Perplexity's decider is open-weight and accepts images, and GPT-6 Luna's entry is a harness around a general model rather than a model built to decide.
For a reader choosing today, the practical split is narrower than the marketing. If your evidence is a sentence or a record and you want the cheapest per-call cost with the least variance in behaviour, Jev 1.13 is the specialist and its $0.042 rate is the lowest of the three published prices we could verify. If your evidence includes a photograph, or you want a decision from a model you already know from ordinary tasks, the Decisions API is the only one of the two closed options that accepts images. The open-weight option answers a different question — control and self-hosting — and we have not tested it.
What we can actually measure, and what nobody has
OpenAI's claim for the endpoint is that it "evaluates text, images, or both and returns typed answers about 10x faster than the Responses API." That is vendor-stated and unreproduced: no region, no input size, no concurrency level, no service-level agreement, and the baseline is the Responses API in general rather than any specific workload. A single speed-up figure is the wrong number to build a deadline on.
What we can put beside it is our own seven-day serving window ending 2026-10-07 on the two models underneath, from our playground traffic — and it is important to say what those numbers are not. They describe ordinary generation requests, not decisions.
• GPT-6 Luna, all request shapes: a median of 1,448 ms and a p95 of 4,912 ms, a 1.31% error rate over 643,394,111 tokens in seven days, which is what throughput of about 125 output tokens a second looks like when the model is writing.
• Jev 1.13: a median of 149 ms and a p95 of 245 ms, a 0.10% error rate over 110,193,080 tokens. That is a genuinely fast endpoint, and it is fast because it does not generate a response — it returns numbers for a state that is ingested once.
Read against each other those two rows are a warning, not a comparison. A decision request emits a handful of tokens, so on the Decisions API the output-throughput figure that dominates Luna's general profile stops being the binding constraint, and the number that starts to matter is how long the model takes to read the evidence. We have not measured that on the decisions endpoint, and as far as we can tell nobody outside OpenAI has published it.
The deeper unmeasured question is calibration, and it is the one that decides whether any of this is usable. A probability of 0.92 for visible damage is only worth routing on if, across your traffic, the photos scored near 0.92 are damaged about 92% of the time. OpenAI's documentation tells you to set thresholds from labelled examples and to choose them by the cost of false positives against false negatives. That is correct advice, and it is also an admission that the calibration of these numbers is something you have to establish yourself. For a first pass, measure the distribution of returned probabilities on a sample you already know the answers to. If everything comes back at 0.99 or 0.01, the threshold is doing no work and the endpoint is a very expensive boolean.
How to try either one without betting a production path on it
Sourcing, plainly: the endpoint contract, the price and the image rules in this piece are from OpenAI's own Decisions API documentation and the plugin's README, both read on 2026-10-07; the Jev 1.13 specification is from TypeSafe's own model documentation; the serving figures are our own playground data over the seven-day window ending 2026-10-07; the package timestamps come from PyPI and the project's Git history. The 10x speed claim is OpenAI's and is labelled as such. The open-weight decider's licence and release date are from its model repository.
The Decisions API itself is OpenAI's own packaging and we do not route it; if you want that specific endpoint, OpenAI is where it lives. The models underneath are a different matter. GPT-6 Luna is a live route on OrcaRouter at OpenAI's list price with zero markup passed through, and typesafe/jev-1.13 at list price is on the same key, which makes the comparison in this article something you can run rather than read: the same state, the same questions, two endpoints, one contract to manage and no second invoice.
That is also the sane way to adopt an unproven surface. Put the decision behind a fallback so that a refusal, a timeout, or a beta limit changing under you degrades into a prompt-based call rather than an outage, and keep that fallback on the same key as the primary so there is nothing to rewire when the endpoint moves toward general availability. OpenAI says GA is expected in the coming weeks and that gpt-6-luna is the only model available in the meantime; both of those are reasons to build against the shape and instrument it now, and neither is a reason to put a payment authorisation behind it yet.

What to watch next
Three things would settle the questions this piece cannot. A published token budget for the decisions endpoint, so a request can be sized rather than guessed. Any GA announcement, which is the point at which the input-only rate stops being a beta promise. And one independent measurement of latency and calibration on the endpoint itself — the first person to put a few thousand labelled pairs through it and publish the reliability curve will do more for the category than either launch post did.
Until then the honest summary is narrow and useful: if the thing you need to decide is text, Jev 1.13 is cheaper and it returns numbers in about 150 milliseconds in our traffic. If the thing you need to decide includes an image, OpenAI's Decisions API is the one that will look at it, at 2.4 times the input price, with a beta label on the box. Both are callable from a command line as of this week, and that is a better position than either was in seven days ago.


