Hero title card reading 'OpenAI's Decisions API', subtitled 'GPT-6 Luna stops chatting and picks one answer', with three rounded cards reading 'One answer from a list you define', 'Vendor-stated ~150 ms' and 'No published price yet', a footer reading 'OpenAI figures vendor-reported; no independent measurement yet.', and the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

OpenAI's Decisions API: GPT-6 Luna Stops Chatting and Picks One Answer

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GPT-6 Luna shipped on September 22, 2026 as the cheap rung of OpenAI's GPT-6 ladder. What is new on September 29 is a job for it that has nothing to do with conversation. OpenAI's Decisions API takes a question you wrote, a finite list of answers you allow, and a piece of context — text or an image — and returns one of those answers. Not prose. Not a paragraph to parse and hope. A single choice, so a support ticket goes to one of five queues, an image lands in one moderation category, or an agent picks its next move out of a policy you defined. It was announced at DevDay 2026 and it is in limited preview now, with OpenAI saying a broad release is planned in the coming days.

Most of what a team needs before it builds on this is not published yet. There is no price per call, no stated limit on how many candidate answers a request can carry, no statement about whether you can tune it on your own data, and — as of September 30 — no entry for it anywhere in OpenAI's own developer documentation index, changelog, or API reference. We checked all three. The capability currently exists in OpenAI's DevDay recap and in what an OpenAI spokesperson told reporters on launch day. So the rest of this article keeps three tiers visibly separate: what OpenAI has confirmed, what one launch-day measurement claims, and what nobody can know yet.

What a constrained decision actually is

Two existing approaches do this job badly, and the Decisions API is aimed at the gap between them. The first is prompting a chat model: you write a careful instruction asking it to choose from a list, it writes you a sentence, and if you are lucky you can read token probabilities out of the response to get something that resembles a confidence score. It works, it burns tokens, and the confidence number is a rough guess. The second is training a small classifier: fast, cheap, and it needs labelled data plus a retraining run every time the label set changes.

A decision model sits between those. It accepts new labels in the prompt — the same way a prompt does — but returns a score or a choice a developer can branch on without parsing, which is the property a classifier had and a chat model never did. OpenAI is not new to the shape: its Moderation API has returned per-category scores rather than text for years, with one difference that matters here. Moderation's categories are OpenAI's. The Decisions API's categories are yours.

The constraint is the product, and it is worth being precise about what it does and does not buy. A finite output space means every permitted value is known in advance, so your application can reject an answer it did not define, attach a different permission level to each branch, and fall back to a human when confidence or context is thin. It also makes offline evaluation tractable, because every test case has a target class and a measurable cost when it is wrong. What it does not buy is correctness. A constrained model can still select the wrong valid answer, be steered by misleading context, or inherit the bias of the categories you wrote.

The 150-millisecond claim, and the number OpenAI did not publish

OpenAI says the Decisions API makes decisions roughly ten times faster than GPT-6 Luna does through the regular API — on the order of 150 milliseconds against about 1.6 seconds. That figure is vendor-stated and unreproduced; there is no independent measurement of it in the two days since launch, and OpenAI's own recap says only "real-time decision-making." No latency distribution, no region, no input size, no concurrency level, and no service-level agreement accompanies the number.

Our own traffic data is not a substitute for that test, but it explains why the missing distribution matters. Across the seven days to September 30, GPT-6 Luna served through OrcaRouter's normal chat-completions route shows a median of 699 milliseconds and a p95 of 7.6 seconds over 3.23 billion tokens of mixed workload — a spread of more than an order of magnitude between the middle and the tail. That is a different request shape from a constrained decision, so it is not a comparison to the 150 ms claim. It is the reason a single headline p50 is the wrong number to design a deadline around. If you test the preview, record median, p95 and p99 separately for text and image inputs, include network time and retries, and measure the deadline your product actually has — a routing decision for an asynchronous queue and one sitting between a click and a payment do not share a latency budget.

A single-column scoreboard titled 'GPT-6 Luna and the Decisions API - the scoreboard' with six rows: Decisions API price 'not published', candidate-answer limit 'not published', fine-tuning on your data 'no statement', stated decision latency '~150 ms vendor-reported', GPT-6 Luna input / output '$0.10 / $0.50 per 1M', and AA Intelligence Index at max effort '37.3', with a footer reading 'OpenAI figures vendor-reported; Index per Artificial Analysis; unpublished means blank, not zero.' and the OrcaRouter logo composited in the bottom-right corner.

Pricing: what the underlying model costs, and why that is not the API's price

OpenAI has published no rate for the Decisions API. It is also not safe to assume the Luna token rates apply, because the Decisions API is its own packaging — a different endpoint with different limits, and OpenAI has said it will share more detail "at broad rollout." Anyone quoting you a per-decision cost right now is inferring it.

What is published is the rate card for the model it is built on, read from OpenAI's pricing page on September 30, 2026:

• Input — $0.10 per million tokens, which becomes $0.20 above 272,000 input tokens
• Output — $0.50 per million tokens, which becomes $0.75 above 272,000 input tokens
• Cached input — $0.01 per million tokens; cache writes bill at $0.125, a 25% premium on the uncached rate
• Batch and Flex processing — 50% of the standard rates; Fast mode — 2x the applicable rates

The long-context line is the one that catches people out, and it works the same way it does across the GPT-6 family: crossing 272,000 input tokens reprices the whole request at the higher tier, not just the overflow. A decision API that hands a long document or a large image set to the model is exactly the kind of call that can cross that line, which is one more thing the preview's unpublished limits leave open.

GPT-6 Luna on its own terms

Strip the API away and the model underneath is a reasoning model at the bottom of a three-rung family — below GPT-6 Sol at $2.00/$10.00 and the flagship GPT-6 Astra at $10.00/$50.00 — with a 1,050,000-token context window, a 128,000-token output ceiling, a May 18, 2026 knowledge cutoff, and reasoning effort settable across six values from none to max. It takes text and image input and returns text, and on OpenAI's own developer documentation the effort default is medium. One practical caveat from that same page, unrelated to the Decisions API but relevant if you are wiring Luna in as a fallback: on Chat Completions, function calling works only when reasoning effort is set to none. Tools on any other effort setting need the Responses API.

On quality, the independent number is Artificial Analysis's Intelligence Index at 37.3 for Luna at maximum effort, against 47.5 for GPT-6 Sol — a gap of about ten points, which is the honest way to read the price gap of twenty times on input. Neither figure is a statement about the Decisions API, whose constrained task is a different measurement entirely and has no published score.

OpenAI's developer documentation page for the gpt-6-luna model, showing the model heading with compare and playground controls, the reasoning, speed and price tiles, the '$0.1 - $0.5' price range, text and image input with text output, a 1,050,000-token context window, a 128,000-token maximum output, a May 18 2026 knowledge cutoff, the note that reasoning.effort supports none, low, medium (default), high, xhigh and max, and the note that Chat Completions supports function calling only with reasoning effort set to none.

Jev, Laya, and the "system one" framing

The Decisions API did not arrive into an empty field. TypeSafe AI's Jev, launched September 15, 2026 by Diogo Almeida and generally available since September 21, is the model that made this category legible to developers: it generates no text at all, returning typed answers with confidence values instead, at $0.042 per million input tokens with output billed at zero. The first independent testing of Jev, run by Every, put 777 judgments across 37 documents in under 0.7 seconds for roughly a quarter of a cent, and a second test timed it at a median 0.35 seconds per passage against 8.83 seconds for Claude Fable 5.1 at high effort — about 25 times faster at roughly 1/580th the cost, while catching six of seven deliberately planted defects where Fable 5 caught all seven. "Good but not perfect" was Every's verdict, and it is the right one-line reading. Laya is the third entrant in the same shape, a decision model that answers without writing a token.

Whether OpenAI built the Decisions API because of Jev is speculation. The New Stack, reporting it on launch day, called a reaction to TypeSafe "likely" and noted OpenAI probably rushed the announcement ahead of DevDay; OpenAI has said nothing of the kind, and the timing — Jev's general availability on September 21, DevDay on September 29 — is suggestive but not evidence. What is not speculation is that the two are not yet comparable on the axis buyers care about. Jev publishes a price, a per-call shape and a GA date. The Decisions API publishes none of those three.

Where it runs, and what to keep warm beside it

The Decisions API is OpenAI's own packaging. It is not a model on our catalogue and we do not route it, so if you want this specific endpoint, OpenAI is where it lives. The model it is built on is a different matter: GPT-6 Luna is a live route on OrcaRouter at OpenAI's list price with zero markup passed through — $0.10 input and $0.50 output per million tokens, with the >272K tier carried at $0.20/$0.75 — and over the seven days to September 30 it has served 3.23 billion tokens through us at a 0.18% error rate.

That matters for how you de-risk a preview. The sane way to test a constrained-decision endpoint is to keep the prompt-based path warm as the fallback, because a preview can change limits or pricing without notice and a decision sitting on a critical path needs somewhere to go. Keeping that fallback on the same key as your other models means no second contract and no code change to the fallback path: one API for 200+ models, with automatic failover across providers when one wobbles. And if your decision is one step of a longer chain, the routing DSL composes models into a single call, which is exactly the orchestrator-plus-decider pattern the Jev coverage popularised — a larger model planning, a small model choosing each step. That pattern is constructible today out of models we serve; the Decisions API itself would sit outside it until OpenAI opens it up.

Who should move, and who should wait

If your problem is a high-volume classification or routing job where a wrong branch is cheap and reversible, the preview is worth a request shape and a labelled set this week — the capability is real, the constraint is a genuine improvement over parsing prose, and a preview is the cheapest possible time to learn where its error costs land. If your decision authorises money, messages a customer, or sits between a user and a payment, nothing about this week changes your calculus: there is no published price to model, no published limit to design around, no SLA, and no word on whether you can tune the thing on your own data. Those four blanks are what "broad rollout" has to fill, and until OpenAI names them, the honest reading of the Decisions API is a promising interface with a fast vendor-stated clock and no bill attached.

OrcaRouter's model page for openai/gpt-6-luna, showing the model name and OpenAI attribution dated 2026-09-22, capability pills for vision, tools, JSON and reasoning, the description of GPT-6 Luna as the fast cost-efficient model below GPT-6 Sol, the /v1/chat/completions route, a $0.10 input and $0.50 output rate per million tokens with a 699 ms p50 time to first token, a 1M-token context with 128K maximum output, and 3234.7M tokens of traffic.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily