A generated hero card titled 'Introducing Astra for Law', subtitled 'OpenAI's legal configuration of GPT-6 Astra', with a date line reading 'Astra for Law announced September 17, 2026 · GPT-6 Astra released September 3, 2026', above three cards labelled 'Legal Search Index — 230M+ U.S. law URLs', 'Corpus — CourtListener, Free Law Project' and 'API — gpt-6-astra-law, coming soon', with the footer 'Benchmark figures vendor-reported; not independently reproduced.' and the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

Introducing Astra for Law: OpenAI Puts a Legal Search Index Behind GPT-6 Astra

Author

Elias Hawthorne

Date Published

Back to all posts

Astra for Law was announced on September 17, 2026 — a legal-work configuration built on GPT-6 Astra, the flagship model the company shipped two weeks earlier, on September 3. It is not a new model. It is the same weights behind a different front door: a dedicated legal search index, legal-specific instructions, and a set of tools aimed at United States case-law research. That distinction matters more than the launch framing suggests, because every number in the announcement measures the configuration against the base model rather than a new capability — and the configuration is doing nearly all of the work.

The headline claim is that Astra for Law passed the overall correctness check on 54.0% of 200 U.S. legal research questions drawn from the private validation set of Vals AI's Legal Research Bench, against 38.7% for GPT-6 Astra using web search alone — a 40% relative improvement, as OpenAI tells it. Those are vendor-reported figures on a private split, and no independent lab has reproduced them. What is independently visible is more useful: Vals AI's own public comparison lists GPT-6 Astra at 39.42% ±3.40 on that benchmark, which is essentially the baseline OpenAI is measuring against. The gap is coming from the legal search index and the instructions, not from the model having become a better lawyer.

Here is what shipped, what has actually been measured, and what is still missing.

What Astra for Law actually is

Three things were added on top of GPT-6 Astra, and none of them is a weight change. In a media briefing, Jason Boehmig — the Ironclad cofounder OpenAI hired in June 2026 — and engineer Sherwin Wu described the release as a configuration of domain instructions, settings such as response length, and legal-industry tooling that is expected to expand over time.

Legal Search Index — a retrieval layer over U.S. case law, statutes, regulations, court rules and administrative decisions spanning more than 230 million URLs, with sources added daily.

Corpus provenance — the index draws on CourtListener, the library maintained by the nonprofit Free Law Project, which OpenAI says covers more than 99.9% of published U.S. precedential case law. It complements licensed content from providers including Thomson Reuters.

Legal instructions — a set of domain instructions for applying research to facts, developing arguments or deal terms, and surfacing weaknesses or uncertainty in a position.

A generated single-column scoreboard titled 'Astra for Law — the scoreboard' with six rows reading 'Underlying model: GPT-6 Astra', 'Legal Search Index: 230M+ U.S. law URLs', 'Corpus: CourtListener, 99.9% of precedential case law', 'Availability: Trusted Access, Am Law 200 only', 'API model id: gpt-6-astra-law, not yet open' and 'Headline benchmark: 54.0% vs 38.7% vendor-reported', above the footer 'All figures OpenAI-reported on a private validation set; no independent reproduction.' with the OrcaRouter logo composited in the bottom-right corner.

The demonstration in the briefing was concrete rather than abstract: Wu drafted a motion to dismiss for a hospital facing a disability-discrimination and retaliation claim, worked a stockholder challenge to an asset sale over a holiday weekend, and handled a sales-commitment dispute. None of those is a new capability for a frontier model. What is new is that the retrieval half of the job no longer routes through general web search.

The benchmark, read carefully

Four figures came out of the launch, all attributed to OpenAI, all on the same private validation set, and all worth separating from one another:

Overall correctness — 54.0% for Astra for Law versus 38.7% for GPT-6 Astra with web search, at the highest reasoning effort.

Reference cases found — 24% more on case-law-focused questions, again at highest reasoning effort.

Relevant passages retrieved — up to 54% more from correct court opinions, at the same reasoning effort as the baseline.

Answer length — roughly twice as long on average across the tested reasoning settings.

That last figure is the one to hold next to the first. An overall-correctness check that rewards coverage is easier to pass with a longer answer, and part of the 15.3-point gain is plausibly the retrieval layer simply putting more material in front of the model. This is not a criticism of the result — finding 24% more reference cases is a real, separately stated improvement — but it does mean the composite number and the retrieval numbers should be read together rather than the composite alone.

The private-split problem is the bigger one. A validation set that OpenAI holds and does not publish cannot be re-run by anyone, and the same benchmark's public leaderboard tells a flatter story: Vals AI's own comparison page lists GPT-6 Astra at 39.42% ±3.40 on Legal Research Bench, with Gemini 3.8 Flash at 38.94% ±3.39. Whatever produced the jump to 54.0% is not visible on the public board, because the public board is scoring the base model without the legal index attached.

A screenshot of the Vals AI 'Compare Models' leaderboard captured September 18, 2026, comparing Gemini 3.8 Flash and GPT-6 Astra: Legal Research Bench (Agentic US legal research) shows Gemini 3.8 Flash 38.94% ±3.39 against GPT-6 Astra 39.42% ±3.40; Vals Index shows 62.25% against 66.61%; and the table also lists Finance Agent (V2) 61.44% vs 53.54%, Tax Agent Bench 66.77% vs 63.34%, MedCode 48.13% vs 48.49%, Terminal-Bench Science 8.57% vs 65.71%, Code Migration 36.55% vs 67.74%, Terminal-Bench 4.0 13.13% vs 57.07% and Vibe Code Bench v1.1 78.65% vs 89.59%, with a Vals Index date of 2026-09-11.

The Legora tie-out, and the number inside it

The most substantive third-party evidence in the launch coverage is a workflow test from Legora, whose legal engineer Percevale Perks ran a financial-statement tie-out — matching draft-account figures against a trial balance, a consolidation schedule and prior-year accounts. Legora's agent completed the tie-out across 41 documents in a single run, in minutes, logging each check and producing a line-by-line record for review. It found all four errors Legora had planted in the accounts, including a £500,000 gap hidden in the revenue note, kept every check the earlier model had passed, and added roughly fifty more.

That is a genuinely good result on a genuinely hard document set. It is also where the most useful number in the entire launch cycle is buried. On Legora's Benchmark for Agentic Reasoning, which covers real end-to-end legal tasks, the improvement on the financial-statement workflow was about 40% — and the average improvement across all BAR tasks was about 3%. Both numbers are Legora's, both describe the same model, and only one of them travelled. If you are evaluating this for a firm, the 3% is the figure to plan against and the 40% is the figure to hope for.

Independent testing points at a different failure mode

Two independent evaluations are worth putting next to the launch numbers, with the caveat that both are small and single-source.

HAQQ ran GPT-6 Astra against 41 real legal questions across 20 practice areas at medium reasoning. Astra led on legal substance at 17.63 out of 20 — by less than a point. It cost $0.2622 per answer, roughly eleven times the cost per answer of GPT-5.6 Luna Pro for under a point of composite gain. HAQQ's own read is sharp: the advantage is not that the model is smarter, it is that it stops writing. The same test found that turning reasoning effort from low to high multiplied cost by about 1.5× and latency by about 1.8× while reducing judged quality by 0.17 points — a cost lever dressed as a quality lever, and a reason to pin the effort setting rather than leave it on max.

The finding that should travel furthest is not about cost. Across non-U.S. jurisdictions, all three models tested cited the correct provision while misstating what it says. Astra was legally correct on 66.7% of those items; Claude Opus 5 was correct on 100%. A citation checker passes that answer, because the citation exists and is the right one. A client does not. This is the failure mode the legal search index is least equipped to address, since the index is U.S.-only and the error is in the reading, not the retrieval.

What you can actually get, and when

Astra for Law is not generally available. Access starts through a new Trusted Access program aimed at Am Law 200 firms, delivered through ChatGPT and Codex. The API is described as coming soon under the model id gpt-6-astra-law, appearing in the model picker as "GPT-6 Astra Law"; Harvey and Legora are named as API customers who will build on it. No price for the legal configuration has been published, which is the single largest open question for anyone budgeting around it.

Trusted Access carries two data terms: Zero Data Retention on the API, and exclusion of ChatGPT Enterprise usage from human review by default. Asked in the briefing about exceptions, Boehmig did not claim the policy was absolute: "It's definitely more complicated than the one sentence," he said, noting the full agreement runs about 30 pages, and adding that OpenAI is confident those 30 pages meet the need. OpenAI says it is working with Latham & Watkins on information permissions, ethical walls, client instructions and firm oversight.

The ecosystem piece is larger than the model piece: 26 vendor plugins including Thomson Reuters, Harvey, Legora, iManage, Relativity, Clio, Intapp and DeepJudge; nine community plugins from legal-engineering groups; and 47 custom skills built by legal power users. ChatGPT for Word also reached general availability for proofreading, suggested edits and formatting flags. Forward-deployed OpenAI engineers worked with Sullivan & Cromwell on an agreement analyzer, Ropes & Gray on M&A diligence, and Cooley on GO Public.

For context on how crowded this lane now is: Anthropic shipped Claude for Legal in May 2026, three months before Astra for Law existed.

The risks no benchmark on that page captures

Nothing in the launch materials addresses the two governance questions a firm's general counsel will ask first. As of September 2026 no SOC 2, ISO or HIPAA certification had been published for GPT-6 Astra, and there is no published detail on firm-specific data siloing, local deployment or data residency. The absence is not evidence of failure — it is simply undocumented, which is a different problem and one that falls on the firm to resolve before privileged material goes in.

On the ethics side, the operative rules are ABA Model Rule 1.6 on confidentiality, which governs what a firm may share with a third-party provider and may require updated disclosures and informed consent, and Rule 5.3 on supervision of nonlawyer assistance, which is where AI output lands. The reason both are live rather than theoretical is Mata v. Avianca, where attorneys were sanctioned for filing nonexistent AI-generated cases. A legal search index over real opinions makes that specific failure less likely. It does not make it impossible, and it does nothing about the misread-provision failure HAQQ documented.

If you want to build on this before the API opens

The honest position today is that gpt-6-astra-law is not in anyone's API, ours included — OpenAI has said coming soon and named no date. What is live is the base model: openai/gpt-6-astra sits on OrcaRouter at 0% markup, meaning OpenAI's provider list price is passed through unchanged, so a vendor price change on that model is reflected the same day. One key reaches more than 200 models, with automatic failover across providers and a routing DSL for composing several models into a single call.

A screenshot of OpenAI's developer documentation page for GPT-6 Astra, captured September 18, 2026, showing the model badge labelled 'Default' with the line 'Our most capable model, built for the hardest end-to-end work', a 1,050,000 context window, 128,000 max output tokens, an Apr 30 2026 knowledge cutoff, reasoning token support, and the pricing card reading Input $10.00, Cached input $1.00, Cache writes $12.50, Output $50.00 per 1M tokens, with the notes that prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request, that cache writes are billed at 1.25x the uncached input token rate, and that Batch and Flex are priced at 50% of Standard rates while Fast mode is priced at 2x.

The practical consequence for a legal-engineering team is a split. Everything in the Astra for Law workflow that does not depend on the Legal Search Index — drafting from facts you supply, restructuring a clause set, formatting to a firm template, generating a review checklist — you can prototype today on base GPT-6 Astra, and swap the model id when the law variant opens without changing your integration. Everything that depends on the index you cannot prototype, because the index is the part that is not for sale yet. Two caveats worth stating plainly: a routed request is not automatically covered by OpenAI's Zero Data Retention approval, which is granted per organisation, so a firm that needs ZDR should confirm coverage on the specific route it uses; and a router is an extra hop in the data path, which is a real cost for privileged material even when it buys you failover.

What to watch

Three things, in order of how much they will change the picture. The API date for gpt-6-astra-law, since that is the difference between a pilot and a product. The price, since the base model's $10 per million input and $50 per million output tokens carry a long-context reprice above 272,000 input tokens that legal document sets will cross routinely — and if the legal configuration prices at a premium to that, the economics of a 41-document tie-out change materially. And an independent re-run of those 200 questions on a split somebody else controls. Until that exists, 54.0% is OpenAI's number about OpenAI's configuration, and the defensible claim is the narrower one: a U.S. case-law index plus legal instructions produces materially better retrieval than general web search on U.S. legal research questions. That claim is worth acting on. The rest is worth waiting for.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily