A generated hero card titled 'Astra for Law vs GPT-6 Astra', subtitled 'Same weights — one adds a legal search index', with a date line reading 'Astra for Law announced September 17, 2026 · GPT-6 Astra released September 3, 2026', above two cards: left 'Astra for Law — legal index + instructions, price not disclosed' and right 'GPT-6 Astra — web search, $10 in / $50 out per 1M', with the footer 'Vendor figures unreproduced; GPT-6 Astra public board score per Vals AI.' and the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

Astra for Law vs GPT-6 Astra: Same Weights, One Extra Search Index

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Astra for Law and GPT-6 Astra are not two models, and framing the choice as a head-to-head is the first thing to get right. Astra for Law is GPT-6 Astra with a United States legal search index bolted to its retrieval path, a set of legal drafting instructions layered on top, and legal tooling attached. The weights are identical. So the real question is not which model is smarter — it is whether the legal search index is worth paying a premium for, and on the evidence available so far the answer depends almost entirely on whether your work is case-law retrieval or document review.

That matters because OpenAI has not published a price for the legal configuration, the API for it is not open, and the only head-to-head measurement against the base model comes from OpenAI itself on a private validation set. This page works through what is actually known, what the independent numbers say, and which side of the comparison a given legal workload belongs on.

What separates them, precisely

Three things were added, and they are not equally hard to replicate.

Legal Search Index — retrieval over U.S. case law, statutes, regulations, court rules and administrative decisions, more than 230 million URLs, sources added daily. This is the part you cannot build from a prompt.

Corpus — built on CourtListener, the Free Law Project library covering more than 99.9% of published U.S. precedential case law, complemented by licensed content from providers including Thomson Reuters. The public half is free; the licensed half and the daily refresh are not something an individual firm reproduces.

Instructions and settings — legal-analysis and legal-writing instructions, plus settings such as response length. These are prompt-level and configuration-level. A competent legal-engineering team can approximate a meaningful fraction of them against the base model today.

A generated two-column scoreboard titled 'Astra for Law vs GPT-6 Astra — the scoreboard'. Left column 'Astra for Law' rows: 'Retrieval: Legal Search Index, 230M+ URLs', 'Instructions: legal-specific', 'Headline score: 54.0% vendor-reported', 'Availability: Trusted Access, Am Law 200', 'API model id: gpt-6-astra-law, coming soon', 'Price: not disclosed'. Right column 'GPT-6 Astra' rows: 'Retrieval: general web search', 'Instructions: general', 'Public board score: 39.42% on Vals AI', 'Availability: generally available now', 'API model id: gpt-6-astra', 'Price: $10.00 in / $50.00 out per 1M'. Footer reads 'Astra for Law figures OpenAI-reported on a private split; GPT-6 Astra public board figure per Vals AI.', with the OrcaRouter logo composited in the bottom-right corner.

On GPT-6 Astra alone, retrieval goes through general web search. That is the whole of the difference in the research path, and it is the axis every number below is measuring.

The only head-to-head that exists

OpenAI tested both sides on 200 U.S. legal research questions from the private validation set of Vals AI's Legal Research Bench. At the highest reasoning effort, Astra for Law passed the overall correctness check on 54.0% of questions against 38.7% for GPT-6 Astra with web search — 15.3 points, described as a 40% relative improvement. Supporting figures from the same run: 24% more reference cases found on case-law questions, up to 54% more relevant passages retrieved from correct court opinions at equal reasoning effort, and answers averaging roughly twice the length.

Three cautions belong on those numbers, and none of them is a reason to dismiss them. First, they are OpenAI's, on a split OpenAI holds, and nothing independent has reproduced them. Second, the doubled answer length is worth holding beside the composite score, because a correctness check that rewards coverage is easier to pass with a longer answer — the retrieval figures (24% more cases, 54% more passages) are the cleaner evidence of what the index actually adds. Third, the public version of the same benchmark does not show the jump: Vals AI's own comparison page lists GPT-6 Astra at 39.42% ±3.40 on Legal Research Bench, with Gemini 3.8 Flash at 38.94% ±3.39. That is the base model, without the index, landing essentially where OpenAI's baseline sits.

A screenshot of the Vals AI 'Compare Models' leaderboard captured September 18, 2026, comparing Gemini 3.8 Flash and GPT-6 Astra: Legal Research Bench (Agentic US legal research) shows Gemini 3.8 Flash 38.94% ±3.39 against GPT-6 Astra 39.42% ±3.40; Vals Index shows 62.25% against 66.61%; and the table also lists Finance Agent (V2) 61.44% vs 53.54%, Tax Agent Bench 66.77% vs 63.34%, MedCode 48.13% vs 48.49%, Terminal-Bench Science 8.57% vs 65.71%, Code Migration 36.55% vs 67.74%, Terminal-Bench 4.0 13.13% vs 57.07% and Vibe Code Bench v1.1 78.65% vs 89.59%, with a Vals Index date of 2026-09-11.

Two independent readings complicate the picture further. Legora ran a financial-statement tie-out across 41 documents in one pass and found all four planted errors, including a £500,000 gap in the revenue note, adding roughly fifty checks the earlier model had missed — about a 40% improvement on that workflow. On Legora's Benchmark for Agentic Reasoning overall, however, the average improvement across all tasks was about 3%. Meanwhile HAQQ's 41-question test across 20 practice areas put Astra's legal-substance lead at 17.63 out of 20, less than a point clear of cheaper models, at $0.2622 per answer. Read together, the honest summary is that the configuration's edge is large and repeatable on retrieval-heavy U.S. case-law work, thin on general legal reasoning, and largely untested on non-U.S. law — where the same test found all three models citing the correct provision while misstating what it says.

Astra for Law has no published price. GPT-6 Astra's rate card is public, and it is the number the comparison has to be built on, because the legal configuration is the same model plus retrieval and cannot plausibly cost less than the model it wraps.

Standard — $10.00 per million input tokens, $50.00 per million output tokens.

Cached input — $1.00 per million, a 90% discount, and the lever that matters most for a firm reusing standing instructions and precedent boilerplate.

Cache writes — $12.50 per million, billed at 1.25× the uncached input rate; caching pays from the second reuse onward.

Above 272,000 input tokens — 2× input and cache rates and 1.5× output, applied to the entire request, not just the tokens above the threshold. In practice that means $20.00 input and $75.00 output per million for any long request.

Batch — 50% of Standard. Fast mode — 2× Standard, and unavailable for GPT-6 Astra with EU data residency.

That 272K threshold is not a footnote for this comparison; it is the centre of it. A 41-document tie-out, a full deposition transcript set, or a multi-year deal file crosses 272,000 tokens routinely, which means the realistic legal rate is the long-context row — $20.00 in and $75.00 out — not the headline $10/$50. For a firm sizing a pilot, the difference between those two rows is the difference between a project that fits a budget and one that does not, and it is a reason to measure your actual per-matter token volume before committing rather than after.

Two more cost notes from independent testing. HAQQ measured $0.2622 per answer on Astra, roughly eleven times the cost per answer of GPT-5.6 Luna Pro for under a point of composite gain — but also noted that the gap is partly because Astra writes about 3.5× fewer tokens than Claude Opus 5, making it cheaper per answer than Opus 5 for that workload. And that same test found the reasoning-effort dial behaves as a cost lever rather than a quality lever: low to high multiplied cost by about 1.5× and latency by about 1.8× while judged quality fell 0.17 points. Pin the effort setting; do not leave it at maximum by default.

When the base model is the better buy

The case for GPT-6 Astra alone is stronger than the launch framing implies, and it rests on three observations.

• Your work is not U.S. case-law retrieval. The index is U.S.-only. For contract drafting, clause restructuring, intake, chronology building, or summarising documents you already have in hand, the thing Astra for Law adds is largely inert.

• You already pay for primary legal research. A firm with Westlaw or LexisNexis subscriptions is buying the same precedential coverage twice if it also pays a premium for a retrieval layer it cannot audit.

• You want the retrieval path you control. Feeding a curated document set into base GPT-6 Astra gives you citation provenance you can inspect, and no dependency on a vendor's index refresh.

There is third-party evidence that the base model is not the weak link. On Mercor's APEX-Agents corporate-lawyer leaderboard, a GPT-6 Astra configuration scored 73.4% Pass@1 across 160 tasks — a third-party result on a legal agent workload, sitting alongside the 3% average gain Legora measured across its own benchmark suite. Neither number is about the search index, and both suggest the underlying model is already doing most of the legal work.

When the configuration earns its premium

The case for Astra for Law is narrower and, where it applies, decisive. If your workload is finding and applying U.S. authority — building a brief's case list, checking whether a position survives a line of decisions, running a fifty-state regulatory survey — then 24% more reference cases and up to 54% more relevant passages is not a marginal improvement, it is the difference between an assistant that misses authority and one that does not. The Legora tie-out is the best available illustration of the same effect on documents rather than case law: cross-referencing 41 files in one pass is exactly the long-context, high-volume comparison work where more retrieved context compounds.

The honest caveat on both: pricing is undisclosed, access is gated to Am Law 200 firms through the Trusted Access program, and no independent lab has re-run the head-to-head. You are being asked to commit to a premium on the strength of a vendor benchmark and one partner's workflow test.

Running both through one endpoint

The routing question here is a sequencing problem rather than a choice. gpt-6-astra-law is not in any API yet, ours included — OpenAI has said coming soon and named no date — so today the only half of this comparison you can actually call is the base model. What is available now is openai/gpt-6-astra on OrcaRouter at 0% markup, with the provider's list price passed through unchanged, so the $10/$50 and $20/$75 rows above are what you pay and a vendor price change lands the same day. More than 200 models sit behind one key, with automatic failover across providers.

A screenshot of the OrcaRouter catalogue page for GPT-6 Astra, captured September 18, 2026, showing the listing openai/gpt-6-astra from provider OpenAI with a 1M token context, 128K max output, input of text + image + file with text output, p50 time to first token of 6.01 s and p95 of 10.00 s, $10.00 input and $50.00 output per 1M tokens, 162.0M tokens of routed traffic over seven days, and an OpenAI-compatible Python sample using base_url https://api.orcarouter.ai/v1.

The routing DSL is the piece that maps onto this specific decision. Because the delta between the two products is a retrieval step, you can compose it yourself — route the base model with your own document retrieval in one call and compare the output against the same question sent with web search only. That is a cheap way to measure how much of the 54.0%-versus-38.7% gap your own corpus reproduces, before you commit to a gated premium product. Model fusion, which runs a panel of models over one prompt, covers the other half: if your concern is that a single model misreads a provision, as HAQQ found on non-U.S. law, a panel surfaces the disagreement instead of hiding it. Because routing keeps both sides on one key, switching the model id when gpt-6-astra-law opens is a config change, not a re-integration.

One caveat worth stating rather than burying: a routed request is not automatically covered by OpenAI's Zero Data Retention approval, which is granted per organisation, so a firm that needs ZDR should confirm coverage on the route it uses. And a router is one more hop in the data path — a real cost for privileged material, even when it buys failover.

Where this lands

If you are choosing today, choose the base model. Astra for Law cannot be bought yet, its premium is unknown, and the independent evidence says GPT-6 Astra already performs at a high level on legal agent workloads without the index. Build on openai/gpt-6-astra now, measure your own retrieval gap, and let the token accounting tell you whether you cross 272K often enough to care about the long-context row.

Then revisit it when three things are known: the price of gpt-6-astra-law, the shape of its API terms, and an independent re-run of the 200-question head-to-head on a split somebody other than OpenAI controls. If the price lands near the base rate and the retrieval gains hold up outside OpenAI's validation set, the configuration becomes the obvious default for U.S. research-heavy work and the base model remains right for everything else. Until then, the defensible claim is the narrow one — a U.S. case-law index plus legal instructions retrieves materially better than general web search on U.S. legal questions. That is worth paying something for. How much is still an open question.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily