The Best Legal Research AI Gets 54% Right
OpenAI announced Astra for Law on September 17: its flagship model paired with a search index of more than 230 million URLs of United States law and 26 partner plugins, aimed at large firms. In its own benchmark testing the system passed the overall correctness check on 54.0% of questions, against 38.7% for the general model using web search alone. OpenAI published that number in the announcement rather than burying it, which is worth acknowledging before anything else.
What was measured
OpenAI tested the complete Astra for Law configuration on 200 United States legal research questions drawn from the private validation set of Vals AI's Legal Research Bench. The benchmark assesses how well a system finds relevant sources and passages and whether its research answers meet the evaluation criteria.
At the highest reasoning effort for both systems, Astra for Law passed the overall correctness check on 54.0% of questions, compared with 38.7% for GPT-6 Astra using web search alone. OpenAI describes that as a 40% relative improvement. On case-law-focused questions it found 24% more reference cases, and on the audited set of target passages it retrieved up to 54% more relevant passages from the correct court opinions at the same reasoning effort.
These are OpenAI's figures from OpenAI's testing. Independent outlets have repeated them rather than reproducing them, which is normal for a launch and worth stating plainly.
Four things the number does not say
| What people will read | What it actually covers |
|---|---|
| Legal AI is 54% accurate | One system, one benchmark, 200 questions |
| It can do legal work | Research, not drafting or advice |
| Better than a junior lawyer | Better than the same model with web search |
| Useful for my firm | United States law, at the most expensive setting |
The baseline row matters most. Beating a general model with web search by a wide relative margin demonstrates that the dedicated index and tooling do real work. It says nothing about how the system compares to the professional whose time you were considering replacing, because that comparison was not run.
The jurisdiction row matters most to readers here. This is a US legal corpus evaluated on US questions. Canadian firms should read it as evidence about the maturity of the field rather than as a product claim that travels, which is the same caution we set out in contract templates from AI.
Publishing 54% was the right call
A company launching a professional product could have led with the 40% relative improvement, the 24% more reference cases, or the 230 million URLs, and left the absolute number in a footnote. Most do.
Putting 54.0% in the announcement sets a standard the rest of the market now has to answer to. When your next vertical AI vendor declines to give you an absolute accuracy figure, you have a reference point for what disclosure looks like from a company with far more to lose.
That matters more this week than usual, because Epoch published an audit of 15 AI benchmarks and flagged nine as flawed. Benchmark numbers are softer than they look even when honestly reported, which cuts in both directions here.
Four questions for any vertical AI vendor
1. What is the number, on which benchmark, measured when, at what settings? All four parts. Astra for Law's 54% is at the highest reasoning effort, which is also the most expensive way to run it, and a cheaper setting produces a different number.
2. Who chose and ran the benchmark? A vendor testing itself on a third-party bench is better than a vendor testing itself on its own bench, and both are weaker than independent evaluation.
3. What is the comparison baseline? Beating a general model is a product claim. Matching a professional is a business case. Vendors frequently present the first and let buyers hear the second.
4. Does the evaluation cover my jurisdiction and my document types? The question that ends most vertical AI conversations in Canada, and the one least likely to be volunteered.
A vendor who answers all four comfortably is telling you they have done the work. One who cannot has either not measured or would rather not say, which is the distinction at the centre of AI washing.
What 54% means operationally
It means a person checks everything. Not a spot check, not a sample, everything. At roughly half correctness on a research task, unreviewed output is not a work product in any profession where being wrong has consequences.
That sounds like it removes the benefit and does not, for the same reason it did not in document extraction: verifying a found citation takes far less time than finding it. The saving is real, it is smaller than the demo implies, and it belongs in the business case as a review cost rather than being assumed away. The pricing consequence of that compression is in billable hours when AI does part of the work.
The pattern generalises well beyond law. Wherever a vendor sells a vertical AI product into a profession, ask for the absolute number. If they give you one as candid as this one, that is a reason to take them seriously rather than a reason to walk away. The related lesson about accuracy being a property of how you frame the task is in extracting data from PDFs with AI, and the broader benchmark problem in why AI benchmarks mislead.
Frequently Asked Questions
What is Astra for Law?
A legal configuration of GPT-6 Astra that OpenAI announced on September 17, 2026, pairing the model with a search index of more than 230 million URLs of United States law and 26 partner plugins. It is available to API customers, with Harvey and Legora named as builders on it, and to law firms admitted to OpenAI’s Trusted Access Program, which targets selected Am Law 200 firms.
What does the 54% figure actually measure?
OpenAI tested the complete Astra for Law setup on 200 United States legal research questions from the private validation set of Vals AI’s Legal Research Bench. At the highest reasoning effort, it passed the evaluation’s overall correctness check on 54.0% of questions, against 38.7% for GPT-6 Astra using web search alone. The benchmark measures finding relevant sources and passages and meeting the evaluation’s criteria, so it assesses research rather than drafting.
Is 54% good or bad?
Both, depending on the comparison. Against the general model with web search at 38.7%, it is a 40% relative improvement and clear evidence that a dedicated index and tooling help substantially. Against a competent professional doing the same research, it is not close. The number is best read as showing that vertical configuration works and that the field is earlier than the marketing around it suggests.
Does this apply to Canadian legal work?
Directly, no. The index OpenAI describes covers United States law, and the benchmark uses US legal research questions. Canadian sources, provincial variation and different procedural rules are not what this system was built or measured against. A Canadian firm reading the announcement should treat it as evidence about the maturity of legal AI generally rather than as a product claim that transfers.
What should I ask a vertical AI vendor about accuracy?
Four things. What is the number, on which benchmark, measured when, and at what settings. Whether the benchmark was chosen and run by the vendor or by an independent party. What the comparison baseline is, since beating a general model is a different claim from matching a professional. And whether the evaluation covers your jurisdiction and document types. A vendor who cannot answer all four has either not measured or would rather not say.
Ask for the absolute number
We test vertical AI products against your own documents and jurisdiction, establish what the accuracy claim means for your work, and price the review step into the business case.
Related Articles
Tabular AI: When AI Finally Gets Your Spreadsheets
Kimi + OpenClaw: Ultra-Long-Context Workflows for Research & Contracts
Your AI Agent Can Now Log In as You
Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.