skip to content
The Weighted Average

Enterprise AI & Work

Astra for Law's 54% Score Uses a Different Test Split

Astra for Law reports 54% on 200 private-validation questions. That is 48.4% of Vals AI's dataset, not the test split used by its leaderboard.

Pale columns frame the entrance to the Supreme Court of Nevada
Pale columns frame the entrance to the Supreme Court of Nevada. Photograph by Quilia

Law-firm technology teams should evaluate OpenAI’s September 17 Astra for Law launch as a specialized research configuration, not insert its 54% correctness score into an unrelated leaderboard. Its reported evaluation uses the 200-question private-validation split—48.4% of Legal Research Bench’s full dataset, and a different set from the benchmark owner’s published test results.

A familiar benchmark name can hide a different exam

The dataset distinction is directly auditable. SiliconANGLE reports that OpenAI evaluated 200 private-validation questions. Vals AI describes 413 questions in total: five public samples, 200 private-validation questions, and 208 test questions. Dividing the launch’s 200 by the benchmark owner’s 413 gives 48.4% after rounding. That measures the share of the dataset represented by the chosen split, not the share of legal work the system can handle.

Vals says the results on its own page use only the 208-question Test set. Its private-validation set is available for license. There is nothing inherently improper about evaluating on that split, and the announcement identifies it. The purchasing error would be to compare scores across the two sets without acknowledging the changed questions, tool configuration, and evaluation procedure. Similar-sized exams are not the same exam.

Within OpenAI’s reported comparison, the specialized configuration passes overall correctness on 54%, versus 38.7% for base GPT-6 Astra with web search, both at the highest reasoning effort. LawSites’ account of the product briefing also reports those figures and the configuration’s longer answers. That supports a narrower conclusion: the added legal context and tools improved the vendor’s result on its selected evaluation. It does not establish superiority over a different provider’s published test score.

The product distinction matters as much as the benchmark distinction. Astra for Law is a configuration of GPT-6 Astra with a legal search index and legal-specific instructions, not a separately announced new base model. Launch coverage describes an index spanning more than 230 million URLs across U.S. case law, statutes, regulations, court rules, and administrative decisions. That is an access-and-retrieval proposition. A larger searchable corpus still needs a test of whether the relevant authority is found and correctly applied.

There is independently attributable infrastructure behind the launch. Free Law Project confirms that its CourtListener plugin is available in ChatGPT, exposing case law, federal filings, citation networks, transcripts, and alerts. It says results link back to underlying documents and that the service works with MCP-speaking clients. This verifies a concrete route to inspect primary legal material; it does not independently reproduce Astra for Law’s benchmark result.

Free Law Project is also temporarily doubling API rate limits through October 1. Treat that as a trial condition, not a permanent capacity entitlement. If an evaluation succeeds during the promotion, confirm that its sustained request pattern remains feasible afterward or determine the membership terms needed for higher limits. A pilot can otherwise mistake temporary data access headroom for the normal operating envelope.

Buy the review process before trusting the score

The appropriate initial buyer is a legal team with access to the program and a defined research workflow, not anyone seeking an autonomous substitute for professional judgment. Selected firms receive access through ChatGPT and Codex under Trusted Access; launch reporting says an API version is coming later without a disclosed date or price. A platform team cannot responsibly commit a production integration timetable or a per-matter saving from an API whose commercial terms remain unresolved.

Budget the work the configuration does not remove: permission design, source inspection, lawyer review, and the handling of uncertain or conflicting authority. Start with matters for which reviewers can establish the expected reasoning and supporting documents. Assess whether the system misses controlling authority, applies the wrong jurisdiction, or omits a condition that changes the answer. These are recommended test dimensions, not a claim that this publication has measured the product’s legal accuracy.

The benchmark’s own methodology explains why partial success is inadequate. Vals distinguishes weighted partial credit from all-pass correctness, where every required rubric item must pass. It says its grading uses an LLM judge and reports human-validation work. Buyers should therefore retain the rubric and examine disagreements rather than use a single percentage as a release gate. A fluent answer can preserve most of an argument while losing the condition that matters to the client.

Our MLPerf analysis separates benchmark qualification from application acceptance. The same separation applies here, with an additional split warning. Do not pool results from different question sets or treat a vendor’s configured workflow as a test of the base model alone. If a firm’s own evaluation changes retrieval, instructions, or available documents, record those changes so later comparisons remain interpretable.

Data terms require a full reading too. Launch accounts describe zero data retention for eligible API use and exclusion of ChatGPT Enterprise usage from human review by default. LawSites reports that OpenAI characterized the governing agreement as substantially more detailed than that summary. Procurement should resolve client instructions, ethical walls, retention, and permitted access in the actual contract before uploading confidential material. A headline privacy assurance is not an organization-specific authorization.

The strongest case for adoption is that maintaining a reliable legal search layer is itself expensive work, and the product may let firms focus on their own precedents and review standards. The strongest counterpoint is that a stronger vendor evaluation may not improve the firm’s particular practice area, document mix, or reviewer burden. Both possibilities deserve a controlled pilot. Neither is settled by a comparison with leaderboard scores that were produced on another split.

Today’s Crusoe lead argues for acceptance evidence rather than financing headlines. For Astra for Law, the acceptance evidence is a reproducible, firm-relevant research result with inspectable sources and known review cost. Expand after that evidence improves under the required permissions and commercial terms. Until then, treat 54% as an attributed result on a specified evaluation—not a license to skip the remaining judgment.

Sources