Legal AI benchmarks in 2026: what has actually been measured
Every vendor in this market claims accuracy. Almost none publish a methodology. This page collects the studies that exist, names who funded each one, and says plainly what they do and do not support.
The short answer: there is no current, independent, published benchmark for AI contract review. The most rigorous work in adjacent legal AI is a Stanford study of research tools, and its findings are considerably less flattering than vendor marketing.
The Stanford study, and why it is the most useful thing here
In May 2024, researchers at Stanford RegLab and the Institute for Human-Centered AI published the first preregistered empirical evaluation of commercial legal AI tools. Magesh, Surani, Dahl, Suzgun, Manning and Ho tested 202 legal queries, hand-scored by legal experts.
| Tool | Incorrect or misgrounded responses |
|---|---|
| Lexis+ AI | More than 17% of queries |
| Westlaw AI-Assisted Research | Approximately 33% of queries |
| General-purpose chatbots (ChatGPT, Claude, Llama) | 58% to 80% of queries |
Two things make this study worth more than everything else on this page. It was preregistered, which means the researchers committed to what they would measure before they measured it. And it was independent of the vendors it tested, which is the opposite of the usual arrangement.
The researchers noted that LexisNexis had claimed "100% hallucination-free linked legal citations" and that Thomson Reuters said its tools "avoid hallucinations by relying on trusted content". Their conclusion was that providers' claims are overstated. Both tools appear on this site, and both are still worth buying for the right buyer; the point is that the vendor's own accuracy claim is not evidence.
The contract review benchmark everyone cites
The study behind most "AI beats lawyers" headlines is LawGeex's 2018 NDA experiment. Twenty experienced lawyers and the LawGeex system reviewed five NDAs comprising 153 clauses, with input from law professors at Stanford, Duke and the University of Southern California.
| Measure | LawGeex AI | Lawyers |
|---|---|---|
| Average accuracy at surfacing risk | 94% | 85% |
| Best individual result | 94% | 94% |
| Worst individual result | n/a | 67% |
| Time to complete | 26 seconds | 92 minutes on average |
Three caveats, all of which matter more than the numbers. The study is from 2018, which in this field is several product generations ago. It was run and funded by LawGeex, testing LawGeex. And the task was NDA clause identification against a known checklist, which is the most automatable thing a contract lawyer does rather than a representative sample of the job.
None of that makes it worthless. It is a real experiment with a published method, which puts it ahead of most vendor claims. It does mean that citing it in 2026 as evidence that current tools beat lawyers is citing a vendor's own eight-year-old marketing study.
What no study currently supports
- That any specific contract review tool on the market today is more accurate than another. No independent head-to-head exists.
- That AI review is faster in practice by the ratios the 2018 study reported. Twenty-six seconds against 92 minutes measured a clause-spotting task, not a negotiation.
- That purpose-built legal tools are hallucination-free. The Stanford work found the two market leaders in legal research wrong on 17% and 33% of queries respectively.
- That accuracy figures quoted by any vendor are comparable with each other. Without a shared task and a shared scorer they are not measurements, they are claims.
How to read a vendor's accuracy claim
- Ask what was measured. "Accuracy" on clause identification, risk flagging and drafting quality are three different tasks with three different difficulties.
- Ask whose documents. Performance on a vendor's curated set predicts very little about performance on your scanned, oddly drafted, twelve-year-old supplier agreement.
- Ask who scored it, and whether they knew which system produced which answer.
- Ask who funded it. Most published legal AI benchmarks are produced by a party with an interest in the result, including the one above.
- Ask for the methodology in writing. A number without one is not a benchmark.
Why we do not score on accuracy
Our guides score tools on fit, scope and price transparency, and deliberately not on accuracy. Scoring accuracy would mean either repeating vendor claims as though they were measurements, or running our own benchmark. The first is dishonest and the second is beyond what this site can do properly.
If an independent, preregistered contract review benchmark is published, we will cover it and say who paid for it. Until then, the most reliable evidence available to a buyer is a trial on their own worst documents, which is why every guide here says to run one.
Sources
- Magesh, Surani, Dahl, Suzgun, Manning and Ho, "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools", Stanford RegLab and Stanford HAI, May 2024.
- LawGeex, "Comparing the Performance of Artificial Intelligence to Human Lawyers in the Review of Standard Business Contracts", February 2018, produced by LawGeex.