The test that separates found from filing-grade.
Most legal AI comparisons measure the wrong thing. They ask “did the system find a relevant case” and “did it invent a fake citation.” Both questions are now easy to pass.

Not fabricating is not the same as being filing-grade.
Verified means the citation exists and is accurate. It does not mean the cited proposition is still controlling, on-stance, decided at the right stage, or currently citable.
A real, verified, on-point case that was reversed last month, or that controls the wrong department, or that decided the issue at summary judgment when you are at the pleading stage, is not a weaker answer. It is an unusable answer that looks identical to a winning one. ADR-Hard is built to separate those two things, because no human cite-checker and no similarity score can.
The fabrication era is over. The misuse era just started. Hallucinated legal citations were the 2023 problem, and they are largely solved: on this benchmark, everyone ties at zero fabrication. The next malpractice wave is real citations used wrongly. So we built the test that catches it, and we are inviting the field, ourselves first, to be measured by it.
Internal measurement, June 2026. Not independently audited. Methodology published.
Everyone finds the neighborhood. Only one returns the case you can file.
| Measure | Legawrite.AI | Frontier model + agentic scaffolding + case-law access | Strictly grounded topical retrieval | Raw frontier model |
|---|---|---|---|---|
| Dispositive Recall at filing (composite) | 88% | 24% | 19% | 9% |
| Fabrication rate | 0.00% | 2% | 0.00% | 22% |
| Raw topical recall (did it find the neighborhood) | 94% | 89% | 84% | 58% |
| Controlling-status recall (the assassin row) | 96% | 33% | 22% | 12% |
Architectures, not brands. Dispositive Recall requires the right proposition, the right stance, the right procedural stage, confirmed good-law status and current citability, all at once; constraints compose multiplicatively. A system can clear any one filter by luck. Almost none clear all of them together.
How a real, correctly quoted case loses.
The department-split trap
Two intermediate appellate departments in the same state disagree. Topical retrieval returns the better-known rule from the wrong department. Dispositive recall requires the controlling line for your forum.
The controlling-status trap
The case is real and on point. The high court reversed it, or review was granted and the opinion may no longer be cited. The document still ranks first by similarity.
The stage trap
A holding on a summary-judgment record, offered as authority on a motion to dismiss. Same words, wrong standard, motion lost.
How we designed the test, and why.
- Design principle
- Select for architectural failure: every question is one a document-level index cannot answer without luck.
- Answers
- Propositional, not topical. A correct answer is a proposition with the right stance, stage, status and citability.
- Corpus
- Real New York case law, with California depublication and circuit-split variants as illustrations.
- Scoring
- Each dimension maps to a typed field; constraints compose multiplicatively. The rubric and the scoring function are published with the harness.
- Integrity tie
- Fabrication is held constant and conceded: a strictly grounded system also reaches 0.00%. The benchmark measures what comes after.
- Status
- Internal measurement, June 2026. Not independently audited. Question count, run date and confidence intervals ship with the harness.
Methodology first, numbers second.
The rubric, the question set and the scoring function are written to be published with the harness, so the measurement can be checked rather than believed. That paper is in preparation. Until it is out, the figures on this page stay labeled for what they are: an internal run, dated June 2026, not independently audited.
The benchmark re-versions when the law moves.
A question set built on live law goes stale. Each re-version will be written up on the news page, with what changed in the law and what changed in the questions.
Concede first, then turn.
Why we concede the 0.00% fabrication tie.
A strictly grounded retrieval product, built on a public corpus with citations verified against primary sources, has a fabrication rate of 0.00%. So do we. Its citations are real because they are pulled from a real corpus and checked against the source. Saying otherwise would be false, and pretending fabrication is the wedge would be a bet on the reader not checking.
That correction reframed the entire benchmark, and improved it. The wedge is not fabrication. The wedge is this: verified means the citation exists and is accurate. It does not mean the cited proposition is still controlling, on-stance, decided at the right stage, or currently citable. We tie the field at zero fabrication. Not fabricating is the floor now. Being filing-grade is the ceiling, and that gap is the whole game.
Bring a motion you’re nervous about.
In a demo we run it through and walk the five ways it could quietly lose, on your own authorities. The line is our standing invitation, not a product you sign up for here.