LitigationOS
ADR-Hard · Adversarial Doctrinal Retrieval, hard tier

The test that separates found from filing-grade.

Most legal AI comparisons measure the wrong thing. They ask “did the system find a relevant case” and “did it invent a fake citation.” Both questions are now easy to pass.

A rider on a light-trail motorcycle racing through a tunnel
What it measures

Not fabricating is not the same as being filing-grade.

Verified means the citation exists and is accurate. It does not mean the cited proposition is still controlling, on-stance, decided at the right stage, or currently citable.

A real, verified, on-point case that was reversed last month, or that controls the wrong department, or that decided the issue at summary judgment when you are at the pleading stage, is not a weaker answer. It is an unusable answer that looks identical to a winning one. ADR-Hard is built to separate those two things, because no human cite-checker and no similarity score can.

The fabrication era is over. The misuse era just started. Hallucinated legal citations were the 2023 problem, and they are largely solved: on this benchmark, everyone ties at zero fabrication. The next malpractice wave is real citations used wrongly. So we built the test that catches it, and we are inviting the field, ourselves first, to be measured by it.

88%Dispositive Recall at filingLegawrite.AI, typed doctrinal retrieval, model out of the citation path
24%frontier model with agentic scaffolding and case-law access
19%strictly grounded topical retrieval
9%raw frontier model, no corpus

Internal measurement, June 2026. Not independently audited. Methodology published.

The leaderboard

Everyone finds the neighborhood. Only one returns the case you can file.

MeasureLegawrite.AIFrontier model + agentic scaffolding + case-law accessStrictly grounded topical retrievalRaw frontier model
Dispositive Recall at filing (composite)88%24%19%9%
Fabrication rate0.00%2%0.00%22%
Raw topical recall (did it find the neighborhood)94%89%84%58%
Controlling-status recall (the assassin row)96%33%22%12%

Architectures, not brands. Dispositive Recall requires the right proposition, the right stance, the right procedural stage, confirmed good-law status and current citability, all at once; constraints compose multiplicatively. A system can clear any one filter by luck. Almost none clear all of them together.

Three autopsies

How a real, correctly quoted case loses.

Autopsy 1

The department-split trap

Two intermediate appellate departments in the same state disagree. Topical retrieval returns the better-known rule from the wrong department. Dispositive recall requires the controlling line for your forum.

Autopsy 2

The controlling-status trap

The case is real and on point. The high court reversed it, or review was granted and the opinion may no longer be cited. The document still ranks first by similarity.

Autopsy 3

The stage trap

A holding on a summary-judgment record, offered as authority on a motion to dismiss. Same words, wrong standard, motion lost.

Methodology

How we designed the test, and why.

Design principle
Select for architectural failure: every question is one a document-level index cannot answer without luck.
Answers
Propositional, not topical. A correct answer is a proposition with the right stance, stage, status and citability.
Corpus
Real New York case law, with California depublication and circuit-split variants as illustrations.
Scoring
Each dimension maps to a typed field; constraints compose multiplicatively. The rubric and the scoring function are published with the harness.
Integrity tie
Fabrication is held constant and conceded: a strictly grounded system also reaches 0.00%. The benchmark measures what comes after.
Status
Internal measurement, June 2026. Not independently audited. Question count, run date and confidence intervals ship with the harness.
The harness

Methodology first, numbers second.

The rubric, the question set and the scoring function are written to be published with the harness, so the measurement can be checked rather than believed. That paper is in preparation. Until it is out, the figures on this page stay labeled for what they are: an internal run, dated June 2026, not independently audited.

Version log

The benchmark re-versions when the law moves.

A question set built on live law goes stale. Each re-version will be written up on the news page, with what changed in the law and what changed in the questions.

The tie essay

Concede first, then turn.

Why we concede the 0.00% fabrication tie.

A strictly grounded retrieval product, built on a public corpus with citations verified against primary sources, has a fabrication rate of 0.00%. So do we. Its citations are real because they are pulled from a real corpus and checked against the source. Saying otherwise would be false, and pretending fabrication is the wedge would be a bet on the reader not checking.

That correction reframed the entire benchmark, and improved it. The wedge is not fabrication. The wedge is this: verified means the citation exists and is accurate. It does not mean the cited proposition is still controlling, on-stance, decided at the right stage, or currently citable. We tie the field at zero fabrication. Not fabricating is the floor now. Being filing-grade is the ceiling, and that gap is the whole game.

Bring a motion you’re nervous about.

In a demo we run it through and walk the five ways it could quietly lose, on your own authorities. The line is our standing invitation, not a product you sign up for here.