Legal AI Reliability
We publish a public, lawyer-first taxonomy of AI failure modes in contract drafting and editing. This is a reliability library that makes failure modes concrete, shows why they matter, and documents how they are detected.
Why This Exists
Lawyers and clients need reliability before automation. Our experience is that models still violate document invariants that trained lawyers routinely protect: numbering integrity, defined-term consistency, cross-reference accuracy, and minimal-diff editing discipline.
Who This Helps
In-house Teams
A grounded way to explain why AI still needs human review in contracts, without sounding anti-innovation.
Foundation Labs
A structured, reproducible failure-mode library that can be turned into evals, lint rules, and training data.
Failure Mode Taxonomy (Draft)
- Document structure invariants (list nesting, parent-child slicing)
- Defined-term integrity (capitalization, duplication, drift)
- Cross-reference integrity (broken sections, exhibit callouts)
- Redline mechanics (non-minimal diffs, numbering cascades)
- Obligation logic drift (shall/may flips, condition changes)
- Citation and quote verification (false citations, wrong paragraph)
Taxonomy Index
Each item links to the gallery entry with a short, lawyer-first explanation.
| Category | Failure Mode | Severity | Detectable | Status |
|---|---|---|---|---|
| Structure | List Integrity (Orphaning) | High | Yes | Lintable |
| Structure | Renumbering Cascade | High | Yes | Lintable |
| Definitions | Definition Alphabetization | High | Yes | Lintable |
| Definitions | Defined Term Capitalization | High | Yes | Lintable |
| Cross-References | Dirty Cross-References | High | Yes | Lintable |
| Comparisons | Distractor Misattribution | Critical | Architectural | Architectural |
| Citations | Unread / Stale Citations | High | Yes | Lintable |
| Identifiers | Identifier Interpolation | High | Yes | Lintable |
| Parties | Placeholder & Party Consistency | High | Yes | Lintable |
| Semantics | Semantic Overcompression | Medium | Partial | Partial |
| Formatting | Formatting & Context Overhead | Medium | Architectural | Architectural |
| Markup | Visual Reasoning / Redlines | High | Architectural | Architectural |
| Citations | TOC Over-Indexing | Medium | Yes | Lintable |
| Editing | Document vs Code Editing | High | Architectural | Architectural |
What Comes Next
We are building public examples and a private evaluation pack. If you want the dataset, scoring harness, or a lab-facing briefing, contact us.