The debate about AI in litigation document review has been running since early e-discovery tools entered the market, and it keeps resurfacing in a form that obscures more than it reveals: "Can AI replace attorney review?" The question is the wrong frame. The more useful question is: in a document review process, which tasks benefit from AI assistance, and which tasks require attorney judgment? The answers are reasonably clear, and they point to a division of labor that makes review both faster and more reliable than the all-human alternative — without asking AI to do something it's not good at.
What First-Pass Relevance Classification Actually Is
In document-intensive litigation, the standard workflow involves reviewing a population of documents — often tens of thousands in a complex commercial matter — and sorting them into categories: relevant/non-relevant, responsive/non-responsive, privileged, hot documents. The purpose is to get from the raw document population to the subset that actually matters for the case before attorney analysis time is invested.
First-pass relevance review is the initial triage. In a 40,000-document population, first-pass review might identify 8,000 documents that are potentially responsive to the discovery requests, 2,500 that require closer examination for privilege, and 200 that appear to be high-priority or hot documents. Those classifications don't decide anything substantively — they route documents to the right stage of the review process. That routing function is where AI classification provides the most concrete value.
The economics are straightforward. If first-pass review of a 40,000-document set takes attorney reviewers 400 hours at billing rates that make the exercise expensive, and AI-assisted triage can produce classifications that a supervising attorney can validate in 40–60 hours of quality control review, the cost savings are real and the time compression is significant. The question isn't whether the AI's classifications are perfect — they aren't, and no one should expect them to be. The question is whether the classifications are accurate enough to be reliable as a routing mechanism, with appropriate attorney validation.
The Technology's Actual Capability in This Context
Relevance classification is a pattern-matching and semantic similarity task: given a document, does it fall within the scope of what the discovery request describes? For well-defined relevance criteria — documents discussing a specific contract, communications about a specific time period, records referencing specific parties or products — AI classification performs well. Recall rates in the 85–95% range for responsive documents are achievable in typical commercial litigation document sets when the relevance criteria are clearly specified.
The limitations are equally important to understand:
- Poorly specified relevance criteria produce poor results. If the instruction to the classification system is vague — "documents relevant to the dispute" — the model has to interpret what "relevant" means. That interpretation may not match what the reviewing attorney means, and errors compound. First-pass AI review requires investing time upfront in precise relevance criteria definition, which is attorney work.
- Contextual relevance requires human recognition. A document that on its face discusses a routine business matter may be highly relevant because of context the attorney knows — a background fact, a timeline implication, an internal nickname for a product or operation that doesn't appear in the document. The model classifies what's in the text. The attorney knows what the text means in the context of the case.
- Privilege classification has a higher error cost. Misclassifying a relevant document as non-responsive is a recall error; the document can often be caught in sampling. Misclassifying a privileged document as non-privileged and producing it may create a waiver issue that can't be corrected. The risk profile is different, and AI privilege classification requires more conservative supervision than relevance classification.
Where Attorney Judgment Is Not Optional
We want to be clear that we're not making an argument that AI can substitute for the substantive work of document review. The argument is the opposite: AI handles the volume-reduction task precisely so that attorney judgment can be concentrated on the work that actually requires it.
Case theory development from documents
What a set of documents collectively reveals about the facts, the parties' intentions, the timeline of events, and the potential witnesses is attorney work. A classification system can tell you which documents discuss a particular topic. It cannot tell you that the pattern across those documents suggests that a key witness had knowledge before the date they claimed. That inferential synthesis is the core of case preparation and it requires the lawyer who knows the case.
Hot document identification
High-value documents — the ones that will be used in deposition, that potentially shift the case's trajectory, that reveal bad facts or key admissions — aren't always the documents that rank highest on a relevance score. Some of the most important documents in complex litigation look ordinary on their face; their significance comes from what the attorney already knows. The "hot document" call requires the attorney.
Privilege and work product calls on borderline documents
Documents that clearly communicate legal advice from counsel to client are easy privilege calls. Documents that discuss a business matter in which counsel was tangentially involved, or that were generated during an internal investigation, or that mix business and legal advice in a single communication — those calls require judgment. The applicable tests for attorney-client privilege and work product protection have factual nuances that aren't resolved by text classification.
Building a Review Protocol That Uses Both Well
A review protocol designed around this division of labor looks different from the traditional all-attorney first-pass model:
- Invest in relevance criteria specification before upload. The attorney responsible for the matter defines relevance criteria precisely — not just "documents about the contract dispute" but the specific custodians, date range, topics, and document types that fall within scope. This takes an hour or two and determines the quality of everything that follows.
- Run AI first-pass on the full population. The system classifies documents into relevance tiers, with confidence scores. Documents above the relevance threshold go to attorney review; documents clearly below are sampled for quality control, not fully reviewed.
- Attorney review focuses on the flagged population. The attorney review team works the responsive and borderline documents, making substantive calls — relevance judgment, privilege assessment, hot document identification. The volume is a fraction of the total.
- Quality control validates the classification. A sample of non-responsive documents is reviewed by an attorney to check recall. If the sample reveals systematic misclassification of a document type, the relevance criteria are adjusted and the affected population is re-classified. The quality control pass is not optional.
- Privilege review follows its own track. Potentially privileged documents identified in first-pass are reviewed under a more conservative standard, with full attorney review of flagged documents rather than sampling.
Consider how this plays out on a typical commercial contract dispute. A breach-of-contract matter involving a three-year business relationship generates a document set of roughly 28,000 items across email, contracts, internal memos, and financial records. Under a traditional all-attorney first-pass model, that review might take six to eight weeks and a meaningful portion of the case budget before any substantive work begins. Under a hybrid model with AI first-pass and attorney quality control, the triage phase compresses to one to two weeks, and the attorneys' time is concentrated on the roughly 4,000 documents that the system flags as responsive.
The result isn't just faster — it's often more thorough. When attorney time isn't consumed by reading 24,000 clearly non-responsive documents, more of it is available for the close reading of documents that actually matter. The attorneys who have used this model consistently report that they feel more prepared for depositions and motions, not less, because the volume-reduction step put them in the responsive documents earlier.
The firm's role doesn't shrink in this model. It shifts toward higher-value work — case strategy, legal analysis, substantive document interpretation — and away from volume processing. That's not a diminishment of the attorney's contribution. It's a reallocation of where that contribution is most needed.