Prompt20
All posts
legallawcontract-reviewe-discoverylegal-researchcomplianceverticalevergreen

AI in Law: Where It Helps and Where It Hallucinates

A grounded look at AI in legal work. Contract review and drafting, e-discovery, legal research, case summarization, and client intake — set against the real risks: fabricated citations, confidentiality, privilege, and the professional-responsibility rules that make lawyers liable for the model's mistakes. Why 'human in the loop' is a legal requirement, not a nicety.

By Prompt20 Editorial · 26 min read

The honest one-line answer: AI is genuinely useful for legal work that is bounded and checkable — summarizing a document you already have, drafting a first pass, sorting a mountain of email — and genuinely dangerous for legal work that is open-ended and authoritative, like "find me the controlling case," because a language model will invent a plausible-looking citation with the same confidence it uses for a real one. The value is real. The ceiling is set by verification cost, not model quality.

That framing matters more in law than almost anywhere else, because a lawyer does not get to say "the AI made it up." Under the professional-responsibility rules that govern lawyers in most jurisdictions, the duty of competence and the duty of candor to the court belong to the human. The model is a tool, like a junior associate or a Westlaw subscription, and you are on the hook for what it produces. So the useful question isn't "is legal AI good?" It's "which tasks can I verify cheaply enough that the model's mistakes can't reach a filing or a client?"

Key takeaways

  • The dividing line is verification cost, not task difficulty. Summarizing a contract you can re-read is safe; asserting what the law is is not, because checking a legal proposition against real authority is expensive and the model's failure mode is confident fabrication.
  • Hallucinated citations are the signature risk. Models generate case names, docket numbers, and quotes that look real and aren't. This has produced sanctioned filings in real courtrooms. Every citation an AI touches must be pulled and read before it goes anywhere.
  • Confidentiality and privilege set hard limits. Pasting client facts into a consumer chatbot can breach your duty of confidentiality and, in some readings, risk waiving privilege. Where the data goes and who can train on it is a threshold question, not a footnote.
  • Retrieval beats raw generation for anything authoritative. A system that pulls from a real, current database of primary law and cites what it retrieved is categorically safer than a model answering from its weights. Ask what the tool is grounded in.
  • "Human in the loop" is a rule, not a nicety. The lawyer's duties of competence, supervision, and candor don't transfer to software. Treat AI output as a draft from an unsworn, occasionally-lying junior — never as a source.
  • The best ROI is on volume-and-tedium tasks: e-discovery triage, first-draft contracts against your own playbook, intake, and summarization — bounded jobs where a human check is fast.

Table of contents

Why law is a special case

Most "AI for X" pieces treat hallucination as a quality bug that shrinks as models improve. In law it's a structural hazard, for three reasons that don't go away with a better model.

First, the cost of a wrong answer is asymmetric and public. A hallucinated fact in a marketing email is embarrassing. A hallucinated case in a brief is sanctionable, gets your name in an opinion other lawyers will read, and can harm a client's matter. The downside dwarfs the upside of the time you saved.

Second, legal correctness is not a property of fluent text. Language models optimize for plausible next tokens, and legal writing has an extremely learnable style — the citation format, the "Accordingly, the court holds" cadence, the Latin. A model can produce something that reads exactly like binding authority and is entirely fictional. Fluency is actively misleading here in a way it isn't when you ask for a poem. If you want the mechanics, see why AI hallucinations happen: the model isn't looking anything up, it's sampling from a distribution.

Third, the duties are personal and non-delegable. The rules that govern lawyers place competence, confidentiality, and candor on the individual. You can delegate the work to a tool, but not the responsibility. That single fact reframes every efficiency claim: the time you save on drafting has to survive the time you spend verifying, or you haven't saved anything — you've moved risk onto your own license.

There's a fourth reason that gets less attention and matters just as much: law runs on adversarial verification. Almost every other domain that adopts AI is cooperative or at worst indifferent — a marketer's fabricated statistic sits in a slide deck that nobody audits line by line. In litigation, the opposing party is paid to find the weak citation, the misquoted holding, the clause that doesn't say what your brief claims. There is a professional on the other side whose job is to catch exactly the errors a language model is most likely to make. This inverts the usual economics of "good enough" output. In most fields an occasional error is absorbed as noise; in an adversarial one, a single fabricated authority hands your opponent a free credibility attack and can taint everything else you filed. The error rate that a marketing team shrugs off is the error rate that loses a motion.

And a fifth, quieter one: legal language is deceptively self-similar. Two contracts can differ by a single word — "shall" versus "may," "including" versus "including without limitation," an indemnity that runs one way instead of both — and that word is the entire deal. Models are trained to smooth toward the typical phrasing, which is precisely the wrong instinct when the value lives in the atypical clause a party negotiated hard for. A summarizer that "cleans up" a bespoke indemnity into the standard version has not made an approximation error; it has silently reversed the allocation of risk the parties agreed to. Fluency, again, is the trap: the smoother the output reads, the easier it is to miss that it quietly normalized away the one term that mattered.

The map of legal AI: six use categories

"AI in law" is not one thing, and treating it as one thing is how firms make bad procurement decisions. It's at least six distinct categories of tool, each with a different risk profile, a different verification cost, and a different honest assessment of whether it's ready. Mapping them separately is the first discipline.

1. Legal research (finding and stating the law). The task most people picture, and the most dangerous. The goal is to find controlling authority for a proposition and state what it holds. This is where fabricated citations live, because the correct answer is not in any document you control — it's in the open-ended, jurisdiction-specific, constantly-moving body of primary law. A raw chatbot is actively hazardous here. A retrieval-grounded research tool wired to a real, current case database is categorically different, but even then the output is a starting point you verify, never an answer you cite.

2. Contract review and drafting. Reading contracts for risk, drafting from templates, redlining against a standard. The strongest commercial fit after e-discovery, because the correct answer often lives in a document you control — your clause library, your playbook, the four corners of the agreement in front of you. Covered in depth below.

3. E-discovery and document review. Classifying, clustering, and prioritizing large document sets for relevance and privilege in litigation. The oldest and most mature category — machine-assisted review predates the generative-AI wave by more than a decade and is well-accepted in many courts. Bounded task, statistically checkable output.

4. Document automation. Generating routine documents — engagement letters, standard filings, disclosure schedules, form pleadings — from structured inputs. Overlaps with drafting but is more template-and-fields than reasoning. Low risk when the template is vetted and the inputs are validated; the failure mode is a wrong merge field, not an invented legal rule.

5. Litigation analytics and outcome prediction. Estimating how a judge tends to rule, how long a matter will take, or whether to settle, from patterns across historical cases. The most oversold category. Discussed below — the short version is that its outputs are hard to falsify and easy to mistake for insight.

6. Client intake, triage, and knowledge management. Front-door automation: structured Q&A to route matters, answer routine questions, populate forms, and surface a firm's own past work product. Genuinely useful, bounded, and lower-stakes — provided there are clear "this is not legal advice" disclaimers and a human before anything binding.

Notice these do not share a risk profile. E-discovery and intake are near-ready and checkable. Legal research and outcome prediction are where the money and the danger both concentrate. A firm that buys "an AI tool" without knowing which of these six it actually is has already lost the thread. The rest of this piece works through the ones that matter most.

Where AI genuinely helps

Sort legal tasks by one axis — how cheaply a competent human can verify the output — and the picture gets clear fast.

Task Why AI helps Verification cost Verdict
E-discovery triage / review Classifies and clusters huge document sets; surfaces likely-relevant material Low–medium (spot-check + statistical sampling) Strong fit
First-draft contracts from a playbook Fills a known template; flags deviations from your standard clauses Low (you know what the clause should say) Strong fit
Summarizing a document you have Condenses a contract, deposition, or filing you can re-open Low (read the source) Strong fit
Client intake & triage Structured Q&A, form-filling, routing Low–medium Good fit, with disclaimers
Legal research (finding authority) Fast, broad, articulate High (every cite must be pulled and read) Only with grounded retrieval + full check
Predicting outcomes / "should we settle" Pattern over past matters Very high / often unfalsifiable Treat as a prompt, not an answer

E-discovery is the strongest fit and the oldest. Machine classification of documents for relevance and privilege predates the current AI wave, and courts in many jurisdictions have long accepted technology-assisted review. The task is bounded (this document set, these categories), the output is checkable (sample and measure precision/recall), and the alternative — humans reading millions of emails — is worse on every axis. Modern models extend this to nuance a keyword search misses.

Contract review and drafting is the second strong fit, when scoped to your own standards. "Draft an NDA from our template and flag anything the counterparty changed against our playbook" is a bounded, checkable task. "Is this contract enforceable?" is not. The difference is whether the correct answer lives in a document you control or in the open-ended body of the law.

Summarization is safe precisely because verification is trivial: you have the source, you can read it. The failure mode — the model omits or distorts something — is caught by anyone who checks against the original. The risk rises the moment the summary becomes the only thing anyone reads.

Intake and triage works well as structured front-door automation, with clear disclaimers that it isn't legal advice and a human before anything binding.

Under the hood, the tools that do these well aren't a raw chatbot with a legal system prompt. They're retrieval-augmented systems that pull from a curated corpus — your document set, your clause library, a real case database — and constrain the model to work from retrieved text. That architecture is what separates a defensible legal product from a demo.

Contract review, in detail

Contract work deserves its own section because it's where most firms will get their first real value, and where the difference between a safe deployment and a dangerous one is entirely about how the task is framed. The mechanics reward a clear head.

What the model is actually doing. A competent contract-review tool is not "reading the contract and telling you if it's good." It's performing a set of narrow, bounded operations: extracting specific provisions (term, governing law, indemnity, limitation of liability, assignment, termination), comparing them against a reference — your playbook, a prior version, the counterparty's markup — and flagging deviations. Framed this way, the task is checkable, because for each flag you can open the clause and see whether the model read it correctly. The unit of work is small and the source is in front of you.

Where it earns its keep. Three jobs, in rough order of maturity:

  • Redline triage against a playbook. Given your standard positions ("we accept mutual indemnity but not one-way; cap liability at fees paid; no automatic renewal beyond one year"), the model reads a counterparty's draft and surfaces every clause that departs from them. This turns a two-hour first read into a fifteen-minute review of a pre-flagged list — provided the playbook is written down and the flags are traceable to specific text.
  • Extraction across a portfolio. "Across these 400 leases, which ones have a change-of-control clause, and what does each say?" This is due-diligence work that used to consume junior associates for weeks. The output is a structured table, and each cell links to the source clause, so verification is a spot-check plus a sample.
  • First-draft generation from a template. "Draft an NDA on our standard form with these parties, this term, and mutual confidentiality." The template constrains the output; the human confirms the fields and the deviations.

Where the framing goes wrong. The dangerous version of every one of these is the open-ended question dressed up as a bounded one. "Is this contract enforceable?" "Are we protected here?" "What's our exposure?" These sound like contract review, but the answer doesn't live in the document — it lives in the governing law, the parties' course of dealing, facts not in the contract, and how a court in that jurisdiction would read it. That's legal research and judgment wearing a contract's clothing, and the model will answer it with the same fluent confidence it uses to extract a termination date. The discipline is to keep asking: is the correct answer inside a document I control, or out in the open law? Only the first kind is safe to hand a model.

The silent-edit hazard, specifically. Contract drafting has a failure mode that summarization mostly doesn't: the model can produce a clause that reads perfectly and is subtly wrong in a way that favors nobody on purpose. It drops a carve-out, flips the direction of an indemnity, or "tidies" a negotiated exception back to the standard form. Because the output is grammatical and lawyerly, there's no visual signal that anything is off. The only defense is to treat generated clauses as drafts to be diffed against intent, not as finished text — and to diff against what the deal requires, not just against what reads cleanly. A clean-reading clause that reverses the risk allocation is worse than an awkward one that gets it right.

A practical rule of thumb. Contract AI is safest when it points you at text and least safe when it replaces reading text. A tool that says "clause 14.3 departs from your playbook, here it is" is doing the job. A tool that says "this contract is fine" is asking you to trust it in exactly the way you shouldn't. Buy the first; be very careful with anything that sounds like the second.

Where it hallucinates

The signature failure is the fabricated citation. Ask a bare language model for authority supporting a proposition and it will often produce a case name, a reporter citation, a court, a year, a pinpoint page, and a persuasive quote — all invented, or stitched from fragments of real cases into something that never existed. This is not rare or exotic. It has put real lawyers in front of real judges explaining why their brief cited opinions that do not exist, and it keeps happening because the output is indistinguishable from correct until you try to pull the case.

Why does this happen? Because the model is not consulting a database. It's predicting tokens that fit the pattern of a citation. A valid-looking citation is just a well-formed string, and the model is excellent at well-formed strings. It has no built-in sense of "this specific case exists" versus "this is what a case about this topic would be called." The plausibility that makes the output useful is the same plausibility that makes the fabrication dangerous.

Three defenses, in order of importance:

  1. Ground the model in a real database and cite what was retrieved. A system that answers only from documents it actually pulled — and links each claim to a source you can open — moves the failure from "invented a case" to "misread a real one," which is a mistake a human check catches. This is the single biggest architectural lever.
  2. Pull and read every citation, every time. Non-negotiable. If a citation goes into a filing, a human opened the case and confirmed it says what the brief claims. No exceptions for tools that "usually get it right."
  3. Prefer tools that show their work. If you can't see what the model retrieved, you can't check it, and you're trusting weights. Opacity is a red flag in a legal tool.

Even grounded tools misread. Retrieval fixes "the case doesn't exist"; it doesn't fix "the case exists but doesn't hold what the summary says," or a holding that was later narrowed or overturned. Currency is its own hazard: a model's training data has a cutoff, and the law moves. A confidently stated rule may be a year out of date with no signal that it's stale. Only a tool wired to a maintained, current source of primary law can be trusted on what the law is right now — and even then you verify.

There's a subtler variant worth naming: the subtly-wrong-quote. A model asked to support a proposition may retrieve a real, on-point case and then quote it slightly incorrectly — dropping a qualifier, changing "may" to "shall," omitting the "however" that limits the holding. The citation checks out; the case exists; a hurried reviewer confirms the case is real and moves on. But the quoted language has been altered in a way that changes the meaning, and that alteration is now in your brief attributed to a court that never wrote it. This is why "the citation is real" is not the same standard as "the case says what we claim." Verification means reading the cited passage in the source and confirming the words, not confirming the case exists. The adversary will read the passage; so should you.

Litigation analytics and the seduction of prediction

The most oversold corner of legal AI is prediction: tools that promise to tell you how a specific judge tends to rule on a motion type, how long a matter will take, what a case is "worth," or whether you should settle. The pitch is seductive because it dresses pattern-matching over past cases in the language of data science, and because the questions it answers are exactly the ones a nervous client most wants answered. Be the most skeptical here.

The core problem is falsifiability. When a research tool fabricates a citation, you can catch it — the case either exists or it doesn't. When a prediction tool says "this motion has a 70% chance of success," there is no clean way to check it against reality. The motion resolves once. If you win, was the tool right, or was it a coin flip that landed your way? If you lose, was the 70% wrong, or did you draw the 30%? A single outcome can't confirm or refute a probability, and no firm runs the same case a hundred times. The output feels like information but resists the verification that would tell you whether it's worth anything.

Three specific traps:

  • Base rates masquerading as insight. "This judge grants summary judgment 40% of the time" is a historical frequency, not a prediction about your motion, whose facts and briefing are what actually decide it. Presenting the base rate as a case-specific forecast is a category error the interface often encourages.
  • Survivorship and selection in the training data. Litigation outcomes are shaped by which cases settle, which get sealed, and which are appealed. The visible record is a biased sample of what happened, and a model trained on it inherits every one of those biases while presenting a clean number.
  • The anchoring risk. A confident settlement figure or win probability anchors the lawyer's and client's judgment, whether or not it's any good. That's a real harm even when the number is meaningless: it shifts decisions that should turn on legal analysis toward a figure produced by pattern-matching over a biased sample.

None of this makes analytics useless. Aggregate patterns can be a prompt — a reason to look harder at a judge's prior opinions, a sanity check on a settlement range you derived independently. The failure is treating the output as an answer. A win probability is a question in disguise ("why might this be true or false?"), never a conclusion. Use it to direct attention; never let it substitute for the analysis it's pretending to summarize.

Confidentiality, privilege, and where the data goes

Before competence, there's a threshold question most efficiency pitches skip: where does the client's information go, and who can see or train on it?

A lawyer's duty of confidentiality covers essentially everything relating to a representation. Paste a client's facts into a general-purpose consumer chatbot and you may have disclosed confidential information to a third party whose terms let them retain and train on your inputs. That can be a breach independent of whether the output is any good. Enterprise and legal-specific tools typically offer contractual terms — no training on your data, defined retention, tenant isolation — that consumer tiers do not. Read them; don't assume. Our guide to AI chatbot privacy walks through the retention and training questions to ask any vendor.

Privilege is the sharper edge. Attorney-client privilege and work-product protection can, under some theories, be jeopardized by disclosure to a third party that isn't cloaked by the privilege. The law here is still developing and varies by jurisdiction, so treat it conservatively: assume that feeding privileged material into a system you don't control could create a waiver argument an opponent will happily make, and structure your tooling so privileged data stays inside vetted, contractually-bound systems.

Practical posture:

  • Match the tool to the data. Public legal research questions with no client facts are low-stakes. Anything with identifiable client information belongs only in a vetted, contractually-bound system.
  • Know the data path. Cloud API, on-prem, or a locally-run model are very different risk profiles. For the most sensitive work, keeping the model on infrastructure you control is the conservative choice.
  • Get client consent where appropriate. Depending on jurisdiction and sensitivity, informing clients that AI tools may be used — and how their data is handled — is prudent and sometimes required.

The verification burden and the non-delegable duty

Everything above converges on one uncomfortable number that most efficiency pitches never mention: the true cost of AI output is generation plus verification, and in law the second term is often the larger one. If it takes the model thirty seconds to draft a paragraph with three citations and it takes you twenty minutes to pull and read those three cases, the honest accounting is twenty minutes and thirty seconds — not thirty seconds. Any claim of time savings that quietly drops the verification term is measuring the wrong thing.

This reframes the entire value proposition. AI helps in law precisely where the verification term is small relative to what generation replaced. Summarizing a contract you'll read anyway: verification is nearly free, because you were going to read the source regardless. E-discovery triage: verification is statistical sampling over a set you could never have read by hand, so the model isn't adding a verification burden, it's replacing an impossible one. First draft from your own playbook: you know what the clause should say, so checking is fast. In all three, the model wins because the check is cheap. The tasks where AI seems to help but doesn't are the ones where verification quietly costs as much as doing it yourself — asking "what's the controlling law?" and then having to do the full research to confirm the answer is trustworthy. You did the work twice.

So the practical test before adopting any legal AI task is not "can the model do this?" but "is verifying the model's output cheaper than doing the task myself, given that I am legally required to verify it anyway?" That question sorts the ready use cases from the seductive traps more reliably than any benchmark.

Why the duty won't move to the vendor. Firms sometimes hope that a well-drafted vendor contract, an indemnity, or a disclaimer shifts the risk off the lawyer. It doesn't, in the way that matters. A vendor might owe you money if their tool fails, but the professional-responsibility duty of candor to the court runs from you to the tribunal, and no commercial agreement between you and a software company changes what you owe the judge. When a fabricated citation reaches a filing, the court's concern is with the signature on the brief, not the license agreement behind it. This is the sense in which the duty is genuinely non-delegable: you can outsource the labor and even the liability-for-dollars, but not the accountability-to-the-court. That accountability is the whole reason the profession is licensed, and it's exactly what no model can hold.

There's a useful analogy here to a calculator versus a witness. A calculator is a tool you can rely on because its output is verifiable in principle and reliable in practice — you don't re-derive long division, but you could. A witness is a source whose statements you must weigh, corroborate, and cross-examine, because they can be wrong or lying. The category error at the root of most AI legal disasters is treating the model as a calculator when it is, epistemically, a witness — fluent, confident, occasionally fabricating, and never under oath. Tools you verify like a calculator; sources you interrogate like a witness. The model is the second thing wearing the interface of the first.

The rule that ties it together: human in the loop

"Human in the loop" gets said so often it sounds like a slogan. In law it's closer to a description of your legal obligations.

  • Competence means you're responsible for understanding the tools you use well enough to use them safely — including their failure modes. "I didn't know it could hallucinate" is not a defense once hallucination is a known property.
  • Supervision means AI output is a draft from a subordinate you must review, not a source you cite. You'd never let an unsworn junior's first draft go out unchecked. The model is that junior, minus the fear of being fired, plus a talent for confident fabrication.
  • Candor to the court means every representation in a filing is yours. The judge is not interested in your software vendor.
  • Billing raises its own question: if a task took the model ten minutes, billing an hour of "research" is its own ethics problem. Efficiency gains generally belong, at least in part, to the client.

The workable mental model: AI is a force multiplier on a competent lawyer, and a liability multiplier on a careless one. It makes a good lawyer faster at drafting and reviewing. It makes a lawyer who doesn't check faster at generating sanctionable work. Same tool, opposite outcomes, and the variable is the human's discipline about verification.

Regulation, court rules, and the standing-order era

Legal AI is unusual among AI applications in that its users are already governed by a dense, mature body of professional rules — and the institutions that enforce those rules move faster than legislatures. You don't have to wait for an "AI law" to be regulated; the existing duties of competence, confidentiality, supervision, and candor already apply, and courts and bar associations have been busy mapping them onto AI specifically.

Several layers are worth tracking, in rough order of immediacy:

  • Court standing orders and local rules. The most concrete constraint. After the first wave of fabricated-citation incidents, individual judges and some courts began issuing standing orders on AI use — variously requiring disclosure that AI was used in a filing, certification that any AI-generated content was checked by a human, or both. These are specific, enforceable, and judge-by-judge: the operative rule can differ from one courtroom to the next, which means "what are this court's AI rules?" is now a question to ask at the start of any matter, the way you'd check page limits and formatting rules. Do not assume a rule you followed last month in one court applies in the next.
  • Bar association guidance and ethics opinions. Bars in various jurisdictions have issued opinions applying the professional conduct rules to generative AI — typically confirming that use is permitted, that the duty of competence now includes understanding the tools' limitations, that confidentiality constrains what data may be entered, and that supervision duties extend to AI output. These rarely create new obligations; they clarify how existing ones bind. That's the pattern to expect and the reason "it's a new technology" is not a defense.
  • Discovery and evidentiary rules. As AI generates and touches more documents, questions about the discoverability of prompts and outputs, the authentication of AI-assisted work, and the treatment of AI-generated evidence are developing. This area is genuinely unsettled and jurisdiction-specific.
  • Broad AI legislation. Horizontal regimes governing AI systems generally — risk tiers, transparency, data governance — increasingly reach legal tech as a downstream user. The broader regulatory picture sits above the profession-specific rules and is moving in its own right.

The through-line: the regulatory floor is rising, unevenly, and from multiple directions at once. For a practitioner the practical posture is not to memorize a fixed rulebook — there isn't one, and it would be stale by next term — but to build the habit of checking the applicable court's orders and the relevant bar's guidance for each matter, and to keep a verification record that demonstrates compliance if anyone ever asks. The lawyers who get burned are rarely the ones who didn't know the rules existed; they're the ones who assumed last year's rules, or another court's rules, still applied.

The economics and the access-to-justice question

Step back from any single firm and the interesting question is what AI does to the structure of legal work, and there are two honest answers pulling in opposite directions.

The optimistic case is access to justice. Legal help is unaffordable for most people and most small businesses. A large share of people who need a lawyer — for an eviction, a benefits denial, a simple will, a small-claims dispute — never get one, not because the law is beyond them but because an hour of a lawyer's time costs more than the matter is worth to them. If AI can genuinely lower the cost of the routine, high-volume, template-shaped legal work — the intake, the standard document, the first-pass research — it could in principle extend some form of legal help to people currently priced out entirely. That's a real and worthwhile prize, and the categories where AI is safest (bounded, checkable, high-volume) overlap substantially with the categories where the unmet need is largest.

The skeptical case is that the verification burden eats the savings exactly where they'd matter most. The economics of legal AI are only favorable when a competent lawyer verifies the output cheaply. But the access-to-justice scenario often removes the lawyer — a self-represented person using a consumer chatbot has no one to catch the fabricated citation or the silently-reversed indemnity. So the very population that most needs cheap legal help is the least equipped to absorb the tool's failure mode. AI that helps a lawyer serve more clients affordably is a genuine gain. AI that replaces the lawyer for someone who can't tell a real holding from an invented one may simply relocate the harm from "can't afford help" to "got confident, fluent, wrong help." The unauthorized-practice-of-law rules exist in part for this reason, and they sit awkwardly against the access mission.

The honest synthesis: AI most plausibly improves access to justice by making lawyers cheaper, not by removing them. A legal-aid clinic that handles three times the caseload because AI drafts and triages — with lawyers still checking — is the durable win. A chatbot that hands unverified legal conclusions to people with no way to check them is a liability wearing the costume of empowerment. Which one gets built is a choice, not a property of the technology, and the difference between them is once again whether a competent human remains in the loop.

There's also a quieter internal economic story: AI compresses the leverage model that funds law firms. A great deal of junior-associate training happens through the grinding tasks AI is best at — the document review, the first drafts, the research memos. If those tasks shrink, the profession has to answer how the next generation of lawyers develops the judgment that verification depends on. You can't supervise a tool whose work you never learned to do yourself. That's not a reason to avoid the tools; it's a reason to be deliberate about how juniors still learn the underlying craft, because the entire safety model rests on humans who can tell right from plausible.

Hype versus what actually ships today

Cutting through the marketing, here is the grounded state of things — the parts that are real, the parts that are oversold, and the parts that are simply not there yet. This is an evergreen assessment: the specific tools change, but the shape of what works and what doesn't has been stable and is likely to stay that way, because it's set by the structure of the tasks, not the capability of any one model.

Real and working today. E-discovery and technology-assisted review — mature, accepted, genuinely better than manual review. Document summarization where the human keeps the source. First-draft generation from a firm's own templates and playbooks. Extraction across document portfolios for due diligence. Structured intake and triage with disclaimers. Retrieval-grounded research that shows its sources as a starting point for a lawyer who then verifies. The common thread, again: bounded task, cheap verification, human owns the output.

Oversold — real capability wrapped in overreach. "AI legal research" that implies you can trust the answer without pulling the cases — the grounding helps, the verification requirement does not go away. Outcome prediction sold as insight rather than a biased base rate. "Autonomous" contract review that offers a verdict ("this contract is fine") rather than pointing you at the clauses. The pattern is a genuinely useful bounded tool marketed as if it had crossed into the unbounded, authoritative territory where the risk lives.

Not there, and structurally unlikely to arrive soon. An AI that you can trust to state the current law without a human verifying it. An AI that can be accountable — that can sign a brief, answer to a court, hold a duty. That last one isn't a capability gap that a better model closes; it's a category difference. Accountability requires a licensed person who can be sanctioned, and no model is a person. Any pitch that quietly assumes the model can carry responsibility is selling the one thing the technology structurally cannot provide.

The test for any legal AI claim is the same question that runs through this whole piece: does it respect the line between text that reads like law and text that is verified to be law? Tools that keep a human on the right side of that line are among the best productivity gains the profession has seen. Tools that blur it are selling a lawsuit. The marketing rarely tells you which one you're looking at; the architecture and the verification story do.

How to adopt it without getting burned

A pragmatic sequence:

  1. Start where verification is cheap. Summarization, first-draft-from-playbook, e-discovery triage — high volume, low authority, fast to check. Build trust on tasks where mistakes are visible and harmless.
  2. Insist on grounding and traceability. Prefer tools that retrieve from real, current sources and show what they pulled. If you can't see the source, you can't verify — walk away. See how to choose an LLM for your app for the evaluation questions that transfer directly.
  3. Write the confidentiality rules down first. Which tools may touch client data, on what terms, with what consent. Do this before anyone pastes a contract into anything.
  4. Make verification mandatory and logged. Every AI-touched citation gets pulled and read; note it in your workflow. This is your defense if a mistake ever slips through. The general playbook — grounding, forced citations, verification passes — is in how to reduce AI hallucinations.
  5. Keep a human signature on everything that leaves. No AI output reaches a client or a court without a lawyer who read it and owns it.
  6. Watch the regulatory floor rise. Courts and bars are issuing standing orders and guidance on AI disclosure and verification. The broader regulatory picture is moving, and legal practice sits near the front of it.

Do this and AI is one of the better productivity tools to reach the profession in a generation. Skip the verification discipline and it's a fast path to a sanction with your name on it. The technology didn't change which of those you get — you did.

FAQ

Can AI give legal advice? Not in the sense a lawyer does. A model can produce text that reads like legal advice, but it can't take responsibility, doesn't reliably know current law, and will fabricate authority. It's a drafting and research aid for a supervising lawyer, not a substitute for one. Consumer tools should carry — and generally do — clear disclaimers that output isn't legal advice.

Why does AI make up case citations? Because a bare language model predicts plausible text rather than looking anything up. A citation is a well-formed string, and models are excellent at well-formed strings, so they generate real-looking case names, numbers, and quotes for cases that don't exist. The fix is to use tools grounded in a real, current legal database and to pull and read every cited case before relying on it.

Is it safe to put client information into ChatGPT or similar tools? Treat consumer chatbots as unsafe for confidential client data unless the terms explicitly prevent training on and retention of your inputs. Doing otherwise can breach confidentiality and, under some theories, risk privilege. Use enterprise or legal-specific tools with contractual data protections, or a model on infrastructure you control, for anything containing client facts.

Does using AI violate legal ethics rules? Using AI isn't inherently a violation — but the lawyer's duties of competence, confidentiality, supervision, and candor still apply fully. Failing to verify AI output, leaking confidential data, or hiding required disclosure can each breach those rules. The tool is permitted; carelessness with it is not.

What legal tasks is AI actually good at? Bounded, checkable, high-volume work: e-discovery triage, first-draft documents against your own playbook, summarizing documents you already have, and structured client intake. The common thread is that a competent human can verify the output quickly. Open-ended, authoritative questions — "what does the law require here?" — are where the risk lives.

Will AI replace lawyers? Not the responsibility, because someone accountable and licensed has to own the work product, verify it, and answer to the court. It will change the mix of what lawyers spend time on — less first-draft grinding and manual review, more judgment, verification, and client work. The durable skill is being the human who checks, and knowing exactly what to check.

Do I have to tell the court I used AI? It depends entirely on the court, and increasingly the answer is yes. After the first wave of fabricated-citation incidents, individual judges and some courts began issuing standing orders that require disclosure of AI use in filings, certification that AI-generated content was human-verified, or both. These rules vary judge by judge and court by court, so the safe habit is to check the applicable court's standing orders and local rules at the start of every matter rather than assume last month's practice still applies. When in doubt, verifying everything and being prepared to disclose is the conservative posture.

Can AI-generated legal work be discovered by the other side? This area is unsettled and jurisdiction-specific, but treat it cautiously. Questions about whether prompts, drafts, and AI outputs are discoverable — and how AI-assisted work interacts with work-product protection — are actively developing. The prudent assumption is that anything you feed a tool you don't control could surface, which is another reason to keep privileged and sensitive material inside vetted, contractually-bound systems and to be deliberate about what you put into any AI tool in the first place.

Is AI a genuine access-to-justice tool for people who can't afford a lawyer? Cautiously, and mostly indirectly. AI most plausibly improves access by making lawyers cheaper and higher-throughput — a legal-aid clinic handling more cases with AI drafting and triage, lawyers still verifying — rather than by replacing the lawyer. A self-represented person using a consumer chatbot has no one to catch a fabricated citation or a silently-reversed clause, so the population that most needs cheap help is the least equipped to absorb the tool's failure mode. The gain is real when a competent human stays in the loop; it becomes a liability the moment the tool hands unverified legal conclusions to someone with no way to check them.