CapabilityCognitive automation2024-05-30
Output from paid AI legal research tools by LexisNexis and Thomson Reuters contained hallucinations 17% to 33% of the time in a Stanford evaluation
Paralegaloccupation page →Event date / reported
2024-05-30
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Checking what the tools produced
Confirming that a cited case exists, that a summary matches the source, that a generated clause does not contradict another.
New task✓ Evidence-backed
Where this applies
A study that tested retrieval-augmented legal research tools sold by LexisNexis (Lexis+ AI) and Thomson Reuters (Westlaw AI-Assisted Research and Ask Practical Law AI), and found each hallucinated between 17% and 33% of the time, with substantial differences between systems. It tests 2024 versions of the tools, not how lawyers or paralegals use them.
What this means
Checking what the machine produced is real work: even specialised legal tools get a sixth to a third of answers wrong.
What it does not yet show
2024 versions of commercial tools; they may have improved, and the study does not measure use in practice.
What you can check
Open arXiv 2405.20362 and find "between 17% and 33% of the time".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Magesh et al. — Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv 2405.20362, submitted 30 May 2024) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.