CapabilityCognitive automation2024-05-30
Leading AI legal research tools from LexisNexis and Thomson Reuters hallucinated between 17% and 33% of the time, a Stanford study found
Judge / magistrateoccupation page →Event date / reported
2024-05-30
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Researching the law
Finding and reading the statutes, precedents and authorities a case turns on.
Being augmented✓ Evidence-backed
Where this applies
A study that tested retrieval-augmented legal research tools sold by LexisNexis (Lexis+ AI) and Thomson Reuters (Westlaw AI-Assisted Research and Ask Practical Law AI), and found each hallucinated between 17% and 33% of the time, with substantial differences between systems. It tests research tools used across the legal profession, not judges' work, with 2024 versions.
What this means
Even the specialised tools built for legal research invent material often enough that a judge has to check every citation.
What it does not yet show
2024 versions of commercial tools; they may have improved.
What you can check
Open arXiv 2405.20362 and find "between 17% and 33% of the time".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Magesh et al. — Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arXiv 2405.20362, submitted 30 May 2024) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.