Worker adoptionCognitive automation2024-03-11
Between 6.5% and 16.9% of text in peer reviews at several AI conferences could have been substantially modified by language models, a Stanford study estimated
University lectureroccupation page →Event date / reported
2024-03-11
Evidence stage
Worker adoptionMeasured, large-scale use of a tool for real work, where the decision to use it was the worker's rather than an employer's. It is more than a capability record — the work is real, not a demo — and less than a deployment record, because no employer put it into production, required it, or built a process around it. Weighted `cautious`: `automating` means the machine can do the task AND there are adoption signs, and this is an adoption sign — but usage can be experimental, and much of the measurement comes from a party with a stake, so one record is never enough and two independent ones are. Note who is counting. Vendor telemetry sees this directly and sells the tool, so such a record names that stake in its scope; a statistics agency asking firms whether their workers use AI in tasks sees the same channel with no stake at all, and that is the better source where it exists.
Tasks this bears on
Peer review
Reviewing other researchers' papers and proposals.
Still human-led✓ Evidence-backed
Where this applies
Peer reviews at AI conferences (ICLR, NeurIPS, CoRL, EMNLP). Using a population-level estimator, the study reports that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, beyond spell-checking or minor writing updates. It covers AI conferences only and estimates shares of text, not individual reviewers.
What this means
Reviewers are already using language models on their reviews, at least in AI research. The rules that follow are about what they may hand over and what they must still judge themselves.
What it does not yet show
AI conferences only; it estimates text shares, not how review quality or judgement changed.
What you can check
Open arXiv 2403.07183 and find "between 6.5% and 16.9%" in the abstract.
Does it change the assessment?
No. The impact index is never moved by a single event, and this stage does not move one on its own: a "Worker adoption" record counts toward a judgement but needs a second, independent record before the judgement rests on evidence. This one is counted; on its own it changed nothing.
Source
Liang, Izzo, Zhang et al. (Stanford University) — Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews, arXiv 2403.07183 (submitted 11 Mar 2024; ICML 2024) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.