PilotCognitive automation2025-04-13
27% of ICLR 2025 reviewers who received AI feedback on their reviews updated them, in a randomised study of 20,000 reviews
University lectureroccupation page →Event date / reported
2025-04-13
Evidence stage
PilotSmall-scale trial in a real setting. Tells us the deployment conditions are being tested, not that they hold — so one pilot is never enough on its own; two independent ones are.
Tasks this bears on
Peer review
Reviewing other researchers' papers and proposals.
Still human-led✓ Evidence-backed
Where this applies
An AI conference's review process in 2025. In a randomised study, an LLM-based agent gave reviewers feedback on their reviews; the paper reports that 27% of reviewers who received feedback updated their reviews, and that over 12,000 feedback suggestions were incorporated. The authors built the system they evaluate, and it ran at one conference.
What this means
AI is being used to review the reviewers — suggesting where a review is vague or unfair — and reviewers act on it. The judgement on the paper remains the reviewer's.
What it does not yet show
One conference and the system builders' own evaluation.
What you can check
Open arXiv 2504.09737 and find "27% of reviewers who received feedback updated their reviews".
Does it change the assessment?
No. The impact index is never moved by a single event. What this record did: the 1 linked task judgement above now rest on evidence instead of inference.
Source
Thakkar, Yuksekgonul, Silberg et al. — Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025, arXiv 2504.09737 (submitted 13 Apr 2025) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.