CapabilityCognitive automation2024-10-25
AI grading of physics exams reached a coefficient of determination of about 0.91 when handling half of the grading load, with the rest left to people
University lectureroccupation page →Event date / reported
2024-10-25
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Marking and feedback
Marking scripts and coursework and giving students feedback.
Being augmented≈ Platform inference
Where this applies
Switzerland, physics exams at one university. The study uses psychometric thresholds to decide which answers AI grades and which go to people, and reports that AI can achieve a coefficient of determination of about 0.91 against human grades when handling half of the grading load, and about 0.96 for one-fifth of the load. It is one exploratory study of one exam type.
What this means
AI can take a share of the marking where it is confident and hand the rest to people. The split, and the responsibility for the grades, stays with the academic.
What it does not yet show
One exploratory study; it does not measure marking practice or time saved.
What you can check
Open arXiv 2410.19409 and find "when handling half of the grading load".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Kortemeyer, Nöhl (ETH Zurich) — Assessing Confidence in AI-Assisted Grading of Physics Exams through Psychometrics: An Exploratory Study, arXiv 2410.19409 (submitted 25 Oct 2024) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.