CapabilityCognitive automation2023-03-15
ChatGPT produced fewer under-corrections but more over-corrections than dedicated grammar tools in an evaluation on a correction benchmark
Editor and proofreaderoccupation page →Event date / reported
2023-03-15
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Proofreading and correction
Finding and fixing errors of spelling, grammar and consistency.
Being augmented✓ Evidence-backed
Where this applies
A benchmark evaluation of grammatical error correction. The authors report that ChatGPT performs worse than dedicated baselines such as Grammarly on automatic metrics, particularly on long sentences, and that human evaluation suggests it produces fewer under-correction or mis-correction issues but more over-corrections. It used a 2023 model.
What this means
A model rewrites more than a proofreader should — which is why someone still has to decide which changes to keep.
What it does not yet show
A 2023 benchmark; newer models may behave differently.
What you can check
Open the arXiv paper "ChatGPT or Grammarly?" (2303.13648) and find "more over-corrections".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Wu et al. — ChatGPT or Grammarly? Evaluating ChatGPT on Grammatical Error Correction Benchmark, arXiv:2303.13648 (submitted 15 Mar 2023) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.