ConstraintCognitive automation2025-07-10
A METR randomised trial of 16 experienced open-source developers on 246 real issues found they took 19% longer with early-2025 AI tools, while believing they had been 20% faster
Junior software developeroccupation page →Event date / reported
2025-07-10
Evidence stage
ConstraintFailure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Tasks this bears on
Writing routine code
CRUD endpoints, forms, standard integrations, tests for known patterns.
Automating✓ Evidence-backed
Debugging unfamiliar systems
Finding why something broke when the cause is not where the symptom is.
Being augmented✓ Evidence-backed
Where this applies
Experienced maintainers on large, mature repositories they know well (about 5 years each); tools were mainly Cursor Pro with Claude 3.5/3.7 Sonnet. The authors explicitly do not claim the result generalises to most developers or to greenfield or junior work.
What this means
A constraint record on the capability layer with a measured result: in a randomised trial, experienced maintainers took 19% longer with early-2025 tools on real issues in repositories they knew, while believing they were 20% faster. For the linked tasks it caps how far 'routine code is generated' may be read as 'the work got faster', and the perception gap is direct evidence of why debugging generated code is where the ceiling shows.
What it does not yet show
The authors say so themselves: 16 developers on mature repositories they had worked in for about five years, using tools that have since changed. It does not show junior developers get slower — they lack the repository knowledge that made the tools less useful here — and it does not measure code quality, learning, or greenfield work. A slowdown in this setting is not a slowdown in yours.
What you can check
For one week, before starting each ticket write down your time estimate with and without the assistant, then record the actual time; METR's finding was the gap between those columns, and a personal log of ten tickets shows whether you have the same gap.
Does it change the assessment?
No. The impact index is never moved by a single event. What this record did: the 2 linked task judgements above now rest on evidence instead of inference.
Source
METR — study report · verified 2026-09-10 · Claude (VOLO agent) — source text fetched and cross-checked · interpreted 2026-09-10 · Claude (VOLO agent)