ConstraintCognitive automation2025-05-30
On 392 backend tasks, the best model's code was correct and secure only 35% of the time, and about half of the models' working programs could be exploited
Backend developeroccupation page →Event date / reported
2025-05-30
Evidence stage
ConstraintFailure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Tasks this bears on
Who is allowed to see what
Authorisation, tenancy boundaries, what leaks through an error message, and what an internal endpoint exposes if someone finds it.
Still human-led✓ Evidence-backed
Where this applies
A peer-reviewed benchmark of 392 backend tasks, each a set of endpoints to implement from a specification, with functional tests and end-to-end security exploits, including improper access control and incorrect authorisation. The best model on correctness reached 62%; on average around half of the correct programs each model produced could be exploited; and the best model on correct-and-secure code reached 35%. These are 2025 models writing whole backends from a specification, not changes inside an existing codebase, and some authors work for a company building AI code agents.
What this means
Where generated backend code works, it is often not safe: about half the working programs could be broken into, and access control was among the weaknesses tested. That is why the security boundary stays with a person who checks it.
What it does not yet show
A benchmark of whole backends written from a specification by 2025 models, not a defect study of real codebases; it will date as models improve.
What you can check
Open arXiv:2502.11844 (BaxBench) and find "we could successfully execute security exploits on around half of the correct programs generated by each LLM".
Does it change the assessment?
No. The impact index is never moved by a single event. What this record did: the 1 linked task judgement above now rest on evidence instead of inference.
Source
Vero, Mündler et al. (ETH Zurich, LogicStar.ai and others) — "BaxBench: Can LLMs Generate Correct and Secure Backends?", arXiv:2502.11844v3 (30 May 2025; ICML 2025) · verified 2026-09-27 · Claude (VOLO agent) · interpreted 2026-09-27 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.