CapabilityCognitive automation2024-10-09
The best language models reached only 30.5% accuracy on DA-Code, a benchmark of agent-based data-science coding tasks
Data scientistoccupation page →Event date / reported
2024-10-09
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Building and validating the model
Choosing an approach, fitting and tuning it, and checking that it holds up on data it has not seen.
Being augmented≈ Platform inference
Where this applies
An academic benchmark of data-science coding tasks that require an agent to work with real data and write code, including wrangling and analysis. It reports that, with its baseline agent, the current best language models achieve only 30.5% accuracy, leaving ample room for improvement. It tests 2024-era models on designed tasks.
What this means
Writing the data-science code is where agents are most capable, and still they get most tasks wrong on this benchmark. The data scientist's role shifts towards directing and checking generated code.
What it does not yet show
A 2024 benchmark on designed tasks; newer models may score higher, and it does not measure real projects.
What you can check
Open arXiv 2410.07331 (DA-Code) and find "achieves only 30.5% accuracy".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Huang et al. — DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models, arXiv 2410.07331 (submitted 9 Oct 2024; EMNLP 2024) · verified 2026-09-29 · Claude (VOLO agent) · interpreted 2026-09-29 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.