CapabilityCognitive automation2024-09-12
The best AI agent solved only 34.12% of realistic data-analysis tasks on DSBench, a benchmark of 466 analysis and 74 modelling tasks from real competitions
Data scientistoccupation page →Event date / reported
2024-09-12
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Cleaning and exploring the data
Finding what is wrong with the data, deciding what to fix or exclude, and getting a feel for what it can and cannot say.
Being augmented✓ Evidence-backed
Building and validating the model
Choosing an approach, fitting and tuning it, and checking that it holds up on data it has not seen.
Being augmented≈ Platform inference
Where this applies
A benchmark of 466 data-analysis tasks and 74 data-modelling tasks drawn from Eloquence and Kaggle competitions, used to test AI agents built on models such as GPT-4o, Claude and Gemini. It reports that the best agent solved only 34.12% of the data-analysis tasks and reached a 34.74% relative performance gap on modelling tasks. It tests models available in 2024–2025 on competition tasks, not work on a company's data; one author's employer develops AI models.
What this means
Agents can do some of the analysis on their own, but on realistic tasks the best still get about a third right. The work is being assisted, and the checking of what the agent produced stays with the data scientist.
What it does not yet show
A benchmark of competition tasks that dates quickly; it does not measure real projects or how much of the work agents now do.
What you can check
Open arXiv 2409.07703 (DSBench) and find "the best agent solving only 34.12% of data analysis tasks".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 2 linked judgements above stand where they were.
Source
Jing et al. (UT Dallas, Tencent AI Lab, USC) — DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?, arXiv 2409.07703 (submitted 12 Sep 2024; revised 11 Apr 2025) · verified 2026-09-29 · Claude (VOLO agent) · interpreted 2026-09-29 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.