Data scientist — tasks, one by one
The unit of analysis is the task, not the job title. Each one below carries its direction, whether the judgement rests on evidence or on platform inference, the reasoning, and what it does not establish.
Every task on this page#
Framing the question and the metric
Still human-led≈ Platform inferenceWorking out what the decision-maker actually needs to know, and which measurable quantity would answer it without rewarding the wrong thing.
The input is a conversation with people who do not yet know what to ask, and the risk is a metric that is easy to move and wrong to optimise. O*NET's newest candidate task for the occupation is interviewing stakeholders to identify the questions data analysis should address — work being added, not removed.
This rests on the nature of the work and on O*NET's candidate task list; no record here measures how framing is done or who does it.
Cleaning and exploring the data
Being augmented✓ Evidence-backedFinding what is wrong with the data, deciding what to fix or exclude, and getting a feel for what it can and cannot say.
AI agents can write the wrangling code, but on a benchmark of realistic data-analysis tasks the best agent solved only about a third of them. Where the model is high-risk, the EU's AI Act makes the data work itself a documented duty: training, validation and testing data must be subject to data-governance practices covering cleaning and examination for possible biases.
Benchmark tasks are not a company's data, and the EU duties for high-risk systems apply later; nothing here measures time spent on cleaning.
Building and validating the model
Being augmented≈ Platform inferenceChoosing an approach, fitting and tuning it, and checking that it holds up on data it has not seen.
This is where AI agents are most capable and still fall well short: on public benchmarks the best agents solve roughly a third of data-analysis tasks and reach about 30% accuracy on data-science coding tasks. The data scientist increasingly directs and checks generated code rather than writing all of it.
Benchmarks date quickly and test short tasks; they do not show how models are built on real projects or how much of the work agents now do.
Designing experiments and reading causes
Still human-led≈ Platform inferenceSetting up the A/B test or the study so the answer can be trusted, and saying whether a change caused the result or merely came with it.
Causal reasoning is where models are weakest in the evidence here: on a benchmark of statistical and causal questions from textbooks and papers, the strongest model reached 58% accuracy, and its authors found models struggle to use causal knowledge and the provided data at the same time.
The benchmark tested 2024-era models on textbook questions; newer models may do better, and it does not measure experiments on real products.
Explaining what it means, and what it does not
Still human-led≈ Platform inferencePresenting the result so the person deciding understands the uncertainty, the assumptions and what would change the answer.
A decision rests on trusting the person who says what the numbers mean and where they stop. US labour projections expect data scientist employment to grow 35 percent from 2025 to 2035, far faster than average, as organisations bring AI systems into their work.
A projection is not a result, and it counts jobs, not which tasks those jobs involve.
Challenging and auditing models
New task✓ Evidence-backedIndependently testing someone else's model — whether it works, where it fails, whether it treats groups unfairly — and documenting it for a regulator or auditor.
The rules around models are written for people: US banking supervisors expect effective challenge by individuals with the expertise, independence and standing to change a model, and leave generative and agentic AI outside that guidance for now; New York City bars using an automated hiring tool without a bias audit in the past year, by an auditor independent of the tool.
The banking guidance is not enforceable and covers large banks; the New York rule covers hiring tools in one city. Neither measures how many people do this work.