Get told when a verified record lands on this occupation → · Mark which of these tasks are yours (VOLO Pro, free during the launch) →
Data scientist
Turns a business or research question into something data can answer: frames it, cleans and explores the data, builds and validates statistical or predictive models, designs and reads experiments, and explains to decision-makers what the result does and does not show. AI agents can now write much of the analysis code, but on public benchmarks the best still solve about a third of realistic data-analysis tasks, and in causal questions they struggle most. Banking supervisors still expect models to be challenged by independent people, and US projections expect data science jobs to grow.
Take one recent analysis you did and write down whether it showed cause or only association, and what experiment would have settled it.
This is not a probability of losing your job. It combines how much of the role's task load is exposed to automation with how far adoption has actually gone — useful for comparing occupations on one consistent basis, and for nothing else.
Written for data scientists who build models and run experiments to inform decisions — in technology companies, banks and insurers, research and the public sector. Analysts working mainly in SQL and dashboards have their own page, as do engineers who build and run learned systems in production. The evidence is benchmarks of AI agents on data-science tasks, banking supervisors' model-risk guidance, rules on auditing automated decisions, and US labour projections; it establishes what agents can do on test tasks and what rules expect of people, not how data scientists' time has changed.
What is actually changing#
The unit of analysis is the task, not the job title. A role is not replaced — its task mix shifts.
Each tile is one task. Its size is how much of the job it is; its colour is where the task is heading. Click a tile to see what the judgement does not establish.
This rests on the nature of the work and on O*NET's candidate task list; no record here measures how framing is done or who does it.
Benchmark tasks are not a company's data, and the EU duties for high-risk systems apply later; nothing here measures time spent on cleaning.
Benchmarks date quickly and test short tasks; they do not show how models are built on real projects or how much of the work agents now do.
The benchmark tested 2024-era models on textbook questions; newer models may do better, and it does not measure experiments on real products.
A projection is not a result, and it counts jobs, not which tasks those jobs involve.
The banking guidance is not enforceable and covers large banks; the New York rule covers hiring tools in one city. Neither measures how many people do this work.
Is this your job? Say so and this page narrows to your share of it.
A job title is a bundle of tasks bought together, and no two people hold the same bundle. Nothing is sent anywhere — it stays in this browser.
Read all 6 tasks in full — direction, reasoning and limits →
Recent changes#
United States. The statistics bureau's occupational outlook projects that employment of data scientists will grow 35 percent from 2025 to 2035, much faster than the average for all occupations, with about 24,800 openings a year on average over the decade. It is a projection for one country, made while firms integrate AI systems into their work; it counts jobs, not which tasks those jobs involve.
A named person with standing publicly predicted something, on a date, in an attributable statement. It is recorded so that who said what, and when, stays checkable — and it never moves a task's assessment, because a prediction is not an observation. Its value arrives later: the record sits on the same page as the evidence about that occupation, so anyone reading the forecast reads the record of what happened next beside it. That is the reckoning; this site publishes no verdict on whether a forecast came true.
United States. The banking agencies' revised model-risk guidance replaces SR 11-7 of 2011. It says effective challenge is performed by individuals with the appropriate expertise to conduct a critical and objective challenge, sufficient independence to maintain objectivity, and the organisational standing and influence to effect change; it states that generative AI and agentic AI models are not within its scope; and it says the guidance does not set enforceable standards, so non-compliance will not result in supervisory criticism. It is most relevant to banking organisations with over $30 billion in total assets.
Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
An academic benchmark of data-science coding tasks that require an agent to work with real data and write code, including wrangling and analysis. It reports that, with its baseline agent, the current best language models achieve only 30.5% accuracy, leaving ample room for improvement. It tests 2024-era models on designed tasks.
A demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
A benchmark of 466 data-analysis tasks and 74 data-modelling tasks drawn from Eloquence and Kaggle competitions, used to test AI agents built on models such as GPT-4o, Claude and Gemini. It reports that the best agent solved only 34.12% of the data-analysis tasks and reached a 34.74% relative performance gap on modelling tasks. It tests models available in 2024–2025 on competition tasks, not work on a company's data; one author's employer develops AI models.
A demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
European Union. Article 10 of the AI Act says training, validation and testing data sets for high-risk AI systems shall be subject to data governance and management practices, including data preparation such as annotation, labelling and cleaning, and examination in view of possible biases; Article 14 says high-risk systems shall be designed so that they can be effectively overseen by natural persons. The obligations bind providers of high-risk systems; under the 2026 amending regulation, the high-risk obligations for systems listed in Annex III apply from 2 December 2027.
Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
An academic benchmark of 411 questions with data sheets from textbooks, online learning materials and academic papers, testing statistical and causal reasoning. It reports that the strongest model, GPT-4, achieves 58% accuracy, and that models have difficulty with data analysis and causal reasoning and struggle to use causal knowledge and provided data simultaneously. It tests 2024-era models on textbook-style questions, not experiments on real products.
A demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
New York City. The final rule implementing the city's law on automated employment decision tools says an employer or employment agency may not use or continue to use such a tool if more than one year has passed since its most recent bias audit; that a bias audit must at a minimum calculate the selection rate and impact ratio for each category; and that an auditor is not independent if it is or was involved in using, developing or distributing the tool. The tools it covers include machine learning, statistical modelling and data analytics. It applies to hiring and promotion in one city.
Regulation, subsidy or public procurement is requiring or funding adoption — the mirror of a constraint. It shows adoption is being required, not that it has happened, so one mandate is never enough on its own; two independent ones are.
What this means for you#
If you are starting out, the parts of the job an AI agent does best — writing routine analysis and modelling code — are the parts junior roles used to be built on. Build your value where the evidence says agents are weakest: framing the question, designing experiments and reasoning about causes, and explaining results to people who must decide.
Expect to review and direct generated code more than write it, and expect more of your time to go to validation, documentation and audits as rules on automated decisions spread. Your standing to challenge a model — and to be believed when you say where it stops — is the part the rules are written around.
Your options#
Four directions, each with its real constraints and one thing you can test this week. Continuing as you are is a legitimate choice — it just has to be a chosen one.
Stay a data scientist, and move towards experiments and causal work
Causal reasoning and experiment design are where the benchmarks show models struggling most, and decisions depend on them.
Experimentation roles sit mainly in large product companies and research; smaller firms may not run enough tests to need one.
Take one recent analysis you did and write down whether it showed cause or only association, and what experiment would have settled it.
Take on model validation and audit
Banking supervisors expect independent challenge of models and New York City requires independent bias audits of hiring tools; the work needs people who understand models and are not the ones who built them.
Independence means not auditing your own team's models, which may mean moving teams or employers; the banking guidance is not binding.
Find out whether your organisation has a model validation or model risk function, and what it checks before a model goes live.
Move towards decision support: product or policy analytics
Explaining what a result means to the people deciding is the part of the job least touched by automation in the evidence here.
These roles reward domain knowledge and communication over technical depth, and pay can differ.
Ask one decision-maker you work with which of your recent results changed a decision, and why.
Common questions#
Not on present evidence. AI agents can write much of the analysis code, but on public benchmarks the best still solve about a third of realistic data-analysis tasks, and they are weakest at causal reasoning. Rules on models expect independent people to challenge and audit them, and US projections expect data scientist jobs to grow. What changes is the mix: less hand-written code, more framing, checking and explaining.
We do not answer that with a number of years. There is a signal you can watch instead: whether AI agents start matching people on causal and experimental questions in benchmarks, and whether rules on models stop requiring an independent person to challenge them. The first is not true today; the second is the opposite of the current direction.
On DSBench, a benchmark of 466 data-analysis and 74 data-modelling tasks from real competitions, the best agent solved about 34% of the analysis tasks; on DA-Code the best models reached about 30.5% accuracy; and on QRData, a benchmark of statistical and causal questions, the strongest model reached 58%. Benchmarks date quickly, but they show where agents are strongest and where they struggle.
US projections expect data scientist employment to grow 35 percent from 2025 to 2035, far faster than average. The work is shifting towards framing questions, designing experiments, validating and auditing models and explaining results — the parts where agents are weakest and rules expect a person.
What these judgements rest on#
2 of 6 task judgements on this page are backed by a verified event and 4 are platform inference, each labelled where it appears. Behind them sit 2 technology dimensions, a reconstructed trajectory since language models reached the public, and 7 verified events.
See which technologies, how it got here, and the method →
Where it sits in the official classification: skills, knowledge, related jobs →
Other roles in the same function#
A company divides its work into functions before it divides it into jobs. These sit in Technology & data alongside this one — a fact about org charts, not a judgement that they are similar or that they are changing in the same direction.
Junior software developer · Experienced software engineer · Frontend developer · Backend developer · Data engineer · Machine learning engineer · AI researcher · Software tester / QA engineer · Data analyst · Business analyst · Business systems owner · IT support specialist / helpdesk · Security analyst (SOC) · DevOps / platform / SRE engineer · Network engineer · Technical writer / documentation engineer
