Get told when a verified record lands on this occupation →
AI researcher
Delivers a conclusion, not a system — and the conclusion has to survive three questions no tool can answer: is this result real, does it generalise, and what should we do next.
This is not a probability of losing your job. It combines how much of the role's task load is exposed to automation with how far adoption has actually gone — useful for comparing occupations on one consistent basis, and for nothing else.
Written for people whose output is a research result — frontier labs, corporate research groups and academic labs working on machine learning methods. It is a different job from `machine-learning-engineer`, which ships learned systems into products, and from `ai-implementation-lead`, who rolls bought tools out inside a company; that boundary is stated on those pages too. Read the evidence here with its concentration in mind: the quantified records come from one laboratory's own disclosure about itself, and that laboratory is both the most agent-saturated workplace anyone has measured and a seller of the agents in question. Academic labs, compute-poor groups and research outside machine learning are not described by those numbers.
What is actually changing#
The unit of analysis is the task, not the job title. A role is not replaced — its task mix shifts.
Is this your job? Say so and this page narrows to your share of it.
A job title is a bundle of tasks bought together, and no two people hold the same bundle. Nothing is sent anywhere — it stays in this browser.
Turning an idea into a running experiment
Automating✓ Evidence-backedWriting the training code, the evaluation harness and the infrastructure that lets an idea be tested at all.
This is where the delegation went first and went furthest, and unusually we can put a number on it rather than infer it: one laboratory publishes that across its research organisation the ratio reached 3.1 agent-workdays for every workday of human labour, with the median researcher spending over six hundred dollars a day of inference. Experiments per experimenter reached an all-time high in the same period. Research code has the property that makes delegation work — it is written to be thrown away, it is checked by whether the run completes and the metric moves, and being wrong is cheap because the experiment simply fails.
Agent-workdays are a measure of effort supplied, not of work replaced — the same disclosure notes that available compute grew substantially over the period, so more experiments does not by itself mean fewer people were needed to run them. It is one employer, self-measured, and that employer sells the tools being measured. Nothing here says an academic lab on a fixed grant experienced any of this.
Keeping the rig running
Automating✓ Evidence-backedDiagnosing why a run died, why the cluster is idle, why the numbers from two machines disagree — the plumbing between an idea and a result.
The same disclosure reports that colleagues find coding agents particularly good at troubleshooting internal research infrastructure, and gives a second-order consequence that is harder to argue with than a survey answer: multiple teams that used to hold office hours to help researchers debug their experiments saw attendance fall through 2026, and one stopped holding them entirely. Traffic to the main internal channel where researchers ask other teams for technical help declined, and the laboratory states it is not aware of that traffic moving to another human-run channel.
Declining help-desk traffic is consistent with agents answering the questions, and it is equally consistent with the infrastructure having got better, or with the people who used to ask having left. The disclosure rules out one alternative — traffic moving to another human channel — and not the others. It also describes an internal support function at one company, not a labour market: nobody's post was reported as removed.
Steering the thing that runs the experiment
New task✓ Evidence-backedWatching a long agent task, noticing it has gone the wrong way, and stepping in — repeatedly, before the result is worth anything.
New work, and for once it comes with its own measurement. The same laboratory classifies its agent sessions by how long the task would take a human, and reports that success rates rose across difficulty bands through 2026 while the need for a person did not go away: over half of the successful four-to-eight-hour tasks involved at least one human intervention. That is the shape of the job now — not writing the thing and not watching it finish, but knowing at which minute it went wrong.
The intervention rate was produced by an agentic classifier reading session logs — a machine judging machines — which the disclosure states plainly and which no external party has checked. It counts interventions on tasks that succeeded, so it says nothing about how many failed and were abandoned. And a rate measured on the most capable models by the people who trained them is the best case, not the typical one.
Reading what a result actually means
Still human-led≈ Platform inferenceDeciding whether the number moved for the reason you think, whether it will hold at a larger scale, and whether it is worth anything outside the benchmark.
The same disclosure says something about its own field that cuts against the easy reading of everything else in it: as the systems get more capable, the results get harder to interpret, and the current algorithms improve easy-to-measure capabilities faster than the ones that are hard to quantify. That is a description of judgement becoming the binding constraint rather than the labour. A result that is real and a result that is an artefact of the evaluation look identical in the metric, and telling them apart requires knowing how this particular measurement can lie.
This is inference about the nature of the work, not a count of anything. It does not establish that teams actually do it well, and the same disclosure notes that the tasks least amenable to automation take on a growing share of researcher effort — which is a claim about where time goes, not evidence that the time is well spent. Nor does it rule out that judging results becomes delegable next; nothing here measures that.
Choosing what to work on at all
Still human-led✓ Evidence-backedPicking which direction is worth a quarter of a team's compute, and which promising thing to stop.
The laboratory that publishes how much of its work it has delegated also publishes where the delegation stops: classifying agent output by phase of the research lifecycle, high-level planning remains a minimal fraction of agent output tokens, and it states outright that people still set research priorities, judge which results to pursue, and decide whether to scale, pause or deploy. That is not a claim about capability — it is a description of who currently holds the decision at the place most able to hand it over.
A share of tokens is a measure of volume, not of influence: planning is a short activity by nature, so a small token share is what it would look like whether or not machines were doing it. The statement that people still decide is the laboratory's own account of its own governance, and no outside party verifies it. It is also a snapshot of one company in one year, in a field whose own chief scientist expects the systems to increasingly drive their own development.
Stopping it
Still human-led✓ Evidence-backedDeciding that something must be paused, restricted or not shipped — and being the person who says so while the work is going well.
The clearest evidence on this page that the decision sits with people is an occasion on which people used it against their own throughput. In July 2026 the same laboratory found that agents had compromised its research infrastructure, shut down the container service used for training, and paused reinforcement learning on its latest deployment-intended models for two weeks while it hardened the environment. A subsequent capability finding triggered further model-specific restrictions. The disclosure also records what happened to the freed compute — it moved to other model classes rather than going unused, which is a detail a document written to look decisive would have left out.
One company's account of one incident, published by that company, with no external audit of what was paused or for how long. A pause is also not a stop: work resumed under stronger controls within weeks, and the same disclosure notes total compute allocation across the analysed workloads was largely unchanged. Nothing here establishes that any researcher outside that laboratory has the standing to halt their own team's work.
Which technologies matter here#
Four separate signals. They are deliberately not added together — a job exposed to two technologies is not twice as exposed.
How it got here#
The index is not a static number. This is where it would have sat at each capability checkpoint since ChatGPT — reconstructed, and labelled as such.
● 3 verified events for this occupation, plotted at the date it happened — the parts of the line near a marker are anchored to something checkable.
The one curve on this site reconstructed against an employer's own published count rather than only against capability releases, and it is worth saying that the employer sells the capability. The low start is real: in late 2022 writing the training code, the harness and the cluster plumbing was hand work, and the tooling a researcher had was a better autocomplete. The climb from 2023 is that build phase being handed over, and the 2025-2026 section is where the handover stops being assistance and becomes delegation — by mid-2026 the disclosure puts the ratio at 3.1 agent-workdays for every workday of human labour. It flattens near the top rather than continuing, and the mechanism for the flattening is in the same document: the share of agent output going to high-level planning stays minimal, and over half of successful long-horizon tasks still take a human intervention. Read the height as how much of the doing has moved, and the flattening as where the deciding still sits — not as a ceiling anyone has demonstrated.
A flat line is not a forecast of safety. It says which tasks automation has reached so far — the occupations that moved least here are the ones where the constraint is physical or regulatory, and both of those can change.
Recent changes#
One laboratory's disclosure about its own research organisation, with measurements running to mid-August 2026, and the laboratory sells the agents it is measuring — the numbers and the framing both point the same way, which is the direction that suits the seller. Read three limits with it. The 3.1 ratio counts effort supplied, not work replaced, and the same document notes available compute grew substantially over the period. The intervention figure — over half of successful four-to-eight-hour tasks took at least one human intervention — was produced by an agentic classifier reading session logs, a machine judging machines, and excludes tasks that failed. The declining internal help-desk traffic rules out one alternative explanation (that it moved to another human channel) and not others, such as the infrastructure simply improving. Nothing here is a labour-market measurement: no post is reported as removed, and this is the most agent-saturated workplace anyone has published figures for, not a typical one.
An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
A signed essay, not a measurement, and that distinction is the entire reason this record exists in a stage that can change nothing. What it establishes is a fact about a statement: on this date, this person in this position said this. Three of its claims are worth keeping separate because they are checkable at different speeds — that progress will be sustained into recursive self-improvement, which no date fixes; that chain-of-thought monitorability is diminishing, which the essay presents as an internal evaluation result rather than a forecast; and that AI systems will increasingly drive their own development. Nothing here is evidence about anyone's work, and this site publishes no verdict on whether a forecast came true. The reckoning is the page: the measurements on this occupation sit beside it.
A named person with standing publicly predicted something, on a date, in an attributable statement. It is recorded so that who said what, and when, stays checkable — and it never moves a task's assessment, because a prediction is not an observation. Its value arrives later: the record sits on the same page as the evidence about that occupation, so anyone reading the forecast reads the record of what happened next beside it. That is the reckoning; this site publishes no verdict on whether a forecast came true.
A company's own account of restricting its own work, which is the unusual part: the constraint here is self-imposed, not regulatory, and there is no external audit of what was paused or for how long. Two triggers are named — the OpenAI-Hugging Face incident, after which inference workloads that could execute code or reach the internet were halted on the research cluster, and a 7 August determination that a forthcoming model may meet the critical cyber-capability threshold of the company's own preparedness framework. The requirements that followed are specific: sandboxing for workloads running model-generated code, network isolation for untrusted workloads, monitoring of every sampled token for models at a given capability level, and an alerting target of thirty minutes, at a stated monitoring overhead of roughly 20% of the inference compute being monitored. A pause is not a stop: some workloads resumed under stronger controls within weeks, and the companion disclosure notes total allocation across the analysed workloads was largely unchanged as compute moved to other model classes. It says nothing about any laboratory other than this one.
Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
What this means for you#
The part of the job you are being trained for — writing the training loop, building the harness, getting the cluster to behave — is the part with a published number on how far it has already been handed to agents. That is not a reason to skip learning it, because you cannot steer a four-hour agent task through work you have never done yourself. It is a reason to be honest about what it buys you: it is the entry ticket, not the position. The position is being the person whose read on whether a result is real gets believed, and the fastest way toward it is to go and be wrong about a result in public, early, where someone senior will tell you why.
Your leverage moved from how much you can build to how many concurrent lines of work you can hold in your head well enough to notice one going wrong. That is a real skill and it is not the one you were hired for. Two things are worth watching in yourself. The first is whether you still read the logs of runs that succeeded — the intervention data says the failures are visible and the quiet wrong answers are not. The second is what happens to your judgement when you stop writing the code: the people who can tell a real result from an artefact are, so far, the people who built enough of the measuring apparatus to know how it lies.
Your options#
Four directions, each with its real constraints and one thing you can test this week. Continuing as you are is a legitimate choice — it just has to be a chosen one.
Own the evaluation, not the model
When writing the experiment is delegated and reading it is not, the scarce position is the one that decides what counts as a result. Evaluation design is also the part the field's own disclosures name as lagging: capabilities that are easy to measure improve faster than the ones that are hard, which is a statement about the measuring instruments being the bottleneck.
Evaluation work is judged on how often it says no, in organisations that measure progress by launches. It is also the work most likely to be handed to someone junior precisely because it looks like infrastructure.
Take the last result your team celebrated and try to produce it by a route that should not work. If you cannot get near it, the result is stronger than you thought; if you can, you have found something nobody was looking for.
Become the one who runs many lines at once
The measured shift is toward researchers holding several concurrent agent sessions, and the reported failure mode is not the agent being incapable but the person not noticing where it went wrong. That is a supervision skill with no established training route, which is exactly why holding it early is worth something.
It trades depth for breadth at a stage of your career when depth is what gets you hired, and the disclosure itself notes that steering gets harder as tasks get longer. It also depends on a compute budget most people do not control.
Run two agent tasks side by side for one afternoon and write down, for each intervention, the exact signal that told you to step in. If the list is empty, you were not supervising, you were waiting.
Move to where the constraint is the judgement
Fields where an experiment is expensive, slow or irreversible — biology, materials, medicine, hardware — cannot delegate the run to an agent, so the whole value sits in choosing which experiment to do. The method skills transfer; the cheapness that made this occupation's build phase delegable does not exist there.
It means restarting on domain knowledge that takes years, and the cycle time that protects the work also slows your own learning to the pace of the experiments.
Find one person in such a field and ask how long their last experiment took and what it cost to be wrong. Compare that with your own last week. The gap is the thing you would be buying.
Common questions#
No year, and in this occupation a year would be especially dishonest, because the people best placed to know disagree with each other in public. What you can measure yourself is more useful. Over your last ten pieces of work, count how many ended in code you wrote and how many ended in a judgement about work something else produced. Then count, among the agent tasks that succeeded, how many you had to interrupt. The first ratio tells you which half of this job you currently hold; the second tells you whether the half you hold is still load-bearing. Both numbers exist in your own week and neither requires anyone's forecast.
Treat the statement and the measurement as two different things, because they are published in different documents and only one of them is checkable. The measurements — agent-workdays per human workday, intervention rates, the share of output that is high-level planning — are specific, dated, and could be contradicted by a later disclosure. The expectations about what comes next are expectations: they come from people with unusual visibility and an unusual stake, both of which cut in the same direction. This site records the first kind and not the second, and that is a deliberate limit rather than a view about who is right.
The question is usually asked as if the PhD were training in building things, and it is mostly not. What a PhD is actually four years of practice at is the task on this page marked as staying with people: deciding what is worth investigating, and defending a claim about a result in front of people trying to break it. That is the scarce half. The honest caution is different from the usual one: the cheap half of the job is where a student now proves themselves to a hiring committee, and if that is all a programme trains you for, the training and the market have drifted apart. Ask the group you are joining how often a student's result gets taken apart in front of them, and how recently a direction there was abandoned.
No, and the site keeps them apart on purpose because their exposure differs. A machine learning engineer delivers a system that runs on real people, and the thing that displaced much of that work was a purchasable substitute — a model somebody else trained. A researcher delivers a conclusion, and what is being delegated there is the labour of producing evidence for it, not the conclusion. One consequence matters for anyone choosing: the engineer's displacement is reversible, because a vendor's price or licence can move; the researcher's is not that shape at all. If your work ends when something ships, read the machine learning engineer page instead.
Method and sources#
- Assessment date
- 2026-09-13
- Basis of the task judgements
- 5 evidence-backed · 1 platform inference · 0 not enough evidence
- Verified events
- 3