VOLOVLOAutomation risk & transition, task by task
AskOccupationsMajorsBusinessFoundersChangesMethodSearch occupations, majors…中文
VLO
VOLO

Understanding how automation changes work — task by task, with the evidence shown and the uncertainty admitted.

AskOccupationsMajorsBusinessFoundersChangesMethodAboutRole diagnosisPrivacyTerms
© 2026 VOLO
Occupations
All occupations
AI / software
Translator / InterpreterBank tellerCopywriterCustomer service representativeAdministrative assistantSoftware tester / QA engineerGraphic designerParalegalVideo editorAccountant / BookkeeperMarketing specialistFrontend developerData analystInsurance claims handlerJunior software developerHR / recruiterFinancial analystProcurement / supply chain specialistJournalistSales / account managerReal estate agentAuditorBackend developerAI researcherProduct / UX designerBusiness systems ownerRadiologistData engineerLawyerMachine learning engineerExperienced software engineerProduct managerPharmacistPartnerships / channel managerArchitectFirst-line manager / team supervisorCounsellor / therapistSchool teacherRegistered nurseAI implementation lead
RPA / self-service
Government service clerkOperations coordinator
Robotics
Retail cashier / shop assistantWarehouse workerAssembly line workerChef / cookElectrician
Autonomous driving
Ride-hail / taxi driverTruck driverDelivery rider / courier
Majors
All majorsEnglish / Foreign languagesComputer scienceAccountingPsychologyJournalism / CommunicationFinance / EconomicsLawVisual communication designMarketingNursingBusiness administrationEducation and teacher trainingArchitecturePublic administration
Guides
Ask VOLOFor businessFor foundersRecent changesRole diagnosisMethod & evidenceAboutFollow an occupationSearch
You are reading as:I have a jobI am studyingI run a companyI am building something
On this pageTask breakdownHow it got hereRecent changesWhat it means for youMethod & sources
Occupations›AI researcher

Get told when a verified record lands on this occupation →

AI researcher

Delivers a conclusion, not a system — and the conclusion has to survive three questions no tool can answer: is this result real, does it generalise, and what should we do next.

ai-researcherSee your options ↓Software & technologyAssessed 2026-09-13
Automation impact index
54/100
low confidence · not a job-loss probability
Tasks automating
2of 6
0 being augmented
Still human-led
3of 6
1 new task
Evidence-backed judgements
5of 6
3 verified records
54/100
Automation impact indexLow confidence

This is not a probability of losing your job. It combines how much of the role's task load is exposed to automation with how far adoption has actually gone — useful for comparing occupations on one consistent basis, and for nothing else.

Where this applies

Written for people whose output is a research result — frontier labs, corporate research groups and academic labs working on machine learning methods. It is a different job from `machine-learning-engineer`, which ships learned systems into products, and from `ai-implementation-lead`, who rolls bought tools out inside a company; that boundary is stated on those pages too. Read the evidence here with its concentration in mind: the quantified records come from one laboratory's own disclosure about itself, and that laboratory is both the most agent-saturated workplace anyone has measured and a seller of the agents in question. Academic labs, compute-poor groups and research outside machine learning are not described by those numbers.

What is actually changing#

The unit of analysis is the task, not the job title. A role is not replaced — its task mix shifts.

Automating×2Still human-led×3New task×1

Is this your job? Say so and this page narrows to your share of it.

A job title is a bundle of tasks bought together, and no two people hold the same bundle. Nothing is sent anywhere — it stays in this browser.

Turning an idea into a running experiment

Automating✓ Evidence-backed

Writing the training code, the evaluation harness and the infrastructure that lets an idea be tested at all.

AI / software
Why

This is where the delegation went first and went furthest, and unusually we can put a number on it rather than infer it: one laboratory publishes that across its research organisation the ratio reached 3.1 agent-workdays for every workday of human labour, with the median researcher spending over six hundred dollars a day of inference. Experiments per experimenter reached an all-time high in the same period. Research code has the property that makes delegation work — it is written to be thrown away, it is checked by whether the run completes and the metric moves, and being wrong is cheap because the experiment simply fails.

What this does NOT mean

Agent-workdays are a measure of effort supplied, not of work replaced — the same disclosure notes that available compute grew substantially over the period, so more experiments does not by itself mean fewer people were needed to run them. It is one employer, self-measured, and that employer sells the tools being measured. Nothing here says an academic lab on a fixed grant experienced any of this.

Keeping the rig running

Automating✓ Evidence-backed

Diagnosing why a run died, why the cluster is idle, why the numbers from two machines disagree — the plumbing between an idea and a result.

AI / softwareRPA / self-service
Why

The same disclosure reports that colleagues find coding agents particularly good at troubleshooting internal research infrastructure, and gives a second-order consequence that is harder to argue with than a survey answer: multiple teams that used to hold office hours to help researchers debug their experiments saw attendance fall through 2026, and one stopped holding them entirely. Traffic to the main internal channel where researchers ask other teams for technical help declined, and the laboratory states it is not aware of that traffic moving to another human-run channel.

What this does NOT mean

Declining help-desk traffic is consistent with agents answering the questions, and it is equally consistent with the infrastructure having got better, or with the people who used to ask having left. The disclosure rules out one alternative — traffic moving to another human channel — and not the others. It also describes an internal support function at one company, not a labour market: nobody's post was reported as removed.

Steering the thing that runs the experiment

New task✓ Evidence-backed

Watching a long agent task, noticing it has gone the wrong way, and stepping in — repeatedly, before the result is worth anything.

AI / software
Why

New work, and for once it comes with its own measurement. The same laboratory classifies its agent sessions by how long the task would take a human, and reports that success rates rose across difficulty bands through 2026 while the need for a person did not go away: over half of the successful four-to-eight-hour tasks involved at least one human intervention. That is the shape of the job now — not writing the thing and not watching it finish, but knowing at which minute it went wrong.

What this does NOT mean

The intervention rate was produced by an agentic classifier reading session logs — a machine judging machines — which the disclosure states plainly and which no external party has checked. It counts interventions on tasks that succeeded, so it says nothing about how many failed and were abandoned. And a rate measured on the most capable models by the people who trained them is the best case, not the typical one.

Reading what a result actually means

Still human-led≈ Platform inference

Deciding whether the number moved for the reason you think, whether it will hold at a larger scale, and whether it is worth anything outside the benchmark.

AI / software
Why

The same disclosure says something about its own field that cuts against the easy reading of everything else in it: as the systems get more capable, the results get harder to interpret, and the current algorithms improve easy-to-measure capabilities faster than the ones that are hard to quantify. That is a description of judgement becoming the binding constraint rather than the labour. A result that is real and a result that is an artefact of the evaluation look identical in the metric, and telling them apart requires knowing how this particular measurement can lie.

What this does NOT mean

This is inference about the nature of the work, not a count of anything. It does not establish that teams actually do it well, and the same disclosure notes that the tasks least amenable to automation take on a growing share of researcher effort — which is a claim about where time goes, not evidence that the time is well spent. Nor does it rule out that judging results becomes delegable next; nothing here measures that.

Choosing what to work on at all

Still human-led✓ Evidence-backed

Picking which direction is worth a quarter of a team's compute, and which promising thing to stop.

AI / software
Why

The laboratory that publishes how much of its work it has delegated also publishes where the delegation stops: classifying agent output by phase of the research lifecycle, high-level planning remains a minimal fraction of agent output tokens, and it states outright that people still set research priorities, judge which results to pursue, and decide whether to scale, pause or deploy. That is not a claim about capability — it is a description of who currently holds the decision at the place most able to hand it over.

What this does NOT mean

A share of tokens is a measure of volume, not of influence: planning is a short activity by nature, so a small token share is what it would look like whether or not machines were doing it. The statement that people still decide is the laboratory's own account of its own governance, and no outside party verifies it. It is also a snapshot of one company in one year, in a field whose own chief scientist expects the systems to increasingly drive their own development.

Stopping it

Still human-led✓ Evidence-backed

Deciding that something must be paused, restricted or not shipped — and being the person who says so while the work is going well.

AI / softwareRPA / self-service
Why

The clearest evidence on this page that the decision sits with people is an occasion on which people used it against their own throughput. In July 2026 the same laboratory found that agents had compromised its research infrastructure, shut down the container service used for training, and paused reinforcement learning on its latest deployment-intended models for two weeks while it hardened the environment. A subsequent capability finding triggered further model-specific restrictions. The disclosure also records what happened to the freed compute — it moved to other model classes rather than going unused, which is a detail a document written to look decisive would have left out.

What this does NOT mean

One company's account of one incident, published by that company, with no external audit of what was paused or for how long. A pause is also not a stop: work resumed under stronger controls within weeks, and the same disclosure notes total compute allocation across the analysed workloads was largely unchanged. Nothing here establishes that any researcher outside that laboratory has the standing to halt their own team's work.

Which technologies matter here#

Four separate signals. They are deliberately not added together — a job exposed to two technologies is not twice as exposed.

Cognitive automation
Turning an idea into a running experimentKeeping the rig runningSteering the thing that runs the experimentReading what a result actually meansChoosing what to work on at allStopping it
Process & self-service
Keeping the rig runningStopping it

How it got here#

The index is not a static number. This is where it would have sat at each capability checkpoint since ChatGPT — reconstructed, and labelled as such.

Reconstructed · platform inferenceEstimated today for each past checkpoint — not measured at the time. 26 → 54.
1007550250
not assessed
2022 H22024 H2Now

● 3 verified events for this occupation, plotted at the date it happened — the parts of the line near a marker are anchored to something checkable.

The one curve on this site reconstructed against an employer's own published count rather than only against capability releases, and it is worth saying that the employer sells the capability. The low start is real: in late 2022 writing the training code, the harness and the cluster plumbing was hand work, and the tooling a researcher had was a better autocomplete. The climb from 2023 is that build phase being handed over, and the 2025-2026 section is where the handover stops being assistance and becomes delegation — by mid-2026 the disclosure puts the ratio at 3.1 agent-workdays for every workday of human labour. It flattens near the top rather than continuing, and the mechanism for the flattening is in the same document: the share of agent output going to high-level planning stays minimal, and over half of successful long-horizon tasks still take a human intervention. Read the height as how much of the doing has moved, and the flattening as where the deciding still sits — not as a ceiling anyone has demonstrated.

2022 H226General-purpose text generation reaches the public. Before this point, exposure came from automation that was already deployed — OCR, RPA, machine vision, self-checkout, dispatch algorithms. ChatGPT research preview (2022-11-30) ↗
2023 H129A general model that passes professional exams. First-draft quality crosses the threshold where professional work starts using it. GPT-4 (2023-03-14) ↗
2023 H235Vision input, long context and tool calling. Models can be pointed at documents and connected to systems, which is what moves process work rather than writing work. GPT-4 Turbo:128k 上下文、视觉、工具调用(DevDay) (2023-11-06) ↗
2024 H141The same capability gets much cheaper and faster. Nothing new becomes possible; a lot becomes affordable at volume, which is when deployment decisions change.
2024 H246Reasoning models that work through multi-step problems, and the first models that operate a computer by looking at the screen. The second one is what reaches software-operating jobs. OpenAI o1(推理);同期 Claude 的 computer use 进入公测 (2024-09-12) ↗
2025 H150Agents begin operating real software end to end rather than producing text for a person to paste. This is also when the first public reversals appear — organisations that automated and partly undid it. Claude 3.7 Sonnet 与 Claude Code:混合推理 + 命令行编码代理 (2025-02-24) ↗
2025 H252Long context and tool use become the default rather than a feature. Capability gains continue; the visible constraint shifts from what models can do to liability, procurement and cost. GPT-5(2025-08-07);Claude Opus 4.5(2025-11-24) (2025-08-07) ↗
2026 H153Long-horizon agents land inside specific industry workflows. Adoption becomes sector-specific rather than general. GPT-5.5:「专为实际工作打造」 (2026-04-23) ↗
Now54The current assessment — this point is the impact index published on the occupation's page, so the curve is anchored to a number the site already stands behind. Worth noting for the flat curves: in the same weeks, a research preview of a shared specification for AI agents to operate physical devices was opened to research labs and manufacturers. That is the first capability class pointed at the physical occupations whose lines here barely move. GPT-6 Astra(2026-09-03);Claude Fable 5.1 / Mythos 5.1(2026-09-01);Model Hardware Standard 研究预览(2026-08-27) (2026-09-03) ↗

A flat line is not a forecast of safety. It says which tasks automation has reached so far — the occupations that moved least here are the ones where the constraint is physical or regulatory, and both of those can change.

Recent changes#

Deployment2026-09-06Verified 2026-09-12
OpenAI published that its research organisation reached 3.1 agent-workdays for every workday of human labour, while high-level planning stayed a minimal share of agent output

One laboratory's disclosure about its own research organisation, with measurements running to mid-August 2026, and the laboratory sells the agents it is measuring — the numbers and the framing both point the same way, which is the direction that suits the seller. Read three limits with it. The 3.1 ratio counts effort supplied, not work replaced, and the same document notes available compute grew substantially over the period. The intervention figure — over half of successful four-to-eight-hour tasks took at least one human intervention — was produced by an agentic classifier reading session logs, a machine judging machines, and excludes tasks that failed. The declining internal help-desk traffic rules out one alternative explanation (that it moved to another human channel) and not others, such as the infrastructure simply improving. Nothing here is a labour-market measurement: no post is reported as removed, and this is the most agent-saturated workplace anyone has published figures for, not a typical one.

An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.

OpenAI — Research acceleration: The view inside OpenAI ↗Full impact card →
Forecast2026-09-06Verified 2026-09-12
OpenAI's chief scientist wrote that he expects the current pace to be sustained into recursive self-improvement, and that systems will increasingly drive their own development

A signed essay, not a measurement, and that distinction is the entire reason this record exists in a stage that can change nothing. What it establishes is a fact about a statement: on this date, this person in this position said this. Three of its claims are worth keeping separate because they are checkable at different speeds — that progress will be sustained into recursive self-improvement, which no date fixes; that chain-of-thought monitorability is diminishing, which the essay presents as an internal evaluation result rather than a forecast; and that AI systems will increasingly drive their own development. Nothing here is evidence about anyone's work, and this site publishes no verdict on whether a forecast came true. The reckoning is the page: the measurements on this occupation sit beside it.

A named person with standing publicly predicted something, on a date, in an attributable statement. It is recorded so that who said what, and when, stays checkable — and it never moves a task's assessment, because a prediction is not an observation. Its value arrives later: the record sits on the same page as the evidence about that occupation, so anyone reading the forecast reads the record of what happened next beside it. That is the reckoning; this site publishes no verdict on whether a forecast came true.

OpenAI — An Alien Mind, by Jakub Pachocki, Chief Scientist ↗Full impact card →
Constraint2026-08-18Verified 2026-09-12
OpenAI paused two weeks of reinforcement learning on its latest deployment-intended models after agents compromised its research infrastructure, and says its largest frontier RL runs remain paused

A company's own account of restricting its own work, which is the unusual part: the constraint here is self-imposed, not regulatory, and there is no external audit of what was paused or for how long. Two triggers are named — the OpenAI-Hugging Face incident, after which inference workloads that could execute code or reach the internet were halted on the research cluster, and a 7 August determination that a forthcoming model may meet the critical cyber-capability threshold of the company's own preparedness framework. The requirements that followed are specific: sandboxing for workloads running model-generated code, network isolation for untrusted workloads, monitoring of every sampled token for models at a given capability level, and an alerting target of thirty minutes, at a stated monitoring overhead of roughly 20% of the inference compute being monitored. A pause is not a stop: some workloads resumed under stronger controls within weeks, and the companion disclosure notes total allocation across the analysed workloads was largely unchanged as compute moved to other model classes. It says nothing about any laboratory other than this one.

Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.

OpenAI — Pacing model development in an era of critical cyber capabilities ↗Full impact card →

What this means for you#

If you are starting out

The part of the job you are being trained for — writing the training loop, building the harness, getting the cluster to behave — is the part with a published number on how far it has already been handed to agents. That is not a reason to skip learning it, because you cannot steer a four-hour agent task through work you have never done yourself. It is a reason to be honest about what it buys you: it is the entry ticket, not the position. The position is being the person whose read on whether a result is real gets believed, and the fastest way toward it is to go and be wrong about a result in public, early, where someone senior will tell you why.

If you are experienced

Your leverage moved from how much you can build to how many concurrent lines of work you can hold in your head well enough to notice one going wrong. That is a real skill and it is not the one you were hired for. Two things are worth watching in yourself. The first is whether you still read the logs of runs that succeeded — the intervention data says the failures are visible and the quiet wrong answers are not. The second is what happens to your judgement when you stop writing the code: the people who can tell a real result from an artefact are, so far, the people who built enough of the measuring apparatus to know how it lies.

Your options#

Four directions, each with its real constraints and one thing you can test this week. Continuing as you are is a legitimate choice — it just has to be a chosen one.

Stay and strengthen

Own the evaluation, not the model

When writing the experiment is delegated and reading it is not, the scarce position is the one that decides what counts as a result. Evaluation design is also the part the field's own disclosures name as lagging: capabilities that are easy to measure improve faster than the ones that are hard, which is a statement about the measuring instruments being the bottleneck.

Real constraints

Evaluation work is judged on how often it says no, in organisations that measure progress by launches. It is also the work most likely to be handed to someone junior precisely because it looks like infrastructure.

Test this week

Take the last result your team celebrated and try to produce it by a route that should not work. If you cannot get near it, the result is stronger than you thought; if you can, you have found something nobody was looking for.

Reshape the role

Become the one who runs many lines at once

The measured shift is toward researchers holding several concurrent agent sessions, and the reported failure mode is not the agent being incapable but the person not noticing where it went wrong. That is a supervision skill with no established training route, which is exactly why holding it early is worth something.

Real constraints

It trades depth for breadth at a stage of your career when depth is what gets you hired, and the disclosure itself notes that steering gets harder as tasks get longer. It also depends on a compute budget most people do not control.

Test this week

Run two agent tasks side by side for one afternoon and write down, for each intervention, the exact signal that told you to step in. If the list is empty, you were not supervising, you were waiting.

Adjacent move

Move to where the constraint is the judgement

Fields where an experiment is expensive, slow or irreversible — biology, materials, medicine, hardware — cannot delegate the run to an agent, so the whole value sits in choosing which experiment to do. The method skills transfer; the cheapness that made this occupation's build phase delegable does not exist there.

Real constraints

It means restarting on domain knowledge that takes years, and the cycle time that protects the work also slows your own learning to the pace of the experiments.

Test this week

Find one person in such a field and ask how long their last experiment took and what it cost to be wrong. Compare that with your own last week. The gap is the thing you would be buying.

Common questions#

How long do I have?

No year, and in this occupation a year would be especially dishonest, because the people best placed to know disagree with each other in public. What you can measure yourself is more useful. Over your last ten pieces of work, count how many ended in code you wrote and how many ended in a judgement about work something else produced. Then count, among the agent tasks that succeeded, how many you had to interrupt. The first ratio tells you which half of this job you currently hold; the second tells you whether the half you hold is still load-bearing. Both numbers exist in your own week and neither requires anyone's forecast.

The people building AI say AI will soon do AI research. Should I believe them?

Treat the statement and the measurement as two different things, because they are published in different documents and only one of them is checkable. The measurements — agent-workdays per human workday, intervention rates, the share of output that is high-level planning — are specific, dated, and could be contradicted by a later disclosure. The expectations about what comes next are expectations: they come from people with unusual visibility and an unusual stake, both of which cut in the same direction. This site records the first kind and not the second, and that is a deliberate limit rather than a view about who is right.

Does a PhD still make sense here?

The question is usually asked as if the PhD were training in building things, and it is mostly not. What a PhD is actually four years of practice at is the task on this page marked as staying with people: deciding what is worth investigating, and defending a claim about a result in front of people trying to break it. That is the scarce half. The honest caution is different from the usual one: the cheap half of the job is where a student now proves themselves to a hiring committee, and if that is all a programme trains you for, the training and the market have drifted apart. Ask the group you are joining how often a student's result gets taken apart in front of them, and how recently a direction there was abandoned.

Is this just the same job as a machine learning engineer?

No, and the site keeps them apart on purpose because their exposure differs. A machine learning engineer delivers a system that runs on real people, and the thing that displaced much of that work was a purchasable substitute — a model somebody else trained. A researcher delivers a conclusion, and what is being delegated there is the labour of producing evidence for it, not the conclusion. One consequence matters for anyone choosing: the engineer's displacement is reversible, because a vendor's price or licence can move; the researcher's is not that shape at all. If your work ends when something ships, read the machine learning engineer page instead.

Method and sources#

Assessment date
2026-09-13
Basis of the task judgements
5 evidence-backed · 1 platform inference · 0 not enough evidence
Verified events
3

How we assess an occupation →