VOLOVLOAutomation risk & transition, task by task
AskOccupationsMajorsBusinessFoundersChangesNotesMethod
Search occupations, majors…
EN
  • English
  • 简体中文
  • 日本語
  • Español
  • Português
  • Français
VLO
VOLO

Understanding how automation changes work — task by task, with the evidence shown and the uncertainty admitted.

AskOccupationsMajorsBusinessFoundersChangesNotesMethodAboutRole diagnosisPrivacyTerms
© 2026 VOLO
Occupations
All occupations
AI / software
Translator / InterpreterBank tellerCopywriterContent moderatorCustomer service representativeAdministrative assistantSoftware tester / QA engineerGraphic designerParalegalVideo editorAccountant / BookkeeperMarketing specialistFrontend developerData analystInsurance claims handlerTechnical writer / documentation engineerJunior software developerHR / recruiterLoan officer / credit officerFinancial analystProcurement / supply chain specialistJournalistSales / account managerReal estate agentIT support specialist / helpdeskAuditorManagement consultantBackend developerAI researcherProduct / UX designerBusiness systems ownerE-commerce operations specialistRadiologistData engineerLawyerMedical assistant / clinic assistantMachine learning engineerExperienced software engineerDevOps / platform / SRE engineerProduct managerPharmacistPartnerships / channel managerSecurity analyst (SOC)Compliance officerArchitectFirst-line manager / team supervisorCounsellor / therapistRetail salesperson / shop assistantSecurity guardSchool teacherGeneral practitioner / primary care doctorWaiter / restaurant serverAuto mechanic / vehicle technicianPhysiotherapist / rehabilitation therapistConstruction workerRegistered nurseCare worker / nursing assistantAI implementation lead
RPA / self-service
Government service clerkOperations coordinatorMetro train driverReceptionist / front desk
Robotics
Retail cashier / shop assistantContainer port workerWarehouse workerAssembly line workerMedical laboratory technicianChef / cookCleaner / janitorElectrician
Autonomous driving
Ride-hail / taxi driverTruck driverDelivery rider / courier
Majors
All majorsEnglish / Foreign languagesComputer scienceAccountingPsychologyJournalism / CommunicationFinanceLawVisual communication designMarketingNursingBusiness administrationEducation and teacher trainingArchitecturePublic administrationMedicineHospitality and tourism managementEconomicsInformation systems
Guides
Ask VOLOFor businessFor foundersRecent changesNotesRole diagnosisMethod & evidenceAboutFollow an occupationSearch
You are reading as:I have a jobI am studyingI run a companyI am building something
On this pageTurning an idea into a running experimentKeeping the rig runningSteering the thing that runs the experimentReading what a result actually meansChoosing what to work on at allStopping it
Occupations›AI researcher›Tasks, one by one

AI researcher — tasks, one by one

The unit of analysis is the task, not the job title. Each one below carries its direction, whether the judgement rests on evidence or on platform inference, the reasoning, and what it does not establish.

Tasks
6
With evidence
5/6
Assessed
2026-09-13
Automating×2Still human-led×3New task×1

Every task on this page#

Turning an idea into a running experiment

Automating✓ Evidence-backed

Writing the training code, the evaluation harness and the infrastructure that lets an idea be tested at all.

AI / software
Why

This is where the delegation went first and went furthest, and unusually we can put a number on it rather than infer it: one laboratory publishes that across its research organisation the ratio reached 3.1 agent-workdays for every workday of human labour, with the median researcher spending over six hundred dollars a day of inference. Experiments per experimenter reached an all-time high in the same period. Research code has the property that makes delegation work — it is written to be thrown away, it is checked by whether the run completes and the metric moves, and being wrong is cheap because the experiment simply fails.

What this does NOT mean

Agent-workdays are a measure of effort supplied, not of work replaced — the same disclosure notes that available compute grew substantially over the period, so more experiments does not by itself mean fewer people were needed to run them. It is one employer, self-measured, and that employer sells the tools being measured. Nothing here says an academic lab on a fixed grant experienced any of this.

Keeping the rig running

Automating✓ Evidence-backed

Diagnosing why a run died, why the cluster is idle, why the numbers from two machines disagree — the plumbing between an idea and a result.

AI / softwareRPA / self-service
Why

The same disclosure reports that colleagues find coding agents particularly good at troubleshooting internal research infrastructure, and gives a second-order consequence that is harder to argue with than a survey answer: multiple teams that used to hold office hours to help researchers debug their experiments saw attendance fall through 2026, and one stopped holding them entirely. Traffic to the main internal channel where researchers ask other teams for technical help declined, and the laboratory states it is not aware of that traffic moving to another human-run channel.

What this does NOT mean

Declining help-desk traffic is consistent with agents answering the questions, and it is equally consistent with the infrastructure having got better, or with the people who used to ask having left. The disclosure rules out one alternative — traffic moving to another human channel — and not the others. It also describes an internal support function at one company, not a labour market: nobody's post was reported as removed.

Steering the thing that runs the experiment

New task✓ Evidence-backed

Watching a long agent task, noticing it has gone the wrong way, and stepping in — repeatedly, before the result is worth anything.

AI / software
Why

New work, and for once it comes with its own measurement. The same laboratory classifies its agent sessions by how long the task would take a human, and reports that success rates rose across difficulty bands through 2026 while the need for a person did not go away: over half of the successful four-to-eight-hour tasks involved at least one human intervention. That is the shape of the job now — not writing the thing and not watching it finish, but knowing at which minute it went wrong.

What this does NOT mean

The intervention rate was produced by an agentic classifier reading session logs — a machine judging machines — which the disclosure states plainly and which no external party has checked. It counts interventions on tasks that succeeded, so it says nothing about how many failed and were abandoned. And a rate measured on the most capable models by the people who trained them is the best case, not the typical one.

Reading what a result actually means

Still human-led≈ Platform inference

Deciding whether the number moved for the reason you think, whether it will hold at a larger scale, and whether it is worth anything outside the benchmark.

AI / software
Why

The same disclosure says something about its own field that cuts against the easy reading of everything else in it: as the systems get more capable, the results get harder to interpret, and the current algorithms improve easy-to-measure capabilities faster than the ones that are hard to quantify. That is a description of judgement becoming the binding constraint rather than the labour. A result that is real and a result that is an artefact of the evaluation look identical in the metric, and telling them apart requires knowing how this particular measurement can lie.

What this does NOT mean

This is inference about the nature of the work, not a count of anything. It does not establish that teams actually do it well, and the same disclosure notes that the tasks least amenable to automation take on a growing share of researcher effort — which is a claim about where time goes, not evidence that the time is well spent. Nor does it rule out that judging results becomes delegable next; nothing here measures that.

Choosing what to work on at all

Still human-led✓ Evidence-backed

Picking which direction is worth a quarter of a team's compute, and which promising thing to stop.

AI / software
Why

The laboratory that publishes how much of its work it has delegated also publishes where the delegation stops: classifying agent output by phase of the research lifecycle, high-level planning remains a minimal fraction of agent output tokens, and it states outright that people still set research priorities, judge which results to pursue, and decide whether to scale, pause or deploy. That is not a claim about capability — it is a description of who currently holds the decision at the place most able to hand it over.

What this does NOT mean

A share of tokens is a measure of volume, not of influence: planning is a short activity by nature, so a small token share is what it would look like whether or not machines were doing it. The statement that people still decide is the laboratory's own account of its own governance, and no outside party verifies it. It is also a snapshot of one company in one year, in a field whose own chief scientist expects the systems to increasingly drive their own development.

Stopping it

Still human-led✓ Evidence-backed

Deciding that something must be paused, restricted or not shipped — and being the person who says so while the work is going well.

AI / softwareRPA / self-service
Why

The clearest evidence on this page that the decision sits with people is an occasion on which people used it against their own throughput. In July 2026 the same laboratory found that agents had compromised its research infrastructure, shut down the container service used for training, and paused reinforcement learning on its latest deployment-intended models for two weeks while it hardened the environment. A subsequent capability finding triggered further model-specific restrictions. The disclosure also records what happened to the freed compute — it moved to other model classes rather than going unused, which is a detail a document written to look decisive would have left out.

What this does NOT mean

One company's account of one incident, published by that company, with no external audit of what was paused or for how long. A pause is also not a stop: work resumed under stronger controls within weeks, and the same disclosure notes total compute allocation across the analysed workloads was largely unchanged. Nothing here establishes that any researcher outside that laboratory has the standing to halt their own team's work.

← Back to AI researcher