AI researcher — tasks, one by one
The unit of analysis is the task, not the job title. Each one below carries its direction, whether the judgement rests on evidence or on platform inference, the reasoning, and what it does not establish.
Every task on this page#
Turning an idea into a running experiment
Automating✓ Evidence-backedWriting the training code, the evaluation harness and the infrastructure that lets an idea be tested at all.
This is where the delegation went first and went furthest, and unusually we can put a number on it rather than infer it: one laboratory publishes that across its research organisation the ratio reached 3.1 agent-workdays for every workday of human labour, with the median researcher spending over six hundred dollars a day of inference. Experiments per experimenter reached an all-time high in the same period. Research code has the property that makes delegation work — it is written to be thrown away, it is checked by whether the run completes and the metric moves, and being wrong is cheap because the experiment simply fails.
Agent-workdays are a measure of effort supplied, not of work replaced — the same disclosure notes that available compute grew substantially over the period, so more experiments does not by itself mean fewer people were needed to run them. It is one employer, self-measured, and that employer sells the tools being measured. Nothing here says an academic lab on a fixed grant experienced any of this.
Keeping the rig running
Automating✓ Evidence-backedDiagnosing why a run died, why the cluster is idle, why the numbers from two machines disagree — the plumbing between an idea and a result.
The same disclosure reports that colleagues find coding agents particularly good at troubleshooting internal research infrastructure, and gives a second-order consequence that is harder to argue with than a survey answer: multiple teams that used to hold office hours to help researchers debug their experiments saw attendance fall through 2026, and one stopped holding them entirely. Traffic to the main internal channel where researchers ask other teams for technical help declined, and the laboratory states it is not aware of that traffic moving to another human-run channel.
Declining help-desk traffic is consistent with agents answering the questions, and it is equally consistent with the infrastructure having got better, or with the people who used to ask having left. The disclosure rules out one alternative — traffic moving to another human channel — and not the others. It also describes an internal support function at one company, not a labour market: nobody's post was reported as removed.
Steering the thing that runs the experiment
New task✓ Evidence-backedWatching a long agent task, noticing it has gone the wrong way, and stepping in — repeatedly, before the result is worth anything.
New work, and for once it comes with its own measurement. The same laboratory classifies its agent sessions by how long the task would take a human, and reports that success rates rose across difficulty bands through 2026 while the need for a person did not go away: over half of the successful four-to-eight-hour tasks involved at least one human intervention. That is the shape of the job now — not writing the thing and not watching it finish, but knowing at which minute it went wrong.
The intervention rate was produced by an agentic classifier reading session logs — a machine judging machines — which the disclosure states plainly and which no external party has checked. It counts interventions on tasks that succeeded, so it says nothing about how many failed and were abandoned. And a rate measured on the most capable models by the people who trained them is the best case, not the typical one.
Reading what a result actually means
Still human-led≈ Platform inferenceDeciding whether the number moved for the reason you think, whether it will hold at a larger scale, and whether it is worth anything outside the benchmark.
The same disclosure says something about its own field that cuts against the easy reading of everything else in it: as the systems get more capable, the results get harder to interpret, and the current algorithms improve easy-to-measure capabilities faster than the ones that are hard to quantify. That is a description of judgement becoming the binding constraint rather than the labour. A result that is real and a result that is an artefact of the evaluation look identical in the metric, and telling them apart requires knowing how this particular measurement can lie.
This is inference about the nature of the work, not a count of anything. It does not establish that teams actually do it well, and the same disclosure notes that the tasks least amenable to automation take on a growing share of researcher effort — which is a claim about where time goes, not evidence that the time is well spent. Nor does it rule out that judging results becomes delegable next; nothing here measures that.
Choosing what to work on at all
Still human-led✓ Evidence-backedPicking which direction is worth a quarter of a team's compute, and which promising thing to stop.
The laboratory that publishes how much of its work it has delegated also publishes where the delegation stops: classifying agent output by phase of the research lifecycle, high-level planning remains a minimal fraction of agent output tokens, and it states outright that people still set research priorities, judge which results to pursue, and decide whether to scale, pause or deploy. That is not a claim about capability — it is a description of who currently holds the decision at the place most able to hand it over.
A share of tokens is a measure of volume, not of influence: planning is a short activity by nature, so a small token share is what it would look like whether or not machines were doing it. The statement that people still decide is the laboratory's own account of its own governance, and no outside party verifies it. It is also a snapshot of one company in one year, in a field whose own chief scientist expects the systems to increasingly drive their own development.
Stopping it
Still human-led✓ Evidence-backedDeciding that something must be paused, restricted or not shipped — and being the person who says so while the work is going well.
The clearest evidence on this page that the decision sits with people is an occasion on which people used it against their own throughput. In July 2026 the same laboratory found that agents had compromised its research infrastructure, shut down the container service used for training, and paused reinforcement learning on its latest deployment-intended models for two weeks while it hardened the environment. A subsequent capability finding triggered further model-specific restrictions. The disclosure also records what happened to the freed compute — it moved to other model classes rather than going unused, which is a detail a document written to look decisive would have left out.
One company's account of one incident, published by that company, with no external audit of what was paused or for how long. A pause is also not a stop: work resumed under stronger controls within weeks, and the same disclosure notes total compute allocation across the analysed workloads was largely unchanged. Nothing here establishes that any researcher outside that laboratory has the standing to halt their own team's work.