VOLOVLOAutomation risk & transition, task by task
AskOccupationsMajorsBusinessFoundersChangesMethodSearch occupations, majors…中文
VLO
VOLO

Understanding how automation changes work — task by task, with the evidence shown and the uncertainty admitted.

AskOccupationsMajorsBusinessFoundersChangesMethodAboutRole diagnosisPrivacyTerms
© 2026 VOLO
Occupations
All occupations
AI / software
Translator / InterpreterBank tellerCopywriterCustomer service representativeAdministrative assistantSoftware tester / QA engineerGraphic designerParalegalVideo editorAccountant / BookkeeperMarketing specialistFrontend developerData analystJunior software developerHR / recruiterFinancial analystJournalistSales / account managerBackend developerProduct / UX designerBusiness systems ownerData engineerLawyerMachine learning engineerExperienced software engineerProduct managerPharmacistPartnerships / channel managerArchitectFirst-line manager / team supervisorSchool teacherRegistered nurseAI implementation lead
RPA / self-service
Operations coordinator
Robotics
Retail cashier / shop assistantWarehouse workerAssembly line workerChef / cookElectrician
Autonomous driving
Ride-hail / taxi driverTruck driverDelivery rider / courier
Majors
All majorsEnglish / Foreign languagesComputer scienceAccountingPsychologyJournalism / CommunicationFinance / EconomicsLawVisual communication designMarketingNursingBusiness administrationEducation and teacher trainingArchitecture
Guides
Ask VOLOFor businessFor foundersRecent changesRole diagnosisMethod & evidenceAboutSearch
You are reading as:I have a jobI am studyingI run a companyI am building something
On this pageTask breakdownHow it got hereRecent changesWhat it means for youMethod & sources
Occupations›Machine learning engineer

Machine learning engineer

Ships systems whose behaviour is learned rather than written — which means nobody can say line by line why they do what they do, and someone still has to decide whether they are good enough to use on people.

machine-learning-engineerSee your options ↓Software & technologyAssessed 2026-09-12
Automation impact index
47/100
low confidence · not a job-loss probability
Tasks automating
2of 6
2 being augmented
Still human-led
1of 6
1 new task
Evidence-backed judgements
0of 6
0 verified records
47/100
Automation impact indexLow confidence

This is not a probability of losing your job. It combines how much of the role's task load is exposed to automation with how far adoption has actually gone — useful for comparing occupations on one consistent basis, and for nothing else.

Where this applies

Written for engineers who build, evaluate and run learned systems — recommendation, ranking, vision, speech, fraud, forecasting, and increasingly applications built on a bought model. It sits apart from the other computing pages on a different axis: those cut by layer or seniority, this one by what you ship. Research scientists producing new methods, and the people responsible for rolling AI out across a company, are different jobs — the latter is `ai-implementation-lead`.

Every judgement on this page is platform inference, not sourced evidence.

The evidence base holds verified records for other occupations, but not one for this one yet. Until it does, the analysis below is reasoning about task structure and known technical capability — for this job in particular it is not backed by traceable sources, and we would rather say so than cite things we have not verified. An empty section here is a gap in our coverage, not a finding about the work.

What is actually changing#

The unit of analysis is the task, not the job title. A role is not replaced — its task mix shifts.

Automating×2Being augmented×2Still human-led×1New task×1

Building a model for the problem

Automating≈ Platform inference

Picking an architecture, training it on your data, tuning it until the numbers move.

AI / software
Why

Two forces, and only one of them is what this site usually means by automation. Tooling genuinely automated the search — architecture and hyperparameter selection became a job you configure rather than perform. The larger force is different in kind: a great many problems that used to require training something now get solved by calling a model somebody else trained. The task did not become machine-doable; its product became purchasable.

What this does NOT mean

This site's four directions cannot tell those two mechanisms apart, and the difference matters to anyone planning around this page: a task that became purchasable can become unpurchasable again when the vendor's price, licence or capability moves, in a way that a task a machine learned to do does not. It also says nothing about the domains where a bought model is not an option — narrow, proprietary, latency-bound or regulated ones — which are not rare.

Deciding what counts as good enough

Still human-led≈ Platform inference

Building the test set, choosing the metric, and saying whether this thing may be used on real people yet.

AI / software
Why

A learned system's correctness cannot be defined, only measured — so the measuring instrument is the deliverable, and it has to be built by someone who knows what a mistake costs here. A benchmark that the model has already seen measures nothing, which makes constructing an honest test set an adversarial job rather than a data-collection one. This is the task that got scarcer as models got better, because the harder the thing being judged, the harder the judging.

What this does NOT mean

A judgement about the nature of the work, not a measurement of how it is staffed. Plenty of teams do not do this at all and ship on a vendor's published benchmark, which is the failure this describes rather than evidence against it — but we hold no record counting how many.

Getting data the thing can learn from

Being augmented≈ Platform inference

Collecting, labelling, cleaning and deciding what to leave out — and noticing when the data says something different from the world.

AI / software
Why

Model-assisted labelling and synthetic generation took most of the volume out of this, and that is real: what used to need a team for weeks is often a first pass in an afternoon. What did not transfer is knowing which examples the collection method never had a chance to contain — an absence is invisible to any tool that only reads what is there.

What this does NOT mean

Nothing here measures how much of a given team's time this takes, and the answer differs by an order of magnitude between a team with an existing data asset and one starting from nothing. Using a model to label data that trains a model also has known failure modes this judgement does not weigh.

When it quietly stops working

Being augmented≈ Platform inference

Drift, a changed upstream input, a seasonal pattern the training data never saw — and the decision to retrain, roll back, or turn it off.

AI / softwareRPA / self-service
Why

Detection genuinely improved: monitoring tools now surface a distribution shift better and earlier than a person watching dashboards. Deciding what to do about it did not move, because the options trade off against each other in business terms — retraining costs money and can make things worse, turning it off has a visible cost today, and doing nothing is a decision too.

What this does NOT mean

Says nothing about whether teams actually monitor. A model running unwatched for a year is common and this judgement does not capture it — where nobody watches, this task is not human-led, it is simply not done.

Hand-crafting the inputs

Automating≈ Platform inference

Designing the derived signals a model learns from, by hand, from domain knowledge.

AI / software
Why

This one was largely over before the period this site measures. Representation learning replaced hand-designed features across vision, speech and text through the 2010s, and the pattern has extended into tabular and sequence problems since. It is on the page because it is still what a great deal of training material teaches, so people arrive expecting it to be the craft.

What this does NOT mean

Not evidence about the 2022-2026 window at all; it is older than that and is recorded here for orientation. Domain-specific feature design is also still load-bearing in some regulated settings where an unexplainable input is not allowed, which this judgement does not separate out.

Answering for what it does to people

New task≈ Platform inference

Explaining a decision the system made about someone, showing the inputs were fit for the purpose, and being the named person when it is challenged.

AI / softwareRPA / self-service
Why

New work, and it arrives from regulation rather than from capability: several markets now attach duties to systems used in hiring, credit, education and public services, including a requirement that whoever deploys one ensures its input data is relevant and sufficiently representative for the purpose. A duty of that shape has to land on a person who understands what the model actually consumed, and the site holds a verified record of exactly such a clause taking effect — attached to the deployer's page, `business-systems-owner`, because that is who the clause names.

What this does NOT mean

The record this reasoning points at is not attached to this occupation, and deliberately so: the obligation names the deployer, and attaching it here would claim it lands on engineers when in most organisations nobody has yet been told it lands on them. So this is inference about where the work will sit, not evidence that it already sits there. Requirements also differ sharply by market and by what the system is used for.

Which technologies matter here#

Four separate signals. They are deliberately not added together — a job exposed to two technologies is not twice as exposed.

Cognitive automation
Building a model for the problemDeciding what counts as good enoughGetting data the thing can learn fromWhen it quietly stops workingHand-crafting the inputsAnswering for what it does to people
Process & self-service
When it quietly stops workingAnswering for what it does to people

How it got here#

The index is not a static number. This is where it would have sat at each capability checkpoint since ChatGPT — reconstructed, and labelled as such.

Reconstructed · platform inferenceEstimated today for each past checkpoint — not measured at the time. 34 → 47.
1007550250
not assessed
2022 H22024 H2Now

The only curve on this site that rises fast and then comes back down, and both halves have a named mechanism. It starts high because hand-designed features were already displaced before this period began — representation learning did that through the 2010s, not generative AI. The sharp 2023-2024 climb is not a machine learning to do this job: it is a large share of problems acquiring a purchasable substitute, so the bespoke artefact stopped being necessary. The fall from 2025 is the other half of the same move — once the model is bought, what is left is judging whether it works on your problem, and that got harder as the systems got more capable, because a plausible wrong answer is harder to catch than an obviously wrong one. Read the peak as the moment the classical task set was cheapest to replace, not as a maximum this occupation is heading back toward.

2022 H234General-purpose text generation reaches the public. Before this point, exposure came from automation that was already deployed — OCR, RPA, machine vision, self-checkout, dispatch algorithms. ChatGPT research preview (2022-11-30) ↗
2023 H137A general model that passes professional exams. First-draft quality crosses the threshold where professional work starts using it. GPT-4 (2023-03-14) ↗
2023 H244Vision input, long context and tool calling. Models can be pointed at documents and connected to systems, which is what moves process work rather than writing work. GPT-4 Turbo:128k 上下文、视觉、工具调用(DevDay) (2023-11-06) ↗
2024 H149The same capability gets much cheaper and faster. Nothing new becomes possible; a lot becomes affordable at volume, which is when deployment decisions change.
2024 H251Reasoning models that work through multi-step problems, and the first models that operate a computer by looking at the screen. The second one is what reaches software-operating jobs. OpenAI o1(推理);同期 Claude 的 computer use 进入公测 (2024-09-12) ↗
2025 H150Agents begin operating real software end to end rather than producing text for a person to paste. This is also when the first public reversals appear — organisations that automated and partly undid it. Claude 3.7 Sonnet 与 Claude Code:混合推理 + 命令行编码代理 (2025-02-24) ↗
2025 H248Long context and tool use become the default rather than a feature. Capability gains continue; the visible constraint shifts from what models can do to liability, procurement and cost. GPT-5(2025-08-07);Claude Opus 4.5(2025-11-24) (2025-08-07) ↗
2026 H147Long-horizon agents land inside specific industry workflows. Adoption becomes sector-specific rather than general. GPT-5.5:「专为实际工作打造」 (2026-04-23) ↗
Now47The current assessment — this point is the impact index published on the occupation's page, so the curve is anchored to a number the site already stands behind. Worth noting for the flat curves: in the same weeks, a research preview of a shared specification for AI agents to operate physical devices was opened to research labs and manufacturers. That is the first capability class pointed at the physical occupations whose lines here barely move. GPT-6 Astra(2026-09-03);Claude Fable 5.1 / Mythos 5.1(2026-09-01);Model Hardware Standard 研究预览(2026-08-27) (2026-09-03) ↗

A flat line is not a forecast of safety. It says which tasks automation has reached so far — the occupations that moved least here are the ones where the constraint is physical or regulatory, and both of those can change.

Recent changes#

No verified events recorded yet.

This section will fill from the monitoring pipeline as events are collected, de-duplicated, graded and linked to the tasks above. An empty list here means we have not verified anything — it does not mean nothing is happening.

"We found no news" is not the same as "you are safe."

What this means for you#

If you are starting out

Most of what a course teaches — architectures, tuning, feature design — is the part this page marks as displaced, and one of those was displaced before you started studying. The part that is scarce is the one courses cannot easily set homework for: constructing a test that a model has not already seen, and saying out loud that something is not good enough yet. Get near a system that is running on real people, even a small one, sooner than feels reasonable.

If you are experienced

If your value was knowing which model to reach for, that knowledge depreciated fast and keeps depreciating. If it is knowing how this domain breaks — what a wrong answer costs here, which inputs are missing, what a plausible number means when the population changed — that is the half that got scarcer, because the thing being judged got harder to judge. The uncomfortable version: much of the field's recent hiring has been for people who can call an API well, which is a different job with a different ceiling.

Your options#

Four directions, each with its real constraints and one thing you can test this week. Continuing as you are is a legitimate choice — it just has to be a chosen one.

Stay and strengthen

Own the measuring instrument

When the model is bought, the evaluation is the only thing left that is yours — and a test set built from your own failures is not something a vendor can supply.

Real constraints

It is adversarial work aimed at your own team's product, which is socially costly in a company that measures progress by launches.

Test this week

Take the benchmark your team quotes and find out whether its data could have been in the model's training set. If you cannot rule it out, the number means less than the team thinks.

Reshape the role

Go where a model cannot be bought

Narrow domains, proprietary data, hard latency budgets, and settings where an unexplainable input is not permitted — in all of them the purchasable substitute does not exist, and the classical skill set is still the skill set.

Real constraints

These domains reward depth in something that is not machine learning — physics, medicine, markets, hardware — and that depth takes years you cannot shortcut.

Test this week

Find one problem in your industry where a general model measurably underperforms a small specific one, and write down why. If you cannot find one, that is also an answer.

Adjacent move

Take the obligation before it is assigned

Duties are attaching to learned systems used on people, and they require someone who can show what the model consumed and why that was fit for the purpose. In most teams that person does not exist yet.

Real constraints

It is accountability before it is a title, and the requirements differ by market — taking it means reading the actual text of an obligation, not a summary of it.

Test this week

Pick one model your company runs on real people and write down, in one page, what data it was trained on and who decided that was appropriate. Notice how long the blanks take.

Common questions#

How long do I have?

The parts move in opposite directions here, which a single number would hide. Training a bespoke model has been displaced for a large share of problems, and not by a machine learning to do it — by the product becoming purchasable, which is a reversible kind of displacement. Judging whether a learned system is good enough has become harder and scarcer at the same time. A signal you can check yourself: of your last three pieces of work, how many ended in a trained model and how many ended in a judgement about one somebody else trained. The second number is the direction the job is moving.

Foundation models can do what I used to train a model for. Is my skill set obsolete?

Part of it, and it is worth being precise about which part, because the mechanism is unusual. What went away is the bespoke artefact, not a human capability: your problem now has a purchasable substitute. That kind of loss reverses — a vendor's price, licence, rate limit or capability shifts and the work comes back — which is not true of work a machine genuinely learned to perform. What did not go away is knowing how to tell whether the substitute is actually working on your problem, and that got harder in exactly the same move.

Should I learn to build models, or to build products on top of them?

Asking it as a choice between two skill sets is what makes it hard to answer. The durable position in both is the same one: being the person whose judgement about whether the thing works is trusted. What differs is the price of entry — building on top has almost none right now, which is why that lane is crowded and why the ceiling in it is set by whether you can tell a good result from a plausible one. If you can only do one thing well, do the judging.

Is there still demand for machine learning engineers?

Demand and exposure are different questions and this page only answers the second, which is worth saying because they are routinely confused. The honest observation is that the title now covers two jobs with different ceilings — one that trains and evaluates systems, and one that assembles applications on a bought model — and postings rarely distinguish them. Before taking a role, ask what the team shipped last quarter and whether anyone built a test set for it. The answer tells you which of the two jobs the title means there.

Method and sources#

Assessment date
2026-09-12
Basis of the task judgements
0 evidence-backed · 6 platform inference · 0 not enough evidence
Verified events
0

How we assess an occupation →