VOLOVLOAutomation risk & transition, task by task
AskOccupationsMajorsBusinessFoundersChangesNotesMethod
Search occupations, majors…
EN
  • English
  • 简体中文
  • 日本語
  • Español
  • Português
  • Français
VLO
VOLO

Understanding how automation changes work — task by task, with the evidence shown and the uncertainty admitted.

AskOccupationsMajorsBusinessFoundersChangesNotesMethodAboutRole diagnosisPrivacyTerms
© 2026 VOLO
Occupations
All occupations
AI / software
Translator / InterpreterBank tellerCopywriterCustomer service representativeAdministrative assistantSoftware tester / QA engineerGraphic designerParalegalVideo editorAccountant / BookkeeperMarketing specialistFrontend developerData analystInsurance claims handlerTechnical writer / documentation engineerJunior software developerHR / recruiterLoan officer / credit officerFinancial analystProcurement / supply chain specialistJournalistSales / account managerReal estate agentIT support specialist / helpdeskAuditorManagement consultantBackend developerAI researcherProduct / UX designerBusiness systems ownerE-commerce operations specialistRadiologistData engineerLawyerMedical assistant / clinic assistantMachine learning engineerExperienced software engineerDevOps / platform / SRE engineerProduct managerPharmacistPartnerships / channel managerSecurity analyst (SOC)Compliance officerArchitectFirst-line manager / team supervisorCounsellor / therapistRetail salesperson / shop assistantSecurity guardSchool teacherGeneral practitioner / primary care doctorWaiter / restaurant serverAuto mechanic / vehicle technicianPhysiotherapist / rehabilitation therapistConstruction workerRegistered nurseCare worker / nursing assistantAI implementation lead
RPA / self-service
Government service clerkOperations coordinatorMetro train driverReceptionist / front desk
Robotics
Retail cashier / shop assistantContainer port workerWarehouse workerAssembly line workerMedical laboratory technicianChef / cookCleaner / janitorElectrician
Autonomous driving
Ride-hail / taxi driverTruck driverDelivery rider / courier
Majors
All majorsEnglish / Foreign languagesComputer scienceAccountingPsychologyJournalism / CommunicationFinanceLawVisual communication designMarketingNursingBusiness administrationEducation and teacher trainingArchitecturePublic administrationMedicineHospitality and tourism managementEconomicsInformation systems
Guides
Ask VOLOFor businessFor foundersRecent changesNotesRole diagnosisMethod & evidenceAboutFollow an occupationSearch
You are reading as:I have a jobI am studyingI run a companyI am building something
On this pageBuilding a model for the problemDeciding what counts as good enoughGetting data the thing can learn fromWhen it quietly stops workingHand-crafting the inputsAnswering for what it does to people
Occupations›Machine learning engineer›Tasks, one by one

Machine learning engineer — tasks, one by one

The unit of analysis is the task, not the job title. Each one below carries its direction, whether the judgement rests on evidence or on platform inference, the reasoning, and what it does not establish.

Tasks
6
With evidence
2/6
Assessed
2026-09-12
Automating×2Being augmented×2Still human-led×1New task×1

Every task on this page#

Building a model for the problem

Automating≈ Platform inference

Picking an architecture, training it on your data, tuning it until the numbers move.

AI / software
Why

Two forces, and only one of them is what this site usually means by automation. Tooling genuinely automated the search — architecture and hyperparameter selection became a job you configure rather than perform. The larger force is different in kind: a great many problems that used to require training something now get solved by calling a model somebody else trained. The task did not become machine-doable; its product became purchasable.

What this does NOT mean

This site's four directions cannot tell those two mechanisms apart, and the difference matters to anyone planning around this page: a task that became purchasable can become unpurchasable again when the vendor's price, licence or capability moves, in a way that a task a machine learned to do does not. It also says nothing about the domains where a bought model is not an option — narrow, proprietary, latency-bound or regulated ones — which are not rare.

Deciding what counts as good enough

Still human-led✓ Evidence-backed

Building the test set, choosing the metric, and saying whether this thing may be used on real people yet.

AI / software
Why

A learned system's correctness cannot be defined, only measured — so the measuring instrument is the deliverable, and it has to be built by someone who knows what a mistake costs here. A benchmark that the model has already seen measures nothing, which makes constructing an honest test set an adversarial job rather than a data-collection one. This is the task that got scarcer as models got better, because the harder the thing being judged, the harder the judging.

What this does NOT mean

A judgement about the nature of the work, not a measurement of how it is staffed. Plenty of teams do not do this at all and ship on a vendor's published benchmark, which is the failure this describes rather than evidence against it — but how many is not counted anywhere.

Getting data the thing can learn from

Being augmented≈ Platform inference

Collecting, labelling, cleaning and deciding what to leave out — and noticing when the data says something different from the world.

AI / software
Why

Model-assisted labelling and synthetic generation took most of the volume out of this, and that is real: what used to need a team for weeks is often a first pass in an afternoon. What did not transfer is knowing which examples the collection method never had a chance to contain — an absence is invisible to any tool that only reads what is there.

What this does NOT mean

Nothing here measures how much of a given team's time this takes, and the answer differs by an order of magnitude between a team with an existing data asset and one starting from nothing. Using a model to label data that trains a model also has known failure modes this judgement does not weigh.

When it quietly stops working

Being augmented≈ Platform inference

Drift, a changed upstream input, a seasonal pattern the training data never saw — and the decision to retrain, roll back, or turn it off.

AI / softwareRPA / self-service
Why

Detection genuinely improved: monitoring tools now surface a distribution shift better and earlier than a person watching dashboards. Deciding what to do about it did not move, because the options trade off against each other in business terms — retraining costs money and can make things worse, turning it off has a visible cost today, and doing nothing is a decision too.

What this does NOT mean

Says nothing about whether teams actually monitor. A model running unwatched for a year is common and this judgement does not capture it — where nobody watches, this task is not human-led, it is simply not done.

Hand-crafting the inputs

Automating≈ Platform inference

Designing the derived signals a model learns from, by hand, from domain knowledge.

AI / software
Why

This one was largely over before the period this site measures. Representation learning replaced hand-designed features across vision, speech and text through the 2010s, and the pattern has extended into tabular and sequence problems since. It is on the page because it is still what a great deal of training material teaches, so people arrive expecting it to be the craft.

What this does NOT mean

Not evidence about the 2022-2026 window at all; it is older than that and is recorded here for orientation. Domain-specific feature design is also still load-bearing in some regulated settings where an unexplainable input is not allowed, which this judgement does not separate out.

Answering for what it does to people

New task✓ Evidence-backed

Explaining a decision the system made about someone, showing the inputs were fit for the purpose, and being the named person when it is challenged.

AI / softwareRPA / self-service
Why

New work, and it arrives from regulation rather than from capability: several markets now attach duties to systems used in hiring, credit, education and public services, including a requirement that whoever deploys one ensures its input data is relevant and sufficiently representative for the purpose. A duty of that shape has to land on a person who understands what the model actually consumed, and the site holds a verified record of exactly such a clause taking effect — attached to the deployer's page, `business-systems-owner`, because that is who the clause names.

What this does NOT mean

The record this reasoning points at is not attached to this occupation, and deliberately so: the obligation names the deployer, and attaching it here would claim it lands on engineers when in most organisations nobody has yet been told it lands on them. So this is inference about where the work will sit, not evidence that it already sits there. Requirements also differ sharply by market and by what the system is used for.

← Back to Machine learning engineer