Machine learning engineer
Ships systems whose behaviour is learned rather than written — which means nobody can say line by line why they do what they do, and someone still has to decide whether they are good enough to use on people.
This is not a probability of losing your job. It combines how much of the role's task load is exposed to automation with how far adoption has actually gone — useful for comparing occupations on one consistent basis, and for nothing else.
Written for engineers who build, evaluate and run learned systems — recommendation, ranking, vision, speech, fraud, forecasting, and increasingly applications built on a bought model. It sits apart from the other computing pages on a different axis: those cut by layer or seniority, this one by what you ship. Research scientists producing new methods, and the people responsible for rolling AI out across a company, are different jobs — the latter is `ai-implementation-lead`.
The evidence base holds verified records for other occupations, but not one for this one yet. Until it does, the analysis below is reasoning about task structure and known technical capability — for this job in particular it is not backed by traceable sources, and we would rather say so than cite things we have not verified. An empty section here is a gap in our coverage, not a finding about the work.
What is actually changing#
The unit of analysis is the task, not the job title. A role is not replaced — its task mix shifts.
Is this your job? Say so and this page narrows to your share of it.
A job title is a bundle of tasks bought together, and no two people hold the same bundle. Nothing is sent anywhere — it stays in this browser.
Building a model for the problem
Automating≈ Platform inferencePicking an architecture, training it on your data, tuning it until the numbers move.
Two forces, and only one of them is what this site usually means by automation. Tooling genuinely automated the search — architecture and hyperparameter selection became a job you configure rather than perform. The larger force is different in kind: a great many problems that used to require training something now get solved by calling a model somebody else trained. The task did not become machine-doable; its product became purchasable.
This site's four directions cannot tell those two mechanisms apart, and the difference matters to anyone planning around this page: a task that became purchasable can become unpurchasable again when the vendor's price, licence or capability moves, in a way that a task a machine learned to do does not. It also says nothing about the domains where a bought model is not an option — narrow, proprietary, latency-bound or regulated ones — which are not rare.
Deciding what counts as good enough
Still human-led≈ Platform inferenceBuilding the test set, choosing the metric, and saying whether this thing may be used on real people yet.
A learned system's correctness cannot be defined, only measured — so the measuring instrument is the deliverable, and it has to be built by someone who knows what a mistake costs here. A benchmark that the model has already seen measures nothing, which makes constructing an honest test set an adversarial job rather than a data-collection one. This is the task that got scarcer as models got better, because the harder the thing being judged, the harder the judging.
A judgement about the nature of the work, not a measurement of how it is staffed. Plenty of teams do not do this at all and ship on a vendor's published benchmark, which is the failure this describes rather than evidence against it — but we hold no record counting how many.
Getting data the thing can learn from
Being augmented≈ Platform inferenceCollecting, labelling, cleaning and deciding what to leave out — and noticing when the data says something different from the world.
Model-assisted labelling and synthetic generation took most of the volume out of this, and that is real: what used to need a team for weeks is often a first pass in an afternoon. What did not transfer is knowing which examples the collection method never had a chance to contain — an absence is invisible to any tool that only reads what is there.
Nothing here measures how much of a given team's time this takes, and the answer differs by an order of magnitude between a team with an existing data asset and one starting from nothing. Using a model to label data that trains a model also has known failure modes this judgement does not weigh.
When it quietly stops working
Being augmented≈ Platform inferenceDrift, a changed upstream input, a seasonal pattern the training data never saw — and the decision to retrain, roll back, or turn it off.
Detection genuinely improved: monitoring tools now surface a distribution shift better and earlier than a person watching dashboards. Deciding what to do about it did not move, because the options trade off against each other in business terms — retraining costs money and can make things worse, turning it off has a visible cost today, and doing nothing is a decision too.
Says nothing about whether teams actually monitor. A model running unwatched for a year is common and this judgement does not capture it — where nobody watches, this task is not human-led, it is simply not done.
Hand-crafting the inputs
Automating≈ Platform inferenceDesigning the derived signals a model learns from, by hand, from domain knowledge.
This one was largely over before the period this site measures. Representation learning replaced hand-designed features across vision, speech and text through the 2010s, and the pattern has extended into tabular and sequence problems since. It is on the page because it is still what a great deal of training material teaches, so people arrive expecting it to be the craft.
Not evidence about the 2022-2026 window at all; it is older than that and is recorded here for orientation. Domain-specific feature design is also still load-bearing in some regulated settings where an unexplainable input is not allowed, which this judgement does not separate out.
Answering for what it does to people
New task≈ Platform inferenceExplaining a decision the system made about someone, showing the inputs were fit for the purpose, and being the named person when it is challenged.
New work, and it arrives from regulation rather than from capability: several markets now attach duties to systems used in hiring, credit, education and public services, including a requirement that whoever deploys one ensures its input data is relevant and sufficiently representative for the purpose. A duty of that shape has to land on a person who understands what the model actually consumed, and the site holds a verified record of exactly such a clause taking effect — attached to the deployer's page, `business-systems-owner`, because that is who the clause names.
The record this reasoning points at is not attached to this occupation, and deliberately so: the obligation names the deployer, and attaching it here would claim it lands on engineers when in most organisations nobody has yet been told it lands on them. So this is inference about where the work will sit, not evidence that it already sits there. Requirements also differ sharply by market and by what the system is used for.
Which technologies matter here#
Four separate signals. They are deliberately not added together — a job exposed to two technologies is not twice as exposed.
How it got here#
The index is not a static number. This is where it would have sat at each capability checkpoint since ChatGPT — reconstructed, and labelled as such.
The only curve on this site that rises fast and then comes back down, and both halves have a named mechanism. It starts high because hand-designed features were already displaced before this period began — representation learning did that through the 2010s, not generative AI. The sharp 2023-2024 climb is not a machine learning to do this job: it is a large share of problems acquiring a purchasable substitute, so the bespoke artefact stopped being necessary. The fall from 2025 is the other half of the same move — once the model is bought, what is left is judging whether it works on your problem, and that got harder as the systems got more capable, because a plausible wrong answer is harder to catch than an obviously wrong one. Read the peak as the moment the classical task set was cheapest to replace, not as a maximum this occupation is heading back toward.
A flat line is not a forecast of safety. It says which tasks automation has reached so far — the occupations that moved least here are the ones where the constraint is physical or regulatory, and both of those can change.
Recent changes#
No verified events recorded yet.
This section will fill from the monitoring pipeline as events are collected, de-duplicated, graded and linked to the tasks above. An empty list here means we have not verified anything — it does not mean nothing is happening.
"We found no news" is not the same as "you are safe."
What this means for you#
Most of what a course teaches — architectures, tuning, feature design — is the part this page marks as displaced, and one of those was displaced before you started studying. The part that is scarce is the one courses cannot easily set homework for: constructing a test that a model has not already seen, and saying out loud that something is not good enough yet. Get near a system that is running on real people, even a small one, sooner than feels reasonable.
If your value was knowing which model to reach for, that knowledge depreciated fast and keeps depreciating. If it is knowing how this domain breaks — what a wrong answer costs here, which inputs are missing, what a plausible number means when the population changed — that is the half that got scarcer, because the thing being judged got harder to judge. The uncomfortable version: much of the field's recent hiring has been for people who can call an API well, which is a different job with a different ceiling.
Your options#
Four directions, each with its real constraints and one thing you can test this week. Continuing as you are is a legitimate choice — it just has to be a chosen one.
Own the measuring instrument
When the model is bought, the evaluation is the only thing left that is yours — and a test set built from your own failures is not something a vendor can supply.
It is adversarial work aimed at your own team's product, which is socially costly in a company that measures progress by launches.
Take the benchmark your team quotes and find out whether its data could have been in the model's training set. If you cannot rule it out, the number means less than the team thinks.
Go where a model cannot be bought
Narrow domains, proprietary data, hard latency budgets, and settings where an unexplainable input is not permitted — in all of them the purchasable substitute does not exist, and the classical skill set is still the skill set.
These domains reward depth in something that is not machine learning — physics, medicine, markets, hardware — and that depth takes years you cannot shortcut.
Find one problem in your industry where a general model measurably underperforms a small specific one, and write down why. If you cannot find one, that is also an answer.
Take the obligation before it is assigned
Duties are attaching to learned systems used on people, and they require someone who can show what the model consumed and why that was fit for the purpose. In most teams that person does not exist yet.
It is accountability before it is a title, and the requirements differ by market — taking it means reading the actual text of an obligation, not a summary of it.
Pick one model your company runs on real people and write down, in one page, what data it was trained on and who decided that was appropriate. Notice how long the blanks take.
Common questions#
The parts move in opposite directions here, which a single number would hide. Training a bespoke model has been displaced for a large share of problems, and not by a machine learning to do it — by the product becoming purchasable, which is a reversible kind of displacement. Judging whether a learned system is good enough has become harder and scarcer at the same time. A signal you can check yourself: of your last three pieces of work, how many ended in a trained model and how many ended in a judgement about one somebody else trained. The second number is the direction the job is moving.
Part of it, and it is worth being precise about which part, because the mechanism is unusual. What went away is the bespoke artefact, not a human capability: your problem now has a purchasable substitute. That kind of loss reverses — a vendor's price, licence, rate limit or capability shifts and the work comes back — which is not true of work a machine genuinely learned to perform. What did not go away is knowing how to tell whether the substitute is actually working on your problem, and that got harder in exactly the same move.
Asking it as a choice between two skill sets is what makes it hard to answer. The durable position in both is the same one: being the person whose judgement about whether the thing works is trusted. What differs is the price of entry — building on top has almost none right now, which is why that lane is crowded and why the ceiling in it is set by whether you can tell a good result from a plausible one. If you can only do one thing well, do the judging.
Demand and exposure are different questions and this page only answers the second, which is worth saying because they are routinely confused. The honest observation is that the title now covers two jobs with different ceilings — one that trains and evaluates systems, and one that assembles applications on a bought model — and postings rarely distinguish them. Before taking a role, ask what the team shipped last quarter and whether anyone built a test set for it. The answer tells you which of the two jobs the title means there.
Method and sources#
- Assessment date
- 2026-09-12
- Basis of the task judgements
- 0 evidence-backed · 6 platform inference · 0 not enough evidence
- Verified events
- 0