Data engineer — how we know
The page itself gives the judgements. This one gives what they rest on: which technologies bear on the work, how the estimate moved since language models reached the public, and the method behind both.
Which technologies matter here#
Four separate signals. They are deliberately not added together — a job exposed to two technologies is not twice as exposed.
How it got here#
The index is not a static number. This is where it would have sat at each capability checkpoint since ChatGPT — reconstructed, and labelled as such.
—— this stretch contains a verified event- - - no event in this stretch — reconstruction only0 = no task exposed, 100 = every task exposed
● 3 verified events for this occupation, plotted at the date it happened — the parts of the line near a marker are anchored to something checkable.
The highest starting point of the three engineering layers, because this work was already being commoditised before generative tools existed — bought connectors and managed warehouses had been eating the building half for years. The middle rise is transformation code becoming cheap to draft and natural-language querying arriving. It flattens earliest and stops rising at all after 2025, and the reason is on the page: cheap querying makes definitions more load-bearing, not less. The one benchmark we hold that was built from real corporate warehouses scores a plain model at zero where public benchmarks read 80 to 90 — the gap is meaning, and meaning lives in people.
A flat line is not a forecast of safety. It says which tasks automation has reached so far — the occupations that moved least here are the ones where the constraint is physical or regulatory, and both of those can change.
Written about this#
These pieces argue from the same records this page holds, and each of their sections names what it rests on.
Method and sources#
- Assessment date
- 2026-09-12
- Basis of the task judgements
- 4 evidence-backed · 2 platform inference · 0 not enough evidence
- Verified events
- 3