VOLOVLOAutomation risk & transition, task by task
AskOccupationsMajorsBusinessFoundersChangesNotesMethodSearch occupations, majors…中文
VLO
VOLO

Understanding how automation changes work — task by task, with the evidence shown and the uncertainty admitted.

AskOccupationsMajorsBusinessFoundersChangesNotesMethodAboutRole diagnosisPrivacyTerms
© 2026 VOLO
Occupations
All occupations
AI / software
Translator / InterpreterBank tellerCopywriterCustomer service representativeAdministrative assistantSoftware tester / QA engineerGraphic designerParalegalVideo editorAccountant / BookkeeperMarketing specialistFrontend developerData analystInsurance claims handlerJunior software developerHR / recruiterFinancial analystProcurement / supply chain specialistJournalistSales / account managerReal estate agentAuditorBackend developerAI researcherProduct / UX designerBusiness systems ownerRadiologistData engineerLawyerMachine learning engineerExperienced software engineerProduct managerPharmacistPartnerships / channel managerArchitectFirst-line manager / team supervisorCounsellor / therapistSchool teacherRegistered nurseAI implementation lead
RPA / self-service
Government service clerkOperations coordinator
Robotics
Retail cashier / shop assistantWarehouse workerAssembly line workerChef / cookElectrician
Autonomous driving
Ride-hail / taxi driverTruck driverDelivery rider / courier
Majors
All majorsEnglish / Foreign languagesComputer scienceAccountingPsychologyJournalism / CommunicationFinance / EconomicsLawVisual communication designMarketingNursingBusiness administrationEducation and teacher trainingArchitecturePublic administration
Guides
Ask VOLOFor businessFor foundersRecent changesNotesRole diagnosisMethod & evidenceAboutFollow an occupationSearch
Enter as:I have a jobI am studyingI run a companyI am building something
Recent changes›Data engineer›2024-11-12
ConstraintCognitive automation2024-11-12

On 632 enterprise data workflows drawn from real warehouses, a code agent solved 21.3% — against 91.2% on the academic version of the same task

Data engineeroccupation page →
Event date / reported
2024-11-12
Evidence stage
ConstraintFailure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Tasks this bears on
Building the pipeline
Extracting from a source, reshaping it, loading it somewhere queryable, and scheduling the whole thing.
Automating✓ Evidence-backed
Owning what a number means
Defining active user, revenue, churn — and holding that definition when two teams want it to mean different things.
Still human-led✓ Evidence-backed
Where this applies
632 problems built from real data applications on BigQuery and Snowflake, with databases often exceeding 1,000 columns. What makes them hard is stated by the authors and is exactly this occupation's daily environment: the answer requires searching database metadata and dialect documentation, reading the project's own codebase, holding very long context, and emitting several queries in different dialects that frequently run past 100 lines. The 21.3% is one agent framework on o1-preview at the end of 2024 and will be out of date; the durable finding is the gap to 91.2% on Spider 1.0 and 73.0% on BIRD, which is a measurement of how much of the difficulty lives in the warehouse rather than in the SQL. It says nothing about employment, and nothing about pipelines that are built rather than queried - it tests producing the query, not operating the system that keeps producing it.
What this means
The difficulty in this job is measurable and most of it is not in the SQL. When the same class of task is posed over a real warehouse — a thousand columns, several dialects, documentation to search, a codebase to read — the success rate falls from roughly nine in ten to roughly two in ten. That gap is a number for the thing data engineers say and are rarely believed about: the query is the easy part, and knowing which table is the real one is the job.
What it does not yet show
A benchmark score is not an employer's decision, and a low one is not safety. The number is from one agent framework on one model generation at the end of 2024 and should be assumed to have moved; what does not move as fast is why it was low. It also tests producing a correct query, which is only one of this occupation's tasks — nothing here bears on building a pipeline, on what happens when the upstream schema changes, or on who is called when the numbers stop arriving.
What you can check
Pick one number your company steers by and trace it back to the tables it is computed from. Count how many of the decisions along the way are written down anywhere. That count, not the SQL, is what a tool would have to reproduce.
Does it change the assessment?
No. The impact index is never moved by a single event. What this record did: the 2 linked task judgements above now rest on evidence instead of inference.
Source
Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows (arXiv:2411.07763) · verified 2026-09-12 · Claude (VOLO agent) — arXiv abstract page read; the 632-problem count, the 1,000-column note and the 21.3% / 91.2% / 73.0% comparison taken from the authors' own abstract · interpreted 2026-09-12 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.

This record is cited in

  • What actually gets automated, and how to tell in advance
All changes for Data engineer →All recent changes →How events become evidence →