DeploymentCognitive automation2026-07-30
Uber says its AI mobile-testing system runs 1,013 end-to-end tests continuously in CI and has turned its test engineers from "clickers" into people who build its context
Software tester / QA engineeroccupation page →Event date / reported
2026-07-30
Evidence stage
DeploymentAn employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Tasks this bears on
Manual regression testing
Clicking through the same flows every release to confirm nothing that used to work has broken.
Automating✓ Evidence-backed
Where this applies
A paper by Uber engineers about a system running in Uber's own pipelines. It says DragonCrawl achieves a 91.6% pass rate on iOS and 92.2% on Android across 1,013 automated tests running continuously in CI/CD, cut test onboarding from 96–120 hours to under 4, and saved an estimated 27 developer years of test maintenance; the test engineers who used to execute scripts were repositioned as 'Context Engineers' rather than 'Clickers', and Uber says it achieved scalable quality assurance without proportional headcount growth. The time savings are Uber's own estimates, and it covers mobile end-to-end tests at one company.
What this means
Clicking through the same app flows every release is being taken over by a model that drives the app itself, at one large company. The testers who did the clicking now write what the system needs to know — the role changes from executing to specifying.
What it does not yet show
One company's own account and estimates; it says headcount did not grow proportionally, not that it fell, and covers mobile tests only.
What you can check
Open arXiv:2607.28750 (DragonCrawl) and find "across 1,013 automated tests running continuously in CI/CD pipelines".
Does it change the assessment?
No. The impact index is never moved by a single event. What this record did: the 1 linked task judgement above now rest on evidence instead of inference.
Source
Uber — "DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing", arXiv:2607.28750 (30 July 2026; accepted at ASE-Industry 2026) · verified 2026-09-27 · Claude (VOLO agent) · interpreted 2026-09-27 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.