Recent changes
Every entry here is a verified event tied to a specific occupation and task. This is the evidence engine's public face — not a news feed: an event only appears once someone has read the source and connected it to a task judgement.
37 verified records across 30 occupations. The stage label says how much an event can move an assessment; nothing here changes the impact index automatically.
Waymo opened fully autonomous paid rides to the public in Denver, San Diego and Tampa, bringing the number of US cities where it carries riders with no driver to 14
Scope: US only; service areas are geofenced within each city and exclude some conditions. Company blog; datePublished 2026-09-01.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Writing roles fell from about one in six of UK agency job postings to about one in thirty by July 2026, while overall agency postings rose 24 percent
Scope: Agency job postings in the UK, measured as a share of all agency postings — not employment, and not in-house or freelance copywriting. A share can fall because other roles grew; here total postings rose while writing rose far less.
What this stage can show: Verifiable change in hiring, headcount, hours or job scope. Highest weight — but causal attribution still has to be argued, not assumed.
A New York school district paused an approved 60,000 dollar humanoid teaching robot after state officials, the teachers union and parents objected
Scope: One rural district, one stationary robot bought as a teaching aid for a high-school robotics course — not a system that teaches a class. What stopped it was procurement, student-data agreements and union objection, not any limit on what the device can do.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
PepsiCo and Gatik began a driverless freight partnership with no safety driver in the cab, serving about 250 retail locations across three US states
Scope: Fixed middle-mile routes on highways and surface streets in Texas, Arizona and Arkansas, serving retail distribution; one remote supervisor oversees several trucks. Says nothing about long-haul, unmapped routes, or the yard and dock work that stays with people.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Abridge made its ambient documentation tool available to nurses across all its health-system clients, more than 250, with several named systems already using it
Scope: US health systems, and the nursing documentation task only. Availability across 250+ clients is not adoption by them — the article names a handful already using it. Says nothing about hands-on care, which is most of a shift.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Estonia's AI Leap pilot reached all 154 upper-secondary schools (~20,000 students, ~4,900 teachers); teachers got ChatGPT and Gemini plus training from Aug 2025, and over 60% use them weekly
Scope: One small country, grades 10–11, a three-year pilot; a Tartu/Stanford/OpenAI impact study is still pending. The student-facing tutor app (Estonian-language, built with OpenAI) only opened in late January 2026, and 47% of activated accounts used it weekly.
What this stage can show: Small-scale trial in a real setting. Tells us the deployment conditions are being tested, not that they hold.
Sweetgreen agreed to sell Spyce, the unit behind its Infinite Kitchen automated makeline, to Wonder for $186.4 million, keeping the technology in its 20+ automated restaurants under a supply deal
Scope: One US fast-casual chain (bowls and salads): assembly-line automation of a fixed menu, with 38 Spyce staff moving to the buyer. Company release; earlier earnings calls cited labour savings, this release gives none.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Mercy said Microsoft Dragon Copilot's ambient nursing tool is in use on inpatient units at hospitals in St. Louis, Springfield and Fort Smith, turning narrated care into Epic flowsheet entries
Scope: One US health system, inpatient units, early rollout; Mercy reports 8–24 minutes saved per shift for high-use nurses and a 29% cut in incremental overtime. Mercy co-developed the tool with Microsoft, so the figures come from an interested party. Nurses review and edit before filing; documentation only, not clinical decisions.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Deloitte agreed to repay the last instalment of a US$290,000 Australian government report after fabricated references and a misattributed court quote were found; it acknowledged using Azure OpenAI
Scope: An assurance review for Australia's Department of Employment and Workplace Relations by a Big Four firm's consulting practice — professional-services output, not bookkeeping. Corrected report issued 26 Sep 2025; the repaid instalment was later put at about A$97,000 (CFO Dive, 21 Oct 2025). Bears on who owns machine-drafted output, not on accounting tasks directly.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
IFR's World Robotics 2025 counted 542,000 industrial robots installed in 2024 and 4.66 million in operation worldwide (+9%); Asia took 74% of new installations and China alone 54%
Scope: Global industry statistics from the robot manufacturers' federation, based on member and national-association data. Robot stock, not job counts; automotive and electronics dominate the installations.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
A METR randomised trial of 16 experienced open-source developers on 246 real issues found they took 19% longer with early-2025 AI tools, while believing they had been 20% faster
Scope: Experienced maintainers on large, mature repositories they know well (about 5 years each); tools were mainly Cursor Pro with Claude 3.5/3.7 Sonnet. The authors explicitly do not claim the result generalises to most developers or to greenfield or junior work.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Crunchyroll's German subtitles for an anime premiere contained the line 'ChatGPT said…'; the company said a third-party vendor had used AI-generated subtitles in violation of its agreement
Scope: Entertainment subtitling, one streaming platform, one episode (Necronomico and the Cosmic Horror Show, ep. 1, 1 July 2025). The company's president had said in 2024 it was testing generative AI for subtitling; the failure was unreviewed vendor output.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Amazon deployed its one-millionth robot across 300+ facilities; its Shreveport next-generation site needs 30% more staff in reliability, maintenance and engineering roles
Scope: Amazon's global fulfilment network; company-reported. The robot count is fleet size, not a statement about total headcount.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Goldman Sachs made its in-house GS AI Assistant available to all employees after use by thousands of staff, for summarising complex documents, drafting initial content and data analysis
Scope: One US investment bank; a firmwide general-purpose assistant with versions for bankers, research analysts and wealth staff, announced in a memo by CIO Marco Argenti. Reported from the internal memo (Reuters first); no headcount statement.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Walgreens opened its 12th robotic micro-fulfilment centre; the network fills 3.5 million+ prescriptions a week for 5,000+ stores, about 40% of a supported store's prescription volume
Scope: One US chain; central fill of a share of routine prescriptions, the rest stays in-store. Company-reported; the release says freed pharmacy time goes to vaccinations and adherence support. Trade press notes this was the first opening since the chain paused expansion in autumn 2024.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
The Chicago Sun-Times and Philadelphia Inquirer printed a syndicated summer reading list with books that do not exist; the freelancer said he used an AI tool and the section was not reviewed
Scope: US newspapers; a syndicated special section produced by King Features (Hearst), not by newsroom staff. King Features said it ended its relationship with the writer; the paper's CEO called it a failure of review, not of reporting.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Klarna's CEO said the company would again hire humans for customer service, saying cost had been 'too predominant' in its AI-first approach and quality had suffered
Scope: Same company as the February 2024 deployment; the assistant kept handling roughly two-thirds of inquiries. Secondary source quoting a Bloomberg interview of 8 May 2025.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Aurora began regular driverless commercial freight deliveries between Dallas and Houston with no human on board, for Uber Freight and Hirschbach
Scope: US, Texas, one interstate corridor, heavy-duty trucks; hub-to-hub with human-handled first and last segments. Expansion to El Paso and Phoenix announced for end-2025.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Meituan received China's first nationwide low-altitude logistics operating certificate from the CAAC for its delivery drones, after 450,000+ drone orders on 53 routes by end-2024
Scope: China; routes concentrated in Shenzhen, Beijing, Shanghai, Guangzhou and Nanjing. Volume is a small fraction of platform orders; state-media source reporting company figures.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
US contractor Rosendin reported field trials of a three-robot solar module installer: robots plus a two-person crew set 350–400 modules per 8-hour shift; electricians do the grid connections
Scope: Utility-scale solar fields in Texas, module placement only (developed with ULC Technologies); demonstrated on site on 17 April 2025. Not building wiring or residential work; contractor-reported rates with no independent measurement.
What this stage can show: Small-scale trial in a real setting. Tells us the deployment conditions are being tested, not that they hold.
EY announced an initial deployment of 150 AI agents to support 80,000 of its tax professionals in data collection, document review and income and indirect tax compliance, built with NVIDIA
Scope: One Big Four firm's own tax practice, global; a launch announcement whose scale figures (3 million tax deliverables, 30 million processes 'over the coming year') are forward-looking targets, not measured results. No headcount statement.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Lloyds Banking Group announced 136 UK branch closures for May 2025–March 2026, citing 10 million fewer branch visits in 2024 and a 48% five-year fall in branch transactions
Scope: UK retail banking (Lloyds, Halifax, Bank of Scotland). Affected staff were offered other roles: a change in where and what branch staff do, not announced redundancies.
What this stage can show: Verifiable change in hiring, headcount, hours or job scope. Highest weight — but causal attribution still has to be argued, not assumed.
A UK cross-government trial gave 20,000 civil servants in 12 organisations Microsoft 365 Copilot for three months; users self-reported saving 26 minutes a day on average, mostly on drafting
Scope: UK central government, 30 Sep–31 Dec 2024, self-reported time savings with no control group; scheduling meetings saved about 9 minutes, drafting documents 24. The report notes the tool struggled with complex data and that human oversight was needed throughout.
What this stage can show: Small-scale trial in a real setting. Tells us the deployment conditions are being tested, not that they hold.
Coca-Cola released 2024 holiday ads made by three AI studios with generative video models, which it called 'a collaboration of human storytellers and the power of generative AI'
Scope: One global brand's broadcast campaign; the AI spots ran alongside conventionally produced ads and drew criticism from creative professionals. Release 12 Nov 2024 per TODAY; NBC report 18 Nov 2024.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Alphabet's CEO said on the Q3 2024 earnings call that more than a quarter of all new code at Google is generated by AI, then reviewed and accepted by engineers
Scope: One large technology company, company-reported, with no definition of how 'generated' is measured (autocomplete versus whole functions). Says nothing about hiring; the review step stayed with engineers.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Uber put QueryGPT, a natural-language-to-SQL tool, into production for operations and support teams, reporting query authoring time down from about 10 to about 3 minutes (~300 daily users)
Scope: One company, limited release; Uber's platform runs about 1.2 million interactive queries a month, so this is a small share. The post itself flags hallucinated tables and columns and prompt quality as open problems.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
The EU AI Act (in force 1 August 2024) lists AI used to recruit or select people — placing targeted job ads, filtering applications, evaluating candidates — as high-risk under Annex III, point 4
Scope: EU market. Annex III obligations (risk management, human oversight, transparency) were scheduled to apply from 2 August 2026 and are subject to amendment. Source mirrors the regulation text; the Official Journal text is Regulation (EU) 2024/1689.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Figma disabled its Make Designs prompt-to-UI feature a week after launch when 'weather app' prompts produced screens closely resembling Apple's; the CEO cited insufficient QA; relaunched Sept 2024
Scope: One design-tool vendor's generative feature (GPT-4o and Amazon Titan on hand-built component design systems), limited beta. A vendor rollback, not a customer deployment; the feature returned within three months as First Draft.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Klarna said it generated 1,000+ marketing images with generative AI in Q1 2024, cutting the image development cycle from six weeks to seven days and image production costs by $6 million a year
Scope: In-house marketing imagery at one fintech (tools named: Midjourney, DALL-E, Firefly, Topaz Gigapixel, Photoroom). Campaign assets, not brand identity or bespoke illustration; company-reported.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Klarna cut sales and marketing spend 11% in Q1 2024 while running more campaigns, attributing 37% of the savings (about $10 million annualised) to generative AI for images, copy and translation
Scope: One fintech's in-house marketing team; external agency spend fell 25%. Company-reported; the release gives no headcount figures.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Klarna said generative AI now handles 80% of its copywriting through an internal 'Copy Assistant', and that AI accounted for about $10 million of annualised marketing savings in Q1 2024
Scope: One global consumer fintech's in-house marketing; company-reported figures with no independent audit. Applies to volume marketing copy, not to positioning work or regulated claims.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Klarna's AI assistant handled 2.3 million conversations — two-thirds of its customer-service chats — in its first month; resolution time fell from 11 to under 2 minutes
Scope: Global consumer fintech, chat channel, 23 markets; refunds, returns, payments, disputes. Company-reported figures; customers could still choose a human agent.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Meta reported deploying TestGen-LLM at Instagram and Facebook test-a-thons: it improved 11.5% of the classes it was applied to and engineers accepted 73% of its recommended test cases into production
Scope: One company, improvement of existing unit-test classes with build/pass/coverage filters; 75% of generated cases built, 57% passed reliably, 25% added coverage. Company-authored paper; a time-boxed test-a-thon setting rather than routine pipeline use.
What this stage can show: Small-scale trial in a real setting. Tells us the deployment conditions are being tested, not that they hold.
Singapore's BCA soft-launched CORENET X on 18 Dec 2023: one coordinated BIM model replaces separate agency submissions, ~20 approval stages become 3 gateways; automated geometric checks in development
Scope: Singapore regulatory submissions only; mandatory for new projects of 30,000 m² GFA or more from 1 Oct 2025 (info.corenet.gov.sg). Automated checks cover straightforward geometric and spatial rules across architecture, C&S and M&E; judgement-based compliance stays with the qualified person.
What this stage can show: An employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
UK supermarket chain Booths removed self-checkouts from 26 of its 28 stores and returned to staffed tills, citing machines that were slow, unreliable and impersonal and trouble with loose produce
Scope: One regional UK chain (28 stores); a reversal, not the industry norm. Some larger chains later trimmed self-checkout too; others kept expanding it.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
A US federal court fined two attorneys and their firm $5,000 for a brief citing six non-existent cases generated by ChatGPT, noting that using a reliable AI tool is not itself improper
Scope: US federal court (S.D.N.Y.). Establishes that the signing lawyer, not the tool, carries responsibility for machine-assisted filings; several bars issued guidance afterwards.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
A US federal court fined two attorneys and their firm $5,000 for a brief citing six non-existent cases generated by ChatGPT, noting that using a reliable AI tool is not itself improper
Scope: US federal court (S.D.N.Y., Judge Castel). The sanction rested on failing to verify and on standing by the citations once challenged. Widely cited in later professional-conduct guidance.
What this stage can show: Failure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
How events become evidence, and why most of the site is still labelled inference: Method and evidence →