DeploymentCognitive automation2024-10-31
Microsoft's AI Red Team says it has run over 80 operations on more than 100 generative AI products, and that automation should not be used to take the human out of the loop
Software tester / QA engineeroccupation page →Event date / reported
2024-10-31 · reported 2025-01-13
Evidence stage
DeploymentAn employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Tasks this bears on
Testing systems with a model inside
Evaluating software whose output is not deterministic — building evaluation sets, catching regressions in behaviour, testing for harmful outputs.
New task✓ Evidence-backed
Where this applies
A paper by Microsoft's AI Red Team about its own work testing generative AI products for safety and security failures. It says that as of October 2024 the team had conducted over 80 operations covering more than 100 products, that the volume made fully manual testing impractical so it built automation, and that such tools should not be used with the intention of taking the human out of the loop, since prioritising risks, designing attacks and defining new categories of harm need human judgement. This is a dedicated red team rather than product QA, and Microsoft sells the models and promotes its open-source tool.
What this means
Testing software with a model inside has become an operation of its own at one large company, run by specialists across a hundred products, with automation doing volume and people doing judgement. That is new testing work, and it sits with people.
What it does not yet show
One company's specialist team; it does not show how product QA teams test model-based features.
What you can check
Open arXiv:2501.07238 and find "should not be used with the intention of taking the human out of the loop".
Does it change the assessment?
No. The impact index is never moved by a single event. What this record did: the 1 linked task judgement above now rest on evidence instead of inference.
Source
Bullwinkel, Minnich et al. (Microsoft) — "Lessons From Red Teaming 100 Generative AI Products", arXiv:2501.07238 (13 January 2025) · verified 2026-09-27 · Claude (VOLO agent) · interpreted 2026-09-27 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.