DeploymentCognitive automation2024-05-16
Microsoft's Bing team reports using LLMs with expert human labellers for most of its offline relevance labels since late 2022, finding them more accurate than third-party labellers
Machine learning engineeroccupation page →Event date / reported
2024-05-16
Evidence stage
DeploymentAn employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Tasks this bears on
Getting data the thing can learn from
Collecting, labelling, cleaning and deciding what to leave out — and noticing when the data says something different from the world.
Being augmented✓ Evidence-backed
Where this applies
Four Microsoft researchers describing Bing's own labelling practice in a peer-reviewed conference paper. They write that Bing has been using LLMs, in conjunction with expert human labellers, for most of its offline metrics since late 2022; that in their experience LLM labels proved more accurate than any third-party labeller, including staff, much faster and many times cheaper; and that they have found it necessary to build harder gold sets over time to keep telling labellers and prompts apart. The labels mostly feed evaluation metrics rather than training data, and Microsoft builds and sells the models used, so it has a stake. It says nothing about how many labellers or engineers the work employs.
What this means
Getting data fit to learn from — here, relevance labels — moved at a major search engine from crowds of human labellers to models checked by experts. The engineer's work shifts to building harder test sets that can still tell good labels from bad.
What it does not yet show
It describes one team's practice, mostly for evaluation labels, by the company that sells the models; it does not show labelling for training data moving the same way or any change in staffing.
What you can check
Open arXiv 2309.10621 (v3), "Large Language Models can Accurately Predict Searcher Preferences", and find "We have been using LLMs, in conjunction with expert human labellers, for most of our offline metrics since late 2022".
Does it change the assessment?
No. The impact index is never moved by a single event. What this record did: the 1 linked task judgement above now rest on evidence instead of inference.
Source
Paul Thomas, Seth Spielman, Nick Craswell and Bhaskar Mitra (Microsoft) — "Large Language Models can Accurately Predict Searcher Preferences", SIGIR '24 (arXiv 2309.10621v3, 16 May 2024) · verified 2026-09-27 · Claude (VOLO agent) · interpreted 2026-09-27 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.