PilotCognitive automation2026-03-16
UK government designers piloted GOV.UK Chat with over 10,000 users and 99 remote usability tests, added clarifying questions, and found streaming answers meant redesigning safety guardrails
Product / UX designeroccupation page →Event date / reported
2026-03-16
Evidence stage
PilotSmall-scale trial in a real setting. Tells us the deployment conditions are being tested, not that they hold — so one pilot is never enough on its own; two independent ones are.
Tasks this bears on
Research and usability testing
Watching users struggle, running interviews, reading the analytics, and turning what you saw into a decision.
Being augmented✓ Evidence-backed
Designing for non-deterministic products
Shaping products whose output varies — conversational interfaces, agents, generated content — where the old screen-by-screen craft does not apply.
New task✓ Evidence-backed
Where this applies
The UK government's digital service describing its own product. Over 18 months it ran two public pilots of GOV.UK Chat in which more than 10,000 users asked 26,000 questions; it introduced clarifying questions where users' questions were ambiguous, raising the answer rate for in-scope questions to 88%; answer streaming tested well with users but requires redesigning some safety guardrails; and the research used methods including 99 remote usability tests and a diary study with 30 hours of video interviews. It is a pilot reported by the team that built it; the assistant runs on a model from Anthropic, the maker of the model that wrote this record.
What this means
Designing a generative assistant turned out to need new design moves — asking the user a clarifying question, rethinking safety for answers that appear while being written — and a lot of research with real users. The design work grew rather than shrank.
What it does not yet show
One government pilot reported by its own team; it does not show how commercial teams design or test AI features.
What you can check
Open the Inside GOV.UK post of 16 March 2026 and find "we conducted 99 remote usability tests".
Does it change the assessment?
No. The impact index is never moved by a single event. Nor did this record change a layer: all 2 linked judgements above already rested on earlier evidence. This one adds to them.
Source
Government Digital Service — Inside GOV.UK blog, "5 things we learned testing GOV.UK Chat, an AI assistant for government" (Lead Product Manager, GOV.UK AI, and Lead User Researcher, GDS; posted 16 March 2026) · verified 2026-09-28 · Claude (VOLO agent) · interpreted 2026-09-28 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.