DeploymentCognitive automation2026-05-28
Two Google engineers described the agents their own SRE organisation runs in incident response: alert grouping, handoff documents, postmortem drafts, and in some cases autonomous mitigation
DevOps / platform / SRE engineeroccupation page →Event date / reported
2026-05-28
Evidence stage
DeploymentAn employer has put it into production. Can move the baseline — weighted by scale and how similar the setting is.
Tasks this bears on
Being woken up
Deciding at 3 a.m., under time pressure and with partial information, what to roll back, what to degrade and what to tell people while it is still broken.
Still human-led✓ Evidence-backed
Where this applies
Google describing its own operations rather than selling into somebody else, written by a Distinguished Software Engineer and a Distinguished Site Reliability Engineer and published on the company blog. What it states in the past or present tense, which is the half worth recording: the SRE team has developed agents that monitor and improve playbooks from their use during incidents and can generate new playbooks from incidents; an alerting agent that groups, pre-processes and enriches alerts before autonomous handlers take them; agents that consolidate the chat spaces, videos and tracking documents used during an incident, create handoff documents between SREs, draft postmortems, and manage internal and external incident communications; agents created to investigate incidents and, in the words of the post, in some cases to autonomously mitigate issues; and AI Insights, a system that reviews past incidents and feeds what it extracts back to those agents. The rest of the post is plans, and is not recorded. Three limits, and the first is the largest. There is not one number in it: no share of incidents touched, no count of agents running, no before and after, and nothing about how many people are on the rota. Second, the stake is direct and is not hidden by the first-party framing — every component named underneath is a Google product (Gemini, the Agent Development Kit, the Gemini Enterprise Agent Platform, MCP on Google API infrastructure), so this post doubles as a reference architecture for things the publisher sells. Third, it is one operator, and an unusual one: an organisation with twenty years of its own reliability discipline is not a fair stand-in for a team of four keeping a pager. What makes it evidence about this task rather than about tooling is where the agents sit: the work around the decision at three in the morning, assembling the context, handing over, telling people what is happening, writing it up afterwards, is the part that moved, and the post says an agentic approach does not necessarily imply removing people from the process for higher-risk services.
What this means
At one very large operator, the parts of a three-in-the-morning incident that went to software first are the ones around the decision rather than the decision itself: assembling the context, grouping the alerts, handing over to the next engineer, telling people what is happening, and writing it up afterwards. The post also says agents in some cases mitigate autonomously, and that for higher-risk services the approach does not necessarily take people out.
What it does not yet show
There is not one number in the post. It does not say what share of incidents these agents touch, how many are running, what changed before and after, or how many people are on the rota now. It also does not say whether anyone is woken up less often, which is the question the task is actually about.
What you can check
Look at your own last three incidents and mark which minutes went to finding out what was happening, which to deciding, and which to telling people. The first and third are the parts this employer handed over; the second is the part it says it kept. If your own split is nothing like that, the record does not transfer to you, and that is worth knowing too.
Does it change the assessment?
No. The impact index is never moved by a single event. What this record did: the 1 linked task judgement above now rest on evidence instead of inference.
Source
Google Cloud — Stevan Malesevic and Christopher Heiser, AI in SRE: where and how Google is deploying agentic AI to improve operations (28 May 2026) · verified 2026-09-22 · VOLO agent · interpreted 2026-09-22 · VOLO agent
Primary source — published by the party that did this, or the authority of record. No co-signature needed.