Every vendor is making the same pitch for an on-call agent. An alert fires at 3am, the agent reads the runbook, checks the dashboards and finds the cause, then either fixes it or wakes a human with a diagnosis instead of a bare pager message. Nobody wants the 3am page, and everybody wants the diagnosis.
The fear is just as easy to picture. The agent misreads a graph, decides the fix is to restart the database, and it's 3am and nobody's watching.
From March to August we worked out where the truth sits between those two. The agent is on the rotation now, in a limited way, and I wouldn't go back. But it got there in four stages, each with a gate, and the gates are the part of the story worth telling.
Stage one: read everything, touch nothing
For the first six weeks the agent could read logs, metrics, traces, the runbooks, the deploy history and the incident history, and couldn't write to anything. When an alert fired, the agent was triggered alongside the human page. It investigated, wrote up what it found and posted that in the incident channel. The human on call did their normal job and read the write-up at some point.
To move on, on-call engineers had to rate the write-up useful, on a two point scale, at least 80 percent of the time across 30 incidents.
It failed the first time. The write-ups were long, restated the alert and hedged, and by week two engineers had stopped reading them. The fix was changing what we asked the agent to produce: one line for the most likely cause, one for the evidence, one for the recommended action, and everything else in a thread.1 After that the rating went to 86 percent, and by week six the on-call engineers were opening the write-up before the dashboard. That was how we knew it was ready.
We learned two things in this stage that shaped everything after it. The agent was better than the median on-call engineer at linking a deploy to an alert, because it always checked the deploy history and humans at 3am don't. And it was worse at telling when a graph was normal, because it had no memory of what the graph usually looked like. We fixed that by giving it a tool that returns seven days of baseline for any metric it looks at, and that one tool cut its false "this is anomalous" rate roughly in half.
Stage two: propose the action, a human clicks
In stage two the agent could offer an action from a fixed list, as a button in the incident channel: restart this service, scale this deployment from 3 to 6, roll back this deploy, flip this feature flag off. The human on call clicked it or didn't. The agent couldn't click.
The list was short on purpose. It had eight actions, all reversible, all things a runbook already told a human to do in that situation. To move on, the human had to click the proposed button unchanged at least 75 percent of the time over 40 proposals, and no proposal could have made things worse if it had been executed.
That second condition needed a review. Every proposal, clicked or not, was looked at the next day by someone who wasn't on call, and they answered one question: if this had run automatically, would it have been correct? Of the first 40, 34 would have been correct, 4 harmless but useless, and 2 wrong. Both wrong ones were rollbacks of a deploy that lined up in time with the alert but didn't cause it. It was the same failure twice.2
We fixed it with a rule instead of a prompt change. A rollback proposal has to point to something in the deployed change that could cause the failing behaviour, on top of the timing. The agent had to read the diff and say what in it could produce the symptom. If it couldn't, it could still propose the rollback with a "timing only" label, and the human treated that label as a reason to look harder. In the next 40, wrong proposals went to zero.
Stage three: act on the short list, tell everyone
Stage three is what people mean when they say "the agent is on call". For four of the eight actions, the agent could go ahead without a click: restart, scale up, flag off, and the rollback with a diff citation. Scaling down, rolling back on timing alone and anything involving data stayed off the list.
Every automatic action followed three rules.
It announces before it acts, in the channel, with a 60 second window where a human can type stop.3
It takes one action per incident. If that action doesn't clear the alert, the agent goes back to proposing and a human decides. The rule exists because autonomous systems rarely fail with one bad action. They fail with a chain of reasonable-looking actions that add up: restart didn't help, so scale up, which didn't help, so roll back, and now three things have changed and nobody knows which one mattered. So it gets one action, then a human.
It records everything as an incident timeline entry, in the same format as a human's actions, with its reasoning attached. The post-incident review reads the agent's actions exactly the way it reads a person's.
The gate here was stage two's second condition again, over 30 automatic actions: none could make things worse. Getting to 30 took eleven weeks, because most incidents in that time were resolved at the proposal stage or were things the agent rightly declined to touch. None made things worse.4
Stage four, which we are in
There are eleven automatic actions now. They were added one at a time, each after a stretch as a proposal with a perfect record, and the stage three rules still apply to all of them.
What changed most in stage four was the rotation, more than the agent. The human on call now gets paged for about 40 percent of the alerts they used to get. The agent either resolves the other 60 percent or diagnoses them far enough that the page says "this is the known flaky job, agent has restarted it, no action needed", and the engineer acknowledges from bed without opening a laptop. Mean time to a first diagnosis, from the alert to the first timeline entry that names a cause, went from around 14 minutes to under 3.
The pages that are left are the real ones, the ones that should wake somebody up.
The rules, on one page
Read before you propose, and propose before you act. Each stage is gated on the record of the one before, reviewed by someone not on call.
Actions come from a fixed list. The agent doesn't make one up, and adding to the list is a decision made in daylight.
Every automatic action is reversible, and announced with a window to stop it.
One action per incident, then a human.
Timing isn't causation. A rollback needs a cited diff.
Give the agent a baseline for every metric, or it'll think everything is an anomaly.
The output is three lines, and everything else goes in the thread.
Record the agent's actions the way you record a person's, and review them the same way.
The one incident
There was one, in stage three. It's worth telling because it wasn't the agent's fault and the rules still mattered. A restart proposal ran automatically, was announced and completed correctly. At that moment a human, paged for a different alert on the same service, was halfway through restarting the same pod by hand. The two restarts overlapped, and the service was down for 40 seconds longer than it needed to be. The fix was a lock: before acting, the agent checks for an active human session on the target and backs off if there is one. The daylight review caught it the next morning, which is exactly why the review exists.
What I would not do
I wouldn't start at stage three because a vendor's demo did. The demo is stage three on a scripted incident, and your incidents aren't scripted.
I wouldn't give the agent any action that touches data: no migration, no cache flush, no queue purge. Those stay with humans, maybe forever, because for those actions "reversible" is a matter of opinion.
I wouldn't skip the daylight review of proposals. It's tedious and takes ten minutes a day, and it's the only thing that turns the agent's mistakes into rules before they turn into incidents.
And I wouldn't measure it by pages avoided. Measure time to first correct diagnosis, and actions that made things worse. The first is what you get. The second is what you pay, and if it isn't zero the agent isn't ready for the next stage, however quiet the pager has got.
Originally published at zeybek.dev.
-
The investigation stayed the same, only the output changed shape. ↩
-
Reading the deploy history for correlation had been the agent's strength in stage one, and in stage two it turned into over-confidence. ↩
-
At 3am nobody will, and that's fine, but during the day someone often does, usually with "wait, I'm already on it". ↩
-
Two were unnecessary, restarts of a service that would have recovered by itself within a minute, and those were fine. ↩
Top comments (1)
the gate in stage one is what i liked: u measured whether people actually read the write-up. the 1-line cause, 1-line evidence, 1-line action format is a nice detail. did the engineers trust it more once it showed the evidence, or was it just easier to scan?