DEV Community

Cover image for Alerts Worth Waking Up For
Mansur Fattakhov
Mansur Fattakhov

Posted on Originally published at mind.mansur.expert

Alerts Worth Waking Up For

The most honest way to test your monitoring is to wait until everything goes down and then see who tells you about it first: the system or a real person. I once had a project where that examination arrived without warning and ended badly. The machines sat with a single provider, Kubernetes with Prometheus and Grafana lived right there among them, and nobody checked anything from the outside — why bother, it seemed, when everything is plainly visible from within. On the day the provider took every machine down at once, not a single message arrived. The graphs did not turn red, because they had stopped updating; the rules did not fire, because there was nothing left to fire; the notification channel held its ordinary workday silence, the very silence that for all the preceding months had meant that things were going well.

We found out about the outage not from monitoring and not straight away. And this was not a story about poorly chosen thresholds or unfortunate rules — it is a story about how monitoring that runs on the same servers as the services falls silent together with them, and monitoring that has fallen silent looks exactly like monitoring for which everything is fine.

Ever since, I have designed alerts not from the question "what else could we measure" but from the question "and if this does not work, who will notice, and when". What follows is what comes of that: which messages are worth sending to a human being, where thresholds come from, and why a good share of the alerts in an average project ought not to exist at all.

An alert is a promise that someone will get up

It all starts with a single distinction that spares more nerves than any technique. An alert is neither an observation nor a curiosity, but a promise: when it arrives, a particular person will take a particular action. If no action is implied, what you have in front of you is not an alert but a graph, and its place is on a dashboard people visit of their own free will, not in a channel that wakes them.

The easiest way to tell the two apart is a thought experiment: imagine the message arriving at three in the morning. If the only thing you will do is glance at it and lie back down, then it has no business arriving at night; and if you would do nothing about it in the daytime either, then it has no business arriving at all. I have written about this before in the piece on KPIs, where the share of alerts that people actually respond to turned out to be a far more honest metric than their number: an alert everyone ignores does more harm than a missing one, because it does not merely keep quiet about a problem — it trains people not to look at the rest of them either.

This rule has a reverse side that gets remembered less often. If action is implied and there is no alert, then you are seriously counting on someone noticing the breakage by accident. Usually it is the users who notice, and usually later than you would have liked.

Silence that looks like order

The breakage hardest to see does not look like a red graph but like the absence of a graph. A process died, an agent dropped off, the network cut away half the environment — and the panel does not flush crimson, it simply freezes. The last green dot is indistinguishable from a normal one, and an eye trained to hunt for red slides calmly past.

The rule "the service is not answering the scraper" is no salvation here, because a dead host fails to answer not out of pain but because, as far as the scraper is concerned, it no longer exists. What is needed is a separate class of rules — ones that watch not for bad values but for the very fact that data is arriving, and that fire when a metric stays quiet longer than is reasonable. I keep such a rule for every host, and it is the only thing that catches "everything has vanished" rather than "everything has broken".

Out of the same story with the provider grows a second conclusion: part of your checks must live on the outside. Internal monitoring knows everything about the services but dies along with them; an external check knows shamefully little — whether the domain answers, whether the certificate is valid — yet it outlives the collapse of everything else and asks precisely what a user would ask. Two loops cost pennies and cover different classes of outage: from the inside you see that things have gone bad, from the outside that you are no longer there.

And the last layer of the same onion is a watchman over the watchman. Rules may fire faithfully while the messages never arrive: the bot token went stale, the notification service went down, someone renamed the channel. The silence that results is indistinguishable from the silence of a calm day, and here the circle closes exactly where it began. So I keep alerts for notifications failing to be delivered and for the external monitor being unreachable — and yes, I understand that this watchman could in theory fall silent too; it is simply that with every layer the probability of a simultaneous failure grows less and less interesting.

Outages known about in advance

There is a separate breed of breakage that it is plainly shameful to learn about after the fact, because the date is known in advance and written into the object itself: a certificate expires, a domain runs out, an integration token goes stale. All of it happens on schedule and, at the appointed hour, brings the service down without fail.

On a busy website such an outage does not live long: the certificate expired, the page stopped opening, a quarter of an hour later somebody wrote in the chat, another quarter and it was mended. But on a website that gets only a few visits a day there is nobody to be first — and that same expired certificate quietly holds the page shut for hours, while you learn of it when you finally drop by yourself. The most galling part is that it is prevented by a single line: a warning two weeks out, then a week, then a day. I keep such checks in the external loop, the very one that will outlive the fall of the internal one.

Beside them stand the background jobs, and the logic is the same, only turned inside out. A backup, a nightly export, a cleanup of old data do not break loudly: they do not fall over with an error, they simply stop happening. So what has to be watched is not the failure but the absence of a successful run: my blog's nightly backup knocks on the monitor itself after every successful attempt, and if no knock has come by morning, a message arrives — regardless of whether the process died, never started on schedule, or quietly hung halfway through.

Where thresholds come from

A threshold taken from thin air is the chief producer of pointless nighttime wake-ups. "Let us say eighty percent" sounds judicious right up to the day it turns out that the service lives perfectly well at eighty-five and wakes you every night until somebody thinks to look at what those numbers have been over the past six months. The order ought to be the reverse: first the actual distribution, then a threshold slightly above the usual maximum, and after the very first false firing — higher still, because reality always proves noisier than it looks on a week's worth of graph.

The second thing without which a threshold is useless is the "for" duration. The rule must require the condition to hold for some time, otherwise any one-second spike will raise the alarm and you will find yourself in front of a laptop at the moment when everything has long since been fine. A couple of nights like that are enough for a person to start opening the messages every other time, and from there see above: an alert that gets ignored is worse than no alert at all.

There is a third consideration as well, far less obvious: a threshold lives not in a vacuum but in concert with the rest of the automation, and that automation has to be known. On my server there is a cleanup after the build runner which begins pruning images when about three quarters of the disk is taken. The free-space alert is set not where the number looks prettier but where it falls between the soft and the hard level of that cleanup: let it fire earlier and it will shout every time the system is already tidying after itself; let it fire later and it arrives when it is too late to mend anything and images must be torn out by hand on a live server. A good threshold is always an answer to the question "what is already happening without me", and not a round number out of one's head.

Severity as a promise, not as ornament

Severity levels make sense only if behind each of them stands an obligation that can be checked. I have three, and they read literally: critical wakes me at night, warning I look at today, info does not reach the main channel at all. The moment a fourth level appears, of the "important, but not that important" sort, the system begins to come apart, because the boundary between neighbouring levels ceases to be checkable and therefore ceases to be observed.

Hence the most reliable symptom of ailing monitoring: if everything of yours is critical, then nothing is critical. Marking by severity says nothing about how dear a metric is to you; it says what a person is obliged to do on seeing the message — and if the answer is the same for every level, the levels can safely be thrown out.

One cause, one message

When a node goes down, a conscientiously configured system sends two dozen messages: about the node itself and about every container that lived on it. Formally it is all true; in essence it is noise, and the single line that matters drowns in it. So messages have to be grouped by cause, and repeats sent no more often than once every few hours while the problem is alive: a person needs one signal and one place to look, not a chronicle of disintegration.

And without fail — a separate message about recovery. Without it there is no telling whether the outage has ended or you have merely turned away from the channel; and the habit of assuming that things will somehow sort themselves out will one day coincide with the case where they did not.

Alerts on meaning, not only on mechanics

Technical rules answer honestly the question of whether the system is working, but say nothing about whether it is doing the thing it exists for. A service can return an impeccable HTTP 200 to every request, all the graphs will be green and even, and inside it is already empty: the data supplier broke, the integration dropped off, an update changed the format, and the page now honestly and quickly shows nothing at all.

So on top of the technical ones I keep a few rules that watch for business-level signals: click-throughs have fallen to zero, the data sources have stopped answering, a regular job has not refreshed the data mart by its appointed hour. Such alerts catch precisely what the mechanics miss, and it is usually they that save your reputation, because a user does not see your uptime — they see an empty page and draw a conclusion about you as a whole.

The rules live in the repository

Last in order but not in importance: alerts are code, and they are worth handling as code. The rules lie next to the configuration, pass a check before rollout, and change by the ordinary working process rather than by mouse in an interface on the fly. The reason is not fastidiousness: monitoring is almost always patched in a hurry and in a foul mood, right after an outage, and it is exactly those changes that matter most to keep in the history — so that six months later one can see not only what the threshold is, but what kind of night gave birth to it.

Checklist

⬜ For every alert it is clear who does what when it arrives

⬜ There are rules for the absence of data, not only for bad data

⬜ There is an external loop that will outlive the fall of the internal one

⬜ There is an alert for the failure of notification delivery itself

⬜ Certificate, domain and token expiry dates are checked in advance

⬜ Background jobs alert on the absence of a successful run, not only on failure

⬜ Thresholds are taken from actual values and have a "for" duration

⬜ Severity levels mean a concrete promise, and critical is not on everything

⬜ One cause yields one message, and recovery arrives separately

⬜ There is at least one alert on business meaning, not only on mechanics

⬜ The rules lie in the repository and pass a check before rollout

Top comments (0)