DEV Community

Mayowa O.
Mayowa O.

Posted on Originally published at Medium

Monitoring engines should know when to stay quiet

A monitoring engine's job is to raise the alarm when something breaks. Most engines do that part well.

The part they mostly leave as invisible plumbing, and the part I've noticed deserves more attention, is knowing when to stay quiet.

I've spent the past few months building a monitoring product called Mycellis. Most of what I learned about staying quiet came from getting it wrong first.

A user adds a new endpoint to their monitoring dashboard. Their staging environment is spinning up, or DNS hasn't finished propagating to their region, or Let's Encrypt is still issuing the cert. One of the ordinary early-life reasons things fail before they work.

The first pulse hits at zero seconds. Fails. The second, two minutes later. Fails. The third, at four minutes. Fails.

At five minutes and one second, the monitoring engine fires an alert. Your endpoint is down.

This can be the failure mode of a monitoring engine. Gather a few data points, hit a threshold, fire. The engine had never seen the endpoint succeed. Not once. It had three data points, all failures, gathered before the endpoint had a chance to exist. And it made a confident claim anyway.

I want to propose a different approach, borrowed from the plant that named the product I built.

In its biennial form, Mycelis muralis, wall lettuce, spends its entire first year as a low rosette. No stalk, no flower, no seed. It builds root and leaf. It commits to visible growth only in year two, once it has stored the energy to survive committing.
Monitoring should work the same way.

Mycelis muralis plant with yellow flowers growing in forest floor

Photo by Robert Flogaus-Faust, CC BY 4.0 via Wikimedia Commons.

How existing tools handle this

Credit up front. This is not a new problem, and I'm not the first person to think about it.

Datadog ships new_group_delay, which defaults to 60 seconds. Any new monitor group skips evaluation for its first minute of life. Datadog Synthetics layers on min_failure_duration, a knob that requires a test to be failing continuously for up to two hours before it alerts.

Better Stack marks fresh heartbeat monitors as "Pending," but only until the first heartbeat arrives. UptimeRobot and Cronitor use "grace periods," which sound related but actually govern how long a late heartbeat is tolerated before it's declared missed. Different problem.

Rob Ewaschuk's My Philosophy on Alerting, which heavily influenced Google's SRE book, argues that "over-monitoring is a harder problem to solve than under-monitoring." Philosophically aligned. But I couldn't find prior art naming new-monitor warm-up as a formal pattern.

The observation isn't that this problem is unsolved. It's that it's usually solved by a config parameter the user never sees, keyed to wall-clock time rather than to the thing you actually care about: whether the monitor has learned anything yet.

Those are the two design decisions I want to defend.

Design decision one: count over time

Time and pulse-count are both proxies for signal maturity. Neither is universally better. Each fails at a different edge of the interval range.

Time-based warm-up works well at fast polling frequencies. new_group_delay = 60s gives a typical monitor plenty of pulses before it starts evaluating. But it degrades as intervals lengthen. A monitor polled every two minutes sees exactly one pulse during that 60-second window, then starts alerting on whatever that single data point happened to say.

Count-based warm-up degrades in the opposite direction. A monitor with a 24-hour interval and a five-pulse threshold sits in warm-up for five days.

I chose count for Mycellis because our polling intervals are user-configurable across four orders of magnitude, from 10 seconds to 24 hours. Time becomes a leaky abstraction across that range. Counting pulses ties the gate to what actually matters: has this monitor learned anything yet? Your product may have different needs.

In Mycellis, N defaults to three. I started with five, watched real stalks in beta sit silent longer than felt right, and pulled it down. The value is arbitrary in the same way any threshold is arbitrary. Three is enough to catch a pattern, small enough that a working monitor doesn't stay silent for long. Tuned to what worked, not to what sounded elegant.

Design decision two: a state, not a config

The second decision is less about mechanics and more about honesty.

A new_group_delay = 60 config parameter is invisible. The user never learns their monitor is in a warm-up phase. It simply doesn't alert for a while. If they add a monitor and don't get paged for an hour, they might reasonably conclude the tool is broken, or that their endpoint is fine when it isn't. The behavior is happening. The user has no vocabulary for it.

In Mycellis, a stalk is what we call a single monitored endpoint. A URL the engine pulses on a schedule to check whether it's alive and fast. I made warm-up a formal state in the stalk lifecycle: AWAKENING. When a stalk is in AWAKENING, the dashboard shows it. The badge is different. The tooltip explains what it means. The user can see that their monitor is learning what normal looks like, and that alerts will begin once it has.

This is a small UX decision that turns into a larger product philosophy. The system is being honest about what it knows. It doesn't pretend to have opinions it hasn't earned.

How it actually works

Mermaid diagram showing the reliability state machine

The reliability state machine has four states: AWAKENING, HEALTHY, DEGRADED, DORMANT. Every new stalk starts in AWAKENING. HEALTHY and DEGRADED are bidirectional based on the rolling health index. AWAKENING is one-way. Once a stalk has earned an opinion, it doesn't unlearn it. DORMANT is designed for user-paused stalks; the pause endpoint is future work.

Stalk stalk = Stalk.builder()
        .reliabilityState(ReliabilityState.AWAKENING)
        .lastActivatedAt(Instant.now())
        // ...
        .build();
Enter fullscreen mode Exit fullscreen mode

lastActivatedAt becomes the anchor for the pulse count. Every completed pulse triggers a state re-evaluation:

private ReliabilityState evaluateReliability(
        double healthIndex, long totalCount, long pulsesSinceActivation) {

    if (pulsesSinceActivation < monitoringProperties.getAwakeningPulseThreshold()) {
        return ReliabilityState.AWAKENING;
    }

    if (healthIndex >= monitoringProperties.getHealthyThreshold()) {
        return ReliabilityState.HEALTHY;
    }
    return ReliabilityState.DEGRADED;
}
Enter fullscreen mode Exit fullscreen mode

The gate is a single condition. Has the stalk seen enough pulses to have an opinion? If not, it stays in AWAKENING regardless of what those pulses actually said, 100% failures included. Once the threshold clears, the same call picks HEALTHY or DEGRADED from the accumulated health index. There is no separate graduation step. The pulse that clears the threshold is the pulse that produces the first real verdict.

That's the state machine. Here's the guard that connects it to the alert pipeline:

private void evaluateStalk(Stalk stalk) {
    if (stalk.getReliabilityState() == ReliabilityState.AWAKENING) {
        return;
    }

    // ...normal evaluation: walk recent pulses, decide DOWN or RECOVERY...
}
Enter fullscreen mode Exit fullscreen mode

Two lines. Every 60 seconds, the alert engine walks every active stalk. Stalks in AWAKENING short-circuit before any pulse history is fetched. No DB read, no decision, no email. Once the guard clears, the normal alert path takes over. Five continuous minutes of failing pulses trigger DOWN. A single successful pulse after DOWN triggers RECOVERY.

The important thing about this guard is not the code. It's how it got there.

What writing this post taught me

The guard shipped six weeks after I designed AWAKENING.

I built the state machine in early July. AWAKENING was the differentiating idea, the whole reason the lifecycle model existed in the first place. I built the state, wired it into the evaluation logic, exposed it on the dashboard, mentally drafted a blog post about how elegant it was, and moved on to the next thing.

Two nights ago, I sat down to actually write this post. Before I could describe how AWAKENING gated the alert pipeline, I did what any honest writer should do. I re-read the code I was about to describe.

AlertEngine.java. I searched for AWAKENING. Zero matches. I searched for ReliabilityState. Also zero. And there in the class-level Javadoc, six weeks ago, past-me had left a note in plain English: out of scope for this commit.

I remembered writing it. I had shipped the state. I had shipped the UI. I had shipped the mental model. I had not shipped the actual guard. At any polling interval slower than about 100 seconds, a brand-new stalk pointed at a failing endpoint would receive a DOWN alert before AWAKENING had a chance to clear. The exact scenario I opened this post with, the one I'm arguing my design prevents, my code did not, in fact, prevent.

I would have published this post claiming the guard existed. It did not.

I added the two-line guard. I wrote the test. I deployed it. I smoke-tested it against https://httpbin.org/status/500, a URL that returns 500 on every request, guaranteed to fail every pulse. Watched the stalk sit in AWAKENING through five minutes of continuous failure with zero alerts fired. Watched it transition to DEGRADED at the three-pulse mark, and the first alert appear seconds later. The pattern I had been claiming to have implemented for six weeks was now, finally, actually implemented.

Then I came back to finish writing this post.

The lesson isn't that I made a mistake. Engineers ship half-finished ideas every day. The lesson is that writing publicly about your own systems is a debugging tool. Not marketing. Not portfolio work. An actual debugging tool that surfaces the gap between what you believe your code does and what your code actually does. You cannot describe what you have not examined, and you cannot examine something without discovering what you assumed but never checked.

If you build things, write about them. Not because it makes you look smart. It will occasionally make you look the opposite. Because it makes the things better.

Trade-offs

"Count-based warm-up is fragile too. A monitor with a 24-hour polling interval sits in AWAKENING for three days." True. Mycellis has no time ceiling on AWAKENING today. A future version probably should. Exit on N pulses OR M hours, whichever comes first. Every threshold has edge cases. Pick the ones you'd rather live with.

"This is new_group_delay with a different name." Same underlying goal, different mechanism (count over time), different UX (visible state over invisible config). Whether either of those differences matters is what this post is arguing. The claim isn't invention. It's design.

"Why not let users configure this themselves?" Because defaults are what most users experience. Formal states enforce the right default while leaving the threshold configurable for the users who want to tune it.

Closing

A stalk in AWAKENING isn't broken. It isn't healthy either. It just hasn't lived long enough to say.

Building products with epistemic honesty means giving the user a language for that state. Not silencing the uncertainty behind a config parameter that pretends the monitor already has an opinion. The plant does not flower in year one. It builds. It commits to visible growth only when it has earned the energy to survive committing.

Mycellis is in closed beta. If you want to see AWAKENING in action, sign up at https://mycellis.dev/signup.

This post was originally published on Medium.

Top comments (0)