DEV Community

Sara
Sara

Posted on

The Alert Fired. It Wasn't the Root Cause.

The alert fires.

You open the dashboard, find the service that triggered it, and start investigating.

Except the thing that alerted isn't always the thing that broke.

Sometimes it's just where the problem became visible.

During a recent panel on production alerting, engineering leaders from Apple, Cars.com, and SLB shared incidents where the first signal pointed in one direction while the root cause was somewhere else.

Three incidents stood out.

1. The recovery hid the failure

At Cars.com, Heather Osborn's team dealt with intermittent application timeouts for months.

The service would time out, restart, and recover.

timeout → restart → healthy → timeout → restart → healthy
Enter fullscreen mode Exit fullscreen mode

The automation was working.

That was also the problem.

Each restart removed the immediate symptom without removing the underlying failure. Eventually, the issue was traced back to a minor PostgreSQL update.

Automated recovery can make a system look healthy while the same problem keeps happening underneath.

A recurring signal may not deserve a page every time, but that doesn't mean you should stop tracking it.

2. One failure created alerts everywhere

At SLB, Barnadeep Bhowmik saw almost the opposite problem.

429s appeared. Storage failures followed. Downstream services started alerting.

Eventually, three separate war rooms were investigating what looked like different problems.

They weren't.

The failures shared an upstream cause, ultimately traced back to an Active Directory server with a patch pending a restart.

             upstream failure
                   │
        ┌──────────┼──────────┐
        ↓          ↓          ↓
      429s      storage     service
                 errors     failures
Enter fullscreen mode Exit fullscreen mode

Every alert could be technically correct and still not tell you anything new.

When alerts suddenly appear across multiple services, the better question may not be:

What is broken?

It may be:

What do these failures have in common?

3. Sometimes the signal is what's missing

Sarah Kaplan from Apple shared an incident that started with 503s.

Capacity looked suspicious. Kubernetes nodes were intermittently going offline. DNS queries were failing.

The investigation eventually reached an Unbound DNS issue involving larger DNS payloads.

But one of the useful clues was much simpler:

Some nodes stopped logging.

We usually monitor things that happen:

errors increased
latency spiked
CPU crossed 80%
requests failed
Enter fullscreen mode Exit fullscreen mode

But absence can be a signal too.

A node normally sends logs. Then it doesn't.

A job normally completes. Then it doesn't.

A service normally receives traffic. Then it stops.

Sometimes what stopped happening tells you more than another threshold crossing.

Why doesn't the first alert always reveal the root cause?

Because alerts detect conditions, not causality.

A downstream service can cross a threshold because an upstream dependency failed. Automated remediation can hide a recurring failure. And sometimes the most useful clue doesn't trigger an alert at all.

The first alert is often an entry point into the incident, not a diagnosis.

That's why I like separating three things:

The question
Observable Could this signal help us understand what happened?
Alertable Does someone need to investigate this?
Pageable Does someone need to act right now?

Those aren't the same bar.

You want plenty of useful telemetry.

You want fewer alerts.

And you want even fewer things waking someone up at 2 a.m.

What should you check after an alert fires?

Before assuming the alerting service is the root cause, ask:

  1. Impact: Are users actually affected?
  2. Correlation: What else changed at the same time?
  3. Dependencies: What do the affected systems share?
  4. Recovery: Is automation repeatedly hiding the symptom?
  5. Absence: What normally happens that suddenly stopped?

The alert doesn't need to know the root cause.

It needs to tell you that something deserves your attention.

Your telemetry needs to help you figure out why.

So the next time an alert fires, don't just ask:

What triggered this?

Ask:

Is this where the problem started, or just where it became visible?


These incidents came from our panel From Alert Noise to Actionable Signals: Lessons from Production, with Sarah Kaplan, Heather Osborn, and Barnadeep Bhowmik.

What's the most misleading alert you've chased in production?

Top comments (0)