The signal problem
I was trying to understand how mature systems separate signal from noise when many alerts arrive at once. I had been working through a related question in TruthPipe, a project I built to represent uncertain data: once uncertainty is detected, how should the system deliver it?
Alertmanager gave me a mature system for examining part of that delivery question. So I followed one alert through the system.
Prometheus helped me understand when a condition becomes an alert. Alertmanager handled what happened afterward—how alerts were grouped, suppressed, routed, and eventually delivered.
The model I thought I understood
The model I was working from was already detailed. The README architecture diagram showed the API, the component that groups alerts and matches them to routes, the cluster, the path toward a notification, and a labeled box the diagram did not explain. That grouping component is the Dispatcher. The path toward a notification is the notification pipeline. The labeled box was Gossip Settle.
The diagram labels Gossip Settle and the major stores and notification components. It does not, by itself, verify their runtime order.
An incoming alert was stored in two places. One was the store that holds active alerts, the Alert Provider. The other was the separate store that holds silence rules, the Silence Provider. The Dispatcher subscribed to new alerts, formed a group, and held the group for a short wait. Deduplication, the check of whether a notification had already been sent, sat outside the Dispatcher, on each receiver.
After the group, the remaining work sat in one sequence: settle the cluster, suppress an alert that is only a symptom of another, apply silence rules, choose the route, wait, skip a notification already sent, retry, and send. The stage list that builds this pipeline, PipelineBuilder.New, lined up with that sequence in the account I was using. That account also said the list included two time checks the diagram had not shown.
Figure 2. The starting hypothesis. The two-store alert claim and the single-chain shape were the claims under test; this is deliberately not a diagram of actual behavior.
The source trace that broke the linear picture
An assistant told me an incoming alert is not written into the silence-rule store. A silence is a different object. The alert goes into the Alert Provider. The Silence Provider is consulted later, when notification processing decides whether a rule mutes the alert. I asked to record that split as likely until the write itself was checked.
The trace prompt asked for one alert from the HTTP API toward a receiver, with no code changes. It asked for ordinary calls to be distinguished from a channel handoff, where one part of the program passes work to another, and from a goroutine, where work continues on its own rather than through one straight sequence of calls. Timers were to be marked separately.
Cursor traced the source. The summary reached me before any instrumentation and before either runtime run. The summary said the alert-ingestion request writes only to the Alert Provider. A silence is created on its own request, through postSilencesHandler into Silences.Set, and read later.
The same summary split the single sequence. After the alert is stored, two parts of the system receive it. The Inhibitor is a background process that tracks whether one alert is only a symptom of another. It keeps that state continuously. The Dispatcher also receives the alert, and it matches a route before the group exists. A later step only selects the pipeline for a route that has already been chosen.
That summary put the group on a timer. Notification processing starts when the timer flushes. The insert puts the alert into the group and leaves it there until that flush. The wait at the start of that processing, the cluster settle from the diagram, and the later wait on each integration are different waits.
I asked what had happened, and why the model had not shown this. I thought the earlier hypothesis had included a static trace, and I had forgotten whether that earlier work had traced code.
What had to be instrumented
About six to ten logging points had been suggested, so the log would not be sprayed with a line at every function. After the source trace, I asked why that many were needed. The answer was that each important boundary needed its own mark, and the log still had to stay readable. Fewer marks were fine if they still separated the boundaries. More were fine if a gap between asynchronous steps would otherwise disappear.
Those boundaries were five questions. Where was the alert stored, and which consumers received it? Did the group wait before the pipeline began? Which stages ran, and which integration branch fired? Did deduplication allow the notification? Did the send happen, and was it recorded?
Cursor then produced a ten-point instrumentation map. The header said the files had already been modified. Each line carried the alert’s fingerprint, the identifying value used to connect its log lines, together with the group, the receiver, the integration, and the flush, so a later line could be tied to the same alert. The inner readiness check inside the cluster-settle stage was left unmarked, so the probe would not change that wait. I answered, “wait what happened.” The files had changed before I understood the map.
Cursor wrote the test file from a configuration I was given, and I pasted it back. The group wait was one second. The interval between later flushes of the same group was five seconds. The only receiver was one local webhook. Those durations came with the configuration. The short waits made the group timer and the repeated flush visible.
Delivery without understanding
The process came up after a first start failed during gossip initialization.
The first test used an alert named TraceProbeTest. A firing webhook arrived. The receiver accepted it, and the recorded reason was the first notification for that alert. The output I had did not show the stage order for that delivery. I asked where the trace stood.
The clean run
The next run was set up on fresh storage, with the log kept, so this alert would not share notification history with the one already delivered. I sent that alert under the name TraceProbeTestRun2. The order below is the correlated extract of that log.
The extract ties the lines to one alert and one flush by the fingerprint, the group, the receiver, the integration, and the flush. The alert is stored in the Alert Provider. The Inhibitor and the Dispatcher both receive it. The Dispatcher matches a route and inserts the group. About one second later the group flushes, and notification processing starts on that flush. The cluster settles, the suppression checks run, and the webhook branch is taken. That branch waits, deduplication allows the first notification, the webhook is called, and the send is recorded.
About five seconds later the group flushes again. The same checks run. Deduplication refuses. No send is attempted, and nothing new is recorded.
**
22:51:59.345 aggregation_group_flush flush_id=1 alert_fingerprints=[70ec33155ecbb17b]
22:51:59.347 dedup_decision flush_id=1 notify=true reason="first notification"
22:51:59.365 integration_notify_success flush_id=1 integration=webhook[0]
22:51:59.367 nflog_write flush_id=1 firing_count=1
22:52:04.353 aggregation_group_flush flush_id=2 alert_fingerprints=[70ec33155ecbb17b]
22:52:04.354 dedup_decision flush_id=2 notify=false reason=none
**
Excerpt. From trace-run.log, Run 2 (TraceProbeTestRun2, fingerprint 70ec33155ecbb17b). The lines above condense the raw log's metadata for reading. figures/trace-log-excerpt.txt contains the six original lines without alteration. There is no webhook attempt or notification-log write for flush 2 in that log. The later resolved notification at flush 61 is outside this essay's firing path.
The model I started from had already placed deduplication outside the Dispatcher. This run showed that check as the gate on whether the send and the record still run.
The corrected model and what changed
Figure 3. The corrected model, based on the Alertmanager source excerpts and Run 2. The two consumers receive the stored alert independently. The group timer separates insertion from pipeline execution. The single webhook branch reaches Dedup before a send and notification-log write.
A silence stays a separate write, on its own request. Later that day the group’s own store was named store.Alerts. It holds the alerts in that group. It is separate from the Alert Provider. Naming it did not move the probes.
This was one webhook on one machine. The lines start when the alert is stored.
OpenBB had taught me to follow function calls once I knew which code was running. Alertmanager made me mark the places where that sequence broke: the handoff to a background consumer, the group timer, and the branch to the webhook. Carrying the same alert identity across those marks let me account for both the notification that went out and the next flush that stopped.



Top comments (0)