I went looking for why my homelab kept telling me the same thing eight times.
The answer was structural, and it is the kind of thing that hides in a system for years because every individual piece of it is reasonable. AEGIS — my self-hosted agent platform — had no concept of a problem. It had notifications. One probe failed on eight consecutive days and wrote eight Todoist tasks, because the only thing it could ask was "did I already write a task about this?" — and it asked by looking at its own most recent log line, which by then said something else.
The Todoist task was the identity of the incident. And that is the bug.
This is the first of three posts about three lanes of AEGIS I rebuilt between 5 and 13 September. The other two — the money lane and the knowledge lane — landed on the same rule from completely different directions, which is the only reason I trust it.
What the measurement said
I don't rebuild anything on a hunch any more. Before touching the code I measured what was actually there on 2026-09-07:
-
Three alerting systems that could not see each other. The investigation flow (a
#alerttask plus a chat card), a set of watchdogs deduping against the audit log, and domain tables with their own private keys for certificate expiry and config drift. The resolved-awareNOT EXISTSdedupe SQL was copy-pasted in three files. -
Thirteen fingerprint or dedupe-key schemes, mutually incompatible. A vendor fingerprint. A synthesised
alertmanager:{alertname}:{instance}. Asentry:{issue_id}. Three separate signature classes. Three mute-key namespaces sharing one primary key. A day-bucketed drift key. Plus three informal links — aLIKE '%' || task_id || '%'against workflow ids, title substring matching, and a comment footer used as an authorship marker. -
Dedupe keyed on the task, not the problem. The dedupe index's
task_idcolumn wasNOT NULLand joined to the task table. One flow refused to use it at all and built a fourth ledger inside a settings row. Because the signature was the primary key, recreating a task resetfirst_seen_atand the occurrence count — recurrence history was destroyed on every recreate. Only 12 of 42 open alert tasks had a signature row at all. - One producer bypassed the capture path entirely, hand-building its command and deduping against the newest audit row. The close-on-resolve function structurally could not reach it. That is where the eight copies came from.
- No service state. No maintenance window, no deploy suppression. The GitHub webhook claimed deployment events and dropped them. Suppression existed only as incidental delays scattered across three files.
And the part that made it expensive to fix: of 72,499 source lines, the four files carrying this logic held 11,212 of them. The code doing all this was the code hardest to change safely.
The primitive: a problem is a row, a task is a projection
The replacement is deliberately small. Every operational signal — an Alertmanager webhook, a Sentry issue, a swarm heartbeat, a certificate about to expire, a stuck social post, a hand-written task — becomes an Event and goes through one function, ingest_event.
That function owns three things and nothing else:
-
Identity. One open
problemsrow percorrelation_key. If a row for that key is already open, this is another occurrence of the same problem, not a new one. -
History. Every occurrence, every resolution, every report is a
problem_eventsrow, idempotent on(source, external_id). Recurrence survives everything, including deleting the ticket. -
The decision. The return value carries
investigate: trueorfalse. The producer starts the investigation workflow only when told to. The flow itself never dedupes and never captures a task.
The Todoist task is created by a projector, from the problem. It is a projection, never the identity. The rule I wrote into the docs, because I knew I'd be tempted to break it: do not look a problem up by its task title or fingerprint.
This is the first part. The full post — including the rest of the working details — is on my site: Your Alerts Have No Identity
Top comments (0)