DEV Community

Anand Kumar
Anand Kumar

Posted on

Incident Deep Dip

I asked Hindsight what keeps breaking — it named the deploy

By Anand

Most incident tools are reactive. Something breaks, you investigate, you fix it, you write a postmortem nobody reads, and then the same class of thing breaks again three weeks later. I got tired of that loop, so when I built IncidentDeepDip I gave it one job beyond responding to incidents: notice what keeps happening. The moment the agent had durable memory of past incidents, it could answer a question no single incident ever can — what is our recurring failure, and what triggers it? The answer, in our case, was blunt: the deploy.

This is a write-up about turning incident memory into pattern discovery, using Hindsight as the memory layer and a small amount of aggregation on top.

From recall to patterns

The core agent already used Hindsight to recall incidents similar to a new one — that's the reactive path. But once incidents live in durable memory as structured records, each carrying a service, a root-cause category, a severity, and whether a deployment preceded it, you can stop asking "what's like this one?" and start asking "what's the shape of everything we've seen?"

The pattern discovery is deliberately simple — it groups incidents by root-cause category and measures how many followed a deployment:

for category, incs in by_category.items():
    count = len(incs)
    deploy_related = sum(1 for i in incs if i.get("deployment"))
    deploy_pct = round(100 * deploy_related / count)

    if deploy_pct >= 60:
        insight = (f"{deploy_pct}% of '{category}' incidents occurred shortly "
                   f"after a deployment — strongly deployment-correlated.")
Enter fullscreen mode Exit fullscreen mode

No machine learning, no clustering model. Just honest counting over memory that Hindsight makes durable and queryable. The intelligence isn't in the algorithm; it's in the fact that the incidents are remembered at all, in a structured form, instead of evaporating into closed tickets. That's the argument Vectorize makes about agent memory — memory is the substrate that makes higher-order reasoning possible — and pattern discovery is a concrete example of it.

What it found

When I ran this over our incident history, the output wasn't subtle:

  • 63% of all incidents followed a deployment. Not a hunch — a computed share across every stored incident.
  • One category dominated: Kafka consumer rebalance, nine occurrences, 100% deployment-correlated. Every single time that class of outage happened, a deploy had just gone out.
  • The rest of the pattern list ranked the next-most-common failure families and their own deploy correlation, so I could see the whole landscape at once.

I want to be precise here: I did not type "63%" anywhere. It's derived from the stored incidents at request time. If we resolve more incidents and retain them, the number updates itself. The tool's opinion about what keeps breaking is always a reflection of what actually broke — which is exactly the property you want and almost never get from a static runbook.

The before/after that made it click

The reactive agent already showed a strong before/after: with memory off, a new payment incident gets a generic "could be a few things" and a confidence around 35%; with Hindsight memory on, it recalls three near-identical past incidents, names the consumer-rebalance root cause, and confidence rises to about 85%.

Pattern discovery extends that from a single incident to the whole system. Before, "we seem to have a lot of Kafka issues" was a feeling someone voiced in a retro. After, it's a ranked, quantified pattern with a deploy-correlation percentage attached — and it points at a specific, testable hypothesis: our rolling deployments are triggering consumer-group rebalances. That's not a vibe you argue about. That's a checklist item you add before the next release.

The honest limitation

I'll be direct about what this is and isn't. Correlation is not causation, and my deploy-correlation metric is exactly that — a correlation. A category being "100% deployment-correlated" means every stored incident of that type had a recent deploy, not that the deploy provably caused it. The tool is good at pointing a human at the right hypothesis fast; it is not a root-cause oracle.

I also had to resist the urge to over-engineer this. My first instinct was to reach for clustering and embeddings to "discover" categories automatically. I stopped, because the honest version — group by the category we already record, count deploy correlation — was more trustworthy and infinitely more explainable. When the tool says "63% followed a deploy," I can point at the exact incidents behind that number. An engineer will believe a number they can audit; they won't act on a black-box "risk score." Keeping it countable was the right call.

Where this goes

The natural next step writes itself: if the memory knows that a failure class is deployment-correlated, the agent can move from response to prevention — flag a risky deploy before it ships, not after it pages someone. The foundation is already there, because every resolved incident gets retained back into memory:

await hindsight.aretain(bank_id=BANK_ID, content=content,
                        context="Incident resolution feedback")
Enter fullscreen mode Exit fullscreen mode

Every confirmed resolution makes the next pattern sharper. The system gets more useful the longer it runs, without anyone retraining anything.

Lessons I'd reuse

  • Durable, structured memory turns retros into queries. The reason I could compute recurring patterns at all is that Hindsight kept incidents around as real records instead of letting them die in a ticket tracker.
  • Count before you cluster. The explainable, auditable metric beat the fancy one. Engineers act on numbers they can trace.
  • Compute insights from memory at request time. Don't cache a "risk score." Derive it from what's actually stored, so it's always current.
  • Be honest about correlation. Point humans at hypotheses; don't claim causation you can't back.
  • Response is step one; prevention is the payoff. Once memory knows the pattern, warning before a deploy is a small step, not a new system.

The single most valuable thing the tool ever told me wasn't how to fix an incident. It was which incident we were going to have again — and that our own deploys were the trigger. Memory made that visible. Nothing else did.

Top comments (0)