I asked Hindsight what keeps breaking — it named the deploy
By Anand
Most incident tools are reactive. Something breaks, you investigate, you fix it, you write a postmortem nobody reads, and then the same class of thing breaks again three weeks later. I got tired of that loop, so when I built IncidentDeepDip I gave it one job beyond responding to incidents: notice what keeps happening. The moment the agent had durable memory of past incidents, it could answer a question no single incident ever can — what is our recurring failure, and what triggers it? The answer, in our case, was blunt: the deploy.
This is a write-up about turning incident memory into pattern discovery, using Hindsight as the memory layer and a small amount of aggregation on top.
From recall to patterns
The core agent already used Hindsight to recall incidents similar to a new one — that's the reactive path. But once incidents live in durable memory as structured records, each carrying a service, a root-cause category, a severity, and whether a deployment preceded it, you can stop asking "what's like this one?" and start asking "what's the shape of everything we've seen?"
The pattern discovery is deliberately simple — it groups incidents by root-cause category and measures how many followed a deployment:
for category, incs in by_category.items():
count = len(incs)
deploy_related = sum(1 for i in incs if i.get("deployment"))
deploy_pct = round(100 * deploy_related / count)
if deploy_pct >= 60:
insight = (f"{deploy_pct}% of '{category}' incidents occurred shortly "
f"after a deployment — strongly deployment-correlated.")
No machine learning, no clustering model. Just honest counting over memory that Hindsight makes durable and queryable. The intelligence isn't in the algorithm; it's in the fact that the incidents are remembered at all, in a structured form, instead of evaporating into closed tickets. That's the argument Vectorize makes about agent memory — memory is the substrate that makes higher-order reasoning possible — and pattern discovery is a concrete example of it.
What it found
When I ran this over our incident history, the output wasn't subtle:
- 63% of all incidents followed a deployment. Not a hunch — a computed share across every stored incident.
- One category dominated: Kafka consumer rebalance, nine occurrences, 100% deployment-correlated. Every single time that class of outage happened, a deploy had just gone out.
- The rest of the pattern list ranked the next-most-common failure families and their own deploy correlation, so I could see the whole landscape at once.
I want to be precise here: I did not type "63%" anywhere. It's derived from the stored incidents at request time. If we resolve more incidents and retain them, the number updates itself. The tool's opinion about what keeps breaking is always a reflection of what actually broke — which is exactly the property you want and almost never get from a static runbook.
The before/after that made it click
The reactive agent already showed a strong before/after: with memory off, a new payment incident gets a generic "could be a few things" and a confidence around 35%; with Hindsight memory on, it recalls three near-identical past incidents, names the consumer-rebalance root cause, and confidence rises to about 85%.
Pattern discovery extends that from a single incident to the whole system. Before, "we seem to have a lot of Kafka issues" was a feeling someone voiced in a retro. After, it's a ranked, quantified pattern with a deploy-correlation percentage attached — and it points at a specific, testable hypothesis: our rolling deployments are triggering consumer-group rebalances. That's not a vibe you argue about. That's a checklist item you add before the next release.
The honest limitation
I'll be direct about what this is and isn't. Correlation is not causation, and my deploy-correlation metric is exactly that — a correlation. A category being "100% deployment-correlated" means every stored incident of that type had a recent deploy, not that the deploy provably caused it. The tool is good at pointing a human at the right hypothesis fast; it is not a root-cause oracle.
I also had to resist the urge to over-engineer this. My first instinct was to reach for clustering and embeddings to "discover" categories automatically. I stopped, because the honest version — group by the category we already record, count deploy correlation — was more trustworthy and infinitely more explainable. When the tool says "63% followed a deploy," I can point at the exact incidents behind that number. An engineer will believe a number they can audit; they won't act on a black-box "risk score." Keeping it countable was the right call.
Where this goes
The natural next step writes itself: if the memory knows that a failure class is deployment-correlated, the agent can move from response to prevention — flag a risky deploy before it ships, not after it pages someone. The foundation is already there, because every resolved incident gets retained back into memory:
await hindsight.aretain(bank_id=BANK_ID, content=content,
context="Incident resolution feedback")
Every confirmed resolution makes the next pattern sharper. The system gets more useful the longer it runs, without anyone retraining anything.
Lessons I'd reuse
- Durable, structured memory turns retros into queries. The reason I could compute recurring patterns at all is that Hindsight kept incidents around as real records instead of letting them die in a ticket tracker.
- Count before you cluster. The explainable, auditable metric beat the fancy one. Engineers act on numbers they can trace.
- Compute insights from memory at request time. Don't cache a "risk score." Derive it from what's actually stored, so it's always current.
- Be honest about correlation. Point humans at hypotheses; don't claim causation you can't back.
- Response is step one; prevention is the payoff. Once memory knows the pattern, warning before a deploy is a small step, not a new system.
The single most valuable thing the tool ever told me wasn't how to fix an incident. It was which incident we were going to have again — and that our own deploys were the trigger. Memory made that visible. Nothing else did.
Top comments (0)