<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Anand Kumar</title>
    <description>The latest articles on DEV Community by Anand Kumar (@anand_kumar_fdbc93f9a8e8a).</description>
    <link>https://dev.to/anand_kumar_fdbc93f9a8e8a</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147992%2F397404fe-47b5-41cd-99e2-85270606491f.png</url>
      <title>DEV Community: Anand Kumar</title>
      <link>https://dev.to/anand_kumar_fdbc93f9a8e8a</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/anand_kumar_fdbc93f9a8e8a"/>
    <language>en</language>
    <item>
      <title>Incident Deep Dip</title>
      <dc:creator>Anand Kumar</dc:creator>
      <pubDate>Mon, 28 Sep 2026 19:58:45 +0000</pubDate>
      <link>https://dev.to/anand_kumar_fdbc93f9a8e8a/incident-deep-dip-4729</link>
      <guid>https://dev.to/anand_kumar_fdbc93f9a8e8a/incident-deep-dip-4729</guid>
      <description>&lt;h1&gt;
  
  
  I asked Hindsight what keeps breaking — it named the deploy
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;By Anand&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Most incident tools are reactive. Something breaks, you investigate, you fix it, you write a postmortem nobody reads, and then the same class of thing breaks again three weeks later. I got tired of that loop, so when I built &lt;strong&gt;IncidentDeepDip&lt;/strong&gt; I gave it one job beyond responding to incidents: notice what keeps happening. The moment the agent had durable memory of past incidents, it could answer a question no single incident ever can — &lt;em&gt;what is our recurring failure, and what triggers it?&lt;/em&gt; The answer, in our case, was blunt: the deploy.&lt;/p&gt;

&lt;p&gt;This is a write-up about turning incident memory into pattern discovery, using &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; as the memory layer and a small amount of aggregation on top.&lt;/p&gt;

&lt;h2&gt;
  
  
  From recall to patterns
&lt;/h2&gt;

&lt;p&gt;The core agent already used &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; to recall incidents similar to a new one — that's the reactive path. But once incidents live in durable memory as structured records, each carrying a service, a root-cause category, a severity, and whether a deployment preceded it, you can stop asking "what's like this one?" and start asking "what's the shape of everything we've seen?"&lt;/p&gt;

&lt;p&gt;The pattern discovery is deliberately simple — it groups incidents by root-cause category and measures how many followed a deployment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;incs&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;by_category&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;incs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;deploy_related&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;incs&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deployment&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;deploy_pct&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;deploy_related&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;deploy_pct&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;insight&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;deploy_pct&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;% of &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; incidents occurred shortly &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                   &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;after a deployment — strongly deployment-correlated.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No machine learning, no clustering model. Just honest counting over memory that Hindsight makes durable and queryable. The intelligence isn't in the algorithm; it's in the fact that the incidents are &lt;em&gt;remembered&lt;/em&gt; at all, in a structured form, instead of evaporating into closed tickets. That's the argument Vectorize makes about &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;agent memory&lt;/a&gt; — memory is the substrate that makes higher-order reasoning possible — and pattern discovery is a concrete example of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it found
&lt;/h2&gt;

&lt;p&gt;When I ran this over our incident history, the output wasn't subtle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;63% of all incidents followed a deployment.&lt;/strong&gt; Not a hunch — a computed share across every stored incident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One category dominated: Kafka consumer rebalance, nine occurrences, 100% deployment-correlated.&lt;/strong&gt; Every single time that class of outage happened, a deploy had just gone out.&lt;/li&gt;
&lt;li&gt;The rest of the pattern list ranked the next-most-common failure families and their own deploy correlation, so I could see the whole landscape at once.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I want to be precise here: I did not type "63%" anywhere. It's derived from the stored incidents at request time. If we resolve more incidents and retain them, the number updates itself. The tool's opinion about what keeps breaking is always a reflection of what actually broke — which is exactly the property you want and almost never get from a static runbook.&lt;/p&gt;

&lt;h2&gt;
  
  
  The before/after that made it click
&lt;/h2&gt;

&lt;p&gt;The reactive agent already showed a strong before/after: with memory off, a new payment incident gets a generic "could be a few things" and a confidence around 35%; with &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; memory on, it recalls three near-identical past incidents, names the consumer-rebalance root cause, and confidence rises to about 85%.&lt;/p&gt;

&lt;p&gt;Pattern discovery extends that from a single incident to the whole system. Before, "we seem to have a lot of Kafka issues" was a feeling someone voiced in a retro. After, it's a ranked, quantified pattern with a deploy-correlation percentage attached — and it points at a specific, testable hypothesis: our rolling deployments are triggering consumer-group rebalances. That's not a vibe you argue about. That's a checklist item you add before the next release.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest limitation
&lt;/h2&gt;

&lt;p&gt;I'll be direct about what this is and isn't. Correlation is not causation, and my deploy-correlation metric is exactly that — a correlation. A category being "100% deployment-correlated" means every stored incident of that type had a recent deploy, not that the deploy provably caused it. The tool is good at pointing a human at the right hypothesis fast; it is not a root-cause oracle.&lt;/p&gt;

&lt;p&gt;I also had to resist the urge to over-engineer this. My first instinct was to reach for clustering and embeddings to "discover" categories automatically. I stopped, because the honest version — group by the category we already record, count deploy correlation — was more trustworthy and infinitely more explainable. When the tool says "63% followed a deploy," I can point at the exact incidents behind that number. An engineer will believe a number they can audit; they won't act on a black-box "risk score." Keeping it countable was the right call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this goes
&lt;/h2&gt;

&lt;p&gt;The natural next step writes itself: if the memory knows that a failure class is deployment-correlated, the agent can move from &lt;em&gt;response&lt;/em&gt; to &lt;em&gt;prevention&lt;/em&gt; — flag a risky deploy before it ships, not after it pages someone. The foundation is already there, because every resolved incident gets retained back into memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;hindsight&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;aretain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                        &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Incident resolution feedback&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every confirmed resolution makes the next pattern sharper. The system gets more useful the longer it runs, without anyone retraining anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons I'd reuse
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Durable, structured memory turns retros into queries.&lt;/strong&gt; The reason I could compute recurring patterns at all is that &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; kept incidents around as real records instead of letting them die in a ticket tracker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count before you cluster.&lt;/strong&gt; The explainable, auditable metric beat the fancy one. Engineers act on numbers they can trace.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compute insights from memory at request time.&lt;/strong&gt; Don't cache a "risk score." Derive it from what's actually stored, so it's always current.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Be honest about correlation.&lt;/strong&gt; Point humans at hypotheses; don't claim causation you can't back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Response is step one; prevention is the payoff.&lt;/strong&gt; Once memory knows the pattern, warning before a deploy is a small step, not a new system.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The single most valuable thing the tool ever told me wasn't how to fix an incident. It was which incident we were going to have again — and that our own deploys were the trigger. Memory made that visible. Nothing else did.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
