<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gayathri Reddy Lenkala</title>
    <description>The latest articles on DEV Community by Gayathri Reddy Lenkala (@gayathri_reddylenkala_3b).</description>
    <link>https://dev.to/gayathri_reddylenkala_3b</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150623%2F6151040f-eb6c-4618-b313-00b6a9f769d2.png</url>
      <title>DEV Community: Gayathri Reddy Lenkala</title>
      <link>https://dev.to/gayathri_reddylenkala_3b</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gayathri_reddylenkala_3b"/>
    <language>en</language>
    <item>
      <title>When My Incident Bot Found the Same Failure Twice</title>
      <dc:creator>Gayathri Reddy Lenkala</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:04:58 +0000</pubDate>
      <link>https://dev.to/gayathri_reddylenkala_3b/when-my-incident-bot-found-the-same-failure-twice-4kbb</link>
      <guid>https://dev.to/gayathri_reddylenkala_3b/when-my-incident-bot-found-the-same-failure-twice-4kbb</guid>
      <description>&lt;p&gt;An incident rarely arrives with a message saying, "You've seen me before."&lt;/p&gt;

&lt;p&gt;It arrives as a symptom.&lt;/p&gt;

&lt;p&gt;A 502.&lt;/p&gt;

&lt;p&gt;A growing queue.&lt;/p&gt;

&lt;p&gt;A sudden latency spike.&lt;/p&gt;

&lt;p&gt;A service that worked yesterday and doesn't work today.&lt;/p&gt;

&lt;p&gt;I built RETRACE around a simple idea: before the incident assistant starts inventing explanations, let it look at what happened the last time.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Incidents leave clues behind
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;I started with structured incident records.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
{&lt;br&gt;
  "id": "INC-1042",&lt;br&gt;
  "service": "checkout-api",&lt;br&gt;
  "severity": "SEV-1",&lt;br&gt;
  "symptoms": "POST /checkout returns 502",&lt;br&gt;
  "root_cause": "Database connection pool exhausted",&lt;br&gt;
  "fix": "Reduced retry fan-out and increased the pool",&lt;br&gt;
  "lesson": "Check DB pool saturation during traffic spikes"&lt;br&gt;
}&lt;br&gt;
That record already contains something valuable.&lt;/p&gt;

&lt;p&gt;It tells us:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what broke&lt;/li&gt;
&lt;li&gt;where it broke&lt;/li&gt;
&lt;li&gt;why it broke&lt;/li&gt;
&lt;li&gt;how it was fixed&lt;/li&gt;
&lt;li&gt;what we learned&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The problem is making that information available when another incident arrives weeks later.&lt;/p&gt;

&lt;p&gt;That's where Hindsight comes in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;## RETRACE has two kinds of context&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The assistant has the &lt;strong&gt;current incident&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Hindsight provides access to &lt;strong&gt;historical experience&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So instead of treating every request independently, the workflow becomes:&lt;br&gt;
              Current incident&lt;br&gt;
                     │&lt;br&gt;
                     ▼&lt;br&gt;
               RETRACE agent&lt;br&gt;
                     │&lt;br&gt;
                     ▼&lt;br&gt;
              Search memory&lt;br&gt;
                     │&lt;br&gt;
          ┌──────────┴──────────┐&lt;br&gt;
          │                     │&lt;br&gt;
     Similar incident       No useful match&lt;br&gt;
          │                     │&lt;br&gt;
          ▼                     ▼&lt;br&gt;
   Historical evidence      Normal investigation&lt;br&gt;
          │&lt;br&gt;
          └──────────┬──────────┘&lt;br&gt;
                     ▼&lt;br&gt;
                 Response&lt;br&gt;
Hindsight's memory model is designed around storing information and retrieving relevant memories later.&lt;/p&gt;

&lt;p&gt;That made it a natural fit for incident history.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure pattern I cared about
&lt;/h2&gt;

&lt;p&gt;**&lt;br&gt;
One of the most interesting examples in the data involves checkout-api.&lt;/p&gt;

&lt;p&gt;There are two separate incidents:&lt;br&gt;
INC-1042&lt;br&gt;
SEV-1&lt;br&gt;
checkout-api&lt;br&gt;
502 responses&lt;br&gt;
traffic spike&lt;br&gt;
DB connection pool exhaustion&lt;br&gt;
retry storm&lt;/p&gt;

&lt;p&gt;and :&lt;/p&gt;

&lt;p&gt;INC-1098&lt;br&gt;
SEV-2&lt;br&gt;
checkout-api&lt;br&gt;
intermittent 502s&lt;br&gt;
retry loop&lt;br&gt;
DB connection pool exhaustion&lt;/p&gt;

&lt;p&gt;They aren't duplicates.&lt;/p&gt;

&lt;p&gt;But they're related.&lt;/p&gt;

&lt;p&gt;Both tell us that retries can amplify a checkout failure until the database connection pool becomes the limiting resource.&lt;/p&gt;

&lt;p&gt;That relationship is much more useful than simply searching for the exact phrase "502".&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I didn't want a giant incident prompt
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;One tempting approach would have been to put every incident into the system prompt.&lt;/p&gt;

&lt;p&gt;Something like:&lt;/p&gt;

&lt;p&gt;Here are all 500 incidents we have ever experienced...&lt;/p&gt;

&lt;p&gt;INC-1...&lt;br&gt;
INC-2...&lt;br&gt;
INC-3...&lt;br&gt;
...&lt;/p&gt;

&lt;p&gt;Then ask the model to figure out which ones matter.&lt;/p&gt;

&lt;p&gt;That doesn't scale well conceptually.&lt;/p&gt;

&lt;p&gt;It also makes the current question compete with a huge amount of historical information.&lt;/p&gt;

&lt;p&gt;Instead, I wanted the memory layer to answer a narrower question:&lt;br&gt;
What previous incidents are relevant to THIS problem?&lt;/p&gt;

&lt;p&gt;That is the job I gave Hindsight.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  The investigation becomes more specific
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;Consider this query:&lt;/p&gt;

&lt;p&gt;"Checkout is returning 502s again. Traffic just increased."&lt;/p&gt;

&lt;p&gt;A generic assistant might say:&lt;br&gt;
Check:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;application logs&lt;/li&gt;
&lt;li&gt;load balancer&lt;/li&gt;
&lt;li&gt;database&lt;/li&gt;
&lt;li&gt;network&lt;/li&gt;
&lt;li&gt;dependencies&lt;/li&gt;
&lt;li&gt;CPU&lt;/li&gt;
&lt;li&gt;memory&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;RETRACE has another option.&lt;/p&gt;

&lt;p&gt;It can find previous checkout incidents.&lt;/p&gt;

&lt;p&gt;Then the response can focus on the evidence:&lt;br&gt;
A similar checkout incident occurred previously.&lt;/p&gt;

&lt;p&gt;The previous failure involved:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a traffic spike&lt;/li&gt;
&lt;li&gt;retry amplification&lt;/li&gt;
&lt;li&gt;database connection exhaustion&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Check the DB connection pool and retry behavior first.&lt;br&gt;
That doesn't prove the current incident has the same root cause.&lt;/p&gt;

&lt;p&gt;And that's important.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Historical memory should generate a useful hypothesis, not pretend to be proof.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  This changed how I thought about agent memory
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;Before building RETRACE, I mostly thought about memory as something that helps an agent remember users or conversations.&lt;/p&gt;

&lt;p&gt;Incident response made the use case feel different.&lt;/p&gt;

&lt;p&gt;An engineering agent can remember:&lt;/p&gt;

&lt;p&gt;What happened?&lt;br&gt;
Why did it happen?&lt;br&gt;
What fixed it?&lt;br&gt;
What should we check next time?&lt;/p&gt;

&lt;p&gt;That's operational memory.&lt;/p&gt;

&lt;p&gt;Hindsight provides the infrastructure for retaining and recalling those memories rather than forcing the entire history into every model request.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqglhnwkaf7zqgjsskq9k.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqglhnwkaf7zqgjsskq9k.jpeg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzy16dncwq5uebm48m26v.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzy16dncwq5uebm48m26v.jpeg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
The structured incident records are the source material.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8jcwhzhzu63ukff5wlpl.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8jcwhzhzu63ukff5wlpl.jpeg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
The assistant turns the current symptoms into a question about previous experience.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftsdiy2obrgy3vwuuqza9.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftsdiy2obrgy3vwuuqza9.jpeg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
The memory layer brings back the relevant incident.&lt;br&gt;
The assistant can then combine the current symptoms with the historical evidence.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture is deliberately small
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;The application doesn't need a complicated collection of services to demonstrate the core idea.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fige5kfbz3ze34g0z2q5l.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fige5kfbz3ze34g0z2q5l.jpeg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  The lesson hidden in every incident report
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;The most useful field in the incident data may actually be the lesson.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;When checkout traffic spikes,&lt;br&gt;
check DB pool saturation before changing application logic.&lt;/p&gt;

&lt;p&gt;That's different from the root cause.&lt;/p&gt;

&lt;p&gt;The root cause describes what happened.&lt;br&gt;
The lesson describes what someone should remember next time.&lt;/p&gt;

&lt;p&gt;That's exactly the distinction that makes an incident history valuable to an agent.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would keep if I rebuilt it
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;There are a few things I'd preserve.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep incident history structured&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A useful incident should contain more than a paragraph.&lt;/p&gt;

&lt;p&gt;The combination of:&lt;br&gt;
service&lt;br&gt;
severity&lt;br&gt;
symptoms&lt;br&gt;
root cause&lt;br&gt;
fix&lt;br&gt;
lesson&lt;/p&gt;

&lt;p&gt;makes the memory much more actionable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep memory separate from reasoning&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The LLM should reason about the incident.&lt;/p&gt;

&lt;p&gt;The memory layer should help it find relevant experience.&lt;/p&gt;

&lt;p&gt;Keeping those responsibilities separate makes the architecture easier to understand and debug.&lt;/p&gt;

&lt;p&gt;*&lt;em&gt;Treat historical matches as evidence&lt;br&gt;
*&lt;/em&gt;&lt;br&gt;
A previous incident is not a diagnosis.&lt;/p&gt;

&lt;p&gt;If two incidents look similar, that's a reason to investigate the same failure mode—not a reason to declare that the current incident has the same root cause.&lt;br&gt;
That distinction matters in production.&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I found most useful
&lt;/h2&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;p&gt;The interesting result wasn't that the bot could answer questions about incidents.&lt;/p&gt;

&lt;p&gt;A normal LLM can do that.&lt;/p&gt;

&lt;p&gt;The interesting part was being able to ask:&lt;/p&gt;

&lt;p&gt;"Have we dealt with this before?"&lt;/p&gt;

&lt;p&gt;and have the answer come from the system's accumulated operational history.&lt;/p&gt;

&lt;p&gt;That's what changed RETRACE from a troubleshooting chatbot into something closer to an incident assistant with institutional memory.&lt;/p&gt;

&lt;p&gt;The goal is straightforward:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When production breaks again, I don't want the investigation to start from zero.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;**&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;p&gt;**&lt;br&gt;
&lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt;&lt;br&gt;
&lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;What is agent memory? (Vectorize)&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>fastapi</category>
      <category>python</category>
      <category>hindsight</category>
    </item>
  </channel>
</rss>
