<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mandadi Vennela Naga Sai</title>
    <description>The latest articles on DEV Community by Mandadi Vennela Naga Sai (@mandadi_vennelanagasai_).</description>
    <link>https://dev.to/mandadi_vennelanagasai_</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150495%2Fa53c0f54-542a-449d-9bc9-e439c622f27e.png</url>
      <title>DEV Community: Mandadi Vennela Naga Sai</title>
      <link>https://dev.to/mandadi_vennelanagasai_</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mandadi_vennelanagasai_"/>
    <language>en</language>
    <item>
      <title>I Gave My SRE Agent a Memory With Hindsight</title>
      <dc:creator>Mandadi Vennela Naga Sai</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:07:02 +0000</pubDate>
      <link>https://dev.to/mandadi_vennelanagasai_/i-gave-my-sre-agent-a-memory-with-hindsight-dkb</link>
      <guid>https://dev.to/mandadi_vennelanagasai_/i-gave-my-sre-agent-a-memory-with-hindsight-dkb</guid>
      <description>&lt;p&gt;Most incident-response agents can analyze an outage.&lt;/p&gt;

&lt;p&gt;The harder question is: what happened the last time we tried that fix?&lt;/p&gt;

&lt;p&gt;While building IncidentIQ, I wanted my incident agent to do more than analyze the current logs. I wanted it to remember previous incidents, understand which fixes worked, which failed, and which were only temporary, and use that experience when recommending what to do next.&lt;/p&gt;

&lt;p&gt;That became the reason I added Hindsight as the persistent memory layer.&lt;/p&gt;

&lt;p&gt;Instead of starting every incident from zero, IncidentIQ can look at what happened before and turn those experiences into evidence for the next decision.&lt;/p&gt;

&lt;p&gt;The Problem With Starting Every Incident From Zero&lt;/p&gt;

&lt;p&gt;Imagine a payments-api suddenly starts returning HTTP 503 errors.&lt;/p&gt;

&lt;p&gt;The incident might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Service: payments-api
Severity: CRITICAL

Alert:
HTTP 503 surge

Logs:
Database connection pool exhausted.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;An LLM can look at those logs and suggest several possible fixes.&lt;/p&gt;

&lt;p&gt;Restart the service.&lt;/p&gt;

&lt;p&gt;Increase the connection pool.&lt;/p&gt;

&lt;p&gt;Check database latency.&lt;/p&gt;

&lt;p&gt;Inspect long-running queries.&lt;/p&gt;

&lt;p&gt;But the LLM doesn't automatically know what happened during previous incidents in the same environment.&lt;/p&gt;

&lt;p&gt;What if restarting the service worked once but only provided temporary relief?&lt;/p&gt;

&lt;p&gt;What if increasing the database connection pool had already solved the same problem twice?&lt;/p&gt;

&lt;p&gt;What if connection-timeout monitoring had previously prevented the same issue from recurring?&lt;/p&gt;

&lt;p&gt;That history can change the recommendation.&lt;/p&gt;

&lt;p&gt;This is the problem I wanted IncidentIQ to solve.&lt;/p&gt;

&lt;p&gt;What I Built&lt;/p&gt;

&lt;p&gt;IncidentIQ is an SRE incident-response system that combines a React and TypeScript frontend, a FastAPI backend, Hindsight for persistent memory, and Groq for AI reasoning.&lt;/p&gt;

&lt;p&gt;I wanted memory to participate in the investigation itself, rather than being something separate from the incident workflow.&lt;/p&gt;

&lt;p&gt;The basic architecture is:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;React Frontend
      |
      v
FastAPI Backend
      |
      +----------&amp;gt; Hindsight
      |            Persistent Memory
      |
      +----------&amp;gt; Groq
                   AI Reasoning
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The frontend never talks directly to Hindsight or Groq.&lt;/p&gt;

&lt;p&gt;The backend controls the memory and reasoning pipeline.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F209uiigpgwj0w21drwi8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F209uiigpgwj0w21drwi8.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The main incident workflow is:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Current Incident
       |
       v
Hindsight Recall
       |
       v
Previous Incidents
       |
       v
Previous Actions + Outcomes
       |
       v
LLM Reasoning
       |
       v
Recommendation
       |
       v
Engineer Action
       |
       v
New Outcome
       |
       v
Hindsight Retain
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This creates a feedback loop instead of a one-time AI interaction.&lt;/p&gt;

&lt;p&gt;Starting an Incident Investigation&lt;/p&gt;

&lt;p&gt;The first step is describing the incident.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Service: payments-api
Severity: CRITICAL

Alert:
HTTP 503 surge

Logs:
Database connection pool exhausted.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;IncidentIQ's investigation workspace allows the engineer to provide the service, severity, alert, and logs.&lt;/p&gt;

&lt;p&gt;It also provides quick presets for common operational problems such as:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DB connection pool
Redis failure
Auth failure
Certificate expiry
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Once the incident is submitted, the system doesn't immediately ask the LLM for an answer.&lt;/p&gt;

&lt;p&gt;It first asks a more useful question:&lt;/p&gt;

&lt;p&gt;Have we seen something like this before?&lt;/p&gt;

&lt;p&gt;Giving the Agent a Memory&lt;/p&gt;

&lt;p&gt;This is where Hindsight comes in.&lt;/p&gt;

&lt;p&gt;IncidentIQ has a dedicated Hindsight service responsible for connecting the application to the Hindsight memory bank.&lt;/p&gt;

&lt;p&gt;The actual implementation is small:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dotenv&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;load_dotenv&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;hindsight_client&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Hindsight&lt;/span&gt;


&lt;span class="nf"&gt;load_dotenv&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;HINDSIGHT_API_KEY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;HINDSIGHT_BASE_URL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_BASE_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.hindsight.vectorize.io&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;HINDSIGHT_BANK_ID&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_BANK_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incidentiq-sre&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Hindsight&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;HINDSIGHT_BASE_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;HINDSIGHT_API_KEY&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;store_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;HINDSIGHT_BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;search_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;HINDSIGHT_BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;There are two operations that matter here:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;recall()
    |
    v
Retrieve previous operational experience

retain()
    |
    v
Store new experience for future incidents
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;You can explore Hindsight on GitHub and read the Hindsight documentation.&lt;/p&gt;

&lt;p&gt;Recalling Relevant Incidents&lt;/p&gt;

&lt;p&gt;When memory is enabled, IncidentIQ builds a query from the current incident:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
Service: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Severity: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;severity&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Alert: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;alert&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Logs: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;logs&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="n"&gt;hindsight_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;search_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;query&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The recalled memories are then filtered according to the affected service before being used for analysis.&lt;/p&gt;

&lt;p&gt;The idea is straightforward:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Current Incident
       |
       v
Hindsight Recall
       |
       v
Relevant Memories
       |
       v
Historical Facts
       |
       v
Recommendation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This is where the agent starts behaving differently from a stateless incident assistant.&lt;/p&gt;

&lt;p&gt;The Interesting Part: Remembering What Worked&lt;/p&gt;

&lt;p&gt;Retrieving an old incident is useful.&lt;/p&gt;

&lt;p&gt;Retrieving the outcome of the previous solution is much more useful.&lt;/p&gt;

&lt;p&gt;Suppose the historical memory contains:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident: INC-001

Action:
restart service

Outcome:
temporary
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Another memory contains:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident: INC-001

Action:
increase database connection pool size

Outcome:
successful
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;And another contains:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident: INC-001

Action:
add connection timeout monitoring

Outcome:
successful
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Now the agent has more than historical context.&lt;/p&gt;

&lt;p&gt;It has evidence about what actually happened.&lt;/p&gt;

&lt;p&gt;The interface makes this pattern visible by connecting previous actions with their outcomes.&lt;/p&gt;

&lt;p&gt;For the payments-api example, the system can identify that restarting the service was not the same as permanently resolving the underlying connection-pool problem.&lt;/p&gt;

&lt;p&gt;Turning Memory Into Evidence&lt;/p&gt;

&lt;p&gt;This was one of the most important design decisions in IncidentIQ.&lt;/p&gt;

&lt;p&gt;I didn't want the LLM to decide historical success rates on its own.&lt;/p&gt;

&lt;p&gt;The backend extracts historical facts from the recalled memories and calculates action statistics.&lt;/p&gt;

&lt;p&gt;The project explicitly separates successful, failed, and temporary outcomes:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;fact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;outcome&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;successful&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;successful_incidents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;incident_id&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;fact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;outcome&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failed_incidents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;incident_id&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;fact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;outcome&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;temporary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;temporary_incidents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;incident_id&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The system then calculates the historical success rate from those explicit outcomes:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;attempts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="n"&gt;successful&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;successful_incidents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;success_rate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;successful&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;success_rate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This means the application owns the historical evidence.&lt;/p&gt;

&lt;p&gt;The LLM reasons over that evidence.&lt;/p&gt;

&lt;p&gt;That separation matters because a generated explanation should not become the source of truth for historical statistics.&lt;/p&gt;

&lt;p&gt;From Historical Evidence to a Recommendation&lt;/p&gt;

&lt;p&gt;After Hindsight memories have been recalled and processed, IncidentIQ passes the current incident and historical memories to Groq.&lt;/p&gt;

&lt;p&gt;The reasoning prompt explicitly tells the model to stay within the evidence:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
You are an SRE incident response assistant.

Analyze the current incident using the historical
Hindsight memories provided below.

IMPORTANT:
- Base your reasoning on the evidence.
- Do not invent historical incidents.
- Do not invent evidence IDs.
- Recommend practical actions an on-call engineer can take.
- Confidence must reflect how strongly the evidence
  supports the diagnosis.
- Historical success rates should only be estimated
  from explicit outcome information in the memories.
- If there is insufficient historical evidence, say so.

CURRENT INCIDENT:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

HISTORICAL HINDSIGHT MEMORIES:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;memories&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The recommendation is also represented using a structured model:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Recommendation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;rationale&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;historical_success_rate&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;default&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;ge&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;le&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;drift_detected&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="n"&gt;evidence_ids&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This gives the frontend structured information rather than one large block of generated text.&lt;/p&gt;

&lt;p&gt;The Recommendation&lt;/p&gt;

&lt;p&gt;For the connection-pool incident, the resulting recommendation can be:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Increase the database connection pool size
(e.g., from 50 to 100) and redeploy the
payments-api service.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The recommendation is displayed together with its historical effectiveness, confidence, and evidence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcnov0xlmpg7n3ct0vm72.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcnov0xlmpg7n3ct0vm72.png" alt=" " width="800" height="424"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For example, the interface shows:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Historical Effectiveness:
100%

2 Successful / 2 Recorded

Confidence:
95%

Evidence:
INC-001
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This changes the interaction from:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI:
"Try increasing the connection pool."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Historical Evidence:
2 successful / 2 recorded

Recommendation:
Increase the database connection pool.

Evidence:
INC-001
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The engineer can inspect the evidence and then decide what action makes sense.&lt;/p&gt;

&lt;p&gt;What Changes With Memory?&lt;/p&gt;

&lt;p&gt;The simplest way to understand the role of Hindsight is to compare the two workflows.&lt;/p&gt;

&lt;p&gt;Without Memory&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident
   |
   v
LLM
   |
   v
Possible Causes
   |
   v
Generic Recommendation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;With Hindsight&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident
   |
   v
Hindsight Recall
   |
   v
Previous Incidents
   |
   v
Previous Actions
   |
   v
Previous Outcomes
   |
   v
LLM Reasoning
   |
   v
Evidence-backed Recommendation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The LLM itself hasn't magically become better at debugging.&lt;/p&gt;

&lt;p&gt;The context available to it has changed.&lt;/p&gt;

&lt;p&gt;Memory doesn't replace reasoning. It gives reasoning something more useful to work with.&lt;/p&gt;

&lt;p&gt;Using the Same Memory Before Deployment&lt;/p&gt;

&lt;p&gt;I also wanted the memory layer to be useful before an incident happens.&lt;/p&gt;

&lt;p&gt;IncidentIQ therefore includes a pre-deployment risk review.&lt;/p&gt;

&lt;p&gt;An engineer can enter a proposed change:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Service:
payments-api

Proposed Change:
Increase connection pool 50 → 100
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;[INSERT SCREENSHOT — Pre-Deploy Risk Review]&lt;/p&gt;

&lt;p&gt;The system evaluates the proposed change using historical operational context.&lt;/p&gt;

&lt;p&gt;The result can include a risk classification, risk score, rationale, safeguards, and related historical incidents.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqr66n9wf9ukukf52tt10.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqr66n9wf9ukukf52tt10.png" alt=" " width="799" height="453"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4lkr3lhefoawclkblxo4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4lkr3lhefoawclkblxo4.png" alt=" " width="799" height="452"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This means the same memory layer can support two different moments in the software lifecycle:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Before Deployment
       |
       v
Historical Experience
       |
       v
Risk Assessment


During Incident
       |
       v
Historical Experience
       |
       v
Incident Recommendation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The memory isn't tied to only one screen or one workflow.&lt;/p&gt;

&lt;p&gt;Exploring What the Agent Remembers&lt;/p&gt;

&lt;p&gt;I didn't want the memory layer to be invisible.&lt;/p&gt;

&lt;p&gt;IncidentIQ includes a Memory Explorer where engineers can inspect the operational memories available to the system.&lt;/p&gt;

&lt;p&gt;For example, the memory view can contain information such as:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident INC-001 was mitigated by restarting
the payments-api service and permanently resolved
by increasing the database connection pool size
and adding connection timeout monitoring.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg4n0sazcwd6sboq00gqi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg4n0sazcwd6sboq00gqi.png" alt=" " width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The memory explorer is useful for another reason: it makes the agent itself easier to inspect.&lt;/p&gt;

&lt;p&gt;If a recommendation seems unexpected, an engineer can look at the memories that were available to the system.&lt;/p&gt;

&lt;p&gt;Memory becomes something that can be investigated instead of a hidden part of the prompt.&lt;/p&gt;

&lt;p&gt;Learning From What Actually Happened&lt;/p&gt;

&lt;p&gt;A recommendation shouldn't automatically become a permanent truth just because an AI generated it.&lt;/p&gt;

&lt;p&gt;The engineer still needs to record what happened.&lt;/p&gt;

&lt;p&gt;IncidentIQ accepts an outcome and turns it into a new memory.&lt;/p&gt;

&lt;p&gt;The backend builds an outcome memory like this:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;outcome_memory&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
Incident &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; affected the
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; service.

Resolution attempt:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Outcome:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;outcome_description&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Engineer notes:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;notes&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Interpretation:
The engineer marked this resolution attempt as
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;outcome_description&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;It then stores that experience:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;store_memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;outcome_memory&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;So the loop becomes:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident
   |
   v
Recall
   |
   v
Historical Evidence
   |
   v
Recommendation
   |
   v
Engineer Action
   |
   v
Real Outcome
   |
   v
Retain
   |
   v
Future Incident
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This is the part of IncidentIQ I find most interesting.&lt;/p&gt;

&lt;p&gt;The agent isn't simply generating answers.&lt;/p&gt;

&lt;p&gt;It has a mechanism for accumulating operational experience.&lt;/p&gt;

&lt;p&gt;When There Isn't Enough Evidence&lt;/p&gt;

&lt;p&gt;IncidentIQ also includes a Fix Drift view.&lt;/p&gt;

&lt;p&gt;The idea is that a fix that worked previously should not automatically be considered reliable forever.&lt;/p&gt;

&lt;p&gt;As more outcomes are recorded, the system can compare newer results with historical results.&lt;/p&gt;

&lt;p&gt;But there is an important limitation.&lt;/p&gt;

&lt;p&gt;The current interface explicitly reports when there aren't enough recorded outcomes to detect fix drift.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fly511zsxkde8j5d22ys7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fly511zsxkde8j5d22ys7.png" alt=" " width="800" height="445"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F98n2ypw51tffiavo2ogb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F98n2ypw51tffiavo2ogb.png" alt=" " width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I prefer this behavior to manufacturing a conclusion.&lt;/p&gt;

&lt;p&gt;If the system doesn't have enough evidence, the answer should be:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Not enough recorded outcomes to detect fix drift.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That is much more useful than pretending the system knows something it doesn't.&lt;/p&gt;

&lt;p&gt;What I Learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Memory Is More Useful When It Includes Outcomes&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Remembering that an incident happened is useful.&lt;/p&gt;

&lt;p&gt;Remembering what was tried and whether it worked is much more useful.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Retrieval Should Be Visible&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Historical context shouldn't disappear inside an LLM prompt.&lt;/p&gt;

&lt;p&gt;Making memories and evidence visible gives engineers a way to inspect why a recommendation was produced.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The LLM Shouldn't Own the Historical Facts&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The application extracts historical outcomes and calculates explicit statistics.&lt;/p&gt;

&lt;p&gt;The LLM reasons over those facts.&lt;/p&gt;

&lt;p&gt;This separation makes the system easier to inspect.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Insufficient Evidence Is Still Information&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If there aren't enough recorded outcomes to detect a pattern, the system should say so.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The Feedback Loop Is the Real Memory&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The interesting architecture isn't:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident → AI → Answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;It's:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident
   |
   v
Memory
   |
   v
Recommendation
   |
   v
Action
   |
   v
Outcome
   |
   v
Memory
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That is what turns previous incidents into experience that can influence future incidents.&lt;/p&gt;

&lt;p&gt;Final Thoughts&lt;/p&gt;

&lt;p&gt;IncidentIQ started from a simple observation:&lt;/p&gt;

&lt;p&gt;SRE teams repeatedly solve problems that have similar histories.&lt;/p&gt;

&lt;p&gt;An LLM can reason about the incident in front of it.&lt;/p&gt;

&lt;p&gt;But persistent memory gives that reasoning access to what happened before.&lt;/p&gt;

&lt;p&gt;That changes the question from:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"What should I do?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"What happened the last time this occurred,
what actually worked,
and what evidence do we have?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That's what I wanted to explore with IncidentIQ: an SRE agent that doesn't just respond to incidents, but can build a persistent memory of operational experience.&lt;/p&gt;

&lt;p&gt;The complete source code for IncidentIQ is available on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Pravallika2789" rel="noopener noreferrer"&gt;
        Pravallika2789
      &lt;/a&gt; / &lt;a href="https://github.com/Pravallika2789/IncidentIQ-SRE-Agent" rel="noopener noreferrer"&gt;
        IncidentIQ-SRE-Agent
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      AI-powered SRE incident response agent that uses Hindsight memory to learn from past incidents, successful and failed fixes, and deployment history.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;IncidentIQ - Intelligent SRE Incident Response&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;IncidentIQ is an intelligent SRE incident-response agent that uses Hindsight persistent memory to help engineers investigate production incidents, recall previous operational experience, learn from successful and failed resolution attempts, and make safer deployment decisions.&lt;/p&gt;

&lt;p&gt;It combines a React frontend, FastAPI backend, Hindsight for persistent operational memory, and Groq-powered AI reasoning to turn past incidents into reusable engineering knowledge.&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Team Code &amp;amp; Chaos&lt;/h2&gt;
&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;1. Reddy Gayatri Satya Sai Pravallika&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Mandha Varshitha&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Chenna Keerthana&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Mandadi Vennela&lt;/strong&gt;&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Video Presentation&lt;/h2&gt;
&lt;/div&gt;

&lt;p&gt;&lt;a href="https://youtu.be/OdgS6RPzdKA?si=XYwCszUG2_4m06cA" rel="nofollow noopener noreferrer"&gt;Watch the IncidentIQ Demo&lt;/a&gt;&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Problem&lt;/h2&gt;

&lt;/div&gt;

&lt;p&gt;When a production incident happens, engineers need more than an analysis of the current logs.&lt;/p&gt;

&lt;p&gt;They need to know:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Has this happened before?&lt;/li&gt;
&lt;li&gt;What caused the previous incident?&lt;/li&gt;
&lt;li&gt;Which fix actually worked?&lt;/li&gt;
&lt;li&gt;Which approaches failed?&lt;/li&gt;
&lt;li&gt;Which runbook was useful?&lt;/li&gt;
&lt;li&gt;What did the team learn from previous incidents?&lt;/li&gt;
&lt;li&gt;Has a previously successful fix become less effective?&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Solution&lt;/h2&gt;

&lt;/div&gt;

&lt;p&gt;IncidentIQ turns previous incidents…&lt;/p&gt;&lt;/div&gt;


&lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Pravallika2789/IncidentIQ-SRE-Agent" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


</description>
      <category>ai</category>
      <category>agents</category>
      <category>devops</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
