<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Praneeth Chelpuri</title>
    <description>The latest articles on DEV Community by Praneeth Chelpuri (@praneeth_chelpuri_6ee9968).</description>
    <link>https://dev.to/praneeth_chelpuri_6ee9968</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147431%2Fe5a179db-cfdf-4f1e-a50f-6c7f3a04d414.png</url>
      <title>DEV Community: Praneeth Chelpuri</title>
      <link>https://dev.to/praneeth_chelpuri_6ee9968</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/praneeth_chelpuri_6ee9968"/>
    <language>en</language>
    <item>
      <title>How Hindsight Turned Incident Resolutions Into Searchable Memory</title>
      <dc:creator>Praneeth Chelpuri</dc:creator>
      <pubDate>Mon, 28 Sep 2026 15:32:01 +0000</pubDate>
      <link>https://dev.to/praneeth_chelpuri_6ee9968/how-hindsight-turned-incident-resolutions-into-searchable-memory-3bl6</link>
      <guid>https://dev.to/praneeth_chelpuri_6ee9968/how-hindsight-turned-incident-resolutions-into-searchable-memory-3bl6</guid>
      <description>&lt;p&gt;An incident report is useful while the team is fixing the problem. It becomes much more valuable when the next investigation can find it, understand what was tried, and see whether that action worked. That is the loop I built into IncidentMind: an incoming incident prompts a Hindsight recall, recalled experience informs Groq’s investigation, an engineer records the resolution, and IncidentMind retains that outcome in Hindsight for future recall.&lt;/p&gt;

&lt;p&gt;The system does not ask a model to remember incidents across unrelated requests. It stores incident records in a Hindsight memory bank and retrieves relevant history when a new investigation arrives.&lt;/p&gt;

&lt;h2&gt;
  
  
  One incident, one memory loop
&lt;/h2&gt;

&lt;p&gt;IncidentMind is a Python application with a FastAPI API and a browser interface in &lt;code&gt;app/static/index.html&lt;/code&gt;. The &lt;code&gt;/investigate&lt;/code&gt; endpoint accepts an incident ID, service, severity, description, and optional report time, then passes those fields to &lt;code&gt;investigate_incident&lt;/code&gt; in &lt;code&gt;app/services/agent_service.py&lt;/code&gt;. That function builds context, calls Hindsight, prepares a prompt, and requests Groq analysis.&lt;/p&gt;

&lt;p&gt;The engineer still owns the resolution. The &lt;code&gt;/resolve&lt;/code&gt; endpoint accepts incident details, the attempted resolution, a success flag, and an optional engineer note. &lt;code&gt;record_resolution&lt;/code&gt; formats these as a learning record and retains it in Hindsight for a later investigation.&lt;/p&gt;

&lt;p&gt;The flow is simple to describe:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;incident → Hindsight recall → historical experience → Groq reasoning → engineer resolution → Hindsight retain → future recall&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Retrieval and learning are distinct operations: during investigation, the system asks what prior material might be relevant; after the engineer records an outcome, it stores that experience. Hindsight is the persistent memory layer between investigations, while Groq reasons over the current incident and recalled context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I put Hindsight in the middle
&lt;/h2&gt;

&lt;p&gt;Incident response narrows hypotheses from incomplete symptoms. A latency spike after a release could come from a query, a connection leak, or another change. A previous case can point the engineer toward useful checks.&lt;/p&gt;

&lt;p&gt;IncidentMind uses &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; to make those records available to later requests. In &lt;code&gt;app/services/agent_service.py&lt;/code&gt;, the Hindsight client reads its URL and API key from &lt;code&gt;HINDSIGHT_API_URL&lt;/code&gt; and &lt;code&gt;HINDSIGHT_API_KEY&lt;/code&gt;; Groq reads &lt;code&gt;GROQ_API_KEY&lt;/code&gt;. The runtime service sets &lt;code&gt;BANK_ID = "incidentmind_final"&lt;/code&gt; and &lt;code&gt;MODEL = "openai/gpt-oss-120b"&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The Hindsight &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;documentation&lt;/a&gt; describes the client and memory API. IncidentMind uses its &lt;code&gt;recall&lt;/code&gt; and &lt;code&gt;retain&lt;/code&gt; operations as parts of one workflow. This is the broader idea of &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;agent memory from Vectorize&lt;/a&gt;: information from earlier interactions can inform a later one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recall starts with the full incident
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;investigate_incident&lt;/code&gt; formats the current event into &lt;code&gt;incident_context&lt;/code&gt;, including the incident ID, service, severity, report time, and description. It queries with all of that context rather than the description alone.&lt;/p&gt;

&lt;p&gt;The call asks Hindsight for up to 4,000 tokens with a &lt;code&gt;mid&lt;/code&gt; budget:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memory_result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hindsight&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;incident_context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;budget&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result texts become &lt;code&gt;memory_context&lt;/code&gt;, separated by blank lines. With no usable results, the prompt includes &lt;code&gt;No relevant historical incidents found.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9k96ktkh3jhgqnaaw17p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9k96ktkh3jhgqnaaw17p.png" alt=" " width="800" height="441"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Deduplication is a small boundary check
&lt;/h2&gt;

&lt;p&gt;I added a short pass between recall and prompt construction. It strips surrounding whitespace, skips empty strings, and preserves the first copy of each exact text in a list, using a set to track what has already appeared:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memories&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="n"&gt;seen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;memory_result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;memories&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Identical returned records add no distinct evidence and consume context. This is exact-text deduplication after trimming whitespace; it will not detect paraphrases or determine whether similar records describe the same event. The filter removes obvious repeats, not memory consolidation.&lt;/p&gt;

&lt;p&gt;The investigation function also returns the filtered list and its length as &lt;code&gt;memory_count&lt;/code&gt;, so the interface can show the historical texts and count beside the analysis.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9nkn4okx52zlpxfaz2gs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9nkn4okx52zlpxfaz2gs.png" alt=" " width="532" height="742"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Groq reasons over history, not instead of it
&lt;/h2&gt;

&lt;p&gt;The prompt contains the current incident and unique recalled memories. It requests six sections: incident summary, relevant historical incidents, possible root causes, recommended investigation and resolution, and why historical memory matters.&lt;/p&gt;

&lt;p&gt;The prompt also sets useful boundaries. It says to distinguish facts from hypotheses, not to claim a root cause without evidence, not to invent logs or metrics, and to treat historical incidents as evidence rather than proof. Then the configured Groq model receives the prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;groq&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a careful production incident response engineer.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_completion_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1500&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Hindsight provides potentially relevant history; the prompt tells Groq to use it cautiously. A past fix can guide investigation, but the current incident still needs current evidence. The function returns the analysis alongside incident fields, memories, and count.&lt;/p&gt;

&lt;h2&gt;
  
  
  The engineer closes the loop
&lt;/h2&gt;

&lt;p&gt;The loop is not complete when Groq writes an investigation. It closes when an engineer records what they tried and whether it worked. &lt;code&gt;record_resolution&lt;/code&gt; creates a structured text record with the incident ID, service, severity, report time, description, resolution attempt, outcome, engineer note, and a general learning statement. If no note is supplied, it records &lt;code&gt;No additional note provided.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The status is converted to &lt;code&gt;SUCCESSFUL&lt;/code&gt; or &lt;code&gt;UNSUCCESSFUL&lt;/code&gt;, and the resulting record is sent to the same Hindsight bank:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;hindsight&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Recording unsuccessful attempts means history can describe outcomes, not only fixes. The code retains both statuses but adds no recall ranking by success; the prompt only says to prefer successful resolutions when evidence supports them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1bi4cc3d9nlgw7ss5hmk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1bi4cc3d9nlgw7ss5hmk.png" alt=" " width="800" height="322"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Before and after: a representative seeded example
&lt;/h2&gt;

&lt;p&gt;The seed scripts contain representative test incidents, not verified production history. &lt;code&gt;seed_final.py&lt;/code&gt; includes &lt;code&gt;INC-2041&lt;/code&gt;, a critical &lt;code&gt;payments-api&lt;/code&gt; case with 503 errors after deployment, attributed to a connection leak exhausting the pool. The recorded successful resolution is rollback and a fix to connection handling. &lt;code&gt;INC-2315&lt;/code&gt; describes payment latency and intermittent 503s attributed to excessive scans from a new query, with query optimization and an index as the resolution.&lt;/p&gt;

&lt;p&gt;Before memory is available, a report of payment latency and 503s reaches Groq without historical context. Once the seed records are retained in the configured bank, a similar report can cause Hindsight to return examples. IncidentMind filters exact duplicates and asks Groq to suggest checks: connection-pool metrics and query plans, with both causes treated as hypotheses to validate.&lt;/p&gt;

&lt;p&gt;If the engineer finds a different cause, &lt;code&gt;/resolve&lt;/code&gt; retains it too. Later recall can include it alongside the examples. This describes the code path and representative seed data, not a real diagnosis or measured improvement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons and limitations
&lt;/h2&gt;

&lt;p&gt;I took a few practical lessons from building this loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Make memory flow visible.&lt;/strong&gt; The code keeps recall, prompt assembly, reasoning, and retain as separate steps, which makes it easier to inspect what context influenced an analysis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Preserve outcomes with the resolution.&lt;/strong&gt; A fix without its success status or engineer note is less useful to a future investigation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use past incidents as leads.&lt;/strong&gt; Similar symptoms can suggest checks, but they cannot establish today’s root cause.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep duplicate handling conservative.&lt;/strong&gt; Exact text equality after trimming is predictable. More aggressive matching could merge distinct incidents, so it would need evidence and tests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Be explicit when memory is absent.&lt;/strong&gt; A clear fallback lets the reasoning prompt distinguish an empty recall from an omitted section.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The repository has no automated coverage for recall ranking, duplicate handling, or prompt contents. Deduplication misses semantic redundancy, and seed/manual scripts are demo scaffolding, not production evidence. The service has no explicit recovery for Hindsight or Groq failures; &lt;code&gt;/health&lt;/code&gt; returns configured labels rather than testing live connectivity. The runtime service uses &lt;code&gt;incidentmind_final&lt;/code&gt;, while seed scripts also contain other bank IDs, including &lt;code&gt;incidentmind_demo&lt;/code&gt; and &lt;code&gt;incidentmind&lt;/code&gt; in the manual test script. These are repository and test-data configuration details; the example data must use the runtime bank ID to be available to the service.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;IncidentMind turns an engineer’s resolution into searchable memory by retaining a structured outcome in Hindsight, then recalling history for a later investigation. Groq reasons over that context alongside the current incident, with instructions to distinguish evidence from confirmed facts. Exact-text deduplication is a supporting safeguard that removes empty and repeated strings before the model sees them.&lt;/p&gt;

&lt;p&gt;This is a memory workflow with an engineer in the loop, not an autonomous responder or a measured production improvement: retain what happened, recall it when relevant, and keep the evidence visible enough to challenge.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>devops</category>
      <category>hindsight</category>
    </item>
  </channel>
</rss>
