<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: M SAI CHARAN REDDY</title>
    <description>The latest articles on DEV Community by M SAI CHARAN REDDY (@reddysaicharan985).</description>
    <link>https://dev.to/reddysaicharan985</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146162%2F9d9f4ef6-2cfa-4d15-94c5-c80c484e17ad.jpg</url>
      <title>DEV Community: M SAI CHARAN REDDY</title>
      <link>https://dev.to/reddysaicharan985</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/reddysaicharan985"/>
    <language>en</language>
    <item>
      <title>How Hindsight Made Incident Advice Specific</title>
      <dc:creator>M SAI CHARAN REDDY</dc:creator>
      <pubDate>Mon, 28 Sep 2026 01:51:00 +0000</pubDate>
      <link>https://dev.to/reddysaicharan985/how-hindsight-made-incident-advice-specific-447b</link>
      <guid>https://dev.to/reddysaicharan985/how-hindsight-made-incident-advice-specific-447b</guid>
      <description>&lt;p&gt;Most incident assistants can generate a checklist. The harder problem is getting them to remember what actually fixed the last outage without pretending that history has already proved today's root cause.&lt;/p&gt;

&lt;p&gt;I built &lt;a href="https://incident-memory-agent-saicharan.streamlit.app" rel="noopener noreferrer"&gt;Incident Memory Agent&lt;/a&gt; to explore that problem. It is a Streamlit application for engineering and DevOps teams that combines persistent incident memory from &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; with reasoning from Groq.&lt;/p&gt;

&lt;p&gt;The result is not an automatic remediation system. It is an investigation assistant that recalls relevant operational experience, shows the evidence it retrieved, and keeps uncertainty visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the application does
&lt;/h2&gt;

&lt;p&gt;The workflow is intentionally narrow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;An engineer describes the currently known incident facts.&lt;/li&gt;
&lt;li&gt;Hindsight recalls semantically related historical incidents.&lt;/li&gt;
&lt;li&gt;Groq produces a baseline investigation using only the current facts.&lt;/li&gt;
&lt;li&gt;Groq produces a second investigation using the same facts plus the recalled memories.&lt;/li&gt;
&lt;li&gt;The interface displays both answers and the memory candidates side by side.&lt;/li&gt;
&lt;li&gt;After the real cause is confirmed, an engineer records the resolution so it can help during a future incident.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I describe this as &lt;strong&gt;Recall, Reason, and Learn&lt;/strong&gt;. Hindsight handles persistent memory, Groq handles language-model reasoning, and the engineer remains responsible for verification.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Current incident] --&amp;gt; B[Hindsight recall]
    A --&amp;gt; C[Baseline reasoning]
    B --&amp;gt; D[Memory-aware reasoning]
    C --&amp;gt; E[Side-by-side comparison]
    D --&amp;gt; E
    E --&amp;gt; F[Engineer verifies evidence]
    F --&amp;gt; G[Record confirmed resolution]
    G --&amp;gt; H[Hindsight retain]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;This separation matters. Memory should improve the investigation path, but it should not silently become the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure mode I wanted to avoid
&lt;/h2&gt;

&lt;p&gt;Suppose the checkout service starts returning HTTP 500 errors immediately after a deployment. A previous checkout incident had the same symptom and was caused by a missing &lt;code&gt;PAYMENT_API_URL&lt;/code&gt; variable.&lt;/p&gt;

&lt;p&gt;A careless agent might respond: "The environment variable is missing again."&lt;/p&gt;

&lt;p&gt;That is a plausible guess, not a confirmed fact. The current failure could also come from a database migration, dependency outage, invalid secret, code regression, or deployment problem.&lt;/p&gt;

&lt;p&gt;The useful response is more precise:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Confirm that HTTP 500 errors began after the deployment.&lt;/li&gt;
&lt;li&gt;Mark the current root cause as unknown.&lt;/li&gt;
&lt;li&gt;Present the previous missing variable as historical evidence.&lt;/li&gt;
&lt;li&gt;Recommend checking environment configuration early.&lt;/li&gt;
&lt;li&gt;Require logs, metrics, configuration, and deployment data before concluding.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This evidence-versus-proof boundary became the central design decision in the project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Connecting Hindsight
&lt;/h2&gt;

&lt;p&gt;The application reads its service configuration from environment variables. Secrets remain outside the repository.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;load_dotenv&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;hindsight_url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_API_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;hindsight_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;bank_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HINDSIGHT_BANK_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;hindsight&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Hindsight&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;hindsight_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;hindsight_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For every investigation, the current incident description becomes the recall query:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memory_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hindsight&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;current_incident&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;recalled_memories&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;memory_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I used one Hindsight bank named &lt;code&gt;incident-memory-agent&lt;/code&gt;. It contains structured incident narratives describing the service, symptom, impact, cause, resolution, prevention step, and date. I also seeded eight realistic synthetic incidents so the retrieval behavior can be tested across different services.&lt;/p&gt;

&lt;p&gt;The value of semantic recall became obvious quickly. The current description does not need to exactly match the wording of the stored incident. A notification failure can still retrieve a past SMTP credential incident because the operational meaning is similar.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt; was useful for understanding the retain and recall workflow. The broader &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;Vectorize explanation of agent memory&lt;/a&gt; also helped clarify why memory should be treated as a separate system rather than an ever-growing prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Comparing the same incident twice
&lt;/h2&gt;

&lt;p&gt;The most useful interface decision was generating two answers.&lt;/p&gt;

&lt;p&gt;The baseline request receives only the current incident. The memory-aware request receives the same incident plus the recalled memories. Everything else stays as similar as possible.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;groq&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openai/gpt-oss-120b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Use historical memories as evidence, not as proof. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Never invent facts, logs, causes or completed actions.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_completion_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The prompt forces the model to separate its response into confirmed current facts, unknown information, relevant historical evidence, investigation steps, and an uncertainty warning.&lt;/p&gt;

&lt;p&gt;That structure is more important than making the answer sound confident. During an outage, a confidently invented log line is worse than a short list of clearly marked unknowns.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Screenshot:&lt;/strong&gt; New Investigation showing the baseline and Hindsight responses side by side.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A concrete learning-loop test
&lt;/h2&gt;

&lt;p&gt;I tested the complete loop with an invoice-service incident.&lt;/p&gt;

&lt;p&gt;First, I recorded this confirmed resolution:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Symptom: PDF invoices stopped being generated.&lt;/li&gt;
&lt;li&gt;Impact: Customers could not download invoices for 25 minutes.&lt;/li&gt;
&lt;li&gt;Root cause: &lt;code&gt;INVOICE_TEMPLATE_PATH&lt;/code&gt; was missing.&lt;/li&gt;
&lt;li&gt;Resolution: The variable was restored and the service restarted.&lt;/li&gt;
&lt;li&gt;Prevention: Required-variable validation was added.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Record Resolution page retained this structured information in Hindsight. The write action is protected by an admin password so public visitors cannot casually add data to the shared bank.&lt;/p&gt;

&lt;p&gt;Next, I submitted a new current incident:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The invoice service stopped generating PDF invoices after today's deployment. The root cause has not yet been confirmed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The baseline response suggested general checks such as reviewing logs and deployment changes. The Hindsight response retrieved the earlier invoice incident and prioritized validating &lt;code&gt;INVOICE_TEMPLATE_PATH&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Crucially, it still described that variable as a historical clue. It did not claim that the present root cause had been confirmed.&lt;/p&gt;

&lt;p&gt;That single interaction demonstrates the behavior I wanted: the system learned from a confirmed resolution, recalled it later, and used it to improve the investigation without replacing verification.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Screenshot:&lt;/strong&gt; Retrieved Memory Candidates showing the stored invoice-service resolution.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Memory must be visible
&lt;/h3&gt;

&lt;p&gt;If users cannot see what the agent recalled, they cannot judge whether the recommendation is grounded or irrelevant. Showing the retrieved candidates makes the reasoning easier to inspect.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Current facts and historical evidence need separate blocks
&lt;/h3&gt;

&lt;p&gt;Combining them into one prompt paragraph encourages the model to blur the boundary. Explicit labels produced clearer responses and made unsupported claims easier to notice.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Only confirmed outcomes should become durable memory
&lt;/h3&gt;

&lt;p&gt;Early incident hypotheses are often wrong. Saving them as truth would make future investigations worse. The learning workflow therefore asks for the confirmed cause, resolution, and prevention step after the incident is understood.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. A before-and-after comparison explains memory better than a feature list
&lt;/h3&gt;

&lt;p&gt;Saying that an agent has persistent memory is abstract. Showing a generic baseline beside a historically informed response makes the difference concrete.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Retrieval quality depends on memory quality
&lt;/h3&gt;

&lt;p&gt;Duplicated, vague, or incomplete incident records create noisy recall results. Stable document identifiers, structured incident summaries, and careful confirmation are important as the bank grows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Current limitations
&lt;/h2&gt;

&lt;p&gt;The application does not yet connect directly to Kubernetes, Grafana, Datadog, cloud logs, or ticketing systems. Engineers currently enter the incident facts manually. The seeded incidents are synthetic, and the admin password is a lightweight write-protection mechanism rather than full user authentication.&lt;/p&gt;

&lt;p&gt;Those limitations keep the system focused on the memory workflow. A production version should add identity-based access, memory provenance, duplicate control, automated tests, observability, and connectors for live operational evidence.&lt;/p&gt;

&lt;p&gt;The safety rule should remain unchanged: memory can decide what to investigate first, but only current evidence can confirm what is happening now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try the project
&lt;/h2&gt;

&lt;p&gt;You can explore the &lt;a href="https://incident-memory-agent-saicharan.streamlit.app" rel="noopener noreferrer"&gt;live Incident Memory Agent&lt;/a&gt;, review the &lt;a href="https://github.com/reddysaicharan985/incident-memory-agent" rel="noopener noreferrer"&gt;source code on GitHub&lt;/a&gt;, or watch the &lt;a href="https://youtu.be/nkVSZ7bMUVk" rel="noopener noreferrer"&gt;complete video walkthrough&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Building this changed how I think about incident assistants. The model was never the missing piece. The missing piece was a memory layer with a disciplined boundary between experience and evidence. Hindsight made the investigation more specific; the uncertainty rules kept it honest.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fptt5wrbupl6fsqjpzlvu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fptt5wrbupl6fsqjpzlvu.png" alt=" " width="800" height="436"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk7l5hujqpo3batac6hqt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk7l5hujqpo3batac6hqt.png" alt=" " width="800" height="434"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>devops</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
