<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: S.Sahasra</title>
    <description>The latest articles on DEV Community by S.Sahasra (@ssahasra_344feac7913891a).</description>
    <link>https://dev.to/ssahasra_344feac7913891a</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147810%2Fedb73e7c-e954-4bec-8e8b-19e1a4c805f4.png</url>
      <title>DEV Community: S.Sahasra</title>
      <link>https://dev.to/ssahasra_344feac7913891a</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ssahasra_344feac7913891a"/>
    <language>en</language>
    <item>
      <title>Building an Agent That Learns from Every Interaction with Hindsight</title>
      <dc:creator>S.Sahasra</dc:creator>
      <pubDate>Mon, 28 Sep 2026 18:37:29 +0000</pubDate>
      <link>https://dev.to/ssahasra_344feac7913891a/building-an-agent-that-learns-from-every-interaction-with-hindsight-3b5h</link>
      <guid>https://dev.to/ssahasra_344feac7913891a/building-an-agent-that-learns-from-every-interaction-with-hindsight-3b5h</guid>
      <description>&lt;h1&gt;
  
  
  Building an Agent That Learns from Every Interaction with Hindsight
&lt;/h1&gt;

&lt;p&gt;Picture the 3 a.m. version of this. An alert fires, you open your incident response tool, and the assistant says: "This looks like INC-214: connection pool exhaustion, fixed by raising max_connections." You go looking for INC-214 in your issue tracker. It doesn't exist.&lt;/p&gt;

&lt;p&gt;That failure mode is why I built MemoryOps. To be clear, INC-214 is a hypothetical example of an LLM hallucination, not a captured model response: seed data in this repository spans INC-101 through INC-116, so any citation of INC-214 is invented. During an outage, a fabricated citation is worse than a generic answer because a citation reads like verified evidence. You either burn minutes verifying it, or you trust it and apply a fix that was never tested.&lt;/p&gt;

&lt;p&gt;I reduced this risk by building persistent operational memory into the incident response lifecycle. Here is how the architecture and implementation work.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. What MemoryOps Does
&lt;/h2&gt;

&lt;p&gt;MemoryOps is an AI-powered incident response platform for DevOps and SRE teams. The frontend is built with React 19, Vite, and Tailwind CSS (providing Dashboard, Incident Creation, Investigation, and Memory Explorer views). The backend is FastAPI with SQLAlchemy over SQLite (data/incidentiq.db) as the system of record for live incident records.&lt;/p&gt;

&lt;p&gt;Hindsight acts as the long-term persistent memory layer for resolved incident learnings, while Groq Cloud LLM (openai/gpt-oss-20b) serves as the AI reasoning engine. Nothing executes actions autonomously; the human engineer remains in full control.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;React UI ─► FastAPI ─► SQLite     (incidents, ai_recommendation)
              ├──────► Hindsight  (RETAIN on resolve, RECALL on analyze, REFLECT)
              └──────► Groq LLM   (current incident + recalled memories)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcmi1nl290ehrb22eq13h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcmi1nl290ehrb22eq13h.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;MemoryOps system architecture — React frontend, FastAPI backend, SQLite database of record, Hindsight persistent memory layer, and Groq LLM reasoning engine.&lt;/p&gt;

&lt;p&gt;The workflow begins when an engineer declares an incident via POST /api/v1/incidents. Navigating to the investigation page invokes POST /api/v1/incidents/{incident_id}/analyze. The backend recalls relevant past incidents from Hindsight, passes them alongside current symptoms to Groq LLM, and presents evidence-backed recommendations. When the incident is resolved via POST /api/v1/incidents/{incident_id}/resolve, its incident learnings are retained in Hindsight.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Hindsight Memory Lifecycle: RETAIN, RECALL, and REFLECT
&lt;/h2&gt;

&lt;p&gt;Instead of passing massive unstructured log streams to an LLM, MemoryOps uses Hindsight to store structured experience documents. SQLite answers "what is happening now," while Hindsight answers "what did we learn from past outages."&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incident Created ──► Investigation ──► Root Cause &amp;amp; Resolution ──► Hindsight RETAIN ──► Future RECALL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  RETAIN Runs Only on Verified Resolution
&lt;/h3&gt;

&lt;p&gt;An incident is not retained when merely created, during unresolved investigation, or from AI guesses. When an engineer resolves an incident, HindsightService.aretain_incident() stores a complete experience document containing ID, service, error, symptoms, severity, root cause, resolution steps, and post-mortem.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;aretain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;content_text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;document_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;incident_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;severity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To support idempotent retention, document_id is set deterministically to incident.id (for example, INC-101). The Incident database model tracks a memory_retained boolean flag, which is flipped to True only after Hindsight confirms successful retention.&lt;/p&gt;

&lt;h3&gt;
  
  
  RECALL Runs During Incident Investigation
&lt;/h3&gt;

&lt;p&gt;When an investigation is triggered, MemoryOps constructs a semantic search query from the current incident:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;recall_query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Service: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; | Error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; | Symptoms: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;symptoms&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;recalled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;hindsight_service&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;arecall_memories&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;recall_query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2048&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The router parses returned memories using parse_memory_item() and supplies them as grounded context to Groq.&lt;/p&gt;

&lt;h3&gt;
  
  
  REFLECT Synthesizes Patterns
&lt;/h3&gt;

&lt;p&gt;MemoryOps also provides POST /api/v1/incidents/reflect and a Memory Explorer tab so engineers can query cross-incident patterns across historical outages.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8zccpldz0xhctq5kz3l0.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8zccpldz0xhctq5kz3l0.jpeg" alt=" " width="799" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Memory Explorer demonstrating direct Hindsight RECALL vector search and REFLECT pattern synthesis.*&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  3. Separating Evidence from Analysis
&lt;/h2&gt;

&lt;p&gt;In the investigation API response (IncidentInvestigationResponse), historical evidence and AI reasoning are kept strictly separate:&lt;/p&gt;

&lt;p&gt;similar_historical_incidents: populated directly from Hindsight recalled memories.&lt;/p&gt;

&lt;p&gt;ai_analysis: contains the structured JSON output returned by Groq LLM.&lt;br&gt;
The UI renders these inputs as distinct pipeline stages so the engineer can evaluate historical evidence independently from LLM reasoning.&lt;/p&gt;

&lt;p&gt;![MemoryOps investigation view]&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F032c1z7dmxy886amqddn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F032c1z7dmxy886amqddn.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;— MemoryOps investigation view displaying the step-by-step pipeline, explicit memory status banner, recalled historical memories, and Groq AI recommendation.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Explicit Memory Status: No Fake Fallback Memories
&lt;/h2&gt;

&lt;p&gt;MemoryOps explicitly exposes three distinct memory states in the API and UI:&lt;/p&gt;

&lt;p&gt;memory_status = "ok": Hindsight successfully recalled relevant memories (✓ Historical Memory Used).&lt;/p&gt;

&lt;p&gt;memory_status = "empty": Hindsight searched but found no matching memories (○ No Relevant Historical Memory).&lt;/p&gt;

&lt;p&gt;memory_status = "unavailable": Hindsight service was offline or unconfigured (⚠ Historical Memory Unavailable).&lt;/p&gt;

&lt;h2&gt;
  
  
  If Hindsight is unavailable, MemoryOps does not query SQLite incident records and label them as Hindsight memories. Local SQLite records remain system-of-record entries and are never disguised as vector-recalled memories.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  5. Before-and-After Example
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Before (Hypothetical LLM Hallucination)
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Probable cause:&lt;/em&gt; Connection pool exhaustion. Matches INC-214, resolved by increasing max_connections to 100.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Stored Memory (Real Seed Data for INC-101 in &lt;code&gt;seed.py&lt;/code&gt;)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Service:&lt;/strong&gt; Payment API&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Error:&lt;/strong&gt; Database connection timeout&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Symptoms:&lt;/strong&gt; High HTTP 504 Gateway Timeouts on /v1/charge endpoint, elevated API latency&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Root Cause:&lt;/strong&gt; Connection pool exhaustion due to leaked unclosed DB sessions during traffic surge&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resolution:&lt;/strong&gt; Increased connection pool size from 20 to 100 and deployed hotfix for session leak&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Post-mortem:&lt;/strong&gt; Connection pool configuration was insufficient for observed traffic surges. Added automated connection pool utilization alerting at 80% capacity.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  After (Illustrative Response Structure)
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"probable_root_cause"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Database connection pool exhaustion"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"recommended_action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Increase max_connections parameter from 20 to 100 and deploy session leak hotfix"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"high"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"reasoning"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"INC-101 historical incident showed identical Gateway Timeout symptoms and was resolved by expanding the pool."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"supporting_historical_incidents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"INC-101: Payment API database connection timeout"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;System prompt instructions direct Groq to cite only actual recalled memories. However, prompt instructions are guidance rather than mathematical guarantees. Application-side memory-status handling ensures that unavailable or empty Hindsight results are explicitly reported rather than silently represented as historical memory.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Failure Handling and Graceful Degradation
&lt;/h2&gt;

&lt;p&gt;When external services fail, MemoryOps degrades gracefully without crashing.&lt;/p&gt;

&lt;p&gt;If Hindsight returns an error or HTTP 402 insufficient credits, the backend sets memory_status = "unavailable", clears similar_historical_incidents, and continues investigation using current incident details alone. If Groq LLM is unconfigured, _fallback_analysis() applies rule-based heuristic analysis and sets analysis_status = "fallback" with confidence = "low".&lt;/p&gt;

&lt;p&gt;![MemoryOps graceful degradation]&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F032c1z7dmxy886amqddn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F032c1z7dmxy886amqddn.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;MemoryOps graceful degradation view displaying explicit memory status warning and low-confidence fallback heuristic reasoning when external services are unavailable.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Engineering Lessons and Limitations
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Persistent Memory vs Weight Fine-Tuning&lt;/strong&gt;: "Learning" in MemoryOps refers to persistent operational memory through Hindsight RAG, not altering LLM weights.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;System of Record vs Memory Bank&lt;/strong&gt;:SQLite records state ("what happened"), while Hindsight stores reusable operational experience ("what worked").&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explicit State Over Silent Fallbacks&lt;/strong&gt;: Disguising dependency failures with mock memories destroys user trust during active production outages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Idempotent Retention&lt;/strong&gt;: Using deterministic document IDs (document_id = incident.id) is intended to make retries idempotent and avoid duplicate retention, subject to Hindsight's handling of repeated document IDs.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The key lesson is that an incident-response agent does not need to change its model weights to learn from previous incidents. It can improve future investigations by retaining resolved operational experience, recalling relevant evidence when a new incident occurs, and clearly communicating when that memory layer is unavailable&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>softwareengineering</category>
    </item>
  </channel>
</rss>
