<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Devalla Harika</title>
    <description>The latest articles on DEV Community by Devalla Harika (@devalla_harika_ae9342db1a).</description>
    <link>https://dev.to/devalla_harika_ae9342db1a</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146970%2F1419fc71-3bcc-4aac-bb3e-81cf807327f3.png</url>
      <title>DEV Community: Devalla Harika</title>
      <link>https://dev.to/devalla_harika_ae9342db1a</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/devalla_harika_ae9342db1a"/>
    <language>en</language>
    <item>
      <title>Designing Hindsight Retain and Recall for SRE Incidents</title>
      <dc:creator>Devalla Harika</dc:creator>
      <pubDate>Mon, 28 Sep 2026 11:09:51 +0000</pubDate>
      <link>https://dev.to/devalla_harika_ae9342db1a/designing-hindsight-retain-and-recall-for-sre-incidents-24</link>
      <guid>https://dev.to/devalla_harika_ae9342db1a/designing-hindsight-retain-and-recall-for-sre-incidents-24</guid>
      <description>&lt;p&gt;An incident-response system can retrieve the right information and still fail to learn anything.&lt;/p&gt;

&lt;p&gt;That was one of the design problems behind OpsMind. We wanted an AI SRE agent that could not only recall previous incident experiences while investigating a failure, but also retain the outcome of a resolved incident so that the experience could influence future investigations.&lt;/p&gt;

&lt;p&gt;The resulting workflow is built around two operations:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Recall before diagnosis. Retain after resolution.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Hindsight provides the persistent memory layer that connects those two stages.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Role of Memory in OpsMind
&lt;/h2&gt;

&lt;p&gt;OpsMind receives an incident containing information such as its service, severity, symptoms, and current telemetry.&lt;/p&gt;

&lt;p&gt;The agent then gathers current evidence from logs and metrics.&lt;/p&gt;

&lt;p&gt;Only after establishing the current situation does it query Hindsight for historical context.&lt;/p&gt;

&lt;p&gt;The overall flow is:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Incident → Evidence → Hindsight Recall → AI Diagnosis → Human Approval → Resolution → Hindsight Retain&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The important part is the position of memory in this workflow.&lt;/p&gt;

&lt;p&gt;Memory is not the first source the agent consults, and it is not the final authority. It sits between current evidence collection and reasoning as a source of additional operational context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frsiounb4askei7x4nb5x.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frsiounb4askei7x4nb5x.jpeg" alt=" " width="800" height="404"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Figure 1 — Hindsight connects historical incident experience to the current incident-response workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing the Recall Query
&lt;/h2&gt;

&lt;p&gt;A useful memory system depends heavily on what you ask it to retrieve.&lt;/p&gt;

&lt;p&gt;For OpsMind, the recall query includes details from the current incident:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Incident ID&lt;/li&gt;
&lt;li&gt;Service&lt;/li&gt;
&lt;li&gt;Severity&lt;/li&gt;
&lt;li&gt;Symptoms&lt;/li&gt;
&lt;li&gt;Current metrics&lt;/li&gt;
&lt;li&gt;Performance behavior&lt;/li&gt;
&lt;li&gt;Error patterns&lt;/li&gt;
&lt;li&gt;Resource saturation&lt;/li&gt;
&lt;li&gt;Previous remediation experience&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, a payment API incident with high latency, HTTP 500 errors, and high database connection utilization should retrieve previous incidents involving similar operational behavior.&lt;/p&gt;

&lt;p&gt;The purpose is not to search for an identical incident ID.&lt;/p&gt;

&lt;p&gt;The purpose is to find &lt;em&gt;relevant operational experience&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;The query therefore asks Hindsight to identify previous SRE incidents with similar symptoms, service behavior, performance problems, errors, resource saturation, remediation experiences, and outcomes.&lt;/p&gt;

&lt;p&gt;That gives the AI agent historical context without hardcoding a specific previous incident as the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Recall Implementation
&lt;/h2&gt;

&lt;p&gt;The memory layer uses Hindsight's asynchronous API:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
response = await hindsight_client.arecall(&lt;br&gt;
    bank_id=HINDSIGHT_BANK_ID,&lt;br&gt;
    query=query,&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The result is then processed before being passed into the diagnosis pipeline.&lt;/p&gt;

&lt;p&gt;One important step is excluding the current incident from the historical results.&lt;/p&gt;

&lt;p&gt;The current incident should never appear as if it were already historical knowledge.&lt;/p&gt;

&lt;p&gt;OpsMind also extracts incident identifiers from returned memory and removes duplicates. This gives the reasoning layer a cleaner historical context.&lt;/p&gt;

&lt;p&gt;The result is effectively a list of previous incident experiences that may be relevant to the current investigation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Recall Cannot Be the Source of Truth
&lt;/h2&gt;

&lt;p&gt;One of the easiest mistakes in an AI incident-response system would be to let memory override current telemetry.&lt;/p&gt;

&lt;p&gt;Imagine that an old incident had a database connection problem and was resolved by increasing the connection pool.&lt;/p&gt;

&lt;p&gt;A new incident might also mention database latency, but its actual root cause could be completely different.&lt;/p&gt;

&lt;p&gt;If the agent blindly copies the old remediation, persistent memory becomes a source of incorrect automation.&lt;/p&gt;

&lt;p&gt;OpsMind therefore gives current logs and metrics priority.&lt;/p&gt;

&lt;p&gt;The system prompt explicitly instructs the AI agent to use current evidence as the primary basis for claims and historical incidents only as supporting context.&lt;/p&gt;

&lt;p&gt;This creates an important separation:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Current telemetry answers: “What is happening now?”&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Hindsight answers: “What have we experienced before?”&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The AI uses both to reason about the incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retain Comes After Resolution
&lt;/h2&gt;

&lt;p&gt;Recall solves only half of the problem.&lt;/p&gt;

&lt;p&gt;If OpsMind can retrieve historical experiences but never creates new ones, its knowledge becomes static.&lt;/p&gt;

&lt;p&gt;The second half of the workflow is therefore retention.&lt;/p&gt;

&lt;p&gt;After an incident is successfully resolved, OpsMind creates an incident learning record.&lt;/p&gt;

&lt;p&gt;The record contains:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Incident ID&lt;/li&gt;
&lt;li&gt;Service&lt;/li&gt;
&lt;li&gt;AI diagnosis&lt;/li&gt;
&lt;li&gt;Actions taken&lt;/li&gt;
&lt;li&gt;Outcome&lt;/li&gt;
&lt;li&gt;Resolution status&lt;/li&gt;
&lt;li&gt;A learning statement for future incidents&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The retention call looks like this:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
hindsight_client.retain(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    content=learning_record,&lt;br&gt;
    context="OpsMind SRE incident learning",&lt;br&gt;
    metadata={&lt;br&gt;
        "incident_id": str(incident_id),&lt;br&gt;
        "service": str(service),&lt;br&gt;
        "type": "incident_outcome",&lt;br&gt;
        "successful": str(successful).lower(),&lt;br&gt;
    },&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;The metadata is useful because it provides structured information alongside the learning record.&lt;/p&gt;

&lt;h2&gt;
  
  
  INC-008 as a Learning Event
&lt;/h2&gt;

&lt;p&gt;INC-008 provides a concrete example.&lt;/p&gt;

&lt;p&gt;The Payment API experienced:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;6.1-second latency&lt;/li&gt;
&lt;li&gt;26% HTTP 500 errors&lt;/li&gt;
&lt;li&gt;97% database connection utilization&lt;/li&gt;
&lt;li&gt;Requests waiting for database connections&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpsMind diagnosed database connection pool exhaustion.&lt;/p&gt;

&lt;p&gt;After the remediation was approved and simulated successfully, the outcome was retained in Hindsight.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1t4owwg6q0wvxxgh8lhp.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1t4owwg6q0wvxxgh8lhp.jpeg" alt=" " width="800" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Figure 2 — The successful INC-008 resolution is retained as an organizational learning record.&lt;/p&gt;

&lt;p&gt;At that point, INC-008 changed from a current incident into historical organizational knowledge.&lt;/p&gt;

&lt;p&gt;That transition is the core behavior we wanted from persistent memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing the Memory Loop
&lt;/h2&gt;

&lt;p&gt;The strongest test was to investigate another incident after INC-008 had been retained.&lt;/p&gt;

&lt;p&gt;When INC-007 was analyzed, Hindsight returned INC-008 as historical context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz92ytylzlnkk6kjutjdf.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz92ytylzlnkk6kjutjdf.jpeg" alt=" " width="800" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Figure 3 — A later investigation retrieves the newly retained INC-008 experience.&lt;/p&gt;

&lt;p&gt;This gave us a complete memory lifecycle:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;INC-008 → Resolve → Retain → INC-007 → Recall INC-008&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The important observation is that the newly retained experience was not manually copied into the second investigation.&lt;/p&gt;

&lt;p&gt;It was available through the memory layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  An Unexpected Engineering Issue
&lt;/h2&gt;

&lt;p&gt;The memory integration also exposed an asynchronous execution issue.&lt;/p&gt;

&lt;p&gt;The initial implementation produced:&lt;/p&gt;

&lt;p&gt;text&lt;br&gt;
Timeout context manager should be used inside a task&lt;/p&gt;

&lt;p&gt;The problem appeared when the synchronous Hindsight recall path interacted with the asynchronous FastAPI request environment.&lt;/p&gt;

&lt;p&gt;We changed the implementation to use an asynchronous Hindsight client and arecall() inside a dedicated asynchronous function.&lt;/p&gt;

&lt;p&gt;The client is explicitly closed afterward:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
await hindsight_client.aclose()&lt;br&gt;
This fixed the event-loop problem and also avoided leaving the underlying client session open.&lt;/p&gt;

&lt;p&gt;It was a useful reminder that an agent-memory integration has to fit the application's execution model, not just its logical architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  What We Learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Recall and retain are one system
&lt;/h3&gt;

&lt;p&gt;Retrieval without learning produces a static knowledge base.&lt;/p&gt;

&lt;p&gt;Retention without retrieval produces information that cannot influence future reasoning.&lt;/p&gt;

&lt;p&gt;The value comes from connecting both.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Memory quality depends on context
&lt;/h3&gt;

&lt;p&gt;A vague query can return broadly related incidents.&lt;/p&gt;

&lt;p&gt;Including service behavior, symptoms, metrics, errors, and resource saturation gives the memory system more useful retrieval context.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Historical context needs boundaries
&lt;/h3&gt;

&lt;p&gt;Previous incidents should inform diagnosis without becoming unquestioned instructions.&lt;/p&gt;

&lt;p&gt;The current incident remains the primary evidence source.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Successful outcomes are especially valuable
&lt;/h3&gt;

&lt;p&gt;OpsMind retains the outcome of successful remediation because that gives future investigations information about what previously worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The central memory design in OpsMind is simple:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Recall before reasoning. Retain after learning.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Hindsight provides the persistent layer that connects those two operations.&lt;/p&gt;

&lt;p&gt;INC-008 demonstrated the complete lifecycle. Its symptoms were analyzed using current evidence and historical context. After successful resolution, the experience was retained. A later investigation could then recall INC-008 as historical context.&lt;/p&gt;

&lt;p&gt;That makes memory part of the incident-response lifecycle rather than a separate documentation system.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>devops</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
