<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sonasri Gummula</title>
    <description>The latest articles on DEV Community by Sonasri Gummula (@sonasri_gummula_5b44c15c4).</description>
    <link>https://dev.to/sonasri_gummula_5b44c15c4</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146471%2F717a0070-610f-479f-a7f0-3718adbaa70a.png</url>
      <title>DEV Community: Sonasri Gummula</title>
      <link>https://dev.to/sonasri_gummula_5b44c15c4</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sonasri_gummula_5b44c15c4"/>
    <language>en</language>
    <item>
      <title>RecallOps: Building Persistent Memory for Engineering Incident Response</title>
      <dc:creator>Sonasri Gummula</dc:creator>
      <pubDate>Mon, 28 Sep 2026 06:57:45 +0000</pubDate>
      <link>https://dev.to/sonasri_gummula_5b44c15c4/recallops-building-persistent-memory-for-engineering-incident-response-3akl</link>
      <guid>https://dev.to/sonasri_gummula_5b44c15c4/recallops-building-persistent-memory-for-engineering-incident-response-3akl</guid>
      <description>&lt;p&gt;&lt;strong&gt;Introduction&lt;/strong&gt;&lt;br&gt;
Software systems can fail unexpectedly. An API can become slow, a database connection can time out, or a new deployment can introduce an issue. When this happens, engineers need to understand the problem quickly and find an appropriate solution.&lt;br&gt;
One challenge in incident response is that teams often have valuable experience from previous incidents, but that experience may not be available when a similar problem occurs again. Engineers may spend time investigating a problem from the beginning even though a similar incident has already been solved.&lt;br&gt;
RecallOps is designed to address this problem by giving engineering incident response a persistent memory. It uses Hindsight to retain engineering experiences and recall relevant previous incidents when a new incident needs to be investigated.&lt;br&gt;
The main idea is simple: instead of repeatedly starting from zero, engineers can use what the team has already learned.&lt;br&gt;
&lt;strong&gt;The Problem&lt;/strong&gt;&lt;br&gt;
Consider a Payment API experiencing an API latency spike because of a database connection timeout.&lt;br&gt;
An engineer investigating the issue may need to check the database, connection pool, recent deployments, queries, logs, and service configuration.&lt;br&gt;
If a similar problem occurred previously, the engineering team may already have information about its root cause and the solution that worked. However, simply storing old incident records is not enough. The previous experience needs to be available when a new and similar incident occurs.&lt;br&gt;
&lt;strong&gt;This led to the main idea behind RecallOps&lt;/strong&gt;:&lt;br&gt;
How can previous engineering experience be made useful during future incident investigations?&lt;br&gt;
RecallOps addresses this by connecting current incidents with relevant historical experiences through persistent memory.&lt;br&gt;
&lt;strong&gt;Our Solution&lt;/strong&gt;: RecallOps&lt;br&gt;
RecallOps provides an engineering incident command center where an engineer can create and investigate incidents.&lt;br&gt;
An incident can contain information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Incident title&lt;/li&gt;
&lt;li&gt;Description&lt;/li&gt;
&lt;li&gt;Service&lt;/li&gt;
&lt;li&gt;Deployment&lt;/li&gt;
&lt;li&gt;Severity
For example, we can create an incident with the following information:&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;Title: API Latency Spike&lt;/li&gt;
&lt;li&gt;Description: Database connection timeout causing high API latency&lt;/li&gt;
&lt;li&gt;Service: Payment API&lt;/li&gt;
&lt;li&gt;Deployment: v2.4.1&lt;/li&gt;
&lt;li&gt;Severity: Critical
After creating the incident, the engineer can use the Investigate with Recall feature.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F37a2e2vs3c401a1vjal0.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F37a2e2vs3c401a1vjal0.jpg" alt="Figure 1: RecallOps engineering incident command center  " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
The dashboard provides a single place to view the active incident, investigate it using memory, review analysis, and record the final outcome.&lt;br&gt;
How Hindsight Provides Persistent Memory&lt;br&gt;
Hindsight is the central memory layer used by RecallOps.&lt;br&gt;
There are two important operations in our implementation: Recall and Retain.&lt;br&gt;
Recall-When a new incident occurs, RecallOps creates a query using the current incident information.&lt;br&gt;
The query contains the incident's symptoms and context and is sent to Hindsight. Hindsight then searches its persistent memory for relevant previous engineering experiences.&lt;br&gt;
For example, a current database connection timeout may be related to previous incidents involving database failures, similar symptoms, deployment problems, or connection configuration.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxs0quhsixi065trj97pr.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxs0quhsixi065trj97pr.jpg" alt="Figure 2: An incident being prepared for investigation in RecallOps" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When the engineer clicks Investigate with Recall, the system searches Hindsight for relevant memories.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fayn7x39vyfrwwyne802y.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fayn7x39vyfrwwyne802y.jpg" alt="Figure 3: Hindsight recalling previous engineering experiences relevant to the current incident. " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
This is an important part of RecallOps because the system is not simply displaying a static list of old incidents. The current incident is used to retrieve relevant previous experiences.&lt;br&gt;
From Recalled Memory to Investigation&lt;br&gt;
The recalled information is used to support the engineering investigation.&lt;br&gt;
RecallOps presents three important areas:&lt;br&gt;
&lt;strong&gt;Root Cause Pattern&lt;/strong&gt;&lt;br&gt;
The system identifies a possible pattern based on previous incidents.&lt;br&gt;
&lt;strong&gt;Recommended Fix&lt;/strong&gt;&lt;br&gt;
The previous experience can provide a useful direction for resolving the current problem.&lt;br&gt;
&lt;strong&gt;Historical Lesson&lt;/strong&gt;&lt;br&gt;
The system presents lessons from previous incidents so that engineers can understand what happened previously.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5tkabcgg72euvsnxwpkx.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5tkabcgg72euvsnxwpkx.jpg" alt="Figure 4: RecallOps presents investigation insights based on previous engineering experience." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The purpose of this process is not to replace the engineer. Instead, it gives the engineer useful historical context that can make investigation more informed.&lt;br&gt;
&lt;strong&gt;Before and After Persistent Memory&lt;/strong&gt;&lt;br&gt;
The difference becomes clearer when comparing incident response without and with persistent memory.&lt;br&gt;
&lt;strong&gt;Without persistent memory&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;New Incident&lt;br&gt;
     ↓&lt;br&gt;
Start Investigation&lt;br&gt;
     ↓&lt;br&gt;
Check Logs&lt;br&gt;
     ↓&lt;br&gt;
Investigate Database&lt;br&gt;
     ↓&lt;br&gt;
Find Root Cause&lt;br&gt;
     ↓&lt;br&gt;
Find Fix&lt;/p&gt;

&lt;p&gt;The engineer may have to reconstruct the investigation from the beginning.&lt;br&gt;
&lt;strong&gt;With RecallOps and Hindsight&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;New Incident&lt;br&gt;
     ↓&lt;br&gt;
Hindsight Recall&lt;br&gt;
     ↓&lt;br&gt;
Previous Experience&lt;br&gt;
     ↓&lt;br&gt;
Compare Patterns&lt;br&gt;
     ↓&lt;br&gt;
Investigate&lt;br&gt;
     ↓&lt;br&gt;
Apply Fix&lt;/p&gt;

&lt;p&gt;For example, suppose a previous incident involved the same type of database connection problem and increasing the connection pool solved it.&lt;br&gt;
When a similar problem occurs again, RecallOps can retrieve that previous experience. The engineer therefore has a useful starting point instead of investigating with no historical context.&lt;br&gt;
This demonstrates the practical value of persistent memory: previous engineering experience becomes available when it is needed.&lt;br&gt;
&lt;strong&gt;Recording the Outcome&lt;/strong&gt;&lt;br&gt;
RecallOps also includes a developer feedback section where the engineer can record whether the recommended solution worked.&lt;br&gt;
For example, after applying a fix, the engineer could record:&lt;br&gt;
Increased the database connection pool and optimized slow database queries. API latency returned to normal.&lt;br&gt;
The engineer can select Fix Worked or Fix Failed and provide additional notes.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8saj3oij93pcgo68llh6.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8saj3oij93pcgo68llh6.jpg" alt="Figure 4: RecallOps presents investigation insights based on previous engineering experience." width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
Figure 5: Recording the outcome of an engineering fix.&lt;br&gt;
This creates a learning loop rather than stopping after the initial investigation.&lt;/p&gt;

&lt;p&gt;Incident&lt;br&gt;
   ↓&lt;br&gt;
Recall&lt;br&gt;
   ↓&lt;br&gt;
Previous Experience&lt;br&gt;
   ↓&lt;br&gt;
Investigation&lt;br&gt;
   ↓&lt;br&gt;
Fix&lt;br&gt;
   ↓&lt;br&gt;
Outcome&lt;br&gt;
   ↓&lt;br&gt;
Retain&lt;br&gt;
   ↓&lt;br&gt;
Future Incident&lt;/p&gt;

&lt;p&gt;If the solution works, the experience can become useful for future investigations. If it fails, that outcome is also valuable because it provides information about an approach that did not solve the problem.&lt;br&gt;
Technical Implementation&lt;br&gt;
RecallOps is implemented as a web-based engineering incident response application.&lt;/p&gt;

&lt;p&gt;The project uses technologies including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Python&lt;/li&gt;
&lt;li&gt;Flask&lt;/li&gt;
&lt;li&gt;HTML&lt;/li&gt;
&lt;li&gt;CSS&lt;/li&gt;
&lt;li&gt;JavaScript&lt;/li&gt;
&lt;li&gt;SQLite&lt;/li&gt;
&lt;li&gt;Hindsight&lt;/li&gt;
&lt;li&gt;Flask handles the application and API routes, while the frontend provides the incident command-center interface.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Hindsight integration is separated into a service layer. This allows the main application to communicate with Hindsight through dedicated memory operations.&lt;br&gt;
A simplified example of retaining an incident is:&lt;br&gt;
client.retain(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    content=incident_text&lt;br&gt;
)&lt;br&gt;
For retrieving relevant previous experiences, RecallOps uses:&lt;br&gt;
result = client.recall(&lt;br&gt;
    bank_id=BANK_ID,&lt;br&gt;
    query=query&lt;br&gt;
)&lt;br&gt;
The returned memories are processed by the application and displayed in the RecallOps interface.&lt;br&gt;
This integration makes Hindsight an important part of the application's workflow rather than simply using it as another data store.&lt;br&gt;
Why Persistent Engineering Memory Matters&lt;br&gt;
Incident history is valuable, but its value increases when it can actively support future incidents.&lt;br&gt;
Traditional incident records mainly tell engineers what happened in the past. RecallOps takes a step further by connecting previous experiences with a current incident.&lt;br&gt;
The system can recall relevant experiences, help the engineer understand possible patterns, and record the outcome of the current investigation.&lt;br&gt;
&lt;strong&gt;This creates a continuous learning process:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Past experience → Current investigation → New experience → Future investigation&lt;/p&gt;

&lt;p&gt;The result is an incident-response workflow where engineering knowledge can become increasingly useful over time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Conclusion&lt;/strong&gt;&lt;br&gt;
RecallOps gives engineering incident response a persistent memory.&lt;br&gt;
A new incident can be investigated using relevant experiences recalled through Hindsight. The engineer can use those experiences to understand possible root causes and solutions, apply a fix, and record the outcome.&lt;br&gt;
The most important concept behind RecallOps is that engineering teams should not have to repeatedly learn the same lesson.&lt;br&gt;
By connecting Recall, investigation, outcome, and Retain, RecallOps creates a continuous learning loop for engineering incident response.&lt;br&gt;
RecallOps turns incident history into reusable engineering experience, helping future investigations benefit from what the team has already learned.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>softwareengineering</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
