<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Latika Lokrey</title>
    <description>The latest articles on DEV Community by Latika Lokrey (@latika_lokrey).</description>
    <link>https://dev.to/latika_lokrey</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147695%2F2e9aae26-9060-4164-b6d3-c3526fbc8b01.png</url>
      <title>DEV Community: Latika Lokrey</title>
      <link>https://dev.to/latika_lokrey</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/latika_lokrey"/>
    <language>en</language>
    <item>
      <title>Building RecallOps: An Incident Response System That Remembers Past Failures</title>
      <dc:creator>Latika Lokrey</dc:creator>
      <pubDate>Mon, 28 Sep 2026 17:56:23 +0000</pubDate>
      <link>https://dev.to/latika_lokrey/building-recallops-an-incident-response-system-that-remembers-past-failures-1m0k</link>
      <guid>https://dev.to/latika_lokrey/building-recallops-an-incident-response-system-that-remembers-past-failures-1m0k</guid>
      <description>&lt;p&gt;A surprising amount of operational knowledge disappears after an incident is resolved. The ticket gets closed, the engineer moves on, and the next person who encounters the same failure often starts from scratch. Most organizations already collect incident data, but very few systems help engineers actively reuse what was learned from previous investigations.&lt;/p&gt;

&lt;p&gt;We built RecallOps to address that problem. Instead of treating incident records as historical artifacts, we designed a system that retains operational knowledge from past incidents and makes it available during future investigations. The goal was simple: when a similar outage occurs, engineers should be able to benefit from prior experience rather than repeating the same debugging process.&lt;/p&gt;

&lt;p&gt;At the center of that idea is memory. By combining structured incident management, persistent memory through Hindsight, AI-assisted recommendations, and a unified investigation workflow, RecallOps transforms incident history into operational context that can be recalled when it matters most.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem We Wanted to Solve&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Incident management tools excel at recording information. They store tickets, timelines, root-cause analyses, and resolution notes. The problem is that most of this information remains buried inside historical records.&lt;/p&gt;

&lt;p&gt;Consider a common scenario.&lt;/p&gt;

&lt;p&gt;An engineer encounters:&lt;/p&gt;

&lt;p&gt;Database timeout errors in the Payment API&lt;/p&gt;

&lt;p&gt;After several hours of investigation, the team discovers:&lt;/p&gt;

&lt;p&gt;Root Cause:&lt;br&gt;
Connection pool exhaustion&lt;/p&gt;

&lt;p&gt;Resolution:&lt;br&gt;
Restart Service X&lt;br&gt;
Increase connection pool size&lt;/p&gt;

&lt;p&gt;The issue is fixed, documented, and eventually forgotten.&lt;/p&gt;

&lt;p&gt;Months later, another engineer encounters nearly the same failure. The knowledge already exists somewhere, but locating it requires searching through tickets, dashboards, chat messages, and documentation.&lt;/p&gt;

&lt;p&gt;The result is duplicated effort.&lt;/p&gt;

&lt;p&gt;We wanted to build a system capable of remembering operational knowledge and surfacing it at the right moment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What RecallOps Does&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;RecallOps is an incident response platform that captures incident knowledge, stores it as memory, retrieves similar historical incidents, and provides engineers with relevant context during investigations.&lt;/p&gt;

&lt;p&gt;At a high level, the workflow looks like this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjqje6my4cs8dlrubsict.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjqje6my4cs8dlrubsict.png" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When a new incident is reported, the system searches historical memory for related failures. Relevant incidents are retrieved and presented as structured operational context. The recommendation layer then uses that context to assist engineers, while the backend records the final resolution for future recall.&lt;/p&gt;

&lt;p&gt;This creates a feedback loop where every resolved incident strengthens the system's memory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System Architecture&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One of our goals was to keep the architecture modular. Rather than creating a tightly coupled application where every component directly communicates with every other component, we divided responsibilities across dedicated layers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frontend&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The frontend provides the primary interface for engineers.&lt;/p&gt;

&lt;p&gt;It allows users to:&lt;/p&gt;

&lt;p&gt;Report incidents&lt;br&gt;
Review historical context&lt;br&gt;
Analyze recommendations&lt;br&gt;
Record resolutions&lt;br&gt;
Explore previous investigations&lt;/p&gt;

&lt;p&gt;The objective was to keep operational workflows simple while exposing the system's memory capabilities in a clear and understandable way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Backend&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The backend serves as the orchestration layer connecting all major components.&lt;/p&gt;

&lt;p&gt;Its responsibilities include:&lt;/p&gt;

&lt;p&gt;Incident lifecycle management&lt;br&gt;
API contracts&lt;br&gt;
Database operations&lt;br&gt;
Memory retrieval coordination&lt;br&gt;
Integration with recommendation services&lt;/p&gt;

&lt;p&gt;Using FastAPI allowed us to expose stable endpoints while keeping the architecture straightforward.&lt;/p&gt;

&lt;p&gt;Core endpoints include:&lt;/p&gt;

&lt;p&gt;POST /incidents&lt;br&gt;
POST /analyze&lt;br&gt;
POST /resolve&lt;br&gt;
GET /incidents&lt;/p&gt;

&lt;p&gt;By centralizing orchestration inside the backend, the frontend remains independent of implementation details while other components can evolve without requiring API redesign.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory Layer&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The memory layer is built around Hindsight.&lt;/p&gt;

&lt;p&gt;One aspect we found particularly compelling in the Hindsight documentation is its focus on retaining and recalling information over time instead of treating every interaction as isolated.&lt;/p&gt;

&lt;p&gt;Traditional incident systems focus on storage.&lt;/p&gt;

&lt;p&gt;Memory systems focus on recall.&lt;/p&gt;

&lt;p&gt;That distinction became the foundation of RecallOps.&lt;/p&gt;

&lt;p&gt;Historical incidents are retained as operational knowledge that can later be retrieved when similar failures occur.&lt;/p&gt;

&lt;p&gt;The project is designed around the Hindsight GitHub repository, which provides the foundation for memory retention and retrieval.&lt;/p&gt;

&lt;p&gt;For readers interested in the broader concept, Vectorize provides an excellent explanation of persistent memory in its article about agent memory systems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recommendation Layer&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Retrieving historical incidents is only part of the solution.&lt;/p&gt;

&lt;p&gt;Engineers still need context that is easy to interpret.&lt;/p&gt;

&lt;p&gt;The recommendation layer consumes retrieved incident knowledge and generates structured guidance based on previous investigations.&lt;/p&gt;

&lt;p&gt;Instead of presenting raw database records, the system highlights:&lt;/p&gt;

&lt;p&gt;Similar incidents&lt;br&gt;
Previous root causes&lt;br&gt;
Historical resolutions&lt;br&gt;
Relevant investigation outcomes&lt;/p&gt;

&lt;p&gt;This transforms historical data into operational context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Designing Around Incident Knowledge&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One lesson became obvious early in development:&lt;/p&gt;

&lt;p&gt;A newly created incident contains symptoms. A resolved incident contains knowledge.&lt;/p&gt;

&lt;p&gt;That realization influenced our data model.&lt;/p&gt;

&lt;p&gt;A simplified incident record contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Incident&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Base&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;__tablename__&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incidents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="nb"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Integer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;primary_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;service&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;error&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;severity&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;root_cause&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;nullable&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;resolution&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;nullable&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The most valuable fields are not necessarily the ones created when an incident begins.&lt;/p&gt;

&lt;p&gt;They are the fields populated after engineers understand what actually happened.&lt;/p&gt;

&lt;p&gt;Root causes and resolutions are the assets that make memory useful.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Memory Retrieval Matters&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A common assumption is that incident management is primarily a storage problem.&lt;/p&gt;

&lt;p&gt;We discovered that retrieval is significantly more important.&lt;/p&gt;

&lt;p&gt;Imagine an engineer reports:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database timeout
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The memory layer identifies a similar historical incident and returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"memory_found"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"recommendation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Previous resolution: Restart Service X"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"memory"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"root_cause"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Connection pool exhaustion"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"resolution"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Restart Service X"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The engineer still investigates the issue.&lt;/p&gt;

&lt;p&gt;The system does not replace human decision-making.&lt;/p&gt;

&lt;p&gt;Instead, it provides relevant context at the beginning of the investigation rather than after hours of searching.&lt;/p&gt;

&lt;p&gt;That small difference can dramatically improve how quickly operational knowledge is discovered and reused.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example Workflow&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;First Incident&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An engineer reports:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Payment API returning database timeout errors
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Investigation reveals:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Root Cause:
Connection pool exhaustion

Resolution:
Restart Service X
Increase pool size
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The incident is resolved and retained as memory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Future Incident&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Several months later, a similar issue occurs.&lt;/p&gt;

&lt;p&gt;The engineer submits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Payment API database timeout
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;RecallOps retrieves historical knowledge and returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Similar Incident Found

Root Cause:
Connection pool exhaustion

Previous Resolution:
Restart Service X
Increase pool size
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The engineer is not forced to repeat the entire discovery process.&lt;/p&gt;

&lt;p&gt;Instead, they begin with accumulated organizational knowledge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lessons We Learned&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;1. Memory Is More Valuable Than Storage&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Collecting incident data is relatively easy.&lt;/p&gt;

&lt;p&gt;Ensuring that knowledge becomes discoverable during future incidents is far more challenging and significantly more useful.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Stable APIs Simplify Integration&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Clearly defined backend contracts allowed frontend, memory, and recommendation components to evolve independently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Resolved Incidents Are the Real Asset&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The most valuable information is not the outage itself.&lt;/p&gt;

&lt;p&gt;It is the explanation of why the outage occurred and how it was resolved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Retrieval Quality Matters More Than Data Volume&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A smaller collection of highly relevant incidents often provides more value than a large repository of poorly retrievable information.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Operational Knowledge Compounds Over Time&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every resolved incident contributes to the system's ability to assist during future investigations.&lt;/p&gt;

&lt;p&gt;The value of memory increases as more knowledge is retained.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Closing Thoughts&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Building RecallOps changed how we think about incident management systems. The challenge was never collecting incident data. Most organizations already have tools that do that effectively.&lt;/p&gt;

&lt;p&gt;The real challenge was creating a system capable of remembering.&lt;/p&gt;

&lt;p&gt;By combining structured incident records, FastAPI-based orchestration, persistent memory through Hindsight, AI-assisted recommendations, and a unified workflow, RecallOps transforms historical incidents into reusable operational knowledge.&lt;/p&gt;

&lt;p&gt;Every incident teaches something. The goal of RecallOps is to ensure that lesson is not forgotten.&lt;/p&gt;

</description>
      <category>automation</category>
      <category>devops</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Building the Backend for an Incident Memory System with Hindsight.</title>
      <dc:creator>Latika Lokrey</dc:creator>
      <pubDate>Mon, 28 Sep 2026 17:20:41 +0000</pubDate>
      <link>https://dev.to/latika_lokrey/building-the-backend-for-an-incident-memory-system-with-hindsight-5ee3</link>
      <guid>https://dev.to/latika_lokrey/building-the-backend-for-an-incident-memory-system-with-hindsight-5ee3</guid>
      <description>&lt;p&gt;A surprising amount of operational knowledge disappears after an incident is resolved. The ticket gets closed, the engineer moves on, and the next person who encounters the same failure starts from scratch. I built RecallOps to address that problem by treating incident history as something an application can actively remember and reuse rather than merely archive.&lt;br&gt;
The backend turned out to be the most important part of that idea. If incident memory is going to be useful, it needs a reliable way to store incidents, retrieve similar failures, and expose that knowledge to other components. My focus was designing the FastAPI infrastructure that connects incident management, persistent storage, and the project's Hindsight-based memory layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the System Does&lt;/strong&gt;&lt;br&gt;
RecallOps is an incident response system that captures operational knowledge from past outages and makes it available during future investigations.&lt;br&gt;
At a high level, the flow looks like this:&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fed5wa0dpc95gj2kg5rj6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fed5wa0dpc95gj2kg5rj6.png" alt=" " width="145" height="315"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When an engineer reports a new issue, the backend retrieves relevant historical incidents, exposes that context to downstream analysis components, and records the final outcome so the system can learn from future resolutions.&lt;br&gt;
The memory layer is built around Hindsight. I found the ideas described in the Hindsight documentation particularly useful because they focus on retaining and recalling information over time rather than treating every interaction as isolated. The broader concept of persistent memory is also explained well in Vectorize's article on agent memory systems.&lt;br&gt;
The system is designed around Hindsight GitHub repository as the foundation for memory retention and retrieval, while the backend provides the APIs and orchestration layer that connect incident workflows to memory recall.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Backend Problem I Actually Needed to Solve&lt;/strong&gt;&lt;br&gt;
The obvious part of the project was storing incidents in a database.&lt;br&gt;
The harder part was deciding what role the backend should play.&lt;br&gt;
I could have pushed all memory logic into the AI layer, but that creates a system where operational knowledge becomes tightly coupled to model prompts. That approach is difficult to debug and even harder to evolve.&lt;br&gt;
Instead, I wanted the backend to own three responsibilities:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; Incident lifecycle management&lt;/li&gt;
&lt;li&gt; Memory retrieval orchestration&lt;/li&gt;
&lt;li&gt; Stable APIs for the rest of the system&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That decision shaped almost every part of the architecture.&lt;br&gt;
The frontend only talks to APIs.&lt;br&gt;
The memory layer only worries about retention and recall.&lt;br&gt;
The recommendation engine consumes structured context instead of raw incident history.&lt;br&gt;
That separation ended up simplifying integration significantly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Designing Around Incident Workflows&lt;/strong&gt;&lt;br&gt;
I started with a simple incident model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Incident&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Base&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;__tablename__&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;incidents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="nb"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Integer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;primary_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;service&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;error&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;severity&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;root_cause&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;nullable&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;resolution&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Column&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;nullable&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This model captures the minimum information required to support both operational workflows and memory retrieval.&lt;br&gt;
One thing I learned early is that incident records become far more useful after they are resolved.&lt;br&gt;
A newly created incident contains symptoms.&lt;br&gt;
A resolved incident contains knowledge.&lt;br&gt;
That distinction influenced how I designed the APIs.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk0dntru4bmh6iadgmck6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk0dntru4bmh6iadgmck6.png" alt=" " width="427" height="364"&gt;&lt;/a&gt;&lt;br&gt;
Figure 1: Backend project structure used for incident storage, retrieval, and API orchestration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Creating Stable API Contracts&lt;/strong&gt;&lt;br&gt;
One of my goals was enabling parallel development without forcing every component to wait for every other component.&lt;br&gt;
To accomplish that, I standardized a small set of backend endpoints.&lt;br&gt;
Creating an incident:&lt;br&gt;
POST /incidents&lt;br&gt;
Analyzing an incident:&lt;br&gt;
POST /analyze&lt;/p&gt;

&lt;p&gt;Resolving an incident:&lt;br&gt;
POST /resolve&lt;/p&gt;

&lt;p&gt;Listing historical incidents:&lt;br&gt;
GET /incidents&lt;/p&gt;

&lt;p&gt;Those endpoints remained stable even as the implementation evolved.&lt;br&gt;
The frontend team could build dashboards.&lt;br&gt;
The memory layer could improve retrieval quality.&lt;br&gt;
The recommendation engine could change its analysis strategy.&lt;br&gt;
None of those changes required API redesign.&lt;br&gt;
That stability was more valuable than I initially expected.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgjnk64bhm5fmn1b40w6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgjnk64bhm5fmn1b40w6.png" alt=" " width="799" height="466"&gt;&lt;/a&gt;&lt;br&gt;
Figure 2: FastAPI Swagger documentation exposing the incident management and memory retrieval APIs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Memory Retrieval Lives Behind the API&lt;/strong&gt;&lt;br&gt;
One design decision I feel strongly about is hiding memory retrieval behind backend endpoints.&lt;br&gt;
The retrieval process looks conceptually simple:&lt;br&gt;
memory = find_similar_incident(&lt;br&gt;
    db,&lt;br&gt;
    request.error&lt;br&gt;
)&lt;br&gt;
But in practice, retrieval becomes increasingly sophisticated.&lt;br&gt;
Similarity matching evolves.&lt;br&gt;
Metadata expands.&lt;br&gt;
Ranking strategies change.&lt;br&gt;
Storage mechanisms improve.&lt;br&gt;
If every consumer directly interacted with memory infrastructure, those changes would ripple throughout the system.&lt;br&gt;
Instead, the backend acts as a boundary.&lt;br&gt;
Consumers ask questions.&lt;br&gt;
The backend decides how memory is retrieved.&lt;br&gt;
That abstraction allowed me to improve retrieval behavior without forcing changes elsewhere.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Turning Historical Incidents into Operational Context&lt;/strong&gt;&lt;br&gt;
The most interesting part of the project was transforming historical data into something actionable.&lt;br&gt;
Imagine an engineer reports:&lt;br&gt;
Database timeout&lt;br&gt;
The backend searches memory and discovers a previous incident.&lt;br&gt;
The historical record contains:&lt;br&gt;
Root Cause:&lt;br&gt;
Connection pool exhaustion&lt;br&gt;
Resolution:&lt;br&gt;
Restart Service X&lt;br&gt;
Instead of returning raw database records, the backend packages that information into structured context.&lt;br&gt;
A response might look like:&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foupkd28mruku8ksr0q20.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foupkd28mruku8ksr0q20.png" alt=" " width="799" height="173"&gt;&lt;/a&gt;&lt;br&gt;
Figure 3: Memory retrieval returning a previously resolved incident and its resolution.&lt;br&gt;
That may seem straightforward, but it changes how engineers interact with operational knowledge.&lt;br&gt;
The system is no longer acting as a storage layer.&lt;br&gt;
It becomes a retrieval layer.&lt;br&gt;
That distinction matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Hindsight Fits This Problem&lt;/strong&gt;&lt;br&gt;
Many incident management systems focus on recording information.&lt;br&gt;
Fewer systems focus on remembering it.&lt;br&gt;
That is where Hindsight proved useful.&lt;br&gt;
The retain-and-recall model maps naturally onto incident response.&lt;br&gt;
An outage occurs.&lt;br&gt;
The investigation identifies a root cause.&lt;br&gt;
A resolution is applied.&lt;br&gt;
The outcome gets retained.&lt;br&gt;
When a similar failure appears later, that knowledge becomes recallable.&lt;br&gt;
This creates a feedback loop:&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjiwsf6x0896x004ldegs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjiwsf6x0896x004ldegs.png" alt=" " width="459" height="495"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The value compounds over time because every resolved incident expands the system's operational memory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example Interaction&lt;/strong&gt;&lt;br&gt;
Consider two incidents separated by several months.&lt;br&gt;
First Incident&lt;br&gt;
An engineer reports:&lt;br&gt;
Payment API returning database timeout errors&lt;br&gt;
Investigation reveals:&lt;br&gt;
Connection pool exhaustion&lt;br&gt;
Resolution:&lt;br&gt;
Restart Service X&lt;br&gt;
Increase pool size&lt;br&gt;
The incident is resolved and retained.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Later Incident&lt;/strong&gt;&lt;br&gt;
A similar failure appears.&lt;br&gt;
The engineer submits:&lt;br&gt;
Payment API database timeout&lt;br&gt;
The backend retrieves historical context and returns:&lt;br&gt;
Similar incident found.&lt;br&gt;
Root Cause:&lt;br&gt;
Connection pool exhaustion&lt;br&gt;
Previous Resolution:&lt;br&gt;
Restart Service X&lt;br&gt;
Increase pool size&lt;br&gt;
The engineer still makes the final decision, but they begin with accumulated organizational knowledge rather than a blank page.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lessons Learned&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Memory Is More Valuable Than Storage
Storing incidents is easy.
Making historical knowledge discoverable at the right moment is significantly harder and far more useful.&lt;/li&gt;
&lt;li&gt;Stable APIs Reduce Team Friction
A small set of consistent contracts allowed multiple components to evolve independently without constant coordination.&lt;/li&gt;
&lt;li&gt;Retrieval Logic Changes Constantly
Similarity matching and memory ranking evolve quickly.
Keeping them behind backend boundaries prevents unnecessary coupling.&lt;/li&gt;
&lt;li&gt;Resolved Incidents Are the Real Asset
The most valuable information is not the failure itself.
It is the explanation of why the failure occurred and how it was fixed.&lt;/li&gt;
&lt;li&gt;Backend Design Shapes Everything Else
When memory retrieval becomes a first-class concern, the backend stops being a simple CRUD layer and becomes the coordination point for operational knowledge.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Closing Thoughts&lt;/strong&gt;&lt;br&gt;
Building RecallOps changed how I think about incident management systems. The challenge was not collecting incident data. Most organizations already do that. The challenge was creating a backend capable of turning historical incidents into useful context during future failures.&lt;br&gt;
By combining structured incident records, stable APIs, and Hindsight-based memory retrieval, the system moves beyond storing operational history and begins reusing it. Over time, that accumulated knowledge becomes one of the most valuable assets in the platform because every resolved incident improves the system's ability to respond to the next one.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
