DEV Community

Latika Lokrey
Latika Lokrey

Posted on

Building RecallOps: An Incident Response System That Remembers Past Failures

A surprising amount of operational knowledge disappears after an incident is resolved. The ticket gets closed, the engineer moves on, and the next person who encounters the same failure often starts from scratch. Most organizations already collect incident data, but very few systems help engineers actively reuse what was learned from previous investigations.

We built RecallOps to address that problem. Instead of treating incident records as historical artifacts, we designed a system that retains operational knowledge from past incidents and makes it available during future investigations. The goal was simple: when a similar outage occurs, engineers should be able to benefit from prior experience rather than repeating the same debugging process.

At the center of that idea is memory. By combining structured incident management, persistent memory through Hindsight, AI-assisted recommendations, and a unified investigation workflow, RecallOps transforms incident history into operational context that can be recalled when it matters most.

The Problem We Wanted to Solve

Incident management tools excel at recording information. They store tickets, timelines, root-cause analyses, and resolution notes. The problem is that most of this information remains buried inside historical records.

Consider a common scenario.

An engineer encounters:

Database timeout errors in the Payment API

After several hours of investigation, the team discovers:

Root Cause:
Connection pool exhaustion

Resolution:
Restart Service X
Increase connection pool size

The issue is fixed, documented, and eventually forgotten.

Months later, another engineer encounters nearly the same failure. The knowledge already exists somewhere, but locating it requires searching through tickets, dashboards, chat messages, and documentation.

The result is duplicated effort.

We wanted to build a system capable of remembering operational knowledge and surfacing it at the right moment.

What RecallOps Does

RecallOps is an incident response platform that captures incident knowledge, stores it as memory, retrieves similar historical incidents, and provides engineers with relevant context during investigations.

At a high level, the workflow looks like this:

When a new incident is reported, the system searches historical memory for related failures. Relevant incidents are retrieved and presented as structured operational context. The recommendation layer then uses that context to assist engineers, while the backend records the final resolution for future recall.

This creates a feedback loop where every resolved incident strengthens the system's memory.

System Architecture

One of our goals was to keep the architecture modular. Rather than creating a tightly coupled application where every component directly communicates with every other component, we divided responsibilities across dedicated layers.

Frontend

The frontend provides the primary interface for engineers.

It allows users to:

Report incidents
Review historical context
Analyze recommendations
Record resolutions
Explore previous investigations

The objective was to keep operational workflows simple while exposing the system's memory capabilities in a clear and understandable way.

Backend

The backend serves as the orchestration layer connecting all major components.

Its responsibilities include:

Incident lifecycle management
API contracts
Database operations
Memory retrieval coordination
Integration with recommendation services

Using FastAPI allowed us to expose stable endpoints while keeping the architecture straightforward.

Core endpoints include:

POST /incidents
POST /analyze
POST /resolve
GET /incidents

By centralizing orchestration inside the backend, the frontend remains independent of implementation details while other components can evolve without requiring API redesign.

Memory Layer

The memory layer is built around Hindsight.

One aspect we found particularly compelling in the Hindsight documentation is its focus on retaining and recalling information over time instead of treating every interaction as isolated.

Traditional incident systems focus on storage.

Memory systems focus on recall.

That distinction became the foundation of RecallOps.

Historical incidents are retained as operational knowledge that can later be retrieved when similar failures occur.

The project is designed around the Hindsight GitHub repository, which provides the foundation for memory retention and retrieval.

For readers interested in the broader concept, Vectorize provides an excellent explanation of persistent memory in its article about agent memory systems.

Recommendation Layer

Retrieving historical incidents is only part of the solution.

Engineers still need context that is easy to interpret.

The recommendation layer consumes retrieved incident knowledge and generates structured guidance based on previous investigations.

Instead of presenting raw database records, the system highlights:

Similar incidents
Previous root causes
Historical resolutions
Relevant investigation outcomes

This transforms historical data into operational context.

Designing Around Incident Knowledge

One lesson became obvious early in development:

A newly created incident contains symptoms. A resolved incident contains knowledge.

That realization influenced our data model.

A simplified incident record contains:

class Incident(Base):
    __tablename__ = "incidents"

    id = Column(Integer, primary_key=True)
    service = Column(String)
    error = Column(String)
    severity = Column(String)

    root_cause = Column(String, nullable=True)
    resolution = Column(String, nullable=True)
Enter fullscreen mode Exit fullscreen mode

The most valuable fields are not necessarily the ones created when an incident begins.

They are the fields populated after engineers understand what actually happened.

Root causes and resolutions are the assets that make memory useful.

Why Memory Retrieval Matters

A common assumption is that incident management is primarily a storage problem.

We discovered that retrieval is significantly more important.

Imagine an engineer reports:

Database timeout
Enter fullscreen mode Exit fullscreen mode

The memory layer identifies a similar historical incident and returns:

{
  "memory_found": true,
  "recommendation": "Previous resolution: Restart Service X",
  "memory": {
    "root_cause": "Connection pool exhaustion",
    "resolution": "Restart Service X"
  }
}

Enter fullscreen mode Exit fullscreen mode

The engineer still investigates the issue.

The system does not replace human decision-making.

Instead, it provides relevant context at the beginning of the investigation rather than after hours of searching.

That small difference can dramatically improve how quickly operational knowledge is discovered and reused.

Example Workflow
First Incident

An engineer reports:

Payment API returning database timeout errors
Enter fullscreen mode Exit fullscreen mode

Investigation reveals:

Root Cause:
Connection pool exhaustion

Resolution:
Restart Service X
Increase pool size
Enter fullscreen mode Exit fullscreen mode

The incident is resolved and retained as memory.

Future Incident

Several months later, a similar issue occurs.

The engineer submits:

Payment API database timeout
Enter fullscreen mode Exit fullscreen mode

RecallOps retrieves historical knowledge and returns:

Similar Incident Found

Root Cause:
Connection pool exhaustion

Previous Resolution:
Restart Service X
Increase pool size
Enter fullscreen mode Exit fullscreen mode

The engineer is not forced to repeat the entire discovery process.

Instead, they begin with accumulated organizational knowledge.

Lessons We Learned
1. Memory Is More Valuable Than Storage

Collecting incident data is relatively easy.

Ensuring that knowledge becomes discoverable during future incidents is far more challenging and significantly more useful.

2. Stable APIs Simplify Integration

Clearly defined backend contracts allowed frontend, memory, and recommendation components to evolve independently.

3. Resolved Incidents Are the Real Asset

The most valuable information is not the outage itself.

It is the explanation of why the outage occurred and how it was resolved.

4. Retrieval Quality Matters More Than Data Volume

A smaller collection of highly relevant incidents often provides more value than a large repository of poorly retrievable information.

5. Operational Knowledge Compounds Over Time

Every resolved incident contributes to the system's ability to assist during future investigations.

The value of memory increases as more knowledge is retained.

Closing Thoughts

Building RecallOps changed how we think about incident management systems. The challenge was never collecting incident data. Most organizations already have tools that do that effectively.

The real challenge was creating a system capable of remembering.

By combining structured incident records, FastAPI-based orchestration, persistent memory through Hindsight, AI-assisted recommendations, and a unified workflow, RecallOps transforms historical incidents into reusable operational knowledge.

Every incident teaches something. The goal of RecallOps is to ensure that lesson is not forgotten.

Top comments (0)