DEV Community

ZMICS
ZMICS

Posted on

Designing an Incident Response Agent with FastAPI and Persistent Memory

Building an incident-response agent is not just about generating an answer to an alert. The real challenge is connecting incident data, historical context, recommendations, and engineers' actions into a single reliable workflow.

Our project approaches this as a backend engineering problem.

The Architecture

The system uses FastAPI as the API layer, with a database for structured incident data and a frontend that communicates with the backend.

The core workflow is:

Create Incident → Analyze → Retrieve History → Recommend Runbooks → Record Response → Resolve → Postmortem

The backend separates responsibilities across components for incidents, search, runbooks, and postmortems instead of putting the entire workflow into one service.

Finding Relevant Incidents

When a new incident arrives, the system searches historical incidents for relevant matches.

The retrieval layer uses TF-IDF and cosine similarity, combined with additional signals such as:

Service
Error signature
Severity
Recency

This produces an explainable incident fingerprint rather than relying entirely on a single similarity score.

For an engineer, this means the system can provide both the historical incident and context about why it was considered relevant.

Connecting Incidents to Runbooks

Historical incidents become more useful when we know what engineers actually did.

The system records runbook usage and whether the outcome was:

Worked → Partial → Failed

Runbooks also have an Elo-style trust rating that changes based on their historical outcomes.

This creates a feedback loop:

Incident → Runbook → Outcome → Updated Trust → Future Recommendation

Instead of simply recommending a runbook because it exists, the system can use previous response experience as additional context.

Adding Persistent Memory with Hindsight

For the Hindsight-enabled version of the project, Hindsight provides the persistent memory layer.

The application can retain useful incident context and recall it when a new incident requires historical information.

The architecture can therefore be thought of as:

FastAPI → Incident Analysis → Memory → Historical Context → Recommendation

This allows the agent to use previous operational experience while keeping structured incident data and application logic separate from the memory layer.

Keeping the Engineer in Control

Historical context should never be treated as an unquestionable answer.

Two incidents can look similar but have different root causes. Infrastructure also changes, which means an old solution may no longer be appropriate.

The system therefore focuses on providing context and recommendations, while leaving the final decision to the engineer.

What We Learned

The main lesson from building this system is that an incident-response agent is more than an AI model.

It requires several pieces working together:

API design + structured data + retrieval + persistent memory + runbook intelligence + human oversight

FastAPI provides the backend interface, the database preserves structured incident state, retrieval finds relevant incidents, Hindsight provides persistent memory, and the runbook system captures response experience.

The interesting engineering challenge isn't simply making an agent respond to an incident.

It's building the architecture that allows it to remember, retrieve, and use operational experience responsibly.

Top comments (0)