DEV Community

Huda
Huda

Posted on

Incident Response Agent — Intelligent Incident Analysis and Response Assistance

When an incident starts at 3 AM, the first question is rarely “Can we search our database?”

It is usually something closer to:

“Has this happened before, and can I trust what the previous responder did?”

That distinction shaped how I think about our Incident Response Agent.

The project is designed around a simple problem: incident history is valuable, but historical information by itself is not enough. A responder needs useful precedent, some explanation for why that precedent is relevant, and an indication of whether the associated runbook has actually been successful before.

Our implementation approaches that problem without an external LLM or network dependency. The application runs locally using FastAPI, SQLAlchemy, SQLite, scikit-learn's TF-IDF implementation, and a vanilla JavaScript frontend.

The result is less like asking an AI to invent a solution and more like giving an on-call engineer a structured memory of previous incidents.

The problem with ordinary incident history

Imagine that payments-api starts returning upstream 502 errors.

An engineer could search through old incident tickets and perhaps find something similar. But “similar” can mean several different things.

Maybe the old incident affected the same service.

Maybe it had the same error signature.

Maybe it happened at the same severity.

Maybe the description contains similar words.

Or maybe the incident is old enough that the corresponding runbook should no longer be trusted without validation.

A simple keyword search doesn't communicate all of that.

This is why our system doesn't return only a single opaque similarity percentage. It creates an incident fingerprint with five dimensions:

  • text similarity
  • service match
  • error-signature match
  • severity match
  • recency

That makes the result more useful to a human responder because the system exposes some of the reasoning behind its match.

Following the responder's workflow

The application's workflow mirrors what an engineer actually does during an incident.

First, a new incident is created with information such as its description, service, severity, and error signature.

The search endpoint then compares the incident against resolved incidents:

matches = similarity.find_similar_incidents(
    query_text=f"{payload.description}",
    query_service=payload.service,
    query_severity=payload.severity,
    query_error_sig=payload.error_signature,
    candidates=candidates,
    top_k=payload.top_k,
)
Enter fullscreen mode Exit fullscreen mode

The important detail here is that the system isn't searching every record indiscriminately. The search router retrieves incidents whose status is RESOLVED, then sends those candidates into the similarity engine.

For the responder, this creates a much more focused question:

Which resolved incidents look like the one I'm dealing with now?

A concrete example

The repository's tests use a payments incident involving:

502-upstream-timeout

The test query describes payments-api returning upstream 502 errors while the connection pool appears saturated.

The historical candidate contains the same service and error signature.

The similarity test explicitly verifies that this incident ranks first and that both the service and error-signature matches are detected.

That matters because it demonstrates an actual behavior of the implementation rather than a hypothetical example.

The response also carries the fingerprint.

Instead of seeing something like:

Match: 87%

the responder can see evidence such as:

  • same service
  • identical error signature
  • same severity class
  • historical incident title
  • textual similarity
  • recency

The search router even constructs a human-readable explanation:

why="Matched on " + ", ".join(why_bits) + "."
Enter fullscreen mode Exit fullscreen mode

That small design choice is important.

During an incident, a recommendation that cannot explain itself creates another question for the engineer: Why should I trust this?

An explainable match at least gives the responder something concrete to inspect.

Memory is more than storing incidents

One of the interesting parts of this project is that “memory” isn't implemented as a collection of old incident descriptions alone.

The database model connects an incident with:

  • root causes
  • resolution steps
  • runbook usage
  • postmortems

A resolution step records the minute offset, actor, type of action, description, and optionally the runbook that was used.

That means the system can preserve a sequence rather than just an outcome.

For example:

T+0: detection
T+3: diagnosis
T+8: mitigation
T+15: resolution

The frontend describes this as a black-box-style incident timeline.

From an engineer's perspective, that is useful because incident response is inherently temporal. Knowing that a particular fix worked is helpful. Knowing what responders tried before that fix can be even more useful when the next incident is slightly different.

The system also remembers whether a runbook worked

This is probably the most interesting part of the project's approach.

A runbook isn't treated as permanently trustworthy just because someone used it successfully once.

Each runbook starts with an Elo-style rating of 1200.

When a runbook is used, the system records the incident, outcome, previous rating, new rating, and timestamp.

The rating update considers the severity of the incident.

The implementation models SEV1 as a stronger “opponent” than SEV4. Consequently, successfully resolving a high-severity incident provides stronger evidence that the runbook can perform under pressure.

A failure also has context.

The system therefore avoids reducing runbook quality to a simple:

worked 8 / used 10

Instead, it tries to represent the confidence gained from different kinds of incident outcomes.

There is another useful detail: ratings decay toward the 1200 baseline when a runbook hasn't been used recently.

That reflects a practical reality of infrastructure: systems change.

A runbook that successfully handled an incident eight months ago might still be useful, but its historical rating shouldn't necessarily be treated as equally strong forever.

From responder memory to organizational memory

This creates an interesting loop.

A responder receives suggestions from historical incidents.

They use a runbook.

They record what happened.

The incident is resolved.

The system calculates MTTR.

A postmortem can then be generated from the incident's structured timeline and root causes.

Finally, feedback about the runbook changes its rating.

The next responder therefore doesn't start with exactly the same information.

The system has accumulated another piece of evidence.

That is the real value of persistent incident history: not simply remembering that an incident existed, but gradually building a more useful record of what happened and what actually worked.

The postmortem is intentionally not an AI-generated answer

There is an important engineering decision here that I found particularly useful.

The project does not call an external LLM to generate the postmortem.

postmortem_gen.py constructs the draft from the incident's own data.

It uses the recorded severity, duration, error signature, root causes, and response timeline to generate sections such as:

  • summary
  • contributing factors
  • lessons learned
  • action items

The generated postmortem is explicitly treated as a draft.

That is a sensible boundary for incident response. A postmortem contains operational context and requires human judgment, especially around customer impact, blast radius, and blameless framing.

Automating the first draft can remove repetitive writing without pretending that the generated text is the final source of truth.

What happens without this context?

The difference can be thought of as two workflows.

Without historical context:

Incident → investigate → search manually → decide what to try → resolve → document later.

With the project's incident memory:

Incident → retrieve resolved precedents → inspect the similarity fingerprint → review associated runbooks → respond → record timeline → resolve → update runbook evidence → generate postmortem.

The second workflow doesn't remove the engineer from the loop.

That is important.

The system isn't deciding that a particular runbook must be executed. It is organizing evidence so that the engineer can make a faster, more informed decision.

There are real limitations

The implementation is deliberately conservative in several areas.

The biggest limitation is that similarity is lexical rather than semantic. TF-IDF works well when important words overlap, but two incidents describing the same root cause using completely different terminology may not match strongly.

The project itself identifies embeddings as a possible future extension behind the existing similarity interface.

The database is also SQLite. That is practical for the current application and demo, but the README explicitly notes that a larger organizational deployment would require a more scalable database such as PostgreSQL.

And the postmortem generator is template-based. It is deterministic and works offline, but it isn't a substitute for a human review.

These limitations are actually useful because they define where the system's confidence should stop.

One important distinction about Hindsight

The project description that accompanied this work refers to Hindsight, but the supplied repository does not contain evidence of a Hindsight SDK, dependency, configuration, or retain/recall calls.

The persistent-memory behavior demonstrated by this ZIP is implemented using the application's own SQLite/SQLAlchemy data model and deterministic similarity engine.

So I would not describe the current repository as a Hindsight integration without additional project files proving that connection.

That distinction matters in technical writing. A good engineering article should describe what the code actually does, not what we wish the architecture looked like.

What I learned from the responder's perspective

The biggest lesson for me is that incident memory is only useful when it helps answer three questions:

Have we seen something like this before?

Why does the system think it is similar?

Can I trust the historical fix?

Our implementation tries to answer those questions with concrete data rather than an unexplained AI score.

It stores the incident.

It preserves the response timeline.

It connects incidents to runbooks.

It records outcomes.

It adjusts runbook trust.

It accounts for freshness.

And it turns the accumulated information into something the next responder can actually inspect.

That changes the meaning of “we've seen this before.”

Instead of being a vague memory from an old incident ticket, it becomes structured evidence that another engineer can investigate before taking action.

For incident response, that distinction is small in code but significant in practice.

Top comments (0)