DEV Community

Syed Rafi
Syed Rafi

Posted on

Incident Response Agent — Intelligent Incident Analysis and Response Assistance

At 3 AM, Similarity Isn't Enough: Building an Incident-Response Workflow That Explains Its Recommendations
When an incident starts at 3 AM, the difficult part is often not recognizing that something is broken. The difficult part is deciding what to do next.
A familiar error can have several possible causes. A runbook may have worked before, but under different conditions. An incident from six months ago may look relevant while being too stale to trust blindly.
That was the problem we wanted our Incident Response Agent to address.
Instead of treating incident history as a searchable archive, the system turns historical incidents into structured evidence. A new incident can be compared against resolved incidents, the similarity can be broken down into understandable dimensions, previously used runbooks can be surfaced, the current response can be recorded as a timeline, and the eventual result can influence how much trust the system places in that runbook later.
The interesting part is that this does not require a large language model or an external API. The implementation runs locally using deterministic components.
The problem: incident history is only useful if responders can act on it
An incident database can contain years of valuable operational knowledge and still be difficult to use during an outage.
Suppose payments-api starts returning upstream 502 errors. A responder may want answers to several questions immediately:

  • Have we seen this before?
  • Was the previous incident on the same service?
  • Was the error signature identical?
  • Was it the same severity?
  • What runbook did the previous responder use?
  • Did that runbook actually work?
  • How recently was that runbook validated?

A simple full-text search answers only part of this.
Our system instead models the matching process explicitly.
The similarity engine compares a new incident against resolved incidents and produces an incident fingerprint containing five dimensions:

  1. Text similarity
  2. Service match
  3. Error-signature match
  4. Severity match
  5. Recency

This is important because the final score is not presented as the entire explanation.
A responder can see why an incident was considered relevant.
From a new incident to historical precedents
The search flow begins in the /api/search/suggest endpoint.
The endpoint retrieves resolved incidents and passes the new incident's information into find_similar_incidents().
The matching engine uses TF-IDF and cosine similarity over incident information including the title, description, service, error signature, and root-cause summaries.
The core implementation looks like this:

python
WEIGHTS = {
"text_similarity": 0.40,
"service_match": 0.20,
"error_signature_match": 0.20,
"severity_match": 0.10,
"recency": 0.10,
}
This gives text similarity the largest contribution, while still allowing operational attributes to influence the result.
For example, an incident on the same service with the same error signature receives explicit matches for those dimensions instead of depending entirely on vocabulary overlap.
The result is an overall score plus the underlying fingerprint.
That distinction matters.
A score such as 0.82 tells me that the system thinks two incidents are similar. The fingerprint tells me whether that similarity comes from the same service, the same error signature, matching severity, recent history, or simply similar language.
The system deliberately avoids a black-box similarity explanation
One of the implementation decisions I found particularly useful is the decision to keep the similarity dimensions separate.
The fingerprint generated by the matcher contains:

python
fingerprint = {
"text_similarity": round(float(text_sim), 3),
"service_match": service_match,
"error_signature_match": error_match,
"severity_match": severity_match,
"recency": round(recency, 3),
}
The API then uses those values to construct a plain-English explanation.
For example, a recommendation can explain that the historical incident was on the same service, had an identical error signature, had the same severity class, and was previously associated with a particular runbook.
That gives the responder something much more useful than an unexplained ranking.
Similar incidents are only useful if their fixes are useful
Finding an old incident is only half the workflow.
The next question is:
What did the responder actually do?
The project connects historical incidents to resolution steps and runbooks. When a matching incident contains a runbook reference, the search endpoint surfaces that runbook as a recommendation.
The system also sorts recommended runbooks by their current trust rating.
That introduces another interesting design problem.
A runbook used successfully twice is not necessarily equivalent to one that has repeatedly succeeded against serious incidents.
Treating runbook trust like an Elo rating
The project uses an Elo-style rating system rather than a simple success percentage.
Every runbook starts at:

python
STARTING_RATING = 1200.0

Incident severity is treated as the strength of the opponent:

python
SEVERITY_DIFFICULTY = {
"SEV1": 1600,
"SEV2": 1400,
"SEV3": 1200,
"SEV4": 1000,
}

A successful runbook against a SEV1 therefore provides stronger evidence than a successful runbook against a SEV4.
The feedback endpoint records whether the runbook:

  • worked,
  • partially worked, or
  • failed.

The rating is then updated and the usage is persisted.
This creates a feedback loop:
incident → historical match → runbook recommendation → outcome → updated trust → future recommendation
That is more interesting than simply storing whether someone clicked "worked."

The incident itself becomes a timeline
The workflow does not stop after a recommendation.
Responders can record resolution steps through:

text
POST /api/incidents/{incident_id}/steps
Each step can include information such as:

  • minute offset,
  • actor,
  • step type,
  • action text,
  • associated runbook.

This creates what the project describes as a "black-box replay" of the incident.
That matters for two reasons.
First, it gives the responder a structured record of what happened.
Second, it creates information that can be used after the incident is resolved.

The system calculates MTTR when the incident is resolved:
python
delta = incident.resolved_at - incident.started_at
incident.mttr_minutes = max(
1,
int(delta.total_seconds() // 60)
)

The response therefore becomes more than a temporary conversation. It becomes structured incident history.
Closing the loop with a postmortem
Once an incident is resolved, the system can generate a postmortem draft.
Importantly, this is **not an LLM-generated postmortem.
The repository's implementation builds the draft deterministically from the incident's stored information, root causes, and timeline.
For example, the postmortem generator can identify the first mitigation step and turn it into a lesson about whether that action could have happened earlier.
It can also generate action items based on root-cause categories such as deployment, capacity, dependency, configuration, and human error.
This is a deliberate engineering choice.
For incident response, a deterministic first draft has an advantage: its content comes from recorded incident data rather than from an external model inventing plausible-sounding details.
It is still a draft. The repository explicitly positions it as something that should support, rather than replace, human review.
What this looks like during a recurring incident
Consider a new payments-api incident with:

text
Description:
payments-api is returning 502 upstream timeouts,
and the connection pool looks saturated.

Service:
payments-api

Severity:
SEV1

Error signature:
502-upstream-timeout

The project's tests verify that the corresponding historical payments incident ranks first when the same service and error signature are supplied.
The resulting fingerprint can expose:

  • matching service,
  • matching error signature,
  • matching severity,
  • textual similarity,
  • recency.

The associated runbook can then be surfaced with an explanation of why the historical incident matched.
The responder can follow the runbook, record the actions taken, resolve the incident, generate the postmortem, and finally report whether the runbook worked.
That last step changes the future.
The next responder does not inherit exactly the same static recommendation. The runbook's rating has been updated using the new evidence.

What I learned from this implementation:
The biggest lesson is that incident memory is more useful when it is structured around decisions.
Simply remembering that an incident happened is not enough.

The system needs to preserve relationships:
incident → root cause → response step → runbook → outcome**
The second lesson is that explainability can be built into the data model instead of being added later as a presentation layer.
The similarity engine does not only return a number. It returns the dimensions behind the number.
The third lesson is that historical information has a shelf life.
The project implements rating decay toward the 1200 baseline when a runbook has not been used recently. The intention is straightforward: infrastructure changes, dependencies change, and an old successful procedure should not automatically retain the same level of trust forever.

What the project does not solve

There are important limitations.
The similarity system uses TF-IDF rather than semantic embeddings. That means two incidents describing the same underlying problem with very different vocabulary may not match strongly.
The runbook rating is also not a scientific measure of correctness. It is an operational trust signal based on the outcomes recorded by the system.
The postmortem generator is deterministic and template-driven. It cannot independently investigate an incident or discover facts that were never recorded.
Finally, the application currently uses SQLite. The repository describes this as appropriate for a team or demonstration environment, with a production-scale deployment potentially moving to PostgreSQL.
These limitations are useful because they define exactly where the current system ends.

One important architectural clarification

The original project brief called for a Hindsight-based memory story. However, the uploaded repository does not demonstrate an actual Hindsight SDK, retain/recall API, or Hindsight service integration.
The implemented system instead provides its own persistent incident-history mechanism through SQLAlchemy/SQLite, similarity search, runbook usage records, and postmortem data.
That distinction matters.

It would be inaccurate to describe TF-IDF incident retrieval as Hindsight memory simply because both systems involve historical information.
The current implementation is its own deterministic incident-memory workflow.

Conclusion

The most useful part of an incident-response system may not be the ability to find an old incident.
It is the ability to turn that old incident into evidence that can be inspected, acted on, and eventually updated.

Our Incident Response Agent connects those stages:
find a precedent → understand why it matches → inspect its runbook → record the current response → measure the result → update future trust**
For me, that is the core engineering lesson from the project.
During an outage, historical knowledge should not just be available.
It should be structured enough to help someone decide what deserves attention next — while still leaving the final decision with the engineer handling the incident.

Top comments (0)