DEV Community

Sanjana Vysyaraj
Sanjana Vysyaraj

Posted on

Building an Incident Response Agent with Hindsight Memory

The Problem

When a production incident occurs, engineers need to move quickly from symptoms to a reliable root cause and a practical remediation. Logs, error rates, latency, and service information provide important evidence, but current evidence is not always enough. A previous incident may already contain the answer: the same service may have experienced similar symptoms, the same failure mode may have occurred before, and an earlier resolution may provide a useful starting point. Our Incident Response Agent was designed to make that historical context part of the investigation flow. Instead of treating every incident as an isolated event, the system can retrieve relevant previous incidents and provide that information to an AI investigation agent.

Our Solution

The solution combines a FastAPI backend, an AI investigation agent powered by Groq, and Hindsight as the incident-memory layer. The backend receives an incident containing fields such as service, severity, symptoms, and logs. During investigation, the agent retrieves relevant historical memories from Hindsight. Those memories are then supplied to the AI model together with the current incident. The investigation agent produces a structured response containing a summary, root cause, evidence, historical matches, recommended action, and confidence. This structure makes the result easier for another service or frontend to consume than an unstructured natural-language response.

How Hindsight Retain and Recall Are Used

Hindsight provides the memory layer for the system. Resolved incidents can be retained with their incident ID, service, severity, timestamp, symptoms, logs, root cause, resolution, and lessons learned. This creates reusable incident knowledge instead of allowing useful troubleshooting experience to disappear after an incident is closed. For a new incident, the system builds a recall query from the current service, severity, symptoms, and logs. Hindsight searches the incident-memory bank for relevant information. The returned memories are converted into historical matches and passed into the AI investigation agent.

A simplified version of the recall implementation is:

result = hindsight.recall(
bank_id=BANK_ID,
query=query
)

historical_matches = []

for memory in result.results:
historical_matches.append({
"id": memory.id,
"text": memory.text,
"type": memory.type,
"context": memory.context
})



Connecting Memory to the Investigation Agent

The integration is handled in the agent layer. When historical memories are not already provided, the investigation function calls recall_incidents(incident). The Hindsight client is then closed after retrieval. The resulting historical memories are passed into IncidentInvestigationAgent along with the current incident. This creates a complete flow: Current incident → Hindsight Recall → historical memories → Groq investigation → structured response. The historical memory is therefore not just displayed to the user. It becomes part of the context used by the investigation agent when producing its root-cause analysis and recommendation.

if historical_memories is None:
try:
historical_memories = recall_incidents(incident)
finally:
close_hindsight()

return agent.investigate_incident(
incident=incident,
historical_memories=historical_memories,
)

A Concrete Before-and-After Example

Consider a payment-api incident with API latency increasing to 5.1 seconds, HTTP 500 errors rising to 41%, and logs reporting “Database connection pool exhausted” and “ConnectionPoolTimeout”. Before historical-memory integration, the investigation could analyze these current symptoms but would not automatically have access to the team's previous incident experience. After Hindsight integration, the system retrieves a previous payment-api incident, INC-001, with similar symptoms and the same underlying database connection-pool exhaustion problem. The agent can use that historical context when producing its recommendation. In the successful test, the agent identified database connection pool exhaustion as the root cause, matched INC-001, recommended increasing the database connection pool and monitoring pool and timeout behavior, and returned a confidence value of 0.92. The behavior changed from investigating only the current evidence to investigating the current evidence together with relevant previous experience.

Why Historical Matches Matter

The historical_matches field makes the memory contribution explicit. A match can contain an incident identifier and a reason explaining why the previous incident is relevant, for example:

{
"incident_id": "INC-001",
"reason": "Same service and similar symptoms/root cause as the current incident."
}

This gives downstream components a compact representation of which previous incidents influenced the investigation. It also makes the result easier to show in an API response or frontend.

A Practical Lesson and Limitation

Historical memory is only as useful as the information stored and retrieved from it. If an incident is retained without meaningful symptoms, logs, resolution details, or lessons learned, future recall results may provide limited value. The quality of the recall query also matters because it determines which aspects of the current incident are used to search memory. The implementation also depends on access to the Hindsight service and available service credits for live memory operations. This makes graceful error handling and clear configuration important for production use. Nevertheless, the implementation demonstrates the value of giving an investigation agent access to persistent operational experience rather than relying only on the current request.

Conclusion

The Incident Response Agent combines current incident evidence with historical operational knowledge. Hindsight Retain provides a way to preserve resolved incident experience, while Hindsight Recall makes that experience available when a new incident is investigated. The Groq agent then uses the combined context to produce a structured root-cause analysis and recommendation. The result is a more connected investigation workflow: incidents are not treated as isolated events, and previous troubleshooting experience can become useful evidence for future investigations.

References

Hindsight GitHub: https://github.com/vectorize-io/hindsight

Hindsight Documentation: https://hindsight.vectorize.io/

Agent Memory: https://vectorize.io/what-is-agent-memory

Top comments (0)