DEV Community

Cover image for I Used Hindsight to Turn Incident Memory Into Evidence
Chenna Keerthana
Chenna Keerthana

Posted on

I Used Hindsight to Turn Incident Memory Into Evidence

I Used Hindsight to Turn Incident Memory Into Evidence

An incident-response agent can retrieve an old incident. The harder question is what it should actually learn from that incident.

While building IncidentIQ, I wanted Hindsight to do more than provide historical context. I wanted the system to turn past incidents into evidence that could influence what it recommends during the next incident.

That became the core of my implementation.

Current Incident → Hindsight Recall → Historical Evidence → Recommendation → Outcome → New Memory

IncidentIQ landing page


The Problem With Simply Remembering Incidents

An incident usually starts with a few pieces of information:

  • Service
  • Severity
  • Alert
  • Logs and observations

For example:

Service: payments-api

Severity: CRITICAL

Alert: HTTP 503 surge

Logs: Database connection pool exhausted

An LLM can analyze this information and suggest possible actions.

But the current incident is only part of the story.

The same service may have experienced similar failures before. Engineers may already have tried different fixes. One might have worked permanently, another might have provided temporary relief, and another might have failed completely.

If that history is ignored, the agent starts every incident almost from scratch.

I wanted IncidentIQ to follow a different path:

Current Incident → Recall Similar Incidents → Analyze Root Causes → Compare Resolutions → Use Outcomes as Evidence → Generate Recommendation

The important part is the middle: historical memory needs to become useful evidence.

Recent investigations


Where Hindsight Fits

I kept the Hindsight integration deliberately focused.

IncidentIQ uses Hindsight to retain incident outcomes and recall relevant historical memories when a new incident is analyzed.

The memory layer follows a simple pattern:

def store_memory(content: str):
    return client.retain(
        bank_id=HINDSIGHT_BANK_ID,
        content=content
    )


def search_memory(query: str):
    return client.recall(
        bank_id=HINDSIGHT_BANK_ID,
        query=query
    )
Enter fullscreen mode Exit fullscreen mode

For a new incident, the backend builds a query from information such as the service, severity, alert, and logs.
The recalled memories are then filtered against the current service before being passed into the analysis flow.
That filtering matters.
A memory about auth-service should not influence a recommendation for payments-api simply because both incidents contain similar words.
Memory Is Not Evidence Yet
This was the part I found most interesting to implement.
Suppose Hindsight returns memories containing:
Resolution attempt: Increase the database connection pool size.
Outcome: Successful.
And another memory says:
Resolution attempt: Restart the payments-api service.
Outcome: Failed.
Both memories are useful, but the recommendation layer needs more structure than a collection of text snippets.
IncidentIQ extracts historical facts from recalled memories:

  • Incident ID
  • Resolution action
  • Outcome
  • Evidence ID I also normalize different descriptions of the same action. For example: def normalize_action(action: str): action_lower = action.lower().strip() if "connection pool" in action_lower: return "increase database connection pool size" if "restart" in action_lower: return "restart service" if "timeout monitoring" in action_lower: return "add connection timeout monitoring" return action_lower

This allows slightly different descriptions of the same operational action to be treated consistently.

A Concrete Example
Imagine the incident history contains several payments-api incidents involving database connection pools.
For example:
Incident Action Outcome
INC-101 Increased connection pool Successful
INC-102 Increased connection pool Successful
INC-103 Restarted service Failed

The system therefore has two different kinds of evidence:

  • Increase connection pool size → successful historical outcomes
  • Restart service → failed historical outcome Now consider a new payments-api incident with connection pool exhaustion. Instead of treating the LLM's recommendation as an isolated answer, IncidentIQ can attach historical information to matching recommendations. The flow becomes: Current incident → Recall related incidents → Extract resolution attempts → Compare outcomes → Calculate historical statistics → Attach evidence → Generate recommendation

That is the difference between simply recalling an incident and actually using incident memory.

Connecting Evidence to Recommendations
The recommendation layer contains fields for historical success rates and evidence IDs.
After the LLM produces recommendations, the backend checks whether a recommendation matches one of the historical actions.
If a match exists, the backend attaches:

  • historical_success_rate
  • evidence_ids The model therefore doesn't have to invent these values. The backend computes the historical statistics separately and connects them to the recommendation. Conceptually: Recommendation → Historical Action → Historical Outcomes → Evidence IDs

This separation is important because historical statistics should come from the stored incident record rather than being guessed by the language model.

The Same Memory Can Help Before Deployment
Once the historical-memory flow worked for active incidents, the same idea could also be used for pre-deployment analysis.
An engineer can provide a service and a proposed change:
Service: payments-api
Proposed change: Increase database connection pool size
IncidentIQ recalls relevant historical incidents and deployments, then uses that context during the pre-deployment analysis.
The result can include:

  • Risk level
  • Risk score
  • Rationale
  • Related incidents
  • Related deployments
  • Safeguards

So the memory layer can support two stages:
During an incident
Historical memory → evidence-backed recommendation

Before deployment
Historical memory → risk assessment

This makes the memory layer part of the overall incident-response workflow rather than an isolated feature.

Closing the Learning Loop
The final part is making sure today's outcome can become tomorrow's evidence.
IncidentIQ records what happened after a resolution attempt.
For example:
Incident: INC-103
Service: payments-api
Resolution attempt:
Restart the payments-api service.
Outcome:
Failed
Engineer notes:
The restart temporarily cleared existing connections, but connection pool exhaustion returned shortly afterward.
That information can then be retained in Hindsight.
The resulting loop is:
Incident → Recall Historical Memory → Analyze → Recommend → Apply Resolution → Record Outcome → Retain Outcome → Future Incident

This creates a path for new operational experience to become future context.

What I Learned

  1. Retrieval Alone Is Not Enough Adding memory to an agent does not automatically make the agent better. The recalled information needs to be transformed into something the rest of the system can reason about. For IncidentIQ, that meant extracting actions, outcomes, incident IDs, and evidence.
  2. Failed Actions Are Valuable A failed resolution can tell the agent what not to repeat. That makes failed outcomes important historical data rather than noise.
  3. Relevance Needs Explicit Handling Not every recalled memory should influence every incident. Filtering memories by service keeps the evidence focused on the system currently being investigated.
  4. Historical Statistics Should Be Computed Separately I didn't want the language model to estimate historical success rates from a long prompt. The backend extracts explicit outcomes and calculates the statistics itself. That makes recommendations easier to trace back to the underlying incidents.
  5. Memory Becomes More Useful When It Closes the Loop The most useful part of the system isn't simply recalling an old incident. It is recording what happened after today's resolution so that the next incident has more context than the previous one did.

The Bigger Idea
The interesting part of adding Hindsight to IncidentIQ wasn't simply giving an SRE agent memory.
It was deciding what to do with that memory.
Historical incidents can become operational evidence when the system can:

  1. Recall relevant incidents.
  2. Extract their actions and outcomes.
  3. Compare successful and unsuccessful approaches.
  4. Connect those outcomes to current recommendations.
  5. Record today's result.
  6. Feed that result back into future analysis. That changes the question from: "What happened before?"

to:
"What happened before, what actually worked, and what evidence supports the action we're considering now?"

For me, that is where incident memory becomes more than stored history.
It becomes part of the reasoning process.
IncidentIQ
IncidentIQ is an AI-powered incident-response system designed around:

  • Incident investigation
  • Historical recall
  • Evidence-backed recommendations
  • Pre-deployment risk analysis
  • Continuous learning

The complete implementation is available on GitHub:
IncidentIQ GitHub Repository
Built Around One Idea
Don't just remember what happened. Learn from what happened.
This project was completed as part of the Code.in program.

Top comments (0)