Production incidents are rarely completely new.
A service may fail because of a connection pool, an expired certificate, a missing database index, or an upstream dependency. The difficult part is not always identifying a possible fix. It is knowing what has actually worked before — and what engineers should avoid trying first.
That was the problem I worked on while building the recommendation engine for IncidentRecall, an AI incident response agent designed around Hindsight memory
My main contribution was the LLM and recommendation layer: turning retrieved incident memories into structured, evidence-grounded recommendations instead of allowing the model to generate generic troubleshooting advice.
The Problem: AI Shouldn't Start From Zero
A typical incident-response assistant receives the current symptoms and generates a response based on the information in the current incident.
That can be useful, but it misses something important: organizational experience.
If engineers previously faced similar incidents, their successful and failed actions are valuable information.
For example:
A restart may have failed before.
Increasing replicas may have made the situation worse.
Increasing a connection pool may have solved a similar incident.
Updating a certificate may have resolved another incident.
The recommendation engine therefore needed to answer a different question:
What happened when we faced something similar before?
And more importantly:
What should we avoid repeating?
Where My Contribution Fits
IncidentRecall has multiple parts working together.
The incident is first analyzed by the backend. The memory layer provides relevant historical memories, and in the live architecture this layer is powered by Hindsight. My part begins when those memories are available.
The recommendation engine takes:
- The current incident
- Retrieved historical memories
- Successful actions
- Failed actions
- Relevant incident IDs
- Evidence supporting each recommendation
and turns them into a structured recommendation.

Figure 1 — Incident creation screen
This is where a new incident enters the system. The current incident contains information such as the affected service and symptoms, which becomes the basis for memory retrieval and recommendation generation.
Using Hindsight as the Evidence Layer
The most important design decision was to keep the recommendation grounded in retrieved memory.
Instead of allowing the model to invent historical incidents, the recommendation engine works with the memories returned by the memory layer.
The core recall operation is conceptually:
results = _hindsight.recall(
bank_id=BANK_ID,
query=query,
budget="low",
max_tokens=2000,
)
return results.results
These retrieved memories contain the previous incident context and the actions that were attempted.
The recommendation engine can then reason over that evidence.
The Memory Inspector makes the retrieved incident memories visible instead of hiding them behind the AI response. This gives an engineer a direct view of the historical context supporting the recommendation.

Figure 3 — Retrieved memory details
The detailed memory view shows the information available to the recommendation engine, including previous incident IDs and the actions associated with those incidents.
Turning Memories Into Recommendations
The recommendation engine does not simply return the retrieved memories.
Its job is to transform them into something useful during an incident.
The output focuses on questions such as:
What action is recommended?
Which historical incidents support it?
What worked previously?
What failed previously?
What should the engineer avoid trying first?
How strong is the supporting evidence?
This makes the response more actionable than a generic troubleshooting checklist.
For example, if several historical incidents show that restarting a service did not solve the problem, while increasing a connection pool did, the recommendation can surface that distinction.
Teaching the Agent What to Avoid First
One of the most useful parts of the recommendation engine is the “Do Not Try First” section.
Traditional troubleshooting systems usually focus on what to do.
IncidentRecall also learns from what didn't work.

Figure 4 — Agent Recommendation and “Avoid First”
This is one of the most important parts of the interface.
Instead of only saying:
Increase the connection pool.
the agent can also provide evidence that previous restart attempts failed in similar incidents.
That changes the role of memory from simply being a source of suggestions to being a source of decision support.
An engineer does not have to blindly follow the recommendation. They can see why the system suggested it.
Keeping Recommendations Evidence-Grounded
A recommendation is much more useful when the engineer can inspect the evidence behind it.

Figure 5 — Historical matches and recommendation evidence
The historical matches connect the recommendation back to previous incidents.
This was an important design requirement for me: the model should only reference incident IDs and historical outcomes that are supported by the retrieved evidence.
The recommendation therefore separates the current incident from the historical evidence supporting the recommendation.
This also makes the system easier to trust and debug.
What Happens Without Memory?
There is an important difference between an AI assistant that only sees the current incident and one that can use organizational memory.
Without historical memory, the assistant can still produce general troubleshooting advice.
In this demonstration, the historical memory bank is intentionally small and deterministic so that the learning loop can be reproduced consistently. In a live deployment, the memory bank would grow as real incidents and engineer feedback are retained.
But it cannot know whether:
a restart already failed several times,
a particular configuration change worked previously,
scaling caused problems in a similar incident,
or another engineer solved the same type of problem differently.
With memory, the recommendation becomes grounded in previous experience.
That is the main idea behind IncidentRecall:
Don't just ask AI what might work. Ask what worked when we faced something similar before.
The Recommendation Is Not the End of the Loop
Another important part of my work was making the recommendation system capable of learning from engineer feedback.
The engineer can tell the system what actually happened after applying the recommendation.
The feedback interface allows the engineer to record whether the suggested action worked or failed.
This adds the outcome of the current incident to the historical context available to the system.
From Feedback to New Memory
When an engineer marks an action as successful, that outcome can be retained as a new memory.

Figure 7 — Memory Updated / New Incident
This closes the learning loop:
Incident → Recall → Recommendation → Engineer Action → Feedback → New Memory
The next incident can then benefit from what was learned previously.
This feedback loop allows previous incident outcomes to become part of the context used for future recommendations.
In a live deployment, the memory bank can grow as engineers interact with the system and new incident outcomes are retained.
Designing the Recommendation Engine
While building this component, I focused on three principles.
- Evidence before confidence
The model should not sound confident simply because it can generate a convincing answer.
The recommendation should be supported by retrieved incident memories.
- Failed actions are also valuable
A failed action is not useless information.
Knowing that something failed previously can prevent an engineer from wasting time repeating the same step.
That is why the recommendation includes an Avoid First concept.
- Feedback should become useful information
The engineer's final outcome should not disappear after the incident is resolved.
A successful or failed action can become part of the system's future memory.
Demo Mode and Real Integration
For reproducible development and demonstration, the project currently supports a local demo mode.
The demo mode uses a deterministic representation of the historical incident dataset, so the application can be demonstrated without requiring external API credentials.
The architecture still separates the memory and LLM layers so that the recommendation engine can work with the real Hindsight/Groq integration when live credentials are configured.
This distinction is important: the local demo is intentionally deterministic rather than pretending that every demonstration request is using a live external model.
What I Learned
Building the recommendation layer changed how I think about AI agents.
A model generating a technically reasonable answer is not enough for an incident-response system.
The more important question is:
What evidence does the agent have for this recommendation?
Memory makes the answer more useful because it gives the model access to previous experience.
The combination of successful and failed actions is particularly valuable. A previous failure can be just as important as a previous success when an engineer is deciding what to try next.
I also learned that making the evidence visible matters. Showing the historical matches allows the engineer to understand where the recommendation came from instead of treating the model as a black box.
Final Thoughts
The goal of IncidentRecall is not to replace engineers during incidents.
It is to make previous engineering experience easier to retrieve and use.
The recommendation engine sits between memory and action:
Historical experience → Retrieved evidence → Recommendation → Engineer feedback → New memory
That loop gives the agent a way to incorporate previous incident outcomes into future responses.
For me, the most interesting part was not making an AI generate another troubleshooting answer.
It was making the answer say:
“Here is what happened before, here is what worked, here is what failed, and here is what you should consider next.”
That is where an AI assistant starts becoming an experience-driven incident-response system.


Top comments (0)