**
Introduction
**
The Incident Response Agent is a web application that gives an on-call team a memory. It
uses Hindsight as its memory layer: past incidents are retained in a shared memory bank
called sre-incidents and recalled when a new incident looks similar.
The problem is familiar to anyone who has worked on production systems. An alert tells us
something is broken, logs tell us what is happening now, and dashboards show the impact.
But none of those tools necessarily tells us what worked the last time the same failure
happened.
That gap is where our agent operates.
Instead of treating every incident as a completely new problem, the system follows a
simple loop:
Report → Recall → Respond → Resolve → Retain
The important part is that memory is not simply displayed as historical information. It
changes what the agent can recommend when the next incident arrives.
**
Step 1: Report the incident
**
The workflow begins in the Report Incident screen.
An engineer enters a short description of what is happening, including the incident title,
affected service, error description, severity, and detection time.
For example:
Incident: Payment service failure
Service: Payment Gateway
Error: Users are unable to complete payments and receive 504 Gateway Timeout
errors
Severity: Critical
The objective is to give the system enough information to identify the type of failure
without requiring the engineer to write a complete incident report before receiving help.
Once the engineer clicks Analyze Incident, the system begins its analysis workflow.
**
Step 2: Recall from Hindsight
**
This is where the behavior differs from a memoryless assistant.
Instead of treating the current incident as an isolated prompt, the system uses the incident
information as a recall query against the sre-incidents memory bank.
In the example workflow, three memories are returned:
The ranking matters.
The system does not simply return every incident stored in memory. The most relevant
previous experience is surfaced first.
Incident #1042 is particularly useful because its symptoms and service context closely
resemble the current failure.
This gives the agent something a generic troubleshooting assistant does not have: a
concrete example of how the team previously handled a similar incident.
**
Step 3: Turn memory into a
response
**
Finding a previous incident is only useful if the engineer can act on it.
The Incident Memory screen surfaces the important information from incident #1042:
Root cause: Payment Gateway API timeout
Successful resolution: Restart the payment service and verify gateway connectivity
Runbook: Payment Gateway Recovery
Match: 92%
The system then uses Hindsight reflect to turn the recalled information into a
recommended response.
The recommendation in the example workflow is:
- Check gateway connectivity.
- Restart the payment service.
- Verify the API response.
- Monitor the error rate. This is intentionally different from returning a long list of generic troubleshooting possibilities. The recommendation is connected to a previous incident and its observed resolution. **
Step 4: Put the response into a
runbook
**
A recommendation is useful, but during an outage an engineer needs something they can
execute.
The Payment Gateway Recovery runbook turns the response into a checklist.
The engineer can open the runbook and work through the steps while investigating the
current incident.
This creates a useful separation:
Memory provides context.
The runbook provides execution steps.
The recalled incident explains why the response is relevant. The runbook turns that
response into an operational procedure.
That makes the information easier to use when the engineer is working under pressure.
**
Step 5: Resolve the incident
**
The agent does not decide that an incident is resolved simply because its recommendation
was followed.
Once the engineer completes the recovery process, they record the actual outcome.
For example, the engineer can confirm that the payment service was restarted and
gateway connectivity was verified.
This distinction matters because there is a difference between:
What the agent predicted would work
and
What actually worked.
The current incident should only become useful historical knowledge after the engineer
confirms the result.
**
Step 6: Retain the new
experience
**
Once the resolution is confirmed, the incident can be saved back into Hindsight.
The retained experience contains information such as the incident, root cause, successful
fix, observed outcome, and runbook.
This closes the loop.
The next similar incident can now potentially recall not only incident #1042, but also the
newly resolved incident.
The memory bank therefore becomes a growing record of operational experience.
**
Why shared memory matters
**
We designed the memory around the team rather than around one engineer.
A production incident does not necessarily happen when the engineer who solved the
previous incident is online.
If the solution exists only in someone's memory or an old chat thread, the next engineer
may have to rediscover it.
A shared memory bank makes the experience available to the next person handling the
service.
This matters across shift changes and team turnover. The system does not need to know
who originally solved the incident. It needs to know what happened, what was tried, and
what actually worked.
**
Why we show the source
**
We also wanted recommendations to be inspectable.
When the system recommends a response, the engineer can see which previous incident
influenced the recommendation and how strongly the current incident matched it.
Instead of simply seeing:
Restart the payment service.
the engineer can see that the recommendation came from a previous payment gateway
incident with a 92% match.
That does not guarantee that the recommendation is correct. Similar incidents can still
have different causes.
But it gives the engineer evidence they can evaluate before taking action.
**
What changes when memory is
added?
**
The biggest change is not that the agent suddenly knows every possible solution.
The change is that the agent no longer has to start every incident from zero.
A memoryless assistant can provide a reasonable troubleshooting checklist.
A memory-enabled assistant can additionally ask:
Have we seen something like this before, and what actually worked?
That changes the starting point for the investigation.
In our incident scenario, the first occurrence takes 47 minutes because the team has to
work through possible explanations before reaching the useful restart. When a similar
incident is later recalled, the engineer can begin with a previously confirmed recovery
path.
The value of memory is therefore not just information retrieval. It changes the sequence of
decisions made during an incident.
**
Lessons learned
**
1. Recall needs to lead somewhere
Retrieving similar incidents is not enough. The recalled information needs to become an
actionable response.
2. Evidence is more useful than
generic advice
A previous incident provides context that a general troubleshooting checklist cannot
provide.
3. Memory should be shared
Operational knowledge becomes more useful when it is available to the next engineer
rather than remaining with the person who originally solved the problem.
4. Retention closes the loop
If the system only recalls old incidents, its knowledge eventually becomes stale. Retaining
confirmed resolutions allows today's incident to become tomorrow's evidence.
5. Engineers remain responsible
The system proposes a response, but the engineer verifies the incident and its outcome.
The agent supports the decision rather than replacing it.
**
Conclusion
**
- The Incident Response Agent follows a simple idea: production experience should be
- available when it is needed.
- Hindsight provides the memory layer that connects today's incident to yesterday's
- experience. Recall finds relevant incidents, reflect helps turn those experiences into a
- response, and retention makes confirmed outcomes available for future incidents.
- The resulting loop is simple:
- Report → Recall → Respond → Resolve → Retain
- The interesting part is not simply that the agent can remember.
- It is that memory changes what happens the next time the same failure appears
Top comments (0)