The hardest part of incident response is not finding a plausible fix. It
is avoiding a fix that already failed the last time the system broke.
I built an incident-response agent around that problem. It collects
current telemetry, investigates the likely root cause, retrieves
relevant operational history through Hindsight, proposes a recovery
action, requires approval where appropriate, verifies the result, and
retains the outcome so a later incident can use it. The interesting part
is not that the agent can recommend restart\_service. The interesting
part is that, after an engineer has shown that restarting was the wrong
response, the next similar incident can stop making the same suggestion.
The system I built

Figure 1: The Incident Intelligence workspace exposes website auditing,
incidents, errors, services, memory, and runbooks in one operational
view.
I think about the system as an incident loop rather than a chatbot:
Incident
↓
Evidence collection
↓
Investigation
↓
Hindsight recall
↓
Failure + correction history
↓
Recommendation
↓
Human approval
↓
Allowlisted action
↓
Verification
↓
Postmortem
↓
Hindsight retain
The repository separates the memory layer from the incident agent.
HindsightMemoryEngine owns operational memories, while IncidentAgent
uses those memories during investigation and remediation.
The memory model is deliberately more structured than a generic
conversation history. Memories have operational types such as
INCIDENT, RCA, ACTION, FAILURE, CORRECTION, OUTCOME,
VERIFICATION, RUNBOOK, and PREVENTION. That matters because
incident response is not only about remembering what happened. I also
need to remember what I tried, what failed, what a human operator
corrected, and what eventually worked.
That is the role I wanted Hindsight to play: durable operational
experience that can influence a later decision.
For background on the underlying approach, I used Hindsight agent
memory as the mental model
for treating memory as part of an agent's decision process rather than
simply storing old conversations. The Hindsight GitHub
repository and Hindsight
documentation are also useful
references for the memory architecture and APIs.
The incident where a restart was the wrong answer
The clearest example in the system is database connection exhaustion in
the Payment Service.
The current incident contains evidence such as a database connection
pool approaching saturation and elevated Payment Service errors. The
investigation identifies the likely root cause as connection exhaustion
caused by connections not being returned correctly in a Payment Service
release.
A stateless incident agent can make a reasonable recommendation from
that evidence:
Database connections are high.
Payment Service is failing.
Restart the Payment Service.
There is nothing obviously absurd about that recommendation. Restarting
a service is a common operational response.
The problem is that the repository also contains a previous incident
where that class of response was insufficient.
The historical record captures a failed remediation: restarting the
Payment Service did not remove the underlying database connection
pressure. It also captures a human correction explaining why. The safer
response was to reduce connection pressure or roll back the problematic
release, followed by verification.
That distinction is exactly why I did not want memory to be a passive
incident archive.

Figure 2: A website audit result showing concrete findings that feed
incident investigation.
Hindsight sits before the recommendation
The important code path is in the incident investigation. In
memory-enabled mode, the agent queries operational history using the
current scenario and relevant remediation terms.
Conceptually, the flow looks like this:
if self.memory\_mode == "WITH\_MEMORY":
memories = self.hindsight.recall(
f"{scenario\_type} connection leak payment service "
"restart rollback"
)
failures = self.hindsight.search\_failures(
"payment service restart connection"
)
corrections = self.hindsight.search\_corrections(
"payment service restart"
)
The exact implementation in the repository is intentionally focused on
operational retrieval rather than generic chat history. The recall path
scores relevant memories using factors such as content and title
overlap, incident type, and memory category. Failure and correction
memories receive additional importance.
That last part is an important design choice.
If I retrieve ten old incidents and treat all of them equally, I may end
up copying historical actions without understanding their outcomes. A
previous successful action is useful, but a previous failed action can
be more important when deciding what not to do.
The agent therefore builds a past-versus-present view rather than
blindly copying an old incident.
Previous incident
-----------------
DB connections: high
Deployment: Payment Service v2.5
Action: restart
Outcome: failed
Current incident
----------------
DB connections: high
Deployment: Payment Service v2.5
Errors: elevated
The comparison gives the agent a reason to question the obvious
remediation.
The recommendation actually changes
This is where Hindsight stops being a feature on a dashboard and becomes
part of the control flow.
The recommendation logic distinguishes the memory-enabled path from the
baseline path. Without relevant historical memory, the baseline can
select:
action = "restart\_service"
With the relevant failure and human correction available, the
recommendation changes toward:
action = "reduce\_connection\_pool"
The important part is not the literal string. It is the causal chain
attached to the recommendation.
The recommendation records that the decision was influenced by a
recalled failed action and a human correction. It also carries expected
outcome, verification conditions, and a rollback plan.
That gives me an audit trail closer to:
Recommended action:
Reduce database connection pressure
Why:
- Current connection pool is saturated
- Similar incident was previously observed
- Restarting the service previously failed
- An engineer corrected the previous remediation
Expected outcome:
Database connection pressure decreases
Verification:
Check connection utilization and Payment Service error rate
Rollback:
Restore the previous safe configuration if verification fails
For incident automation, this is much more useful than a generic
explanation such as "I recommend restarting the service because the
service is unhealthy."
Memory is useful only if it changes behavior
One of the easiest mistakes with agent memory is to measure the wrong
thing.
A Memory page can show dozens of records and still have no effect on the
agent. A search endpoint can return excellent historical incidents while
the recommendation code ignores them. In that system, memory is
decoration.
I designed the important test around behavior instead.
With memory disabled, the same current evidence can lead to the baseline
remediation:
WITHOUT MEMORY
DB connections: 98%
Payment Service errors: elevated
Recommendation:
Restart Payment Service
With Hindsight available:
WITH HINDSIGHT
DB connections: 98%
Payment Service errors: elevated
Previous failure:
Restart Payment Service did not resolve connection pressure
Human correction:
Restarting alone is insufficient
Recommendation:
Reduce connection pressure
and recycle affected workers
That before-and-after is the actual proof point I care about. The agent
has not magically become infallible. It has gained access to experience
that changes one decision.
The action layer stays deterministic
I also did not want the LLM to turn its recommendation directly into
arbitrary infrastructure commands.
The repository uses an allowlist of supported actions:
SAFE\_ACTIONS = {
"restart\_service",
"rollback\_deployment",
"reduce\_connection\_pool",
"scale\_service",
"disable\_faulty\_dependency",
"clear\_application\_cache",
}
The agent can recommend an action, but execution is constrained by the
backend.
That separation is important. Hindsight can influence the reasoning, but
it should not bypass operational controls.
For higher-risk actions, the workflow includes human approval. After
approval, the action engine executes the known operation and the
verification stage checks actual recovery conditions.
This gives me three distinct boundaries:
Memory
→ informs the decision
Agent
→ proposes the decision
Deterministic action layer
→ controls what can actually execute
I prefer this architecture to letting an LLM generate arbitrary commands
because a memory system can be wrong, stale, incomplete, or irrelevant.
Memory should improve reasoning without becoming an authority that
bypasses safeguards.
Verification closes the loop
The other design decision I consider important is that a successful
action is not the same thing as a successful recovery.
After execution, the agent checks the resulting system state. For the
connection-exhaustion scenario, that means looking at the relevant
service and database metrics rather than accepting an execution response
as proof.
The workflow is effectively:
Action requested
↓
Action executed
↓
Metrics collected
↓
Recovery thresholds evaluated
↓
Verified / failed
If verification succeeds, the postmortem can capture the outcome. If it
fails, the system has evidence that the remediation did not solve the
incident.
That outcome then becomes another piece of operational knowledge.
Retain turns one incident into future context
After recovery, the incident agent builds a structured postmortem
containing information such as the incident summary, impact, timeline,
root cause, evidence, failed actions, corrections, successful
remediation, verification, prevention, and lessons learned.
The important operation is not generating the postmortem itself. It is
retaining the useful parts for future incidents.
self.hindsight.retain(
memory\_type="OUTCOME",
title="Payment Service connection recovery",
content=postmortem\_content,
)
That creates a learning loop:
Incident 1
↓
Failure / correction
↓
Hindsight retain
↓
Incident 2
↓
Hindsight recall
↓
Different recommendation
↓
Verification
I like this pattern because it makes the value of memory concrete. The
system does not need to "learn" in the model-training sense to become
more useful. It can preserve operational experience and retrieve it when
the next decision resembles the old one.
What I learned building it
- Failed actions deserve first-class memory Incident systems naturally collect successful runbooks and root-cause summaries. I found the negative information just as important. "Restart worked" tells me what to try. "Restart failed because the underlying connection pressure remained" tells me what not to repeat. For remediation agents, I would explicitly model failures and corrections instead of burying them inside a long postmortem.
- Retrieval should serve a decision It is tempting to optimize memory around search quality alone. I think the better question is: what decision does the retrieved memory change? In this system, the useful retrieval is not "show me similar incidents." It is "show me similar incidents, especially failed actions and human corrections, before I choose a recovery action." That keeps memory retrieval tied to an operational purpose.
- Human corrections are valuable training data without retraining a model An experienced operator rejecting an action is a high-value event. I do not need to immediately fine-tune a model to preserve that lesson. I can retain the correction as structured operational memory and make it available during the next investigation. That is a much shorter feedback loop.
- Memory should not remove safety boundaries Adding Hindsight made the recommendation better informed, but it did not make the agent trustworthy enough to execute anything it generated. I still want allowlisted actions, approval for risky operations, explicit verification, and rollback behavior. Memory improves context. It does not replace controls.
- The real test is the second incident The first incident proves that the system can investigate. The second similar incident proves whether the memory architecture matters. If the second incident produces exactly the same recommendation despite a retained failure and correction, I have built an incident archive, not a learning system. That distinction shaped the architecture more than any individual UI feature. Where I would take it next The repository's current incident model gives me a foundation for extending memory beyond a single service failure. The same structure can support deployment regressions, dependency failures, latency incidents, memory leaks, and recurring operational patterns. The next step I would prioritize is improving how memories are evaluated for relevance and freshness as the operational history grows. More memory is not automatically better memory. Old incidents can become misleading when services, dependencies, architectures, or runbooks change. I would also keep the boundary between remembered experience and current evidence explicit. Historical memory should influence an investigation, not override what the system is actually reporting now. That is ultimately why Hindsight fits this architecture for me. The useful unit of memory is not "the agent remembers a conversation." It is "the agent remembers that this remediation failed, why an engineer rejected it, what eventually worked, and what evidence proved the recovery." For incident response, that is the difference between an agent that can produce an answer and one that can carry operational experience forward.
Top comments (0)