Introduction
Giving an AI agent long-term memory is powerful, but it introduces a risk that stateless
assistants do not have: the memory itself can be wrong.
In incident response, that matters.
A confident but incorrect recommendation delivered during an outage can waste valuable
time. If that incorrect recommendation is then saved as historical knowledge, the problem
becomes larger: the next incident can retrieve the same mistake and treat it as evidence.
That led us to a simple design principle for our Incident Response Agent:
The agent can suggest a fix, but only an engineer can turn a fix into memory.
The distinction between prediction and confirmed experience is one of the most important
parts of the system.
The risk of learning from guesses
Imagine an incident agent receives a payment gateway timeout.
It analyzes the incident and recommends restarting the payment service.
The engineer tries the restart.
It does not solve the problem.
If the agent automatically stores its original recommendation, the memory bank now contains
a false lesson: restarting the payment service is associated with that failure even though it
did not actually resolve it.
The next time a similar incident occurs, the system recalls that memory.
Because the information came from the team's historical memory, it may appear more
authoritative than a new suggestion.
The engineer follows it again.
Now the system has created a feedback loop in which an unverified suggestion becomes
increasingly difficult to distinguish from an actual operational fact.
We designed our workflow specifically to avoid that.
Our principle: memory comes from
confirmed experience
The agent separates suggestion from fact.
During analysis, the system can recall previous incidents and generate a recommended
response.
That response remains a recommendation.
After the incident is resolved, the engineer gets a separate opportunity to record what
actually happened.
The Record Actual Resolution step asks two required questions:
What actually fixed the incident?
What was the outcome?
Both fields are required before the experience can be saved.
The interface makes the rule explicit:
Only confirmed resolutions become memory.
The agent's own analysis is never automatically stored as fact.
Why this distinction matters
This creates a clean boundary in the system.
Before resolution:
AI-generated information = hypothesis
After engineer confirmation:
Observed resolution = experience
That distinction gives the memory bank a much clearer meaning.
A recalled memory should not mean:
The model once suggested this.
It should mean:
This is what happened during a previous incident, according to the engineer who resolved it.
That is a much stronger foundation for future recommendations.
How the workflow works
The process begins with normal incident analysis.
An engineer reports the incident. Hindsight recalls similar incidents. The agent uses those
memories to produce a recommended response.
At this stage, the current recommendation is not automatically treated as historical truth.
The engineer follows the relevant steps and observes the result.
Once the incident is actually resolved, the engineer records the confirmed resolution and
outcome.
Only then does the system retain the experience in Hindsight.
For example:
Suggested response: Restart the payment service and verify gateway connectivity.
Observed outcome: The payment service restart restored successful payment requests and
gateway connectivity was verified.
The second statement is what becomes reusable experience.
Grounded memory
This design gives us what we call grounded memory.
Every stored record represents something an engineer verified rather than something the
model predicted.
That does not mean a stored record can never become outdated. Production systems
change, dependencies change, and an old fix may eventually stop applying.
But it establishes a much better starting point for future incidents.
When a future incident recalls a previous resolution, the engineer knows that the original
resolution was confirmed during an actual incident.
Recommendations remain
auditable
Another consequence is that recommendations can be connected to their source.
The Incident Response Agent shows the previous incident that influenced a recommendation
and its match score.
For example, a future incident might recall:
1048 — 94% match
1042 — 89% match
Both incidents contain evidence about the payment gateway failure and its resolution.
This lets an engineer inspect the basis for a recommendation instead of treating the output
as an unexplained instruction.
The source incident matters because similarity alone does not prove that two incidents are
identical.
A 94% match is evidence to investigate, not permission to stop thinking.
Repeated confirmation becomes
useful evidence
The design also creates an interesting effect when the same solution is confirmed across
multiple incidents.
Suppose incident #1042 shows that restarting the payment service worked.
Later, incident #1048 produces the same result.
A subsequent incident can now recall both experiences.
For incident #1051, the system can identify that the same fix was confirmed twice, based on
1048 and #1042.
That is more useful than simply having one old recommendation.
The system has accumulated repeated operational evidence.
Importantly, the evidence comes from separate resolved incidents rather than from the agent
repeatedly copying its own suggestion.
Humans stay in control
The confirmation step also keeps the engineer in the loop.
The agent proposes.
The engineer investigates.
The engineer decides what actually fixed the incident.
The engineer records the outcome.
Only then does the system remember it.
This is deliberately different from designing an agent that automatically converts every
output into a permanent instruction.
During an outage, engineers need speed, but speed does not remove the need for
verification.
A short confirmation step at the end of an incident is a small amount of friction compared
with allowing incorrect memories to propagate through future incidents.
Trust under pressure
On-call engineers make decisions with incomplete information and limited time.
They are unlikely to trust an automated recommendation simply because it sounds confident.
The system therefore tries to make recommendations inspectable.
The engineer can see:
which incident was recalled,
how closely it matched,
what the confirmed resolution was,
and which runbook was associated with it.
That evidence gives the engineer something concrete to evaluate.
They can follow the recommendation, adapt it, or reject it.
The system is useful without requiring the engineer to surrender judgment.
What we learned
- Persistent memory changes the safety problem A wrong answer from a stateless assistant disappears when the conversation ends. A wrong answer saved into long-term memory can influence future incidents.
- Separate hypotheses from facts The agent can be generous with suggestions while being conservative about what it stores.
- Confirmation should happen at the right moment The engineer already knows what worked when the incident is resolved. Asking for a short confirmation at that point makes retention practical.
- Memory needs provenance A recommendation becomes easier to evaluate when the system can show where it came from.
- Repeated experience is valuable When separate incidents confirm the same resolution, future recommendations can draw on more than one piece of evidence. Limitations This approach does not guarantee that every memory will remain correct forever. A confirmed resolution can become outdated when infrastructure changes. A similar incident can also have a different root cause. That is why memory should be treated as operational evidence rather than as an unquestionable command. The engineer still needs to evaluate the current incident. Conclusion A memory system should not remember everything an AI says. For our Incident Response Agent, memory is deliberately narrower: it stores confirmed operational experience. Hindsight provides the memory layer, but the workflow determines what is allowed into that memory. The agent proposes. The engineer verifies. The confirmed resolution becomes reusable experience. That boundary is what makes persistent memory useful without turning every model prediction into institutional knowledge
Top comments (0)