At 2:00 AM, an incident rarely starts with a completely new problem. More often, the symptoms look familiar: a service suddenly becomes slow, database connections climb, error rates increase, and someone remembers that something similar happened a few weeks ago.
The difficult part is finding that previous incident quickly enough to make it useful.
I built Resilify.AI around this problem. It is an SRE incident investigation and remediation system that combines live telemetry with persistent organizational memory. Instead of treating every production incident as an isolated prompt, Resilify.AI uses Hindsight to recall relevant incidents, reason over their outcomes, and retain what the team learned after recovery.
The design goal was simple:
When the next incident looks like something we've already experienced, the agent should be able to remember what happened last time.
What Resilify.AI Does
The important design decision here is that telemetry becomes a structured incident before it reaches the reasoning layer.
That gives the rest of the system a consistent representation of an outage.
The Interesting Part: Giving the Agent Organizational Memory
A stateless language model can reason about the current incident, but it does not automatically know how the same organization handled similar incidents in the past.
That was the main reason I introduced Hindsight.
I use Hindsight as the persistent memory layer for Resilify.AI. Its documentation describes the memory primitives I use for recalling previous incidents and retaining the results of resolved incidents. I also found Vectorize's explanation of agent memory useful when thinking about how operational experience should become part of an agent's future context.
I use an incident memory bank:
src-incidents-bank
The memory service supports both cloud-backed Hindsight and a local development fallback.
One design decision I found useful was making the storage mode configurable rather than hard-coding a single environment.
The service checks the environment before deciding which implementation to use:
get mode() {
const envMode =
(process.env.HINDSIGHT_MODE || '').toLowerCase();
if (
envMode === 'cloud' &&
process.env.HINDSIGHT_API_KEY &&
!process.env.HINDSIGHT_API_KEY.startsWith('paste_')
) {
return 'cloud';
}
return 'local';
}
This kept local development simple while allowing the same service to connect to Hindsight Cloud.
Recalling Historical Incidents
When a new incident arrives, the agent searches the historical memory for relevant experiences.
The cloud path sends the incident query to the Hindsight memory bank:
if (this.mode === 'cloud' && this.apiKey) {
try {
const response = await fetch(
${this.apiUrl}/v1/banks/${bank}/recall,
{
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': Bearer ${this.apiKey}
},
body: JSON.stringify({
query,
limit: options.limit || 3
})
}
);
if (response.ok) {
const cloudResult = await response.json();
return {
...cloudResult,
mode: 'cloud'
};
}
} catch (err) {
// fallback handling
}
}
The important point is not simply that old incidents are stored.
The current incident is actively used as the query for finding relevant historical experience.
For example, an incident involving:
payment-gateway
high latency
database connections
recent deployment
can retrieve an earlier incident involving similar symptoms.
Why Memory Changes the Investigation
The important question was not:
"Can I store an incident in memory?"
The more interesting question was:
"Does remembering the incident actually change what the agent does next?"
That became the central design principle.
Suppose the current telemetry shows:
DB Connections: 100/100
Error Rate: 38.7%
P99 Latency: 4820 ms
Recent Deployment: v2.8.1
The agent retrieves a historical incident with a similar pattern.
The previous incident contained:
Root Cause:
Hikari CP connection pool starvation
Remediation:
RB-PAY-04
Outcome:
Recovery verified
The current agent can now use that historical evidence when generating its hypothesis.
Instead of simply saying:
"Check the database."
it can explain why connection-pool exhaustion is a likely cause and show the historical incident supporting that hypothesis.
This is the difference between storing memory and using memory.
Separating Evidence From Diagnosis
I deliberately avoid presenting the model's diagnosis as absolute truth.
The interface uses:
LIKELY ROOT CAUSE
rather than:
ROOT CAUSE
The reasoning layer combines current telemetry with historical evidence.
For example:
Likely Root Cause:
Database connection pool exhaustion
Evidence:
DB connections reached 100/100
P99 latency increased sharply
Error rate increased
Recent deployment occurred
Similar historical incident was resolved using the same remediation pattern
This makes the reasoning inspectable.
An engineer can see both the recommendation and the evidence behind it.
Keeping an Engineer in the Loop
I did not want the system to blindly execute infrastructure changes.
The workflow therefore includes an explicit approval boundary.
The engineer can inspect the recommendation before allowing the remediation to proceed.
For example:
Recommended Runbook: RB-PAY-04
Increase database connection capacity
Restart the connection service
Monitor error rate
Monitor latency
Verify service health
[ APPROVE & EXECUTE ]
[ REJECT ]
Remediation runs through a controlled execution layer rather than directly modifying production infrastructure.
Verifying Recovery Instead of Assuming It
Executing a remediation command does not prove that an incident is fixed.
Resilify.AI therefore checks the telemetry after remediation.
A successful recovery should be visible in the system:
Error Rate
38.7% → 24.2% → 8.1% → 1.4% → 0.3%
P99 Latency
4820 ms → 45 ms
Health probes are also checked before the incident is marked as recovered.
The final state becomes:
RECOVERY VERIFIED
This gives the system an actual outcome to remember.
Closing the Loop
The post-mortem captures what happened and what the system learned.
Lesson:
High database connection utilization combined with increased latency after deployment is a strong indicator of connection pool exhaustion.
The post-mortem is then retained in Hindsight.
The code path that closes the loop is conceptually:
if (this.mode === 'cloud' && this.apiKey) {
try {
const response = await fetch(
${this.apiUrl}/v1/banks/${bank}/memories,
{
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': Bearer ${this.apiKey}
},
body: JSON.stringify({
content,
metadata
})
}
);
if (response.ok) {
const cloudData = await response.json();
this.saveLocalMemory(
bank,
content,
metadata
);
return {
...cloudData,
mode: 'cloud'
};
}
} catch (err) {
// fallback handling
}
}
The important idea is that the resolution becomes part of the agent's future context.
The next time a similar incident occurs, this experience becomes available during investigation.
The Complete Learning Loop
The resulting architecture is more than:
Incident → LLM → Answer
It becomes:
Remember → Investigate → Act → Verify → Learn
The system therefore treats every resolved incident as another piece of operational experience.
That is the part of the design I found most interesting.
The agent does not become useful simply because it can generate troubleshooting text.
It becomes more useful when previous outcomes influence future investigations.
What I Learned
- Memory is useful only when it changes decisions
It is easy to build a system that retrieves old information.
The harder and more important question is whether that information changes the current recommendation.
For Resilify.AI, historical incidents are useful only when they provide evidence that affects the current investigation.
- Retrieval and reasoning are different problems
Finding a similar incident and deciding what that incident means are not the same operation.
That is why the system separates historical retrieval from reasoning over historical experience.
- Evidence makes AI recommendations easier to trust
Instead of presenting a diagnosis as certain, I show the evidence supporting a likely root cause.
This gives an engineer something concrete to inspect before approving remediation.
- The outcome is part of the memory
A remediation without an outcome is incomplete.
The system needs to know whether the recommended action actually worked.
That outcome is what makes the historical incident useful to future investigations.
- The best memory systems close the loop
The most important architectural decision was connecting the beginning and end of the incident lifecycle.
The system therefore has a simple mechanism for accumulating operational experience over time.
LIMITATION
One limitation I ran into was that persistent memory is only useful when the retrieved experience is relevant to the current incident.
A historical incident can provide strong supporting evidence, but it cannot replace current telemetry.
That is why Resilify.AI treats historical memory as evidence for investigation rather than as an unquestionable diagnosis.
The current incident still has to be verified through its own telemetry and recovery checks.
Conclusion
Building Resilify.AI changed how I think about AI agents for operational systems.
A language model can already explain logs, suggest commands, and generate troubleshooting steps.
The harder problem is giving it useful organizational experience.
With Hindsight, a resolved incident can become persistent memory that is available during future investigations.
That changes the role of the agent from a system that simply answers questions into a system that can use the organization's previous experience.
The core loop is simple:
Remember.
Investigate.
Act.
Verify.
Learn.
That is the foundation of Resilify.AI.
AI Assistance Disclosure
I used AI tools to assist with editing and structuring this article. The project implementation, technical decisions, and results described here are based on my own work.










Top comments (1)
hold up. green uptime tiles are not a usage receipt.
1 cut: when the cloud invoice fight opens, can a buyer GET a signed meter tip of what ran, or only another vendor seal?
receipts > seals. #marker1228-obs