What Building an AI SRE Agent Taught Us
The hardest part of building an AI SRE agent was not generating a diagnosis.
It was designing the system around that diagnosis.
OpsMind started from a straightforward problem: incident responders repeatedly encounter similar failures, but the useful experience from one incident is not always available when another incident occurs.
We built the system around persistent agent memory so that previous incident outcomes could become part of future investigations.
The result was a workflow combining current telemetry, Hindsight memory, AI reasoning, human approval, simulated remediation, and retained learning.
The Architecture
OpsMind follows this lifecycle:
Incident → Evidence → Recall → Diagnosis → Approval → Remediation → Retain
Each stage has a different responsibility.
Current logs and metrics provide evidence.
Hindsight provides historical operational context.
The AI SRE agent reasons over the information.
A human approves remediation.
The successful outcome becomes memory.
Figure 1 — OpsMind turns incident resolution into a persistent learning loop.
The architecture looks simple when written as a sequence.
Implementing it exposed several engineering decisions that were less obvious.
Lesson 1: Memory Is Not the Same as Knowledge
Adding a vector or memory database does not automatically make an agent knowledgeable.
The system needs to decide:
- What should be remembered?
- When should it be remembered?
- What context should be stored?
- When should it be retrieved?
- How should retrieved information influence reasoning?
OpsMind retains incident outcomes rather than indiscriminately storing every interaction.
A successful incident creates a learning record containing the incident, diagnosis, actions, outcome, and resolution status.
That makes the memory operationally meaningful.
Lesson 2: Current Evidence Must Remain Primary
Persistent memory can be powerful, but it creates a new failure mode.
An agent may find an old incident that looks similar and assume that the same cause applies again.
OpsMind addresses this by giving current logs and metrics priority.
The AI receives current incident information and evidence before historical context is incorporated into the reasoning process.
For INC-008, current telemetry showed 6.1-second latency, 26% HTTP 500 errors, and 97% database connection utilization.
Hindsight then added historical incidents involving related database behavior.
The model was instructed not to blindly copy previous remediation.
That separation is one of the most important design choices in the system.
Lesson 3: Retrieval Needs to Be Testable
It is easy to say that an agent has memory.
It is harder to demonstrate that the memory actually changes future behavior.
We used INC-008 as a learning event.
The incident was analyzed and resolved.
Its successful outcome was retained.
Then INC-007 was analyzed afterward.
Hindsight returned INC-008 as historical context.
Figure 2 — The later INC-007 investigation recalls the newly retained INC-008 experience.
This gave us a concrete test:
Can a future investigation retrieve an experience created by a previous investigation?
The answer in this workflow was yes.
Lesson 4: Human Approval Changes the Architecture
It would have been possible to connect the AI's recommended actions directly to an execution layer.
We deliberately did not.
OpsMind inserts a human approval gate between diagnosis and remediation.
The agent can produce:
- Root cause
- Evidence
- Reasoning
- Recommended actions
- Confidence
- Runbook
The operator then reviews the proposed action.
Only after approval does the resolution workflow proceed.
The current remediation is simulated rather than connected to production infrastructure.
This keeps the prototype's operational boundary explicit.
Lesson 5: Agent Memory Creates a Feedback Loop
Traditional incident handling often ends with:
Resolved → Closed
The resolution may remain inside a ticketing system or someone's memory.
OpsMind changes the lifecycle to:
Resolved → Retained → Available to Future Investigations
After successful remediation, the incident outcome is retained in Hindsight.
Figure 3 — A successful incident outcome becomes persistent organizational memory.
That creates a feedback loop.
The next incident can use the previous experience.
The next successful resolution can then become another memory.
Over time, the system can accumulate operational experience.
Lesson 6: Integration Problems Are Part of the Work
The Hindsight integration also produced a technical problem that was not visible in the initial architecture.
The first recall implementation encountered:
text
Timeout context manager should be used inside a task
The issue appeared because the synchronous memory client was being used within the asynchronous FastAPI environment.
We changed the implementation to use Hindsight's asynchronous recall operation:
python
response = await hindsight_client.arecall(
bank_id=HINDSIGHT_BANK_ID,
query=query,
)
and explicitly closed the client afterward:
python
await hindsight_client.aclose()
This solved the event-loop problem.
The lesson was broader than this particular error.
When integrating external services into an agent system, application execution models matter just as much as API design.
Lesson 7: A Good Demo Should Prove a Behavior
For an AI agent, showing a generated answer is not enough.
We wanted the workflow to demonstrate a state change.
The sequence is:
INC-008
↓
Diagnosis
↓
Human approval
↓
Resolution
↓
Memory retained
↓
INC-007
↓
INC-008 recalled as historical context
Figure 4 — The memory relationship becomes observable when the later incident retrieves the earlier outcome.
This is more meaningful than simply showing a chatbot response because it demonstrates persistent state across investigations.
What We Would Improve Next
There are several areas that would need additional engineering for a production SRE environment.
First, real observability integrations would be needed instead of controlled incident datasets.
Second, remediation would need stronger safeguards, authorization, audit logging, rollback support, and failure handling.
Third, memory retrieval would need continuous evaluation to determine whether retrieved incidents are genuinely relevant.
Fourth, the system would need mechanisms for handling outdated or contradictory historical knowledge.
Persistent memory introduces a new operational responsibility:
not only deciding what the agent should remember, but also deciding when remembered information should no longer influence decisions.
The Main Takeaway
The most useful lesson from OpsMind was that agent memory is not an isolated feature.
It changes the lifecycle of the application.
Without memory:
Incident → Diagnosis → Resolution
With persistent memory:
Incident → Evidence → Historical Context → Diagnosis → Resolution → Learning → Future Incident
Hindsight provides the persistent layer that makes the second workflow possible.
But the rest of the architecture still matters.
Evidence grounds the diagnosis.
The AI interprets the evidence.
Human approval controls remediation.
Successful outcomes create new memory.
Future investigations can then retrieve those experiences.
Conclusion
Building OpsMind changed our understanding of what an AI SRE agent should do.
The objective is not simply to create an AI that can explain why an incident happened.
The more interesting objective is to build a system that can learn from successful incident resolution and make that experience available when the next problem appears.
INC-008 demonstrated that loop.
It was investigated using current telemetry and historical context, resolved through a controlled workflow, and retained as organizational memory.
Later, INC-007 could retrieve that experience.
That is the behavior we wanted from persistent agent memory:
not a replacement for engineering judgment, but a mechanism for making previous operational experience reusable.
Hindsight on GitHub
Hindsight Documentation
What Is Agent Memory? — Vectorize
OpsMind — GitHub Repository




Top comments (0)