How I Gave My Kubernetes SRE Agent a Persistent Memory
Let me take you back to a Tuesday at 2 AM.
Our primary payment service in production started throwing 500 errors. PagerDuty screamed, Slack channels flooded with frantic messages, and three senior engineers were dragged out of bed to debug a Kubernetes cluster that was actively rejecting traffic.
We spent thirty minutes tailing logs, analyzing Grafana dashboards, and running endless kubectl describe pod commands.
Finally, we found the culprit:
Database connection pool exhaustion caused by a rogue background job.
We killed the job, restarted the pods, and went back to sleep.
The real tragedy?
We had experienced the exact same incident three months earlier.
But that knowledge was buried in a closed Slack thread and forgotten in an outdated wiki. So we were forced to debug the problem from scratch.
This is one of the biggest problems I see in modern Site Reliability Engineering.
We have become incredibly good at building distributed systems, orchestrating containers, monitoring infrastructure, and deploying microservices.
But we are still surprisingly bad at retaining operational knowledge.
Our diagnostic systems are mostly stateless.
When an outage occurs, we often treat it like a brand-new mystery.
That means wasted debugging time, higher MTTR, consumed error budgets, and unnecessary operational cost.
The problem gets worse with AI agents
LLMs and autonomous agents can help with incident response, but they introduce another problem.
A stateless AI agent has to reason through the incident again and again.
It might need to:
- inspect logs
- query Prometheus
- analyze metrics
- formulate hypotheses
- inspect Kubernetes events
- propose a fix
- evaluate whether the fix worked
This ReAct-style loop can be powerful, but it can also be slow, expensive, and limited by the available context.
So I started thinking about a different question:
What if an SRE agent could actually remember?
Not just remember text in a context window.
I wanted it to remember:
- previous incidents
- root causes
- symptoms
- successful remediation steps
- failed approaches
- human modifications
- historical outcomes
That led to KubePilot.
What KubePilot Actually Does
KubePilot is an autonomous SRE platform designed to monitor, diagnose, and remediate Kubernetes incidents.
Instead of building one giant script, I designed it around an event-driven microservices architecture.
The stack includes:
- FastAPI for backend services
- React for the frontend
- Kubernetes for execution
- telemetry and monitoring data for incident detection
- a multi-agent workflow for diagnosis and planning
- Hindsight as the persistent semantic memory layer
The basic workflow looks like this:
Telemetry
↓
Anomaly Detection
↓
Incident Engine
↓
AI Orchestrator
↓
Memory Recall ──────→ Hindsight
↓
Decision Engine
↓
Recovery Playbook
↓
Kubernetes Controller
↓
Incident Resolution
↓
Retain Outcome → Hindsight
The important part is not simply the LLM.
The important part is the memory.
I integrated Hindsight to act as the persistent semantic layer for KubePilot.
Hindsight stores information about previous incidents, including their symptoms, root causes, remediation steps, and outcomes.
So when the next incident happens, KubePilot can ask:
"Have I seen something like this before?"
The Recall → Retain Loop
The architecture is based around a continuous recall and retain loop.
When an anomaly is detected, the incident engine triggers the orchestrator.
Before spending LLM tokens generating new hypotheses, KubePilot performs a semantic search against the Hindsight memory layer.
The simplified workflow is:
Incident Detected
↓
Generate Symptom Fingerprint
↓
Semantic Memory Search
↓
┌───────────────────────────┐
│ Known incident found? │
└───────────────────────────┘
↓
Yes
↓
Retrieve Proven Playbook
↓
Decision Engine
↓
Execute / Request Approval
↓
Verify Recovery
↓
Retain Result in Memory
Here is a simplified example of how the backend-for-frontend layer proxies memory requests from the React UI to the internal backend:
BACKEND_BASE_URL = os.getenv(
"BACKEND_BASE_URL",
"http://host.minikube.internal:8000"
)
@app.get("/api/memory/bank")
async def get_memory_bank():
async with httpx.AsyncClient(timeout=TIMEOUT) as client:
try:
resp = await client.get(
f"{BACKEND_BASE_URL}/memory/bank"
)
return resp.json()
except Exception as e:
logger.error(
f"Failed to fetch from backend: {e}"
)
return {
"bank_id": "error",
"total_memories": 0,
"memories": []
}
When Hindsight returns a high-confidence match, KubePilot retrieves the associated playbook and presents it to the operator.
If the playbook is executed and successfully restores service health, the incident enters the retain phase.
The incident's telemetry signature, selected resolution, and outcome are then written back into Hindsight.
That creates a compounding effect.
Every resolved incident can become operational knowledge for future incidents.
As the memory bank grows, the system can rely more heavily on historically proven remediation paths rather than generating a completely new solution every time.
You can learn more about the concept of agent memory here:
The Short-Circuit: When Memory Beats the LLM
This is where the architecture became especially interesting.
A traditional autonomous agent may need to perform several sequential LLM interactions during an incident.
For example:
Read logs
↓
Ask LLM for hypothesis
↓
Query metrics
↓
Ask LLM to analyze metrics
↓
Inspect Kubernetes events
↓
Generate remediation plan
That reasoning process can work well for novel problems.
But what about an incident we have already solved five times?
Why make the model rediscover the answer?
With semantic memory, KubePilot can bypass the full reasoning loop when it finds a sufficiently strong historical match.
For example, a retained incident might look like this:
{
"incident_id": "INC-7731",
"symptoms": "500 errors from payment-service",
"root_cause": "Database connection pool exhaustion",
"playbook": [
"Increase connection pool size to 50",
"Restart payment-service pods"
],
"success_rate": 0.88,
"outcome": "human_modified",
"human_approved": True
}
Now imagine another payment-service incident produces a very similar symptom fingerprint.
Instead of starting from zero, KubePilot can retrieve the historical incident.
The decision engine can evaluate the match and the historical outcome.
A high-confidence match can then move through the recovery workflow without requiring the entire LLM reasoning cycle.
The UI can also make this explicit to the operator:
MEMORY RECALL
Known Incident Detected
Historical Success Rate: 88%
Recommended Playbook:
1. Increase connection pool size
2. Restart payment-service pods
Source: Historical Incident INC-7731
This distinction matters.
The operator can see whether a recommendation came from historical operational knowledge or from new generative reasoning.
Before vs After
During development, I compared two approaches for recurring incidents.
Without semantic memory
A database timeout triggered a full reasoning workflow:
Incident
↓
Prometheus
↓
Metrics Analysis
↓
LLM Reasoning
↓
Kubernetes Events
↓
Hypothesis
↓
Recovery Plan
In our testing, this workflow took roughly 20–30 seconds and required substantial LLM inference.
With Hindsight memory
The workflow became:
Incident
↓
Symptom Fingerprint
↓
Hindsight Search
↓
Known Playbook
↓
Execution Queue
For recurring incidents, our measured pipeline dropped to approximately 1.5 seconds, with 0 LLM tokens consumed on a recall hit.
The important point is not that every incident will suddenly become a 1.5-second operation.
Novel incidents still require deeper reasoning.
The real advantage is reducing repeated reasoning for known problems.
That is exactly where persistent operational memory becomes valuable.
What I Got Wrong
The transition to a memory-backed SRE architecture was not completely smooth.
My biggest mistake was underestimating the cold-start problem.
When we first deployed the system, Hindsight was effectively a blank slate.
There were no historical incidents available for recall.
That meant every anomaly still triggered the slower reasoning workflow.
I had assumed the memory bank would quickly populate itself.
But there was a problem:
Production incidents are relatively rare.
In a stable environment, you may only get a few significant incidents each week.
That means it can take a long time for the system to accumulate enough high-confidence memories.
So the real lesson was:
Don't wait for the agent to learn everything from scratch.
Seed the memory first.
What I Would Do Differently
If I rebuilt KubePilot today, I would bootstrap the memory layer before production deployment.
I would process existing operational knowledge from sources such as:
- historical incident tickets
- Jira issues
- Slack discussions
- postmortems
- existing runbooks
- internal troubleshooting documentation
Then I would convert that information into a structured incident-memory schema and populate Hindsight before the agent goes live.
Instead of starting with:
Empty Memory
↓
Incident
↓
Reason
↓
Store Memory
I would start with:
Historical Operational Knowledge
↓
Memory Extraction
↓
Hindsight Database
↓
KubePilot
↓
New Incidents
That could dramatically reduce the cold-start period.
What I'd Do Next
The current KubePilot workflow uses Hindsight to retrieve playbooks for recurring incidents.
The next challenge is making the execution layer more autonomous while keeping appropriate safety controls.
Currently, when a high-confidence memory match occurs, the playbook can still be queued for human approval before the Kubernetes controller executes it.
A future version could use a dynamic trust model.
For example:
High-confidence memory match
+
Repeated successful approvals
+
Verified incident signature
↓
Higher automation level
A possible policy could require a memory match above a defined confidence threshold and several previous successful human approvals before allowing automatic execution.
That way, autonomy is earned through evidence rather than being enabled blindly from day one.
I also want to improve telemetry fingerprinting.
Right now, incident matching can rely heavily on symptoms and root-cause descriptions.
A stronger system could incorporate:
- Prometheus metric shapes
- time-series patterns
- Kubernetes events
- distributed trace information
- service dependency graphs
This could allow the memory system to recognize incidents based not only on their textual description, but also on the actual behavior of the system.
Try It
Persistent memory changes the way I think about autonomous agents.
A stateless agent asks:
"What should I do now?"
A memory-backed agent can ask:
"What happened the last time this occurred, what did we do, and did it work?"
That difference is extremely important for production systems.
The context window of an LLM is temporary.
Operational knowledge should not be.
For agents that need to learn from previous actions, a dedicated memory layer can provide a more persistent foundation than repeatedly stuffing historical information into prompts.
You can explore Hindsight here:
You can also explore the project itself here:
Final Thought
The goal of KubePilot isn't to make an LLM magically solve every Kubernetes incident.
The goal is to build an SRE system that gets better at recurring problems.
When an incident is truly new, use reasoning.
When an incident is familiar, use memory.
And when an automated action has been repeatedly validated, gradually increase the level of automation.
That combination of reasoning + memory + deterministic execution is what makes autonomous SRE interesting to me.


Top comments (0)