Why Autonomous SRE Needs Episodic Memory: How We Built Opsight with Hindsight
Every on-call engineer knows the sinking feeling of getting paged at 3:14 AM for a SEV-1 outage that feels eerily familiar, yet nobody can remember which config flag or database lock script resolved it last time. We spend millions instrumenting distributed traces, metrics, and logs, yet when a cascading failure hits production, institutional memory vanishes into a graveyard of disconnected Slack threads and closed Jira postmortems.
Over the past year, we set out to solve this fundamental breakdown in site reliability engineering. We built Opsight, an autonomous incident observation and self-healing system operating across our distributed commerce platform, NEXA. Rather than building another dashboard that turns red and waits for human intervention, or slapping a stateless LLM onto our logs that offers generic advice like "try restarting the container," we anchored the entire architecture around persistent, episodic memory.
By integrating Vectorize agent memory via Hindsight—an open-source framework for giving autonomous agents structured, long-term recall—we gave our reliability agent the ability to correlate real-time telemetry signatures against historical incident postmortems, extract verified code patches, and execute surgical remediations in seconds.
Here is the story of how the system hangs together, why naive RAG failed us, and how episodic memory changed our approach to keeping production systems alive.
1. System Architecture: How Opsight Hangs Together
At its core, Opsight bridges the gap between edge client experience, distributed backend telemetry, and active remediation. The system comprises four tightly coupled layers:
-
The Edge Telemetry Plane (
NEXA Storefront): High-throughput micro-frontends streaming synthetic and user-driven interaction telemetry. When an API call fails—such as a checkout mutation throwing an unexpected500 Internal Server Error—the client runtime captures request timings, trace IDs, and local session state, immediately dispatching a structured diagnostic payload to the orchestrator. -
The Ingress & Topology Stream (
Service Connections): A cluster of streaming workers ingesting metrics from Envoy edge proxies, Kubernetes service meshes, Redis caches, and primary PostgreSQL databases at over 24,000 messages per second. -
The Spatial Visualization Engine (
Opsight 3D Observatory): A hardware-accelerated WebGL scene built on React Three Fiber and Three.js that projects live network graphs, cluster saturation, and memory crystals directly into three-dimensional space for continuous human-in-the-loop oversight. -
The Episodic Memory Core (
Hindsight Agent Memory): The cognitive engine that indexes past incident vectors, root causes, symptoms, and git pull request diffs, allowing our remediation agents to reason across months of operational history.
When an anomaly is flagged at the edge or within the backend service mesh, the incident is not dumped into an unindexed backlog. Instead, the incident signature is immediately passed into our memory pipeline.
2. The Failure of Stateless Agents and Naive RAG
When we first began experimenting with autonomous agents for incident triage, we hit a wall with standard prompt engineering and naive RAG (Retrieval-Augmented Generation).
The industry’s default recipe for building an AI ops agent is straightforward: dump raw CloudWatch or Datadog logs into a prompt, vectorize recent documentation chunks in a vector database, and ask an LLM what broke. In practice, this failed catastrophically during live production incidents for three specific reasons:
The Signal-to-Noise Catastrophe
During a cascading outage—say, a connection pool exhaustion in a payment worker—downstream services begin throwing hundreds of secondary errors per second. Gateways throw 504 Gateway Timeout, ingress workers log retry storms, and authentication caches log socket drops. Ingesting this firehose directly into an LLM context window diluted the root cause. The model spent its attention budget analyzing the symptoms rather than the originating fault.
Amnesia Across Lifecycles
Stateless models have no concept of time or precedent. They don't know that three weeks ago, our senior database engineer discovered that batch checkout bursts require a 3,000ms acquisition timeout on the HikariCP connection pool. A stateless agent repeatedly recommended generic remedies: "Increase CPU limits on the Kubernetes pod," or "Restart the payment service." In production, restarting the payment service during pool starvation only creates a massive spike in connection handshakes, immediately knocking the database completely offline.
The Semantic Disconnect Between Symptoms and Diff Patches
Standard embedding models do a decent job of matching English text, but they fail when correlating numeric telemetry signatures (like Pool active: 98%, Lock waits: 840ms, 5xx rate: 8.4%) to actual structural code patches in infrastructure configurations.
We needed a system that treated incidents not as arbitrary blocks of text, but as structured, episodic experiences containing symptoms, temporal telemetry profiles, verified root causes, and git-level remediation diffs. That led us to Hindsight.
3. The Core Technical Story: Episodic Memory with Hindsight
According to the Hindsight documentation, effective agent memory requires separating short-term working context from long-term episodic retrieval. In Opsight, we designed our data layer around structured incident entities where each past resolution is preserved as an actionable episodic memory.
Here is the TypeScript interface defining our episodic memory schema in src/lib/opsight-data.ts:
export interface PreviousTicket {
id: string;
title: string;
serviceId: ServiceId;
relatedServiceId?: ServiceId;
similarity: number;
date: string;
mttr: string;
severity: "SEV-1" | "SEV-2" | "SEV-3";
status: "RESOLVED";
rootCause: string;
symptoms: string[];
fixSummary: string;
fixSteps: string[];
fixCode?: string;
appliedPR?: string;
engineer: string;
position3D: [number, number, number];
}
Notice what is captured here: this is not merely a summary paragraph. It captures:
- Exact telemetry symptoms (
"504 Gateway Timeouts on POST /v2/charge","DB active connection pool hit 100% saturation"). - Root-cause causality chains.
- Concrete configuration or code patches (e.g., HikariCP pool expansion or concurrent PostgreSQL indexing commands).
- Spatial coordinates (
position3D) to project the memory directly into our live infrastructure topology.
When an active incident emerges, our vector pipeline generates an embedding representing the incident's vector:
$$\vec{V}_{\text{incident}} = \mathcal{E}\Big(\text{Service}, \text{ErrorSignatures}, \Delta\text{Metrics}, \text{CallGraph}\Big)$$
Using Hindsight, we query our historical repository of verified postmortems. Rather than doing a blunt lexical search, Hindsight calculates cosine similarity across the topological graph and symptoms, returning prioritized candidates with confidence scores:
// Sample historical incident ticket retrieved via Hindsight memory
export const previousTickets: PreviousTicket[] = [
{
id: "INC-014",
title: "Database Connection Pool Starvation Post-Deployment",
serviceId: "payment",
relatedServiceId: "database",
similarity: 94,
date: "Aug 14, 2026",
mttr: "12m",
severity: "SEV-1",
status: "RESOLVED",
rootCause:
"HikariCP connection pool was capped at 30 connections with no acquisition timeout during payment peak, causing worker thread exhaustion when checkout batch traffic spiked.",
symptoms: [
"504 Gateway Timeouts on POST /v2/charge",
"DB active connection pool hit 100% saturation",
"PaymentWorker threads locked waiting for free socket",
],
fixSummary:
"Hot-expanded HikariCP maxPoolSize from 30 -> 90, enabled a 3,000ms acquisition timeout, and deployed rate limiting on batch checkout queue.",
fixSteps: [
"Patched ConfigMap to raise hikari.maximum-pool-size from 30 to 90",
"Set hikari.connection-timeout=3000 to prevent infinite thread locks",
"Executed pg_terminate_backend() on stale idle-in-transaction sockets",
"Merged PR #4829 adding backpressure throttling to payment worker pool",
],
fixCode: `// Spring / HikariCP Production Fix (PR #4829)
spring.datasource.hikari.maximum-pool-size=90
spring.datasource.hikari.minimum-idle=25
spring.datasource.hikari.connection-timeout=3000
spring.datasource.hikari.idle-timeout=30000
spring.datasource.hikari.max-lifetime=1200000
spring.datasource.hikari.leak-detection-threshold=2500`,
appliedPR: "PR #4829: Hikari pool resize & timeout guard",
engineer: "Sarah Lin (Staff SRE)",
position3D: [-5.2, 2.2, 4.2],
},
// ... Additional historical incident vectors
];
By presenting this structured memory to the agent, the model skips hours of trial-and-error diagnostics. It verifies whether the live failure matches the preconditions of INC-014 (e.g., database pool utilization at 98%, transaction lock waits rising, payment endpoint throwing 500s). Once verified, it can propose the exact, pre-tested configuration diff.
4. Visualizing Episodic Memory in 3D Space
Most telemetry tools isolate data into disparate silos: Grafana for dashboards, Jaeger for traces, and Confluence for postmortems. We realized that spatial adjacency matters when humans and AI agents collaborate during high-stress incidents.
In Opsight's 3D Environment, every active microservice is rendered as a physical node on an interactive WebGL coordinate plane. But more importantly, Hindsight memories are materialized as holographic memory crystals floating directly adjacent to the afflicted services.
Here is how we implemented the holographic memory crystal in Three.js and React Three Fiber (src/components/opsight/InfrastructureScene.tsx):
function MemoryCrystalNode({
ticket,
selected,
isRelatedToActiveNode,
hide3DLabels = false,
onSelect,
}: {
ticket: PreviousTicket;
selected: boolean;
isRelatedToActiveNode: boolean;
hide3DLabels?: boolean;
onSelect: (ticket: PreviousTicket) => void;
}) {
const group = useRef<THREE.Group>(null);
const ring = useRef<THREE.Mesh>(null);
const [hovered, setHovered] = useState(false);
const color = ticket.similarity >= 90 ? "#38bdf8" : "#c084fc";
useFrame(({ clock }, delta) => {
if (group.current) {
group.current.rotation.y += delta * 0.85;
group.current.position.y =
ticket.position3D[1] + Math.sin(clock.elapsedTime * 1.8 + ticket.similarity) * 0.12;
}
if (ring.current) {
ring.current.rotation.x += delta * 0.5;
ring.current.rotation.z += delta * 0.3;
}
});
return (
<group ref={group} position={ticket.position3D} onClick={() => onSelect(ticket)}>
{/* Holographic 3D Crystal Octahedron */}
<mesh>
<octahedronGeometry args={[0.3, 0]} />
<meshPhysicalMaterial
color={color}
emissive={color}
emissiveIntensity={selected ? 1.8 : isRelatedToActiveNode ? 0.9 : 0.4}
metalness={0.9}
roughness={0.15}
transparent
opacity={0.88}
/>
<Edges color={color} threshold={15} />
</mesh>
{/* Orbiting Halo Ring */}
<mesh ref={ring} rotation-x={Math.PI / 2}>
<torusGeometry args={[0.48, 0.016, 6, 32]} />
<meshBasicMaterial
color={color}
transparent
opacity={selected ? 0.95 : 0.3}
toneMapped={false}
/>
</mesh>
</group>
);
}
When an incident strikes the Payment API, the scene dynamically spotlights the degraded service in crimson and illuminates the corresponding Hindsight memory crystal in cyan. Clicking the crystal or triggering the agent exposes the exact postmortem, telemetry delta, and verified fix.
5. Live Incident Walkthrough: Resolving a Cascading Payment Outage
To understand the real-world impact of this architecture, consider an incident that occurred during high-volume traffic on our NEXA platform.
Step 1: Edge Detection
A user shopping for a smart watch reaches checkout and clicks "Pay." Under the hood, the client issues a POST /api/payment. The payment gateway encounters thread starvation, times out at 4.0 seconds, and returns a 500 Internal Server Error. The client UI displays an alert while the embedded Opsight copilot intercepts the failure and signals the backend orchestrator.
Step 2: Anomaly Ingestion and Spatial Mapping
In the Opsight observatory, the Payment Service flips to DEGRADED (SEV-1), and the Primary Database cluster registers 98% connection pool saturation with lock waits climbing to 840ms. The automated topology engine draws a crimson dependency line from Payment to Database.
Step 3: Hindsight Vector Correlation
Opsight's agent evaluates the telemetry signature against historical vectors stored in Hindsight:
- Metric Signature:
HikariCP Pool > 95%,POST /v2/charge 504,Active DB connections: 498/500. - Memory Match:
INC-014matches with 94% cosine similarity.
The agent does not guess. It notes:
"Incident correlates with INC-014 (Aug 14, 2026), resolved by Sarah Lin. Root cause was HikariCP thread starvation under batch queue bursts. Verified fix required expanding maximum pool size from 30 to 90 and setting connection timeout to 3,000ms."
Step 4: Verification and Remediation
The agent presents the verified configuration patch in our HUD. An engineer can click Apply Fix, or, if operating in fully autonomous mode, Opsight applies the ConfigMap patch via Kubernetes API and reloads the pool gracefully without restarting the pods. Within 90 seconds, connection pool saturation drops from 98% to 32%, lock wait times plummet back to 14ms, and checkout transaction success returns to 100%.
6. Lessons Learned from Building Opsight
Building an autonomous incident management platform taught us several hard truths about distributed systems, WebGL rendering, and agent architectures:
1. Episodic Memory Trumps Raw Model Intelligence
An 8-billion parameter model equipped with structured episodic memory from Hindsight consistently outperforms a 70-billion parameter stateless model relying solely on zero-shot reasoning. Giving an agent specific examples of how your infrastructure failed and how your team fixed it previously eliminates hallucination and drastically reduces mean time to remediation (MTTR).
2. High-Dimensional Telemetry Must Be Distilled Before Vectorization
Raw log lines contain too much non-deterministic noise (timestamps, ephemeral IP addresses, UUIDs) to produce clean vector embeddings. We had to implement a preprocessing step that extracts deterministic features—service dependencies, HTTP status ratios, database wait event types, and metric slopes—before passing them into our vectorizer.
3. Coordinate Spaces Between 3D WebGL and HTML Overlays Require Strict Boundaries
When mixing Three.js with HTML UI overlays (like Drei's <Html> components), managing z-indexes and mouse event propagation can be treacherous. We ran into bugs where 3D labels were bleeding directly through our glassmorphism dialog modals. The fix was ensuring all 3D HTML labels are dynamically unmounted whenever modals are triggered, and strictly scoping pointer events with event.stopPropagation().
4. Code Diffs Are the Ultimate Truth for SRE Agents
Natural language postmortems often drift into vague descriptions: "We adjusted connection parameters." That is useless to an automated agent. Storing exact configuration diffs and SQL scripts inside the episodic memory entity transformed postmortems from passive documentation into executable operational playbooks.
The Path Ahead
The future of site reliability is not about giving human engineers more dashboards to stare at during 3 AM emergencies. It is about building systems that learn from their own operational history.
By coupling real-time distributed telemetry with Vectorize agent memory and Hindsight, we have moved closer to a world where production infrastructure possesses genuine self-healing capabilities—not through brute-force automation, but through institutional memory that never forgets an outage.
References & Resources
- The Hindsight GitHub repository: https://github.com/vectorize-io/hindsight
- The documentation for Hindsight: https://hindsight.vectorize.io/
- The agent memory page on Vectorize: https://vectorize.io/what-is-agent-memory





Top comments (0)