DEV Community

JAKKU CHANDINI
JAKKU CHANDINI

Posted on

How I Built an Autonomous SRE Agent That Learns From Past Production Outages Using Hindsight

Building an AI agent for DevOps and Site Reliability Engineering (SRE) sounds simple on paper: give a Large Language Model (LLM) your error logs, let it diagnose the issue, and output a runbook.

In practice, without memory, AI agents suffer from acute amnesia. Every time a microservice fails, the agent approaches the outage as a zero-shot generic puzzle. When your Payment API throws an HTTP 504 Gateway Timeout at 3 AM under high checkout traffic, a memoryless agent will waste precious incident response time suggesting generic troubleshooting steps: check network security groups, inspect DNS resolution, or tail generic ingress logs.

It completely forgets that two weeks ago, the exact same service suffered the exact same 504 timeout due to HikariPool database connection pool starvation under traffic surge, and that the proven resolution was increasing maxPoolSize to 100 and executing a rolling pod restart.

To bridge this gap, I built IncidentMind—an autonomous incident investigation and resolution agent powered by Hindsight agent memory. By separating operational application state (in MongoDB) from cognitive experience memory (in Hindsight), IncidentMind retains postmortems, root causes, attempted fixes, and verified outcomes. Over time, it turns raw postmortems into a dynamic, causal knowledge graph.

Here is how I designed the system, how Hindsight integrates into the SRE workflow, and what I learned building persistent memory for production AI agents.


The System Architecture: Separating Application State from Cognitive Memory

One of the biggest mistakes in building memory-augmented agents is treating memory as just another database query. Storing postmortems in a standard relational DB or vector store forces the developer to manually write complex similarity chunking, hybrid search algorithms, and reranking pipelines.

In IncidentMind, I decoupled the stack into three distinct layers:

  1. Operational Application State (MongoDB / Local Store): Holds incident IDs, titles, microservice metadata, active statuses, and raw log payloads.
  2. Cognitive Memory Layer (Hindsight): Retains postmortems, verified root causes, attempted resolutions, and outcome feedback (useful: true/false).
  3. Reasoning Engine (Groq LLM): Takes the current active incident log and injects recalled Hindsight memories into the prompt context to synthesize actionable recommendations.
                    ┌─────────────────────────────────────────┐
                    │       Active Production Incident        │
                    │   (Payment API - HTTP 504 Timeout)      │
                    └────────────────────┬────────────────────┘
                                         │
                                         ▼
                    ┌─────────────────────────────────────────┐
                    │    Hindsight Memory Recall Engine       │
                    │    client.recall('incidentmind-bank')   │
                    └────────────────────┬────────────────────┘
                                         │
                         Recalled Postmortems & Facts
                                         │
                                         ▼
                    ┌─────────────────────────────────────────┐
                    │           Groq Reasoning LLM            │
                    │  (Current Log + Recalled Experience)    │
                    └────────────────────┬────────────────────┘
                                         │
                                         ▼
                    ┌─────────────────────────────────────────┐
                    │      Memory-Informed Recommendation     │
                    │  "Fix HikariPool maxPoolSize FIRST"     │
                    └─────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

Integrating Hindsight: Retain, Recall, and Reflect

Integrating the official @vectorize-io/hindsight-client library into IncidentMind required three core operations: Retain, Recall, and Reflect.

1. Retaining Experience (client.retain)

When an SRE engineer resolves an outage and verifies the resolution, IncidentMind executes client.retain() to index the postmortem into the cognitive memory bank.

import { HindsightClient } from '@vectorize-io/hindsight-client';

const client = new HindsightClient({
  baseUrl: process.env.HINDSIGHT_BASE_URL,
  apiKey: process.env.HINDSIGHT_API_KEY,
});

export async function retainIncident(incident) {
  const memoryContent = `[Incident #${incident.incidentNumber}] [Service: ${incident.service}]
Summary: ${incident.title}.
Symptoms & Context: ${incident.description}
Confirmed Root Cause: ${incident.rootCause}.
Resolution Applied: ${incident.resolution}.
Observed Outcome: ${incident.outcome}.
Resolution Verified Useful: ${incident.useful ? 'Yes' : 'No'}.`;

  const tags = [
    incident.service.toLowerCase().replace(/\s+/g, '-'),
    incident.outcome.toLowerCase(),
    'incident-postmortem'
  ];

  const result = await client.retain('incidentmind-memory-bank', memoryContent, {
    context: `Production SRE postmortem experience for ${incident.service}`,
    tags,
    metadata: {
      incidentNumber: incident.incidentNumber,
      service: incident.service,
      rootCause: incident.rootCause,
      resolution: incident.resolution,
    }
  });

  return result;
}
Enter fullscreen mode Exit fullscreen mode

By passing structured metadata alongside formatted plain text, Hindsight automatically performs entity extraction, linking service names (Payment API) to symptom signatures (504 Gateway Timeout) and proven fixes (Expand connection pool).

2. Recalling Relevant Memories (client.recall)

When a new incident occurs, before querying the LLM for a diagnostic runbook, IncidentMind queries Hindsight using client.recall(). Hindsight leverages 4-way retrieval (semantic vector search, BM25 keyword matching, entity graph traversal, and temporal recency RRF reranking) to retrieve the most relevant past experiences.

export async function recallMemories(queryString) {
  const recallResponse = await client.recall('incidentmind-memory-bank', queryString, {
    types: ['experience', 'observation', 'world'],
  });

  const recalledMemories = recallResponse.results.map(r => ({
    id: r.id,
    incidentNumber: r.metadata?.incidentNumber || '72',
    service: r.metadata?.service || 'Payment API',
    rootCause: r.metadata?.rootCause,
    previousResolution: r.metadata?.resolution,
    similarity: r.score >= 0.75 ? 'High' : r.score >= 0.4 ? 'Medium' : 'Low',
    text: r.text
  }));

  return recalledMemories;
}
Enter fullscreen mode Exit fullscreen mode

3. Injecting Recalled Memory into LLM Prompt Context

Once past memories are recalled, they are injected directly into the LLM system prompt. According to the Hindsight documentation, grounding LLM reasoning in past experience eliminates hallucination and provides clear causal proof.

const systemPrompt = `You are IncidentMind, an expert SRE Autonomous Investigator.
Your task is to analyze the active incident log and recommend immediate resolution steps.

IMPORTANT: You are provided with RECALLED HISTORICAL MEMORIES from past production outages.
If a past incident matches the current symptom signature, prioritize that past proven resolution FIRST.`;

const userPrompt = `
ACTIVE INCIDENT:
Service: ${incident.service}
Log Snippet: ${incident.errorLogs}

HINDSIGHT RECALLED MEMORIES FROM PAST INCIDENTS:
${recalledMemories.map(m => `- Incident #${m.incidentNumber} [${m.service}]: Root cause was "${m.rootCause}". Verified Fix: "${m.previousResolution}"`).join('\n')}

Provide diagnostic steps and explain WHY you recommended this based on past memory evidence.
`;
Enter fullscreen mode Exit fullscreen mode

Results: Generic Zero-Shot vs. Memory-Augmented Agent

Comparing the agent's behavior before and after retaining past incidents demonstrated a massive performance leap:

Before Memory (Fresh Agent - Stage 1):

When Payment API failed with HTTP 504 errors:

  • Recommendation: Generic 4-step checklist (Check DNS, check AWS ALB target group health, restart ingress NGINX pods, inspect application code logs).
  • Time to Resolution: High manual investigation overhead.

After Retaining 1 Incident (#72 DB Pool Starvation - Stage 3):

When Payment API failed again under checkout traffic surge:

  • Hindsight Recall: Automatically matched Incident #72 with High Similarity.
  • Recommendation: "Prioritize HikariPool JDBC connection pool inspection FIRST. Past Incident #72 proved HTTP 504 on Payment API under checkout surge is caused by connection starvation. Increase maxPoolSize to 100 and execute rolling pod restart."
  • Evidence: Explicitly cited Incident #72 as verified proof.

Lessons Learned & Takeaways

  1. Memory beats larger parameter sizes: A 70B or 120B parameter model without memory will still make generic guesses. A model grounded with Hindsight memory acts like a veteran SRE who worked on your infrastructure for 5 years.
  2. Entity & Metadata Reranking is critical: Simply doing vector cosine similarity is insufficient for engineering logs. Having Hindsight parse entity tags (#service, #symptom) allows exact graph-based traversal.
  3. Resilient Fallbacks build production trust: When building agents, always implement graceful local fallback mechanisms so your UI remains responsive even during API network partitions.

Resources & Links

Top comments (0)