DEV Community

Duvvi Sai Charan
Duvvi Sai Charan

Posted on

How I Built an Incident Response Agent With Hindsight

Most automated incident management tooling does little more than blast high-priority Slack notifications and wake engineers up at 2 AM. Once an on-call engineer acknowledges the page, they are left to manually scour log aggregators, cross-reference outdated wiki runbooks, and guess whether the current alert looks like last month's database stall.

To solve this, we built a hands-on triage system that does not just alert, but actively retains the institutional memory of past outages to triage active failures in real time.

In this practical, step-by-step walkthrough, I will show you how to build a production-grade incident response agent using Python 3.11, FastAPI, Groq's high-speed Llama-3.3-70b model, and Hindsight for persistent episodic memory.


What the System Does and How It Hangs Together

Our agent continuously automates the incident remediation lifecycle through four synchronized phases:

  1. Alert Ingestion: Ingests raw alerts, system logs, impacted microservice tags, and severity levels via REST endpoints.
  2. Episodic Recall: Queries persistent agent memory to retrieve historically similar outages, past root causes, and previously successful runbooks.

  3. Reasoned Triage: Passes the incoming alert and retrieved historical context to Groq's Llama-3.3 engine to calculate a root-cause hypothesis, an explainable confidence score, and prioritized resolution steps.

  4. Active Learning Loop: Collects operator feedback (thumbs up/down) and writes verified post-mortems back into Hindsight to improve future accuracy.

 [ PagerDuty / Datadog Alert ]
               |
               v
       +---------------+
       | FastAPI API   | <=======> [ SQLite Database ]
       +---------------+           (State & Feedback Weights)
         |           |
         v           v
  +-----------+  +--------------------+
  | Hindsight |  | Groq LLM Engine    |
  | Memory    |  | (Llama-3.3-70b)    |
  +-----------+  +--------------------+
         \           /
          v         v
   [ React Triage Console ]
   (Actionable Runbook & Score)

Enter fullscreen mode Exit fullscreen mode

Core Technical Story: Implementing the Memory Lifecycle

When integrating an LLM into an operational workflow, you quickly discover that static system prompts fail. If you embed standard runbooks directly into prompt templates, you hit token limits, introduce latency, and dilute attention.

The agent requires dynamic, stateful retrieval. We structured our implementation around three core operations using the Hindsight documentation:

  • Retaining Episodic Incidents: Serializing resolved incidents with structured metadata (service, symptoms, resolution steps, runbook IDs).
  • Recalling Operational Context: Querying memory based on runtime symptoms to surface the top-$K$ historical incidents.
  • Recording Efficacy Updates: Registering user feedback to adjust runbook rankings dynamically.

Code-Backed Implementation

Step 1: Environment and Dependencies

Set up your project environment with the required dependencies:

pip install fastapi uvicorn pydantic sqlalchemy httpx groq

Enter fullscreen mode Exit fullscreen mode

Ensure your .env configuration contains your provider credentials:

LLM_PROVIDER=groq
GROQ_API_KEY=gsk_your_groq_api_key_here
GROQ_MODEL=llama-3.3-70b-versatile
MEMORY_BACKEND=hindsight
HINDSIGHT_API_KEY=hsk_your_hindsight_api_key_here
HINDSIGHT_BASE_URL=https://api.hindsight.vectorize.io/v1

Enter fullscreen mode Exit fullscreen mode

Step 2: The Hindsight Memory Client

Create app/memory/hindsight_store.py to handle the retain and recall lifecycle against the Hindsight REST API:

import os
import httpx
from typing import List, Dict, Any

class HindsightMemoryStore:
    def __init__(self):
        self.api_key = os.getenv("HINDSIGHT_API_KEY")
        self.base_url = os.getenv("HINDSIGHT_BASE_URL", "https://api.hindsight.vectorize.io/v1")
        self.headers = {
            "Authorization": f"Bearer {self.api_key}",
            "Content-Type": "application/json"
        }

    async def retain_incident(self, incident: Dict[str, Any]) -> str:
        """Stores a resolved incident and its post-mortem lessons into memory."""
        payload = {
            "document_id": f"inc-{incident['id']}",
            "context_type": "episodic_incident",
            "content": (
                f"Service: {incident['service']}\n"
                f"Symptoms: {incident['symptoms']}\n"
                f"Root Cause: {incident.get('root_cause', '')}\n"
                f"Resolution Steps: {incident.get('resolution_steps', '')}"
            ),
            "metadata": {
                "service": incident["service"],
                "severity": incident["severity"],
                "runbook_used": incident.get("runbook_used", "none")
            }
        }
        async with httpx.AsyncClient(timeout=10.0) as client:
            resp = await client.post(f"{self.base_url}/retain", json=payload, headers=self.headers)
            resp.raise_for_status()
            return resp.json().get("memory_id")

    async def recall_similar(self, service: str, symptoms: str, limit: int = 3) -> List[Dict[str, Any]]:
        """Recalls historical outages sharing symptom semantics."""
        payload = {
            "query": f"Service: {service}. Symptoms: {symptoms}",
            "filter": {"context_type": "episodic_incident"},
            "top_k": limit
        }
        async with httpx.AsyncClient(timeout=10.0) as client:
            resp = await client.post(f"{self.base_url}/recall", json=payload, headers=self.headers)
            resp.raise_for_status()
            return resp.json().get("results", [])

Enter fullscreen mode Exit fullscreen mode

Step 3: FastAPI Ingestion and Triage Endpoints

Implement the core triage and feedback endpoints in app/main.py:

from fastapi import FastAPI, HTTPException, Depends
from pydantic import BaseModel
from typing import Optional, List
from groq import Groq
import os
from app.memory.hindsight_store import HindsightMemoryStore

app = FastAPI(title="Incident Response Agent")
memory_store = HindsightMemoryStore()
groq_client = Groq(api_key=os.getenv("GROQ_API_KEY"))

class IncidentCreate(BaseModel):
    service: str
    severity: str
    symptoms: str
    logs_snippet: str

class ResolveRequest(BaseModel):
    root_cause: str
    resolution_steps: str
    runbook_used: str

@app.post("/incidents/{incident_id}/analyze")
async def analyze_incident(incident_id: str, payload: IncidentCreate):
    # 1. Recall historical incidents from Hindsight
    recalled_matches = await memory_store.recall_similar(
        service=payload.service,
        symptoms=payload.symptoms,
        limit=2
    )

    context_str = ""
    for match in recalled_matches:
        context_str += f"\n- Historical Case: {match['content']} (Score: {match.get('similarity', 0):.2f})"

    # 2. Invoke Groq Llama-3.3 for grounded root-cause analysis
    system_prompt = (
        "You are an on-call triage agent. Analyze the active incident using the provided "
        "historical context. Ground your conclusions in past evidence. Cite past cases explicitly."
    )
    user_prompt = f"""
    Current Incident:
    Service: {payload.service}
    Severity: {payload.severity}
    Symptoms: {payload.symptoms}
    Logs: {payload.logs_snippet}

    Retrieved Historical Context:
    {context_str if context_str else "No matching historical incidents found."}

    Output format:
    1. Probable Root Cause (with % confidence)
    2. Ranked Resolution Steps
    3. Recommended Runbook to execute
    """

    chat_completion = groq_client.chat.completions.create(
        messages=[
            {"role": "system", "content": system_prompt},
            {"role": "user", "content": user_prompt}
        ],
        model=os.getenv("GROQ_MODEL", "llama-3.3-70b-versatile"),
        temperature=0.1
    )

    return {
        "incident_id": incident_id,
        "historical_matches": recalled_matches,
        "analysis": chat_completion.choices[0].message.content
    }

@app.post("/incidents/{incident_id}/resolve")
async def resolve_incident(incident_id: str, data: ResolveRequest, payload: IncidentCreate):
    # Persist the confirmed resolution back into Hindsight episodic memory
    memory_id = await memory_store.retain_incident({
        "id": incident_id,
        "service": payload.service,
        "severity": payload.severity,
        "symptoms": payload.symptoms,
        "root_cause": data.root_cause,
        "resolution_steps": data.resolution_steps,
        "runbook_used": data.runbook_used
    })
    return {"status": "resolved", "memory_id": memory_id}

Enter fullscreen mode Exit fullscreen mode

Results and Behavior Verification

To verify the pipeline end-to-end, we executed our demo scenario: an unexpected database pool exhaustion following a canary deployment.

  1. Ingest Active Failure:
curl -X POST "http://localhost:8000/incidents/inc-902/analyze" \
  -H "Content-Type: application/json" \
  -d '{
    "service": "billing-service",
    "severity": "P1",
    "symptoms": "HTTP 500 error spikes; DB connection timeout during checkout",
    "logs_snippet": "FATAL: remaining connection slots are reserved for non-replication superuser connections"
  }'

Enter fullscreen mode Exit fullscreen mode
  1. Observed System Output: The agent queried Hindsight, identified a previous occurrence in billing-service with 93% semantic similarity, and returned the triage deduction in 1.4 seconds:
1. Probable Root Cause: 
   Database connection pool exhaustion caused by unclosed connections in the 
   canary release (Matches historical incident inc-084, Confidence: 94%).

2. Ranked Resolution Steps:
   - Step 1: Execute Runbook RB-DB-DRAIN-RECYCLE immediately.
   - Step 2: Roll back canary release v2.14.1 to restore baseline connection pool hygiene.
   - Step 3: Monitor connection acquisition latency in Grafana.

3. Recommended Runbook: 
   RB-DB-DRAIN-RECYCLE

Enter fullscreen mode Exit fullscreen mode
  1. Resolve and Update Memory: When the on-call engineer confirmed the fix, calling /incidents/inc-902/resolve wrote the post-mortem metadata into Hindsight. When a similar connection drop occurred on another microservice two days later, the agent immediately surfaced RB-DB-DRAIN-RECYCLE with increased weight.

Lessons Learned

  1. Strict Output Structures Matter: Using low temperature settings (0.1) and structured instructions prevents the model from embellishing operational steps or speculating beyond the evidence in memory.
  2. Asynchronous Memory Retain: Retaining resolved incidents directly in the HTTP request thread can introduce unnecessary latency for the user. Decouple memory persistence into background worker tasks for high-load production environments.
  3. Sanitize Incident Data Before Ingestion: Alert logs frequently contain API keys, authorization tokens, or sensitive payload headers. Always scrub secrets using regex filters before writing logs into external memory layers.
  4. Dynamic Context Trumps Prompt Engineering: Tweaking prompt wording will never solve knowledge gaps. Equipping your agent with persistent episodic memory enables the system to continuously adapt alongside your evolving infrastructure.

By wrapping FastAPI and Groq around Hindsight's memory store, you can build a resilient triage agent that turns isolated post-mortems into an automated operational safety net.
Project Interface & Operational Walkthrough
Here is a look at the live user interface built for on-call engineers to triage incidents in real time:

Figure 1:
 The main triage console showing active alert analysis, root-cause deduction with confidence scoring, and past incident citations.

When an alert triggers, the engineer interacts with three key components:

Explainable Confidence Scores: The composite ranking displaying both semantic vector similarity and historical runbook win-rates directly on screen.

Cited Historical Evidence: Direct references to previous incident post-mortems retrieved from Hindsight memory, removing guesswork during live outages.

One-Click Human Feedback: Thumbs up and thumbs down controls that dynamically adjust runbook effectiveness weights for future triage cycles.

Figure 2:
 The memory explorer and analytics screen displaying runbook success rates, MTTR reduction trends, and episodic memory retention.

Integrating stateful agent memory turned our incident agent from a novelty chatbot into a reliable on-call co-pilot that gets smarter every time production breaks.

Top comments (0)