We subjected 25 real production incident traces to a head-to-head evaluation: an ungrounded, stateless Llama-3 model versus the exact same model backed by Hindsight for persistent incident memory. The test was simple: present identical raw alerts across Redis, Kafka, and Kubernetes, measure the quality of root-cause deductions, and track whether the recommended mitigations actually fixed the problem.
The outcome was definitive. Stateless LLMs consistently produced generic, textbook troubleshooting checklists that wasted critical time during outages. Without historical context, they cannot determine which microservice version introduced a leak, nor do they know which runbooks previously succeeded or failed.
Here are the empirical benchmarks, failure modes, and architectural differences that convinced us to abandon stateless prompts in favor of persistent agent memory.
What the System Does and How It Hangs Together
Our automated triage agent operates in front of our on-call engineers. When an alert fires in monitoring systems like Prometheus or Datadog, it routes directly to a FastAPI service.
Instead of dumping the entire alert payload into an empty LLM conversation window, the agent orchestrates a retrieval pipeline:
- Symptom Extraction: The incoming alert, stack trace, and service identifiers are structured into a uniform query.
- Episodic Memory Query: The agent queries Hindsight to fetch the top matching historical incidents along with their documented post-mortems and resolutions.
- Runbook Correlation: The agent checks the semantic runbook library to find procedures matching the incident profile.
- Grounded Synthesis: Groq's high-speed Llama-3.3-70b engine evaluates the active incident against recalled historical evidence, generating a prioritized action plan that explicitly cites past incident IDs.
+-------------------------------------------------------------------------------+
| BENCHMARK EVALUATION HARNESS |
+-------------------------------------------------------------------------------+
|
+----------------------+----------------------+
| |
v v
[ Pipeline A: Stateless Prompt ] [ Pipeline B: Stateful Agent ]
- Raw Alert + Generic System Prompt - Raw Alert + Hindsight Context
- No Historical Memory - Recalled Past Incidents & Runbooks
- LLM: Groq Llama-3.3-70b - LLM: Groq Llama-3.3-70b
| |
+----------------------+----------------------+
|
v
+---------------------------------------------+
| EVALUATION METRICS |
| - Hallucination Rate |
| - Correct Runbook Selection % |
| - MTTR Reduction Potential |
| - Actionability Score (1-5) |
+---------------------------------------------+
Core Technical Story: The Benchmark Methodology
To run a fair empirical test, we curated an evaluation suite of 25 complex infrastructure incidents from historical post-mortems across our stack:
- Database & Cache: Postgres connection starvation, Redis key eviction cascades, replication lag.
- Orchestration: Kubernetes CrashLoopBackOff due to OOMKills, node disk pressure, broken ConfigMap mounts.
- Streaming & Auth: Kafka consumer rebalance loops, expired TLS certificates on API gateways.
Each test case included the exact raw logs, affected services, error strings, and verified ground-truth fixes documented in historical post-mortems.
We evaluated both pipelines across four objective criteria:
- Specific Root-Cause Identification: Did the model identify the actual failure mechanism or settle for a vague description?
- Hallucination & Speculation: Did the model fabricate configuration flags, nonexistent CLI commands, or imagine historical context?
- Correct Runbook Retrieval: Did the pipeline identify the precise, verified mitigation script?
- Actionable First Step: Was the immediate recommendation safe to execute in production without causing secondary cascading failures?
Code-Backed Implementations: The Two Pipelines
1. The Stateless Baseline Pipeline
The baseline pipeline reflects the typical implementation: passing system instructions and the raw error trace to the LLM.
async def run_stateless_triage(alert_data: dict, llm_client) -> dict:
prompt = f"""
You are an expert SRE triage assistant.
Analyze the following production alert and provide:
1. Probable root cause
2. Ranked troubleshooting steps
3. Recommended remediation
Alert Service: {alert_data['service']}
Severity: {alert_data['severity']}
Logs: {alert_data['logs']}
"""
# Direct execution with no access to past incidents or runbooks
return await llm_client.generate_json(prompt)
2. The Stateful Hindsight Pipeline
The stateful pipeline leverages the Hindsight documentation API to pull verified institutional memory before generating the response.
async def run_stateful_triage(alert_data: dict, memory_client, llm_client) -> dict:
# 1. Recall historical incidents matching symptoms and service
recalled_incidents = await memory_client.recall(
query=f"Service: {alert_data['service']} Symptoms: {alert_data['symptoms']}",
context_type="episodic_incident",
top_k=3
)
# 2. Extract proven fixes and historical post-mortem context
memory_context = "\n---\n".join([
f"Past Incident ID: {item['metadata']['document_id']}\n"
f"Resolution: {item['content']}\n"
f"Runbook Used: {item['metadata'].get('runbook_used')}"
for item in recalled_incidents
])
# 3. Grounded generation prompt
prompt = f"""
You are an expert SRE triage assistant with access to verified past incidents.
Analyze the following alert and ground your diagnosis in the historical context.
Active Alert:
Service: {alert_data['service']}
Logs: {alert_data['logs']}
Verified Past Incidents from Memory:
{memory_context}
Rules:
- Cite the specific Incident ID that matches this failure pattern.
- If a past incident matches, suggest the exact runbook that succeeded previously.
- If no historical match exists, state clearly: 'NO_HISTORICAL_MATCH'.
"""
return await llm_client.generate_json(prompt)
Results & Comparative Evaluation
Across the 25 benchmark scenarios, the stateful agent decisively outperformed the stateless baseline across every core metric:
| Metric | Stateless Baseline (Llama-3.3) | Stateful Agent (Hindsight + Llama-3.3) |
|---|---|---|
| Root Cause Accuracy | 44% (11/25) | 92% (23/25) |
| Hallucination / Speculation Rate | 36% (9/25) | 0% (0/25) |
| Accurate Runbook Selection | 28% (7/25) | 96% (24/25) |
| Dangerous Recommendations | 4 cases (e.g., primary DB reboot) | 0 cases |
| Average Prompt Tokens | 420 tokens | 1,180 tokens |
Concrete Test Case: Kafka Consumer Group Rebalance Storm
-
Alert Payload:
CommitFailedException: Kafka consumer poll timeout exceeded; consumer group rebalancing constantly. -
Ground Truth Fix: Increase
max.poll.interval.msfrom300000to900000and scale batch size down to50to handle long-running payload processing.
Stateless Baseline Response:
Diagnosis: Kafka consumer group is failing heartbeats.
Action Plan:
1. Restart all Kafka broker nodes sequentially.
2. Check network connectivity between consumers and broker cluster.
3. Scale up consumer pods from 5 to 15 replicas.
Impact: Restarting brokers would have aggravated cluster instability, while adding consumer pods directly accelerated rebalancing storms.
Stateful (Hindsight) Response:
Diagnosis: Matches Incident #INC-067 (95% similarity). Consumer thread blocked
during database batch insertion, exceeding max.poll.interval.ms.
Action Plan:
1. Do NOT scale consumer pods (will trigger additional rebalance waves).
2. Execute Runbook RB-KAFKA-CONSUMER-TUNE.
3. Apply hotfix config: max.poll.interval.ms=900000, max.poll.records=50.
4. Verified in INC-067 post-mortem: resolves storm within 90 seconds.
Lessons Learned
- Stateless LLMs Default to Disruptive Fixes: When uncertain, standard models frequently suggest broad operational interventions (rebooting instances, scaling replicas, flushing caches). In distributed systems, these generic fixes often turn minor localized degradations into systemic outages.
- Context Compression Beats Long Context Windows: Dumping raw runbooks and entire log indexes into long context windows degrades attention and inflates inference latency. Recalling targeted episodic chunks via Hindsight kept our prompts lean, focused, and fast.
-
Deterministic Constraint Stops Hallucinations: Constraining the model to cite specific historical incident IDs or explicitly declare
NO_HISTORICAL_MATCHeliminated speculative diagnosis entirely. - Historical Memory Converts On-Call Runbooks into Living Knowledge: Operational runbooks rot when left as static markdown documents. Grounding runtime triage in retained episodic memory ensures that as soon as a post-mortem is resolved, every on-call engineer immediately benefits from the fix.
Persistent agent memory isn't an optional enhancement for incident triage—it is the difference between an LLM guessing in the dark and an agent operating with institutional knowledge.
Project Interface & Operational Walkthrough
Here is a look at the live user interface built for on-call engineers to triage incidents in real time:
Figure 1:

The main triage console showing active alert analysis, root-cause deduction with confidence scoring, and past incident citations.
When an alert triggers, the engineer interacts with three key components:
Explainable Confidence Scores: The composite ranking displaying both semantic vector similarity and historical runbook win-rates directly on screen.
Cited Historical Evidence: Direct references to previous incident post-mortems retrieved from Hindsight memory, removing guesswork during live outages.
One-Click Human Feedback: Thumbs up and thumbs down controls that dynamically adjust runbook effectiveness weights for future triage cycles.
Figure 2:

The memory explorer and analytics screen displaying runbook success rates, MTTR reduction trends, and episodic memory retention.
Integrating stateful agent memory turned our incident agent from a novelty chatbot into a reliable on-call co-pilot that gets smarter every time production breaks.
Top comments (0)