Designing a 3-Tier Agent Memory Engine Using Hindsight
Giving an LLM a massive raw context window during a live infrastructure outage creates noise, latency, and hallucinations—not actionable solutions. When an API gateway starts shedding traffic at 5,000 requests per second, feeding 20 pages of past incident post-mortems into a prompt buffer burns critical seconds while leaving engineers to guess whether the surfaced recommendations actually worked in the past.
To build a production-grade triage system, we had to stop treating agent state as a flat collection of text embeddings. We designed a stateful incident response engine built on FastAPI, Groq's Llama-3.3-70b, and a multi-layer memory architecture powered by Hindsight.
Here is the system architecture, mathematical retrieval model, and state-management implementation behind our 3-tier memory engine.
What the System Does and How It Hangs Together
Our incident response platform automates live triage by pairing deterministic retrieval with bounded generative reasoning.
Rather than relying on a single vector index, the platform divides knowledge retention into three distinct operational domains:
- Episodic Memory: Granular historical records of every incident—timestamps, impacted microservices, severity levels, raw symptom profiles, stack traces, identified root causes, and explicit remediation steps.
- Semantic Memory: Curated operational runbooks, disaster recovery procedures, and verified post-mortem post-incident reviews (PIRs).
- Effectiveness Memory: Dynamic, mutable telemetry tracking how often a specific runbook was recommended, how often it succeeded, its historical Mean Time to Resolution (MTTR), and operator feedback ratings.
+-----------------------------------------------------------------------------------+
| INGESTION & TRIAGE LAYER |
| [ PagerDuty Alert ] ---> [ FastAPI Gateway ] ---> [ Incident Entity Extraction ] |
+-----------------------------------------------------------------------------------+
|
v
+-----------------------------------------------------------------------------------+
| 3-TIER MEMORY RECALL SYSTEM |
| |
| +-----------------------+ +----------------------+ +--------------------+ |
| | Episodic Memory | | Semantic Memory | | Effectiveness Meta | |
| | (Hindsight Store) | | (Runbook Guides) | | (SQLite / Stats) | |
| +-----------------------+ +----------------------+ +--------------------+ |
| \ | / |
| v v v |
| [ Symptom Similarity ] [ Runbook Correlation ] [ Historical Win-Rate ] |
| | |
| v |
| [ Hybrid Scoring & Re-Ranking Engine ] |
+-----------------------------------------------------------------------------------+
|
v
+-----------------------------------------------------------------------------------+
| GENERATION & LEARNING |
| [ Groq Llama-3.3-70b ] ---> Cites Evidence, Confidence %, Ranked Action Plan |
| [ Operator Feedback ] ---> Mutates Weights & Retains Post-Mortem in Hindsight |
+-----------------------------------------------------------------------------------+
FastAPI handles incoming webhooks and exposes REST endpoints. SQLite serves as the transactional source of truth for runtime application states, while Hindsight manages our active episodic and semantic agent context.
Core Technical Story: The Hybrid Scoring Retrieval Engine
Standard vector search prioritizes syntactic similarity over operational truth. In an infrastructure outage, two distinct incidents might share identical generic symptoms—such as connection reset by peer or HTTP 504 Gateway Timeout—while stemming from entirely different origins, such as an expired TLS certificate versus an out-of-memory worker pool.
Relying exclusively on dense vector proximity returns false matches. Conversely, strictly filtering by microservice tags breaks cross-service incident correlation (for instance, a Kafka broker bottleneck presenting as an auth service failure).
To resolve this, we designed a composite retrieval engine that computes an explainable hybrid score:
$$\text{FinalScore} = (w_{\text{sim}} \times S_{\text{vector}}) + (w_{\text{svc}} \times S_{\text{service}}) + (w_{\text{sev}} \times S_{\text{severity}}) + (w_{\text{eff}} \times E_{\text{runbook}}) - (\lambda \times \Delta t)$$
Where:
$S_{\text{vector}}$: Cosine similarity surfaced from the Hindsight documentation API via semantic symptom comparison.
$S_{\text{service}}$: Binary or topological service overlap ($1.0$ for exact match, $0.5$ for upstream/downstream dependencies, $0.0$ for unrelated).
$S_{\text{severity}}$: Distance between alert tiers (e.g., matching a
P1outage with anotherP1rather than a minor warning).$E_{\text{runbook}}$: Empirical success rate of the linked runbook ($\frac{\text{successes}}{\text{total executions}}$).
$\lambda \times \Delta t$: Time decay constant penalizing outdated architecture contexts older than 180 days.
This design ensures our agent memory delivers context that is not just textually similar, but operationally reliable and proven in production.
Code-Backed Architecture
1. Abstracting the Hybrid Memory Engine
The MemoryStore interface exposes both vector persistence and multi-attribute hybrid retrieval methods across our persistent layers.
from abc import ABC, abstractmethod
from typing import List, Dict, Any
from pydantic import BaseModel
class ScoredIncidentMatch(BaseModel):
incident_id: str
content: str
vector_similarity: float
service_match_score: float
runbook_success_rate: float
composite_score: float
explanation: Dict[str, float]
class MemoryEngine(ABC):
@abstractmethod
async def retain_incident_lifecycle(self, incident: Dict[str, Any]) -> str:
"""Persist incident context, root cause, and metadata into Hindsight."""
pass
@abstractmethod
async def recall_hybrid_context(
self,
service: str,
severity: str,
symptoms: str,
limit: int = 5
) -> List[ScoredIncidentMatch]:
"""Execute hybrid re-ranking over episodic context."""
pass
2. Implementing Hybrid Scoring with Hindsight & Relational State
Our hybrid implementation coordinates the vector recall from Hindsight with local effectiveness metrics stored in SQLite.
import os
import httpx
from datetime import datetime
from typing import List, Dict, Any
from app.models.runbook import RunbookStats
class HindsightHybridEngine(MemoryEngine):
def __init__(self, db_session):
self.api_key = os.getenv("HINDSIGHT_API_KEY")
self.base_url = os.getenv("HINDSIGHT_BASE_URL", "https://api.hindsight.vectorize.io/v1")
self.headers = {
"Authorization": f"Bearer {self.api_key}",
"Content-Type": "application/json"
}
self.db = db_session
async def recall_hybrid_context(
self, service: str, severity: str, symptoms: str, limit: int = 5
) -> List[ScoredIncidentMatch]:
# 1. Fetch top semantic candidates from Hindsight
payload = {
"query": f"Service: {service}. Symptoms: {symptoms}",
"filter": {"context_type": "episodic_incident"},
"top_k": limit * 2 # Oversample for downstream re-ranking
}
async with httpx.AsyncClient(timeout=10.0) as client:
resp = await client.post(f"{self.base_url}/recall", json=payload, headers=self.headers)
resp.raise_for_status()
raw_memories = resp.json().get("results", [])
scored_candidates = []
for mem in raw_memories:
meta = mem.get("metadata", {})
vector_sim = float(mem.get("similarity", 0.0))
# Service match component
svc_score = 1.0 if meta.get("service") == service else 0.2
# Severity match component
sev_score = 1.0 if meta.get("severity") == severity else 0.5
# Fetch real-time runbook win-rate from relational storage
runbook_id = meta.get("runbook_used")
stats = self.db.query(RunbookStats).filter_by(runbook_id=runbook_id).first()
eff_score = (stats.successes / stats.total_attempts) if stats and stats.total_attempts > 0 else 0.5
# Compute weighted composite score
composite = (
(vector_sim * 0.40) +
(svc_score * 0.25) +
(eff_score * 0.25) +
(sev_score * 0.10)
)
scored_candidates.append(ScoredIncidentMatch(
incident_id=meta.get("document_id", "unknown"),
content=mem.get("content", ""),
vector_similarity=round(vector_sim, 3),
service_match_score=svc_score,
runbook_success_rate=round(eff_score, 3),
composite_score=round(composite, 3),
explanation={
"vector_weight": round(vector_sim * 0.40, 3),
"service_weight": round(svc_score * 0.25, 3),
"effectiveness_weight": round(eff_score * 0.25, 3)
}
))
# Sort descending by composite score
scored_candidates.sort(key=lambda x: x.composite_score, reverse=True)
return scored_candidates[:limit]
3. Feedback Loop: Dynamic Weight Recalibration
When engineers complete an incident, feedback is recorded. This mutates local effectiveness memory and writes an episodic update back to Hindsight.
async def record_incident_resolution(
db,
memory_engine: MemoryEngine,
incident_id: str,
runbook_id: str,
success: bool,
postmortem_notes: str
):
# 1. Update Effectiveness Memory in Relational Store
stats = db.query(RunbookStats).filter_by(runbook_id=runbook_id).first()
if stats:
stats.total_attempts += 1
if success:
stats.successes += 1
db.commit()
# 2. Retain post-mortem insight into Hindsight episodic memory
await memory_engine.retain_incident_lifecycle({
"id": incident_id,
"runbook_used": runbook_id,
"outcome": "success" if success else "failed",
"notes": postmortem_notes,
"resolved_at": datetime.utcnow().isoformat()
})
Results and Behavior Verification
We verified the multi-tier memory architecture against a classic cascading failure: a Redis cache eviction wave causing downstream relational connection saturation.
When querying the system with the symptom string: Redis connection reset during key eviction spike; auth-service latency > 1200ms:
Standard Dense Vector Retrieval (Single Layer)
- Top Match:
INC-012(Redis cluster network interface partition). - Vector Similarity:
0.89. - Issue: Suggested rebuilding the Redis cluster network topology—a 45-minute invasive procedure that would not solve key eviction storms.
3-Tier Hybrid Retrieval (Hindsight + Effectiveness)
The system evaluated raw vector matches against historical operational win-rates and surfaced the following explainable breakdown:
{
"top_recalled_incident": "INC-044",
"recommended_runbook": "RB-REDIS-VOLATILE-LRU-SCALE",
"vector_similarity": 0.84,
"service_match_score": 1.0,
"runbook_success_rate": 0.94,
"composite_score": 0.891,
"score_breakdown": {
"semantic_similarity_points": 0.336,
"service_points": 0.250,
"historical_win_rate_points": 0.235,
"severity_alignment_points": 0.070
},
"rationale": "INC-044 has slightly lower raw log similarity than INC-012, but its runbook boasts a 94% verified resolution rate for memory-eviction spikes."
}
By surfacing the composite score, the engineer saw exactly why RB-REDIS-VOLATILE-LRU-SCALE was chosen over network restarts. The runbook dynamically adjusted maxmemory-policy and provisioned read replicas, resolving the degradation in under four minutes.
Lessons Learned
- Decouple Dynamic Metrics from Static Vector Embeddings: Never embed mutable stats (like success rates or resolution times) directly into raw vector documents. Keep vector embeddings focused on semantic descriptions, and inject dynamic effectiveness weights at query-time via hybrid re-ranking.
- Explainability Breeds Trust Under Pressure: Engineers ignoring an agent's recommendation usually do so because the suggestion looks like a black-box guess. Exposing the exact mathematical breakdown of vector score versus historical success rate creates immediate operational confidence.
- Write Path Isolation Is Crucial: Post-incident updates and feedback loops must happen asynchronously. If writing a post-mortem to memory blocks the API response during incident closure, network hiccups will cause triage records to drop. Use background tasks or messaging queues to retain context into Hindsight reliably.
- Architect for Graceful Degradation: Always provide fallback logic. In our setup, if the external memory API exceeds strict response budgets, our architecture automatically degrades to local vector search without terminating the triage pipeline.
Separating memory into episodic context, semantic runbooks, and real-time effectiveness telemetry transforms an unpredictable LLM into a disciplined, data-driven systems engineer.
Project Interface & Operational Walkthrough
Here is a look at the live user interface built for on-call engineers to triage incidents in real time:

Figure 1: The main triage console showing active alert analysis, root-cause deduction with confidence scoring, and past incident citations.
When an alert triggers, the engineer interacts with three key components:
Explainable Confidence Scores: The composite ranking displaying both semantic vector similarity and historical runbook win-rates directly on screen.
Cited Historical Evidence: Direct references to previous incident post-mortems retrieved from Hindsight memory, removing guesswork during live outages.
One-Click Human Feedback: Thumbs up and thumbs down controls that dynamically adjust runbook effectiveness weights for future triage cycles.

Figure 2: The memory explorer and analytics screen displaying runbook success rates, MTTR reduction trends, and episodic memory retention.
Integrating stateful agent memory turned our incident agent from a novelty chatbot into a reliable on-call co-pilot that gets smarter every time production breaks.
Top comments (0)