Building Stateful Incident Investigation with Hindsight Agent Memory
When production breaks, every engineer asks the same fundamental question: what changed, and have we solved this before?
Most automated incident response tools treat LLMs as stateless calculators. You feed in a blob of raw logs, prompt the model to diagnose the issue, and display whatever markdown text it generates. Once the HTTP request finishes, everything about that investigation vanishes. The next time the exact same microservice fails with the exact same symptoms, your system starts from complete scratch.
When building DeployLens—an incident investigation agent for our fictional e-commerce SaaS platform, NovaCart—we realized early on that traditional request/response AI reasoning was insufficient. Relational application state and long-term operational memory serve two completely different engineering requirements.
We used SQLite to handle standard relational state such as service registries, deployment records, and alert histories. But for long-term operational experience—such as postmortem findings, verified resolutions, and failed fix attempts—we integrated the Hindsight agent memory framework.
Here is how we designed a stateful investigation architecture that separates relational database state from persistent vector memory.
The Dual-Storage Architecture
An incident investigation agent requires two distinct storage paradigms working in tandem.
Relational databases excel at structured schema constraints, primary key relations, and transactional updates. For example, linking service_id on an alerts table to a services table requires strict foreign keys. However, querying a relational database for past incidents based on ambiguous symptoms like "intermittent Redis timeout during checkout surge" requires rigid SQL LIKE queries or complex full-text search indexes that fail when phrasing differs slightly across postmortems.
This is where Vectorize agent memory fundamentally changes the system design. Hindsight stores operational experiences as semantic memory units inside dedicated memory banks. Rather than searching exact string matches, Hindsight performs vector retrieval over past operational events using semantic similarity, tags, and metadata filters.
[DIAGRAM: DeployLens dual-storage architecture showing SQLite relational DB vs Hindsight memory bank]
In DeployLens, all transactional operational data stays in SQLite, while semantic operational memories live inside a dedicated Hindsight memory bank named novacart-production-memory.
Wrapping the Hindsight Client
To ensure our FastAPI backend remained clean and decoupled, we wrapped the official hindsight-client Python SDK in a dedicated wrapper class, HindsightClientWrapper.
Here is how we initialize the Hindsight client and handle connection status in backend/app/memory/hindsight_client.py:
python
class HindsightClientWrapper:
"""
Wrapper around official hindsight-client Python SDK.
Provides robust status reporting and dynamic connection handling.
"""
def __init__(self):
self.bank_id = settings.HINDSIGHT_BANK_ID
self.base_url = settings.HINDSIGHT_BASE_URL
self.api_key = settings.HINDSIGHT_API_KEY
self.client = None
self._is_connected = False
self._init_client()
def _init_client(self):
self.bank_id = settings.HINDSIGHT_BANK_ID
self.base_url = settings.HINDSIGHT_BASE_URL
self.api_key = settings.HINDSIGHT_API_KEY
try:
from hindsight_client import Hindsight
kwargs = {"base_url": self.base_url}
if self.api_key:
kwargs["api_key"] = self.api_key
self.client = Hindsight(**kwargs)
if (
self.api_key
or "localhost" in self.base_url
or "vectorize.io" in self.base_url
):
self._is_connected = True
else:
self._is_connected = False
except Exception as e:
logger.error(f"Failed to initialize Hindsight SDK client: {e}")
self.client = None
self._is_connected = False
Why This Matters
Wrapping the SDK inside `HindsightClientWrapper` gives the application two essential guarantees:
1. **Dynamic Configuration Reloading:** If environment variables change, the client re-reads configuration settings without forcing a complete application rebuild.
2. **Explicit Health Status:** The backend can report connection status directly to the Next.js frontend, ensuring the user interface displays `Hindsight Connected` or `Hindsight Unavailable` rather than throwing uncaught HTTP exceptions.
Memory Service Abstraction and Local Caching
Above the raw SDK wrapper, we built `MemoryService` in `backend/app/memory/memory_service.py`. This service manages writing memories to both Hindsight and a local in-memory fallback store when running in offline or testing environments.
Here is the implementation of `remember_event` and `add_local_memory`:
python
def add_local_memory(self, content: str, source_type: str, metadata: Dict[str, Any]):
self._local_memories.append({
"id": f"MEM-{len(self._local_memories) + 1001:04d}",
"content": content,
"source_type": source_type,
"metadata": metadata,
"timestamp": datetime.utcnow().isoformat()
})
def remember_event(
self,
content: str,
source_type: str,
metadata: Dict[str, Any],
tags: Optional[List[str]] = None
) -> str:
tag_list = tags or [
source_type,
metadata.get("service", "unknown")
]
mem_id = f"MEM-{len(self._local_memories) + 1001:04d}"
# Save to local fallback store
self.add_local_memory(content, source_type, metadata)
# Save to Hindsight
success = hindsight_client.retain(
content=content,
metadata=metadata,
tags=tag_list
)
logger.info(
f"remember_event [{source_type}] Hindsight status: {success}"
)
return mem_id
Why This Matters
The `remember_event` method standardizes how operational events enter long-term storage.
Every event receives:
* **Semantic Content:** Formatted narrative text summarizing the deployment or resolution.
* **Source Metadata:** Attributes like `service`, `incident_id`, and `version`.
* **Tags:** High-level category markers (`deployment`, `resolution`, `checkout-api`) used by Hindsight to index memory units for precise vector recall.
* **Dual Retention:** Storing events locally in `_local_memories` ensures the system can continue operating smoothly even if external connectivity is temporarily interrupted.
## Visualizing Memory Connection Status
In the frontend, judges and engineers must never have to guess whether the memory layer is active. We surfaced the connection status directly in the primary navigation header.
When Hindsight is active, the UI displays a green indicator reading `Hindsight Connected` pointing to the `novacart-production-memory` bank.
## Before vs After: Stateless RAG vs Hindsight Memory Architecture
To illustrate why this dual-storage architecture matters, consider how an investigation proceeds with and without Hindsight memory.
## WITHOUT HINDSIGHT MEMORY
When `checkout-api` experiences an error rate spike to 18.4% alongside Redis connection acquisition timeout warnings, a traditional stateless agent inspects only current telemetry:
1. System reads current alert: `REDIS_TIMEOUT` on `redis-cluster`.
2. System reads current incident: `INC-2051` on `checkout-api`.
3. Agent output: "Redis connection timed out. Check network connectivity, inspect application logs, or restart the `checkout-api` service."
Because the agent has no cross-session memory, it recommends generic troubleshooting steps and fails to recognize that restarting `checkout-api` was already attempted during previous outages without success.
## WITH HINDSIGHT MEMORY
With the [Hindsight documentation API](https://hindsight.vectorize.io/) integrated into the investigation pipeline:
1. System reads current alert and deployment records (`DEP-1837` updating `payment-service` config `PAYMENT_REDIS_POOL_SIZE=20`).
2. System queries Hindsight memory bank `novacart-production-memory` for matching symptom patterns.
3. Hindsight recalls `INC-1042` (91% similarity match): *"INC-1042 affected checkout-api. Resolved by increasing payment-service Redis pool size from 20 to 50. Restarting checkout-api 3 times provided only temporary relief."*
4. Agent output: "Do NOT restart `checkout-api`. Increasing `PAYMENT_REDIS_POOL_SIZE` from 20 to 50 in `payment-service` environment settings resolved identical outage `INC-1042`."
## Architectural Limitations and Tradeoffs
While this architecture significantly improves investigation accuracy, it introduces explicit tradeoffs:
1. **Network Latency:** Querying an external vector memory bank introduces an HTTP round-trip (typically 120ms–250ms) during investigation context assembly. We mitigated this by running memory recall concurrently with local database queries.
2. **Formatting Overhead:** Operational data must be converted into structured semantic narratives before calling `retain()`. Raw JSON dumps produce inferior vector recall results compared to carefully formatted operational sentences.
3. **Memory Synchronization:** When running multiple instance replicas, in-memory local caches must be synchronized or refreshed periodically to maintain consistency with the remote Hindsight vector index.
## Engineering Takeaways
1. **Separate State from Experience:** Use relational databases for transactional data constraints and Hindsight for semantic operational memories.
2. **Standardize Memory Formatting:** Format operational events into clear semantic prose before storing them.
3. **Decouple via Service Wrappers:** Wrap external memory SDKs behind internal service interfaces to allow graceful degradation when offline.
4. **Surface System Health:** Always display memory layer connectivity status directly in the user interface.
5. **Tag Operational Context Explicitly:** Attach service names, temporal timestamps, and outcome types as metadata during memory retention.


Top comments (0)