Building OpsMemory: Injecting Persistent Memory into Incident Response
If you spend enough time in production engineering, you start to notice a frustrating pattern: production incidents repeat. Not always exactly, but often close enough that the mitigation steps are identical. When alarms go off at 2:00 AM, someone is usually digging through fragmented Slack channels or outdated post-mortem wikis to figure out how the team fixed a database connection pool exhaustion months ago.
When an incident occurs, the context required to solve it is largely tribal knowledge. Organizations rely heavily on the tenure of the on-call engineer. If they haven't seen the specific failure mode before, they start from zero. They rediscover the same dead ends, re-run the same diagnostic commands, and eventually arrive at a resolution that someone else in the organization had already figured out previously.
We built OpsMemory to address this specific bottleneck. OpsMemory is an AI-powered incident response agent that turns production incidents into reusable organizational knowledge.
The central idea driving the project is simple: Every production incident should make the next incident easier to solve.
Why Persistent Memory Matters
There is a growing temptation to solve incident response simply by placing a Large Language Model (LLM) in front of an engineer. But out of the box, an LLM is a stateless reasoning engine. It understands generic software engineering concepts, but it doesn't know your specific architecture, your unique failure modes, or your organization's history.
If you ask a standard AI assistant to help diagnose a payment service timeout, it will give you a generic list of possible network, database, and application issues. It lacks the context required to narrow down the search space.
To make AI genuinely useful during a firefight, it needs persistent organizational memory. It needs to know what broke yesterday, how it was fixed, and what lessons were learned.
How Hindsight Fits into the Architecture
In OpsMemory, we use Groq (specifically the openai/gpt-oss-120b model) as our reasoning layer, and Hindsight as our persistent memory layer.
Hindsight provides persistent memory and semantic recall, allowing OpsMemory to retrieve relevant knowledge from previous incidents and retain verified resolutions for future incidents. For the MVP, Hindsight serves as our persistent memory layer, so we don't need to build a separate incident-knowledge store.
The interaction with Hindsight revolves around two primary operations:
- Recall: When a new incident is reported, we query Hindsight to retrieve similar historical incidents and their outcomes.
- Retain: When an incident is successfully resolved, we push the validated root cause and resolution back into Hindsight, allowing the memory bank to consolidate and improve over time.
From Incident to Recommendation
The core OpsMemory workflow relies on a tight feedback loop: Recall → Reason → Resolve → Retain → Recall again.
- An engineer reports an ongoing incident in the UI.
- OpsMemory queries Hindsight to recall relevant historical incidents.
- The current incident description, along with the recalled historical context, is packaged and sent to the Groq LLM.
- The AI generates a structured incident analysis, identifying a likely root cause, recommended response actions, investigation steps, and prevention measures.
- The engineer investigates, mitigates the issue, and then verifies the actual root cause and resolution.
- The verified resolution is retained in Hindsight.
- Future incidents can benefit from that accumulated knowledge.
Human Verification as a Safety Boundary
Step 5 is a critical design decision in our architecture. The engineer explicitly verifies the AI's hypothesis before the resolution is retained.
An AI-generated diagnosis is a hypothesis, not guaranteed ground truth. We do not claim that OpsMemory automatically fixes production incidents, nor that its initial root cause guesses are guaranteed to be correct. Automatically storing incorrect AI assumptions would contaminate the memory bank, leading to degraded advice in future incidents.
Human verification acts as a necessary quality gate. The AI accelerates the investigation by bringing historical context to the surface, but the human confirms reality before organizational knowledge becomes persistent. The AI produces a likely diagnosis, the engineer verifies the actual root cause, resolution, and outcome, and only that verified information is retained.
The Architecture
To build this, we kept the stack standard and robust.
- Frontend: A React and Vite single-page application, styled with Lucide React for a clean, command-center aesthetic. It is deployed as a static site.
- Backend: A Java 17 Spring Boot API utilizing Spring WebFlux for non-blocking HTTP client calls. It is deployed as a Web Service.
- AI/Memory Integration: The backend orchestrates the Hindsight and Groq integrations.
The API flow is conceptually straightforward:
POST /api/incidents/analyze
Receives the incident description, triggers a Hindsight recall, queries Groq with the combined context, and returns the AI's analysis.
POST /api/incidents/resolve
Receives the engineer's verified resolution and root cause, and sends the verified resolution to Hindsight for retention.
GET /api/incidents/history
Lists historical incidents directly from Hindsight for auditing and dashboard display.
A Concrete Incident
To understand how this works in practice, let's look at the project's demonstration scenario: a simulated payment-service incident backed by historical incident memories stored in Hindsight.
Imagine a scenario where the payment service is timing out and checkout requests are failing intermittently.
BEFORE OpsMemory:
Current incident → generic AI reasoning
An engineer pastes the symptoms into a generic AI assistant. The AI suggests checking network partitions, DNS issues, JVM garbage collection pauses, and database locks. The engineer spends time investigating each possibility before finally identifying the issue.
AFTER OpsMemory:
Current incident → Hindsight historical memory → context-aware LLM reasoning → human verification → Hindsight retention → future incident benefits from accumulated knowledge
The engineer reports the exact same symptoms to OpsMemory.
OpsMemory queries Hindsight, which recalls previous payment incidents involving:
- database connection-pool exhaustion
- long-running transactions
- payment-service failures
- database saturation
- previous mitigation strategies
OpsMemory passes this history to Groq. Groq reasons over the data and generates an analysis: A likely root cause involving database connection-pool exhaustion caused by long-running checkout transactions.
The engineer can use this contextualized analysis to guide the investigation and mitigation.
What Changes After the Incident is Resolved?
Once the payment service is stable, the engineer fills out the resolution form in OpsMemory, noting the verified actions they took. For our example scenario, the resolution is:
- Terminated stuck transactions
- Restarted payment-service
- Increased connection pool limits
- Added saturation alerts
The outcome: Payment service recovered and checkout requests returned to normal.
They click "Retain Memory." The backend sends the verified resolution to Hindsight for retention. The verified resolution is stored as a new memory item in Hindsight. The loop is closed.
Why This is Different From a Normal AI Assistant
A conventional LLM interaction does not automatically provide OpsMemory with persistent, organization-specific incident memory. Without a dedicated memory layer, the agent cannot reliably reuse verified incident knowledge across future interactions.
OpsMemory is explicitly designed for organizational learning. As verified incidents accumulate, the Hindsight memory bank is designed to become a more specialized, context-rich source of knowledge about the organization's infrastructure, failure modes, and successful mitigations.
Current MVP and Future Extensions
It is important to clearly distinguish our current MVP from future extensions.
Today, OpsMemory is a working, deployed MVP.
CURRENT MVP:
- Incident reporting
- Hindsight historical-memory recall
- AI incident analysis
- Likely root-cause identification
- Recommended actions/investigation
- Human verification
- Hindsight retention
- Incident history
- Deployed frontend and backend
In the future, this architecture naturally extends to much deeper integrations.
FUTURE EXTENSIONS:
- Live log ingestion
- Metrics
- Distributed traces
- Deployment-event correlation
- PagerDuty
- Slack/Teams
- Automated incident detection
- Automated low-risk remediation
- Runbook retrieval
- Automated postmortem generation
Today, we are focused entirely on proving the foundational value of the core loop.
Conclusion
Production incidents are expensive, stressful, and inevitable. But paying the operational tax to solve an incident should at least buy you immunity from having to solve it from scratch the next time.
By combining the reasoning capabilities of Groq with the persistent memory of Hindsight, OpsMemory ensures that your organization actually remembers its lessons.
Every production incident should make the next incident easier to solve.
Top comments (0)