RecallOps: Building an AI Incident Response Agent That Learns From Experience
Introduction
Production incidents are unavoidable. The difficult part is solving them quickly and making sure the same problem becomes easier to solve the next time.
Most AI assistants can analyze an incident, but they often lack persistent organizational memory. They may know general troubleshooting techniques, but they don't automatically remember how a specific engineering team solved a similar incident before.
That's the problem we wanted to solve with RecallOps.
What is RecallOps?
RecallOps is an AI-powered incident response platform designed for DevOps and SRE teams.
Its core idea is simple:
Every production incident should become a lesson for the next incident.
RecallOps remembers previous incidents, root causes, successful resolutions, failed attempts, and engineer feedback.
When a new incident occurs, the system recalls relevant historical experiences and uses them during investigation.
The Problem
Consider a Payment API returning 502 Bad Gateway errors.
A traditional AI assistant might recommend:
- Check application logs
- Check recent deployments
- Check service health
- Check dependencies
These are useful, but they are generic.
What if your team had already experienced the exact same problem?
Suppose the previous incident was caused by PostgreSQL connection pool exhaustion and the successful fix was increasing the connection pool from 50 to 100.
RecallOps can remember that experience and use it during the next investigation.
How Hindsight Powers RecallOps
Hindsight is the persistent memory layer in RecallOps.
It stores experiences such as:
- Previous incidents
- Root causes
- Successful resolutions
- Failed attempts
- Engineer observations
- Lessons learned
When a new incident is created, RecallOps searches Hindsight for relevant historical experiences.
The retrieved memories are then provided as context for the AI investigation.
The Learning Loop
The RecallOps workflow is:
Incident
↓
Recall
↓
Investigate
↓
Resolve
↓
Learn
↓
Remember
↓
Improve
This creates a continuous learning cycle.
Example
First Incident
Payment API starts returning 502 Bad Gateway errors.
Investigation discovers:
Root Cause: PostgreSQL connection pool exhaustion
Resolution: Increase connection pool from 50 to 100
Failed Attempt: Increasing the gateway timeout did not solve the underlying database issue.
This experience is saved to Hindsight.
Similar Incident Later
Another Payment API incident produces similar 502 errors.
RecallOps searches its memory and finds the previous incident.
Instead of starting from a generic troubleshooting checklist, it can recommend checking PostgreSQL connection pool utilization first.
This is the key difference:
The system learns from the team's own experience.
Architecture
RecallOps uses:
- React + Tailwind CSS for the frontend
- ASP.NET Core + C# for the backend
- Supabase Authentication
- PostgreSQL for structured application data
- Groq for AI reasoning
- Hindsight for persistent AI memory
The backend orchestrates communication between the frontend, database, AI model, and Hindsight.
Why PostgreSQL and Hindsight?
They have different responsibilities.
PostgreSQL
Stores structured application data:
- Users
- Incidents
- Services
- Status
- Severity
- Timestamps
- Feedback
Hindsight
Stores AI memory:
- Previous experiences
- Root causes
- Resolutions
- Failed attempts
- Lessons
- Relevant patterns
PostgreSQL stores the application's data.
Hindsight stores the agent's experience.
User Experience
RecallOps provides an incident dashboard where engineers can:
- Create an incident
- Start an AI investigation
- View similar historical incidents
- Review recommended checks
- See suggested resolutions
- Resolve the incident
- Save the experience to memory
Technology Stack
Frontend:
React + Tailwind CSS
Backend:
ASP.NET Core + C#
Database:
PostgreSQL / Supabase
AI:
Groq
Memory:
Hindsight by Vectorize
Deployment:
Vercel + Render
What We Learned
The most important lesson from building RecallOps is that AI memory is more than storing previous conversations.
Useful memory needs to become part of the agent's reasoning process.
The goal isn't simply:
"Here are some old incidents."
The goal is:
"Here is what happened before, what worked, what failed, and why that experience matters to the current incident."
Future Improvements
We plan to integrate RecallOps with tools commonly used by DevOps teams, including:
- Grafana
- Prometheus
- Datadog
- Slack
- GitHub
- Jira
- PagerDuty
- AWS
- Azure
- Google Cloud
These integrations could allow RecallOps to automatically collect incident context and continuously learn from real production events.
Conclusion
RecallOps is an AI incident-response system designed around a simple principle:
Don't just solve today's incident. Learn from it.
By combining AI reasoning with persistent Hindsight memory, RecallOps can turn previous incident experiences into useful context for future investigations.
RecallOps — AI Incident Response That Learns From Experience.
Links
GitHub:https://youtu.be/THSTvqjVvq0
Live Demo:https://recallops-ashy.vercel.app/
Demo Video:https://youtu.be/THSTvqjVvq0?si=sYQGela0mi8t70CM
Hindsight: https://hindsight.vectorize.io/
Top comments (0)