DEV Community

Cover image for “RecallOps: Building an AI Incident Response Agent That Learns From Experience”
Narendra Sahu
Narendra Sahu

Posted on

“RecallOps: Building an AI Incident Response Agent That Learns From Experience”

RecallOps: Building an AI Incident Response Agent That Learns From Experience

Introduction

Production incidents are unavoidable. The difficult part is solving them quickly and making sure the same problem becomes easier to solve the next time.

Most AI assistants can analyze an incident, but they often lack persistent organizational memory. They may know general troubleshooting techniques, but they don't automatically remember how a specific engineering team solved a similar incident before.

That's the problem we wanted to solve with RecallOps.

What is RecallOps?

RecallOps is an AI-powered incident response platform designed for DevOps and SRE teams.

Its core idea is simple:

Every production incident should become a lesson for the next incident.

RecallOps remembers previous incidents, root causes, successful resolutions, failed attempts, and engineer feedback.

When a new incident occurs, the system recalls relevant historical experiences and uses them during investigation.

The Problem

Consider a Payment API returning 502 Bad Gateway errors.

A traditional AI assistant might recommend:

  • Check application logs
  • Check recent deployments
  • Check service health
  • Check dependencies

These are useful, but they are generic.

What if your team had already experienced the exact same problem?

Suppose the previous incident was caused by PostgreSQL connection pool exhaustion and the successful fix was increasing the connection pool from 50 to 100.

RecallOps can remember that experience and use it during the next investigation.

How Hindsight Powers RecallOps

Hindsight is the persistent memory layer in RecallOps.

It stores experiences such as:

  • Previous incidents
  • Root causes
  • Successful resolutions
  • Failed attempts
  • Engineer observations
  • Lessons learned

When a new incident is created, RecallOps searches Hindsight for relevant historical experiences.

The retrieved memories are then provided as context for the AI investigation.

The Learning Loop

The RecallOps workflow is:

Incident
↓
Recall
↓
Investigate
↓
Resolve
↓
Learn
↓
Remember
↓
Improve

This creates a continuous learning cycle.

Example

First Incident

Payment API starts returning 502 Bad Gateway errors.

Investigation discovers:

Root Cause: PostgreSQL connection pool exhaustion

Resolution: Increase connection pool from 50 to 100

Failed Attempt: Increasing the gateway timeout did not solve the underlying database issue.

This experience is saved to Hindsight.

Similar Incident Later

Another Payment API incident produces similar 502 errors.

RecallOps searches its memory and finds the previous incident.

Instead of starting from a generic troubleshooting checklist, it can recommend checking PostgreSQL connection pool utilization first.

This is the key difference:

The system learns from the team's own experience.

Architecture

RecallOps uses:

  • React + Tailwind CSS for the frontend
  • ASP.NET Core + C# for the backend
  • Supabase Authentication
  • PostgreSQL for structured application data
  • Groq for AI reasoning
  • Hindsight for persistent AI memory

The backend orchestrates communication between the frontend, database, AI model, and Hindsight.

Why PostgreSQL and Hindsight?

They have different responsibilities.

PostgreSQL

Stores structured application data:

  • Users
  • Incidents
  • Services
  • Status
  • Severity
  • Timestamps
  • Feedback

Hindsight

Stores AI memory:

  • Previous experiences
  • Root causes
  • Resolutions
  • Failed attempts
  • Lessons
  • Relevant patterns

PostgreSQL stores the application's data.

Hindsight stores the agent's experience.

User Experience

RecallOps provides an incident dashboard where engineers can:

  1. Create an incident
  2. Start an AI investigation
  3. View similar historical incidents
  4. Review recommended checks
  5. See suggested resolutions
  6. Resolve the incident
  7. Save the experience to memory

Technology Stack

Frontend:
React + Tailwind CSS

Backend:
ASP.NET Core + C#

Database:
PostgreSQL / Supabase

AI:
Groq

Memory:
Hindsight by Vectorize

Deployment:
Vercel + Render

What We Learned

The most important lesson from building RecallOps is that AI memory is more than storing previous conversations.

Useful memory needs to become part of the agent's reasoning process.

The goal isn't simply:

"Here are some old incidents."

The goal is:

"Here is what happened before, what worked, what failed, and why that experience matters to the current incident."

Future Improvements

We plan to integrate RecallOps with tools commonly used by DevOps teams, including:

  • Grafana
  • Prometheus
  • Datadog
  • Slack
  • GitHub
  • Jira
  • PagerDuty
  • AWS
  • Azure
  • Google Cloud

These integrations could allow RecallOps to automatically collect incident context and continuously learn from real production events.

Conclusion

RecallOps is an AI incident-response system designed around a simple principle:

Don't just solve today's incident. Learn from it.

By combining AI reasoning with persistent Hindsight memory, RecallOps can turn previous incident experiences into useful context for future investigations.

RecallOps — AI Incident Response That Learns From Experience.

Links

GitHub:https://youtu.be/THSTvqjVvq0

Live Demo:https://recallops-ashy.vercel.app/

Demo Video:https://youtu.be/THSTvqjVvq0?si=sYQGela0mi8t70CM

Hindsight: https://hindsight.vectorize.io/

Top comments (0)