RecallOps: An AI Incident Response Agent That Learns From Every Incident Modern incident-response tools can analyze production problems using AI, but there is a major limitation: many of them do not truly learn from what happens after an incident is resolved.
RecallOps, an AI-powered incident response agent built with Hindsight, is designed to solve this problem by creating a continuous learning loop.
The central idea is simple: an incident should not just be solved. It should become reliable knowledge for the next incident.
What Is RecallOps?
RecallOps is a lightweight Python service with a Streamlit interface. An on-call engineer describes a production incident in plain language, selects the affected service and severity, and receives a structured analysis based on relevant incidents stored in Hindsight.
The system follows a five-step cycle:
Incident → Recall → Reason → Resolve → Learn
The first four steps are common in AI-assisted incident management. The important difference is the final step: Learn.
After an engineer resolves an incident, RecallOps stores the confirmed resolution and outcome back into Hindsight. This means that every completed incident can improve the knowledge available for future investigations.
How the Learning Loop Works
RecallOps does not immediately store the AI's analysis when it produces an answer. Instead, it waits until the incident has actually been resolved.
The engineer must provide two pieces of information:
- What actually fixed the incident?
- What was the outcome?
Both fields are required before the information can be stored. This prevents incomplete investigations or unconfirmed guesses from becoming future knowledge.
The information retained in Hindsight contains four parts:
- Production Incident - What happened.
- Actual Resolution - What the engineer did to fix it.
- Outcome - What happened after the fix.
- Lesson Learned - How the experience can help with similar incidents.
One of the most important design decisions is that the agent's own analysis is not stored as memory. The AI's reasoning may contain assumptions or incorrect possibilities. Storing it as a confirmed event could cause those assumptions to influence future investigations.
Instead, RecallOps stores the incident, the real resolution, and the observed outcome.
Why Hindsight Matters
Hindsight acts as the memory layer of RecallOps.
Initial incidents can be loaded into Hindsight using the same structure that is later used for real incidents. This means seeded examples and real production experiences can be recalled in the same way.
For example, an initial incident might describe a payment API experiencing high error rates and database connection problems. The recorded resolution could be increasing the database connection pool and cleaning unused connections, followed by an outcome showing that the service stabilized.
When a similar incident happens later, RecallOps can retrieve this previous experience and use it as context for its investigation.
An Example of Learning From Different Outcomes
Consider a payment service experiencing timeouts.
During one incident, Hindsight retrieves a previous incident involving database connection exhaustion. The engineer investigates the connection pool and confirms that increasing its size resolves the problem. That resolution is then stored.
Several weeks later, another payment timeout occurs. This time, however, the root cause is a slow external fraud-check service rather than database connections.
The second resolution is also stored.
Now the memory bank contains two incidents with similar symptoms but different root causes and solutions. During a future payment timeout, RecallOps can compare both experiences rather than assuming that every timeout has the same cause.
This is an important part of the system: contradicting experiences are valuable because they prevent the AI from blindly repeating an old pattern.
Key Design Principles
- The Write Path Is as Important as the Read Path Simply retrieving old incidents and asking an AI model to reason over them creates a useful assistant, but it does not create a continuously improving system.
The retain step allows the system to accumulate real operational knowledge over time.
- Store Confirmed Results, Not Predictions RecallOps only writes to Hindsight after an engineer provides the actual resolution and outcome.
This creates a clear distinction between what the AI thinks might have happened and what actually happened in production.
That distinction helps maintain the reliability of the memory bank.
- Keep the Read and Write Formats Consistent Both initial seed incidents and newly resolved incidents follow the same structure:
Incident → Resolution → Outcome → Lesson
Because the format remains consistent, Hindsight can retrieve different types of memories without requiring special handling.
- Contradicting Outcomes Are Valuable A previous solution not working in a later incident is not useless information. It can actually make future investigations more accurate.
If the system only stored incidents that confirmed existing expectations, its memory could become increasingly confident while becoming less useful.
Recording different outcomes gives the AI meaningful comparisons when symptoms look similar.
Technology Behind the Project
RecallOps is implemented as a small Python service with a Streamlit front end. Hindsight provides the memory layer for storing and retrieving incident experiences.
The project is organized around a few core components:
- agent.py - Recall and reasoning logic
- app.py - Streamlit interface and resolution capture
- hindsight_memory.py - Memory storage and retrieval
- seed_memory.py - Initial incident data
The reasoning model used in the implementation is openai/gpt-oss-120b through Groq.
However, the main focus of the project is not the specific model. The central concept is the memory loop that allows confirmed incident experiences to be retained and reused.
What Makes the Approach Different?
Traditional AI incident assistance can be thought of as:
Incident → AI Analysis → Answer
RecallOps extends this into:
Incident → Recall → AI Analysis → Human Resolution → Confirmed Outcome → Memory
The second approach creates a feedback loop. Every completed incident can contribute another verified experience to the knowledge base.
This means the system's usefulness can grow with real operational history rather than relying only on a fixed collection of examples.
Future Improvements
The current design intentionally focuses on confirmed incident outcomes, but there is still room for improvement.
One potential extension would be the ability to identify when a previously stored resolution later turns out to be incomplete or incorrect.
Such a mechanism could allow the memory bank to correct itself instead of only accumulating new experiences.
Conclusion
RecallOps demonstrates an important idea for AI-powered incident response: the value of an agent is not only how well it reasons about an incident today, but how accurately it remembers what happened yesterday.
By combining AI reasoning with Hindsight's memory capabilities, RecallOps creates a continuous cycle of recall, reasoning, resolution, and learning.
Instead of treating every incident as an isolated event, the system turns completed incidents into confirmed organizational knowledge.
Over time, that knowledge can provide richer comparisons, preserve valuable operational experience, and help future investigations learn from both successful and unexpected outcomes.
The key principle is straightforward:
Don't just solve the incident. Remember what actually solved it.
Top comments (1)
Built RecallOps to explore how incident response can learn from confirmed production outcomes. What do you think about this approach? I'd love to hear your feedback! 🚀