Production incidents have an annoying habit of feeling new even when they are not.
An API suddenly returns 500 errors. A machine-learning model rejects an input tensor. A deployment that worked yesterday stops after one dependency changes. Someone searches the logs, another developer searches old issues, and eventually somebody says:
“Wait, didn’t we fix something almost exactly like this before?”
That question became the idea behind RecallOps.
RecallOps is an incident-response agent designed around a simple observation: engineering teams already solve many of their future problems in the past. The difficult part is finding that knowledge when another incident happens.
Instead of treating every error as an isolated event, RecallOps uses Hindsight as a persistent memory layer.
When a new incident arrives, RecallOps searches the engineering memory for similar failures. Those memories are supplied to the reasoning layer before it proposes a root cause or resolution.
When engineers finally solve the incident, the important details are retained again.
That creates a loop:
incident → recall → reason → resolve → retain → reuse
The result is not an agent that magically fixes production systems.
It is something more practical:
an agent that does not start every investigation from zero.
Figure 1 — RecallOps prototype overview. A new incident is analyzed using relevant engineering memories, previous resolutions are surfaced, and confirmed fixes become future memory.
The Problem Wasn't Detecting Errors
Modern applications are already extremely good at telling developers when something has gone wrong.
We have application logs, cloud monitoring, telemetry, exception trackers, traces, alerts and dashboards.
Those systems answer an important question:
What happened?
But during debugging, that is only the beginning.
Developers still need to determine:
- Why did it happen?
- Has something similar happened before?
- What was the root cause last time?
- Which fix actually worked?
- Was that fix temporary or permanent?
- Is the current incident really the same problem?
That knowledge often exists somewhere.
It might be inside an old issue.
It might be buried in a Slack conversation.
It might exist in a previous incident report.
Or, most commonly, it might exist only in the memory of the developer who solved the problem.
Consider an error from an image-classification pipeline:
Crop classification failed.
Expected input: 224 x 224
Received input: 37 x 39
At first glance, several parts of the system could be responsible.
The satellite image could be corrupted.
The model could be wrong.
The input might contain the wrong channels.
The inference service could have loaded the wrong model.
In this case, however, the real problem was much simpler.
The image-processing pipeline was sending the raw extracted image directly into a model that expected a fixed 224 × 224 input.
The solution was:
Satellite Image
↓
Extract Region
↓
Resize to 224 × 224
↓
Normalize Values
↓
Convert to Tensor
↓
Model Inference
Once the problem was understood, fixing it was straightforward.
The interesting question was what would happen when a similar error appeared several weeks later.
Would we remember the solution?
Would another developer know about it?
Would we spend another thirty minutes discovering the same preprocessing problem?
That is the information RecallOps is designed to preserve.
Turning Incidents Into Engineering Memory
The first version of the idea was tempting:
ERROR
↓
LLM
↓
SOLUTION
Send the error to a language model and ask what might be wrong.
That works reasonably well for generic programming questions.
But there is a major limitation.
The model might understand Python, C#, Docker, APIs and neural networks, but it does not automatically know the history of our application.
Our system has its own knowledge.
For example:
Service: Crop Classification
Problem:
Model received a satellite crop with incorrect dimensions.
Root cause:
Image resizing was skipped during preprocessing.
Resolution:
Resize to 224×224, normalize pixel values,
then create the model tensor.
That information is far more useful than another generic explanation of tensor shapes.
It is knowledge produced by actually operating the system.
That is where Hindsight's agent memory fits into RecallOps.
Instead of continuously expanding the prompt or manually searching old incident reports, we retain important outcomes as long-term memory.
Then, when a new error occurs, RecallOps queries Hindsight for relevant past incidents.
The Architecture
I deliberately kept the architecture small.
APPLICATION
│
│ Error / Logs
▼
┌────────────────────┐
│ RecallOps │
│ Incident Agent │
└──────────┬─────────┘
│
│ Recall similar incidents
▼
┌────────────────────┐
│ Hindsight │
│ Memory Bank │
└──────────┬─────────┘
│
│ Historical context
▼
┌────────────────────┐
│ Reasoning Layer │
└──────────┬─────────┘
│
▼
Likely Root Cause
Suggested Resolution
Similar Past Incident
│
▼
ENGINEER
│
│ Confirm resolution
▼
┌────────────────────┐
│ Retain Resolution │
└──────────┬─────────┘
│
└──────────────► Hindsight
There are two important Hindsight operations in the workflow:
Recall before investigation.
Retain after resolution.
The full Hindsight documentation describes the memory system in more detail, but these two operations were enough to define the core RecallOps workflow.
Recall Before Reasoning
This became one of the most important decisions in the prototype.
The system should retrieve historical context before asking the reasoning model for an explanation.
A simplified version of the incident-analysis flow looks like this:
public async Task<IncidentAnalysis> AnalyzeIncident(
Incident incident)
{
// 1. Search previous engineering memory
var memories = await RecallRelevantIncidents(
incident.ErrorMessage
);
// 2. Combine the new incident with recalled history
var context = BuildIncidentContext(
incident,
memories
);
// 3. Ask the reasoning layer to analyze everything
var analysis = await AnalyzeWithAI(context);
return analysis;
}
The important part is the order.
Without memory:
New Incident
↓
Reasoning Model
↓
Generic Troubleshooting
RecallOps instead follows:
New Incident
↓
Recall Relevant Memories
↓
Current Evidence + Historical Evidence
↓
Reasoning Model
↓
Suggested Root Cause
↓
Recommended Checks
This means the reasoning layer does not have to rediscover every known fact about the application.
It can start with what previous incidents already taught us.
Building the Recall Query
The error message alone may not contain enough information.
RecallOps can construct a query using the important characteristics of the incident.
For example:
private string BuildRecallQuery(Incident incident)
{
return $"""
Find previous incidents similar to this problem.
Service: {incident.Service}
Error:
{incident.ErrorMessage}
Component:
{incident.Component}
Look for previous root causes,
resolutions and debugging notes.
""";
}
For our example, the query could become:
Service: crop-classification
Error:
Expected image dimensions 224x224.
Received dimensions 37x39.
Component:
model preprocessing
The goal is not exact string matching.
The current error might say:
Received image dimensions: 37x39
while the older incident might say:
Inference tensor shape does not match model input.
They are different strings describing essentially the same class of problem.
The memory layer gives RecallOps a way to retrieve the older engineering context rather than relying entirely on identical wording.
What We Retain
Storing every line of every log would defeat the purpose.
We already have logging systems for that.
The useful memory is the conclusion of the investigation.
A resolved incident can be represented like this:
public class IncidentMemory
{
public string Service { get; set; }
public string Error { get; set; }
public string Component { get; set; }
public string RootCause { get; set; }
public string Resolution { get; set; }
public string Status { get; set; }
}
For the crop-classification example:
{
"service": "crop-classification",
"error": "Expected 224x224 but received 37x39",
"component": "image preprocessing",
"rootCause": "Resize step was skipped before inference",
"resolution": "Resize to 224x224 and normalize before creating the model tensor",
"status": "resolved"
}
Once an engineer confirms the resolution, RecallOps prepares a concise memory:
private string BuildMemory(Incident incident)
{
return $"""
Incident resolved.
Service:
{incident.Service}
Error:
{incident.ErrorMessage}
Root cause:
{incident.RootCause}
Resolution:
{incident.Resolution}
Component:
{incident.Component}
""";
}
That content can then be retained in Hindsight.
Conceptually:
await hindsight.retain(resolvedIncidentMemory);
The objective is not to remember everything.
The objective is to remember the information that makes the next investigation better.
The Second Incident Is Where It Gets Interesting
The first incident only creates memory.
The real value becomes visible when another incident appears.
Imagine that a few weeks later another inference request fails:
Inference failed.
Model input required:
224 x 224
Input received:
48 x 41
It is not the exact same message.
It is not even the same image size.
RecallOps sends the incident context to Hindsight and retrieves the previous preprocessing incident.
The returned context could contain:
Previous relevant incident:
Crop-classification inference previously failed because
a raw satellite crop was passed directly to the model.
Model expected:
224 × 224
Previous input:
37 × 39
Root cause:
The image resize preprocessing step was skipped.
Resolution:
Resize and normalize the image before creating
the inference tensor.
Now the reasoning layer can generate a much more useful response.
For example:
LIKELY ROOT CAUSE
The current failure appears similar to a previously
resolved preprocessing incident.
The model expects 224×224 input, but the current
image is 48×41.
SIMILAR PREVIOUS INCIDENT
A previous crop-classification request failed when
a 37×39 satellite crop was passed directly into
the same model.
PREVIOUS ROOT CAUSE
The image resize step was missing.
RECOMMENDED CHECKS
1. Verify that image resizing runs before inference.
2. Confirm output dimensions are 224×224.
3. Apply the same normalization used during training.
4. Confirm channel ordering.
5. Only then construct the model tensor.
That is very different from asking a model:
“What causes tensor dimension errors?”
RecallOps can instead say:
“We experienced a similar problem before. This was the cause, this was the fix, and these are the checks worth performing first.”
Memory Is Evidence, Not Truth
There is an important limitation.
A similar error does not guarantee an identical root cause.
Two timeout exceptions can come from completely different components.
Two database connection failures can have different causes.
Two model inference errors can look nearly identical while originating at different stages of the preprocessing pipeline.
For that reason, RecallOps does not treat retrieved memory as an automatic command.
It presents it as historical evidence.
The interface should distinguish between:
Similar Past Incident
and:
Confirmed Current Root Cause
Likewise, there is a difference between:
Suggested Resolution
and:
Engineer Confirmed Resolution
This distinction matters because automatically applying an old fix to a new production failure could create an even larger problem.
The agent helps narrow the investigation.
The engineer still makes the final decision.
From a Stateless Agent to a Learning Workflow
Most simple AI applications follow this pattern:
INPUT
↓
MODEL
↓
OUTPUT
Each request is mostly independent.
RecallOps behaves differently:
┌──────── MEMORY ────────┐
│ │
▼ │
INCIDENT → RECALL → REASON → RESOLVE → RETAIN
That final retain step changes what the next execution knows.
Consider a sequence of incidents:
INCIDENT 01
Image preprocessing failure
↓
Resolution retained
INCIDENT 07
Similar model failure
↓
Incident 01 recalled
↓
New resolution retained
INCIDENT 18
Authentication configuration failure
↓
Resolution retained
INCIDENT 42
New production failure
↓
Multiple relevant engineering memories available
The model itself has not necessarily changed.
The context surrounding it has.
That is what makes long-term agent memory interesting.
What I Learned Building RecallOps
The first lesson was that memory becomes much more useful when connected to a specific workflow.
“An AI agent with memory” is a technical capability.
“An incident agent that remembers how we fixed previous failures” is a concrete use case.
The second lesson was that remembering everything is not the goal.
Logs already remember almost everything.
What is frequently missing is the relationship between:
SYMPTOM
+
ROOT CAUSE
+
CONFIRMED FIX
That is the part worth retaining.
The third lesson was recall before reasoning.
The agent should see relevant historical evidence before generating a theory.
Otherwise it can generate a theory first and use later information simply to support what it already decided.
The fourth lesson was that human confirmation matters.
Past solutions are useful evidence, but production systems change.
Dependencies change.
Configuration changes.
Infrastructure changes.
A previous fix should guide an investigation rather than automatically control it.
Finally, I learned that persistent memory changes the way I think about AI applications.
The most interesting improvement does not always come from using a larger model.
Sometimes it comes from giving the existing model access to better context.
The Goal Isn't to Replace Debugging
RecallOps is not trying to remove engineers from incident response.
Debugging still requires understanding the current system.
Engineers still need to inspect logs, reproduce problems, verify assumptions and validate fixes.
The problem RecallOps addresses is more specific:
valuable debugging knowledge disappears too easily.
We solve a production issue.
We close the incident.
We continue building.
Several months later, somebody encounters almost the same problem and starts the investigation again from the beginning.
With persistent memory, that cycle can change.
A resolved incident becomes context for a future incident.
A previous root cause becomes a clue.
A successful resolution becomes a debugging path worth checking.
And every confirmed fix makes the system's engineering memory slightly more useful.
The idea behind RecallOps can therefore be summarized in one question:
Before spending an hour solving this problem, have we already solved something like it?
With Hindsight providing persistent agent memory, RecallOps can actually ask that question every time.
The new error may be unfamiliar.
The fix does not have to be.
Top comments (0)