DEV Community

Cover image for The Error Was New. Hindsight Remembered the Fix
K U S H A L
K U S H A L

Posted on

The Error Was New. Hindsight Remembered the Fix

Production incidents have an annoying habit of feeling new even when they are not.

An API suddenly returns 500 errors. A machine-learning model rejects an input tensor. A deployment that worked yesterday stops after one dependency changes. Someone searches the logs, another developer searches old issues, and eventually somebody says:

“Wait, didn’t we fix something almost exactly like this before?”

That question became the idea behind RecallOps.

RecallOps is an incident-response agent designed around a simple observation: engineering teams already solve many of their future problems in the past. The difficult part is finding that knowledge when another incident happens.

Instead of treating every error as an isolated event, RecallOps uses Hindsight as a persistent memory layer.

When a new incident arrives, RecallOps searches the engineering memory for similar failures. Those memories are supplied to the reasoning layer before it proposes a root cause or resolution.

When engineers finally solve the incident, the important details are retained again.

That creates a loop:

incident → recall → reason → resolve → retain → reuse

The result is not an agent that magically fixes production systems.

It is something more practical:

an agent that does not start every investigation from zero.

Figure 1 — RecallOps prototype overview. A new incident is analyzed using relevant engineering memories, previous resolutions are surfaced, and confirmed fixes become future memory.


The Problem Wasn't Detecting Errors

Modern applications are already extremely good at telling developers when something has gone wrong.

We have application logs, cloud monitoring, telemetry, exception trackers, traces, alerts and dashboards.

Those systems answer an important question:

What happened?

But during debugging, that is only the beginning.

Developers still need to determine:

  • Why did it happen?
  • Has something similar happened before?
  • What was the root cause last time?
  • Which fix actually worked?
  • Was that fix temporary or permanent?
  • Is the current incident really the same problem?

That knowledge often exists somewhere.

It might be inside an old issue.

It might be buried in a Slack conversation.

It might exist in a previous incident report.

Or, most commonly, it might exist only in the memory of the developer who solved the problem.

Consider an error from an image-classification pipeline:

Crop classification failed.

Expected input: 224 x 224
Received input: 37 x 39
Enter fullscreen mode Exit fullscreen mode

At first glance, several parts of the system could be responsible.

The satellite image could be corrupted.

The model could be wrong.

The input might contain the wrong channels.

The inference service could have loaded the wrong model.

In this case, however, the real problem was much simpler.

The image-processing pipeline was sending the raw extracted image directly into a model that expected a fixed 224 × 224 input.

The solution was:

Satellite Image
      ↓
Extract Region
      ↓
Resize to 224 × 224
      ↓
Normalize Values
      ↓
Convert to Tensor
      ↓
Model Inference
Enter fullscreen mode Exit fullscreen mode

Once the problem was understood, fixing it was straightforward.

The interesting question was what would happen when a similar error appeared several weeks later.

Would we remember the solution?

Would another developer know about it?

Would we spend another thirty minutes discovering the same preprocessing problem?

That is the information RecallOps is designed to preserve.


Turning Incidents Into Engineering Memory

The first version of the idea was tempting:

ERROR
  ↓
LLM
  ↓
SOLUTION
Enter fullscreen mode Exit fullscreen mode

Send the error to a language model and ask what might be wrong.

That works reasonably well for generic programming questions.

But there is a major limitation.

The model might understand Python, C#, Docker, APIs and neural networks, but it does not automatically know the history of our application.

Our system has its own knowledge.

For example:

Service: Crop Classification

Problem:
Model received a satellite crop with incorrect dimensions.

Root cause:
Image resizing was skipped during preprocessing.

Resolution:
Resize to 224×224, normalize pixel values,
then create the model tensor.
Enter fullscreen mode Exit fullscreen mode

That information is far more useful than another generic explanation of tensor shapes.

It is knowledge produced by actually operating the system.

That is where Hindsight's agent memory fits into RecallOps.

Instead of continuously expanding the prompt or manually searching old incident reports, we retain important outcomes as long-term memory.

Then, when a new error occurs, RecallOps queries Hindsight for relevant past incidents.


The Architecture

I deliberately kept the architecture small.

             APPLICATION
                  │
                  │ Error / Logs
                  ▼
        ┌────────────────────┐
        │     RecallOps      │
        │   Incident Agent   │
        └──────────┬─────────┘
                   │
                   │ Recall similar incidents
                   ▼
        ┌────────────────────┐
        │     Hindsight      │
        │   Memory Bank      │
        └──────────┬─────────┘
                   │
                   │ Historical context
                   ▼
        ┌────────────────────┐
        │   Reasoning Layer  │
        └──────────┬─────────┘
                   │
                   ▼
        Likely Root Cause
        Suggested Resolution
        Similar Past Incident
                   │
                   ▼
              ENGINEER
                   │
                   │ Confirm resolution
                   ▼
        ┌────────────────────┐
        │ Retain Resolution  │
        └──────────┬─────────┘
                   │
                   └──────────────► Hindsight
Enter fullscreen mode Exit fullscreen mode

There are two important Hindsight operations in the workflow:

Recall before investigation.

Retain after resolution.

The full Hindsight documentation describes the memory system in more detail, but these two operations were enough to define the core RecallOps workflow.


Recall Before Reasoning

This became one of the most important decisions in the prototype.

The system should retrieve historical context before asking the reasoning model for an explanation.

A simplified version of the incident-analysis flow looks like this:

public async Task<IncidentAnalysis> AnalyzeIncident(
    Incident incident)
{
    // 1. Search previous engineering memory
    var memories = await RecallRelevantIncidents(
        incident.ErrorMessage
    );

    // 2. Combine the new incident with recalled history
    var context = BuildIncidentContext(
        incident,
        memories
    );

    // 3. Ask the reasoning layer to analyze everything
    var analysis = await AnalyzeWithAI(context);

    return analysis;
}
Enter fullscreen mode Exit fullscreen mode

The important part is the order.

Without memory:

New Incident
     ↓
Reasoning Model
     ↓
Generic Troubleshooting
Enter fullscreen mode Exit fullscreen mode

RecallOps instead follows:

New Incident
     ↓
Recall Relevant Memories
     ↓
Current Evidence + Historical Evidence
     ↓
Reasoning Model
     ↓
Suggested Root Cause
     ↓
Recommended Checks
Enter fullscreen mode Exit fullscreen mode

This means the reasoning layer does not have to rediscover every known fact about the application.

It can start with what previous incidents already taught us.


Building the Recall Query

The error message alone may not contain enough information.

RecallOps can construct a query using the important characteristics of the incident.

For example:

private string BuildRecallQuery(Incident incident)
{
    return $"""
        Find previous incidents similar to this problem.

        Service: {incident.Service}

        Error:
        {incident.ErrorMessage}

        Component:
        {incident.Component}

        Look for previous root causes,
        resolutions and debugging notes.
        """;
}
Enter fullscreen mode Exit fullscreen mode

For our example, the query could become:

Service: crop-classification

Error:
Expected image dimensions 224x224.
Received dimensions 37x39.

Component:
model preprocessing
Enter fullscreen mode Exit fullscreen mode

The goal is not exact string matching.

The current error might say:

Received image dimensions: 37x39
Enter fullscreen mode Exit fullscreen mode

while the older incident might say:

Inference tensor shape does not match model input.
Enter fullscreen mode Exit fullscreen mode

They are different strings describing essentially the same class of problem.

The memory layer gives RecallOps a way to retrieve the older engineering context rather than relying entirely on identical wording.


What We Retain

Storing every line of every log would defeat the purpose.

We already have logging systems for that.

The useful memory is the conclusion of the investigation.

A resolved incident can be represented like this:

public class IncidentMemory
{
    public string Service { get; set; }

    public string Error { get; set; }

    public string Component { get; set; }

    public string RootCause { get; set; }

    public string Resolution { get; set; }

    public string Status { get; set; }
}
Enter fullscreen mode Exit fullscreen mode

For the crop-classification example:

{
  "service": "crop-classification",
  "error": "Expected 224x224 but received 37x39",
  "component": "image preprocessing",
  "rootCause": "Resize step was skipped before inference",
  "resolution": "Resize to 224x224 and normalize before creating the model tensor",
  "status": "resolved"
}
Enter fullscreen mode Exit fullscreen mode

Once an engineer confirms the resolution, RecallOps prepares a concise memory:

private string BuildMemory(Incident incident)
{
    return $"""
        Incident resolved.

        Service:
        {incident.Service}

        Error:
        {incident.ErrorMessage}

        Root cause:
        {incident.RootCause}

        Resolution:
        {incident.Resolution}

        Component:
        {incident.Component}
        """;
}
Enter fullscreen mode Exit fullscreen mode

That content can then be retained in Hindsight.

Conceptually:

await hindsight.retain(resolvedIncidentMemory);
Enter fullscreen mode Exit fullscreen mode

The objective is not to remember everything.

The objective is to remember the information that makes the next investigation better.


The Second Incident Is Where It Gets Interesting

The first incident only creates memory.

The real value becomes visible when another incident appears.

Imagine that a few weeks later another inference request fails:

Inference failed.

Model input required:
224 x 224

Input received:
48 x 41
Enter fullscreen mode Exit fullscreen mode

It is not the exact same message.

It is not even the same image size.

RecallOps sends the incident context to Hindsight and retrieves the previous preprocessing incident.

The returned context could contain:

Previous relevant incident:

Crop-classification inference previously failed because
a raw satellite crop was passed directly to the model.

Model expected:
224 × 224

Previous input:
37 × 39

Root cause:
The image resize preprocessing step was skipped.

Resolution:
Resize and normalize the image before creating
the inference tensor.
Enter fullscreen mode Exit fullscreen mode

Now the reasoning layer can generate a much more useful response.

For example:

LIKELY ROOT CAUSE

The current failure appears similar to a previously
resolved preprocessing incident.

The model expects 224×224 input, but the current
image is 48×41.

SIMILAR PREVIOUS INCIDENT

A previous crop-classification request failed when
a 37×39 satellite crop was passed directly into
the same model.

PREVIOUS ROOT CAUSE

The image resize step was missing.

RECOMMENDED CHECKS

1. Verify that image resizing runs before inference.
2. Confirm output dimensions are 224×224.
3. Apply the same normalization used during training.
4. Confirm channel ordering.
5. Only then construct the model tensor.
Enter fullscreen mode Exit fullscreen mode

That is very different from asking a model:

“What causes tensor dimension errors?”

RecallOps can instead say:

“We experienced a similar problem before. This was the cause, this was the fix, and these are the checks worth performing first.”


Memory Is Evidence, Not Truth

There is an important limitation.

A similar error does not guarantee an identical root cause.

Two timeout exceptions can come from completely different components.

Two database connection failures can have different causes.

Two model inference errors can look nearly identical while originating at different stages of the preprocessing pipeline.

For that reason, RecallOps does not treat retrieved memory as an automatic command.

It presents it as historical evidence.

The interface should distinguish between:

Similar Past Incident
Enter fullscreen mode Exit fullscreen mode

and:

Confirmed Current Root Cause
Enter fullscreen mode Exit fullscreen mode

Likewise, there is a difference between:

Suggested Resolution
Enter fullscreen mode Exit fullscreen mode

and:

Engineer Confirmed Resolution
Enter fullscreen mode Exit fullscreen mode

This distinction matters because automatically applying an old fix to a new production failure could create an even larger problem.

The agent helps narrow the investigation.

The engineer still makes the final decision.


From a Stateless Agent to a Learning Workflow

Most simple AI applications follow this pattern:

INPUT
  ↓
MODEL
  ↓
OUTPUT
Enter fullscreen mode Exit fullscreen mode

Each request is mostly independent.

RecallOps behaves differently:

                 ┌──────── MEMORY ────────┐
                 │                         │
                 ▼                         │
INCIDENT → RECALL → REASON → RESOLVE → RETAIN
Enter fullscreen mode Exit fullscreen mode

That final retain step changes what the next execution knows.

Consider a sequence of incidents:

INCIDENT 01
Image preprocessing failure
        ↓
Resolution retained


INCIDENT 07
Similar model failure
        ↓
Incident 01 recalled
        ↓
New resolution retained


INCIDENT 18
Authentication configuration failure
        ↓
Resolution retained


INCIDENT 42
New production failure
        ↓
Multiple relevant engineering memories available
Enter fullscreen mode Exit fullscreen mode

The model itself has not necessarily changed.

The context surrounding it has.

That is what makes long-term agent memory interesting.


What I Learned Building RecallOps

The first lesson was that memory becomes much more useful when connected to a specific workflow.

“An AI agent with memory” is a technical capability.

“An incident agent that remembers how we fixed previous failures” is a concrete use case.

The second lesson was that remembering everything is not the goal.

Logs already remember almost everything.

What is frequently missing is the relationship between:

SYMPTOM
   +
ROOT CAUSE
   +
CONFIRMED FIX
Enter fullscreen mode Exit fullscreen mode

That is the part worth retaining.

The third lesson was recall before reasoning.

The agent should see relevant historical evidence before generating a theory.

Otherwise it can generate a theory first and use later information simply to support what it already decided.

The fourth lesson was that human confirmation matters.

Past solutions are useful evidence, but production systems change.

Dependencies change.

Configuration changes.

Infrastructure changes.

A previous fix should guide an investigation rather than automatically control it.

Finally, I learned that persistent memory changes the way I think about AI applications.

The most interesting improvement does not always come from using a larger model.

Sometimes it comes from giving the existing model access to better context.


The Goal Isn't to Replace Debugging

RecallOps is not trying to remove engineers from incident response.

Debugging still requires understanding the current system.

Engineers still need to inspect logs, reproduce problems, verify assumptions and validate fixes.

The problem RecallOps addresses is more specific:

valuable debugging knowledge disappears too easily.

We solve a production issue.

We close the incident.

We continue building.

Several months later, somebody encounters almost the same problem and starts the investigation again from the beginning.

With persistent memory, that cycle can change.

A resolved incident becomes context for a future incident.

A previous root cause becomes a clue.

A successful resolution becomes a debugging path worth checking.

And every confirmed fix makes the system's engineering memory slightly more useful.

The idea behind RecallOps can therefore be summarized in one question:

Before spending an hour solving this problem, have we already solved something like it?

With Hindsight providing persistent agent memory, RecallOps can actually ask that question every time.

The new error may be unfamiliar.

The fix does not have to be.

Top comments (0)