<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: K U S H A L</title>
    <description>The latest articles on DEV Community by K U S H A L (@kushal300).</description>
    <link>https://dev.to/kushal300</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150108%2F908dac2f-9f9c-40b4-8035-caa09aec61b4.jpg</url>
      <title>DEV Community: K U S H A L</title>
      <link>https://dev.to/kushal300</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kushal300"/>
    <language>en</language>
    <item>
      <title>The Error Was New. Hindsight Remembered the Fix</title>
      <dc:creator>K U S H A L</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:24:21 +0000</pubDate>
      <link>https://dev.to/kushal300/the-error-was-new-hindsight-remembered-the-fix-hcm</link>
      <guid>https://dev.to/kushal300/the-error-was-new-hindsight-remembered-the-fix-hcm</guid>
      <description>&lt;p&gt;Production incidents have an annoying habit of feeling new even when they are not.&lt;/p&gt;

&lt;p&gt;An API suddenly returns 500 errors. A machine-learning model rejects an input tensor. A deployment that worked yesterday stops after one dependency changes. Someone searches the logs, another developer searches old issues, and eventually somebody says:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“Wait, didn’t we fix something almost exactly like this before?”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That question became the idea behind &lt;strong&gt;RecallOps&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;RecallOps is an incident-response agent designed around a simple observation: engineering teams already solve many of their future problems in the past. The difficult part is finding that knowledge when another incident happens.&lt;/p&gt;

&lt;p&gt;Instead of treating every error as an isolated event, RecallOps uses &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; as a persistent memory layer.&lt;/p&gt;

&lt;p&gt;When a new incident arrives, RecallOps searches the engineering memory for similar failures. Those memories are supplied to the reasoning layer before it proposes a root cause or resolution.&lt;/p&gt;

&lt;p&gt;When engineers finally solve the incident, the important details are retained again.&lt;/p&gt;

&lt;p&gt;That creates a loop:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;incident → recall → reason → resolve → retain → reuse&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The result is not an agent that magically fixes production systems.&lt;/p&gt;

&lt;p&gt;It is something more practical:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;an agent that does not start every investigation from zero.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1 — RecallOps prototype overview. A new incident is analyzed using relevant engineering memories, previous resolutions are surfaced, and confirmed fixes become future memory.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem Wasn't Detecting Errors
&lt;/h2&gt;

&lt;p&gt;Modern applications are already extremely good at telling developers when something has gone wrong.&lt;/p&gt;

&lt;p&gt;We have application logs, cloud monitoring, telemetry, exception trackers, traces, alerts and dashboards.&lt;/p&gt;

&lt;p&gt;Those systems answer an important question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happened?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But during debugging, that is only the beginning.&lt;/p&gt;

&lt;p&gt;Developers still need to determine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why did it happen?&lt;/li&gt;
&lt;li&gt;Has something similar happened before?&lt;/li&gt;
&lt;li&gt;What was the root cause last time?&lt;/li&gt;
&lt;li&gt;Which fix actually worked?&lt;/li&gt;
&lt;li&gt;Was that fix temporary or permanent?&lt;/li&gt;
&lt;li&gt;Is the current incident really the same problem?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That knowledge often exists somewhere.&lt;/p&gt;

&lt;p&gt;It might be inside an old issue.&lt;/p&gt;

&lt;p&gt;It might be buried in a Slack conversation.&lt;/p&gt;

&lt;p&gt;It might exist in a previous incident report.&lt;/p&gt;

&lt;p&gt;Or, most commonly, it might exist only in the memory of the developer who solved the problem.&lt;/p&gt;

&lt;p&gt;Consider an error from an image-classification pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Crop classification failed.

Expected input: 224 x 224
Received input: 37 x 39
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At first glance, several parts of the system could be responsible.&lt;/p&gt;

&lt;p&gt;The satellite image could be corrupted.&lt;/p&gt;

&lt;p&gt;The model could be wrong.&lt;/p&gt;

&lt;p&gt;The input might contain the wrong channels.&lt;/p&gt;

&lt;p&gt;The inference service could have loaded the wrong model.&lt;/p&gt;

&lt;p&gt;In this case, however, the real problem was much simpler.&lt;/p&gt;

&lt;p&gt;The image-processing pipeline was sending the raw extracted image directly into a model that expected a fixed &lt;strong&gt;224 × 224&lt;/strong&gt; input.&lt;/p&gt;

&lt;p&gt;The solution was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Satellite Image
      ↓
Extract Region
      ↓
Resize to 224 × 224
      ↓
Normalize Values
      ↓
Convert to Tensor
      ↓
Model Inference
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once the problem was understood, fixing it was straightforward.&lt;/p&gt;

&lt;p&gt;The interesting question was what would happen when a similar error appeared several weeks later.&lt;/p&gt;

&lt;p&gt;Would we remember the solution?&lt;/p&gt;

&lt;p&gt;Would another developer know about it?&lt;/p&gt;

&lt;p&gt;Would we spend another thirty minutes discovering the same preprocessing problem?&lt;/p&gt;

&lt;p&gt;That is the information RecallOps is designed to preserve.&lt;/p&gt;




&lt;h2&gt;
  
  
  Turning Incidents Into Engineering Memory
&lt;/h2&gt;

&lt;p&gt;The first version of the idea was tempting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR
  ↓
LLM
  ↓
SOLUTION
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Send the error to a language model and ask what might be wrong.&lt;/p&gt;

&lt;p&gt;That works reasonably well for generic programming questions.&lt;/p&gt;

&lt;p&gt;But there is a major limitation.&lt;/p&gt;

&lt;p&gt;The model might understand Python, C#, Docker, APIs and neural networks, but it does not automatically know the history of &lt;strong&gt;our application&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Our system has its own knowledge.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Service: Crop Classification

Problem:
Model received a satellite crop with incorrect dimensions.

Root cause:
Image resizing was skipped during preprocessing.

Resolution:
Resize to 224×224, normalize pixel values,
then create the model tensor.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That information is far more useful than another generic explanation of tensor shapes.&lt;/p&gt;

&lt;p&gt;It is knowledge produced by actually operating the system.&lt;/p&gt;

&lt;p&gt;That is where &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;Hindsight's agent memory&lt;/a&gt; fits into RecallOps.&lt;/p&gt;

&lt;p&gt;Instead of continuously expanding the prompt or manually searching old incident reports, we retain important outcomes as long-term memory.&lt;/p&gt;

&lt;p&gt;Then, when a new error occurs, RecallOps queries Hindsight for relevant past incidents.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Architecture
&lt;/h2&gt;

&lt;p&gt;I deliberately kept the architecture small.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             APPLICATION
                  │
                  │ Error / Logs
                  ▼
        ┌────────────────────┐
        │     RecallOps      │
        │   Incident Agent   │
        └──────────┬─────────┘
                   │
                   │ Recall similar incidents
                   ▼
        ┌────────────────────┐
        │     Hindsight      │
        │   Memory Bank      │
        └──────────┬─────────┘
                   │
                   │ Historical context
                   ▼
        ┌────────────────────┐
        │   Reasoning Layer  │
        └──────────┬─────────┘
                   │
                   ▼
        Likely Root Cause
        Suggested Resolution
        Similar Past Incident
                   │
                   ▼
              ENGINEER
                   │
                   │ Confirm resolution
                   ▼
        ┌────────────────────┐
        │ Retain Resolution  │
        └──────────┬─────────┘
                   │
                   └──────────────► Hindsight
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There are two important Hindsight operations in the workflow:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recall before investigation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retain after resolution.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The full &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt; describes the memory system in more detail, but these two operations were enough to define the core RecallOps workflow.&lt;/p&gt;




&lt;h2&gt;
  
  
  Recall Before Reasoning
&lt;/h2&gt;

&lt;p&gt;This became one of the most important decisions in the prototype.&lt;/p&gt;

&lt;p&gt;The system should retrieve historical context &lt;strong&gt;before&lt;/strong&gt; asking the reasoning model for an explanation.&lt;/p&gt;

&lt;p&gt;A simplified version of the incident-analysis flow looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;IncidentAnalysis&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;AnalyzeIncident&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;Incident&lt;/span&gt; &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// 1. Search previous engineering memory&lt;/span&gt;
    &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;memories&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;RecallRelevantIncidents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ErrorMessage&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="c1"&gt;// 2. Combine the new incident with recalled history&lt;/span&gt;
    &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;BuildIncidentContext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;memories&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="c1"&gt;// 3. Ask the reasoning layer to analyze everything&lt;/span&gt;
    &lt;span class="kt"&gt;var&lt;/span&gt; &lt;span class="n"&gt;analysis&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;AnalyzeWithAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;analysis&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is the order.&lt;/p&gt;

&lt;p&gt;Without memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;New Incident
     ↓
Reasoning Model
     ↓
Generic Troubleshooting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;RecallOps instead follows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;New Incident
     ↓
Recall Relevant Memories
     ↓
Current Evidence + Historical Evidence
     ↓
Reasoning Model
     ↓
Suggested Root Cause
     ↓
Recommended Checks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This means the reasoning layer does not have to rediscover every known fact about the application.&lt;/p&gt;

&lt;p&gt;It can start with what previous incidents already taught us.&lt;/p&gt;




&lt;h2&gt;
  
  
  Building the Recall Query
&lt;/h2&gt;

&lt;p&gt;The error message alone may not contain enough information.&lt;/p&gt;

&lt;p&gt;RecallOps can construct a query using the important characteristics of the incident.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="nf"&gt;BuildRecallQuery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Incident&lt;/span&gt; &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s"&gt;$"""
&lt;/span&gt;        &lt;span class="n"&gt;Find&lt;/span&gt; &lt;span class="n"&gt;previous&lt;/span&gt; &lt;span class="n"&gt;incidents&lt;/span&gt; &lt;span class="n"&gt;similar&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt; &lt;span class="n"&gt;problem&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

        &lt;span class="n"&gt;Service&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Service&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ErrorMessage&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;Component&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Component&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;Look&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;previous&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt; &lt;span class="n"&gt;causes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;resolutions&lt;/span&gt; &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="n"&gt;debugging&lt;/span&gt; &lt;span class="n"&gt;notes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
        &lt;span class="s"&gt;""";
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For our example, the query could become:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Service: crop-classification

Error:
Expected image dimensions 224x224.
Received dimensions 37x39.

Component:
model preprocessing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not exact string matching.&lt;/p&gt;

&lt;p&gt;The current error might say:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Received image dimensions: 37x39
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while the older incident might say:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Inference tensor shape does not match model input.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They are different strings describing essentially the same class of problem.&lt;/p&gt;

&lt;p&gt;The memory layer gives RecallOps a way to retrieve the older engineering context rather than relying entirely on identical wording.&lt;/p&gt;




&lt;h2&gt;
  
  
  What We Retain
&lt;/h2&gt;

&lt;p&gt;Storing every line of every log would defeat the purpose.&lt;/p&gt;

&lt;p&gt;We already have logging systems for that.&lt;/p&gt;

&lt;p&gt;The useful memory is the conclusion of the investigation.&lt;/p&gt;

&lt;p&gt;A resolved incident can be represented like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;IncidentMemory&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Service&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;set&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Error&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;set&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Component&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;set&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;RootCause&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;set&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Resolution&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;set&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;public&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="n"&gt;Status&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="k"&gt;get&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;set&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the crop-classification example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"service"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"crop-classification"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Expected 224x224 but received 37x39"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"component"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"image preprocessing"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"rootCause"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Resize step was skipped before inference"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"resolution"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Resize to 224x224 and normalize before creating the model tensor"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"resolved"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once an engineer confirms the resolution, RecallOps prepares a concise memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csharp"&gt;&lt;code&gt;&lt;span class="k"&gt;private&lt;/span&gt; &lt;span class="kt"&gt;string&lt;/span&gt; &lt;span class="nf"&gt;BuildMemory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Incident&lt;/span&gt; &lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s"&gt;$"""
&lt;/span&gt;        &lt;span class="n"&gt;Incident&lt;/span&gt; &lt;span class="n"&gt;resolved&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

        &lt;span class="n"&gt;Service&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Service&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ErrorMessage&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;Root&lt;/span&gt; &lt;span class="n"&gt;cause&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RootCause&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;Resolution&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Resolution&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="n"&gt;Component&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;incident&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Component&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="s"&gt;""";
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That content can then be retained in Hindsight.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;await hindsight.retain(resolvedIncidentMemory);
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The objective is not to remember everything.&lt;/p&gt;

&lt;p&gt;The objective is to remember the information that makes the next investigation better.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Second Incident Is Where It Gets Interesting
&lt;/h2&gt;

&lt;p&gt;The first incident only creates memory.&lt;/p&gt;

&lt;p&gt;The real value becomes visible when another incident appears.&lt;/p&gt;

&lt;p&gt;Imagine that a few weeks later another inference request fails:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Inference failed.

Model input required:
224 x 224

Input received:
48 x 41
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is not the exact same message.&lt;/p&gt;

&lt;p&gt;It is not even the same image size.&lt;/p&gt;

&lt;p&gt;RecallOps sends the incident context to Hindsight and retrieves the previous preprocessing incident.&lt;/p&gt;

&lt;p&gt;The returned context could contain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Previous relevant incident:

Crop-classification inference previously failed because
a raw satellite crop was passed directly to the model.

Model expected:
224 × 224

Previous input:
37 × 39

Root cause:
The image resize preprocessing step was skipped.

Resolution:
Resize and normalize the image before creating
the inference tensor.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the reasoning layer can generate a much more useful response.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LIKELY ROOT CAUSE

The current failure appears similar to a previously
resolved preprocessing incident.

The model expects 224×224 input, but the current
image is 48×41.

SIMILAR PREVIOUS INCIDENT

A previous crop-classification request failed when
a 37×39 satellite crop was passed directly into
the same model.

PREVIOUS ROOT CAUSE

The image resize step was missing.

RECOMMENDED CHECKS

1. Verify that image resizing runs before inference.
2. Confirm output dimensions are 224×224.
3. Apply the same normalization used during training.
4. Confirm channel ordering.
5. Only then construct the model tensor.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is very different from asking a model:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“What causes tensor dimension errors?”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;RecallOps can instead say:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“We experienced a similar problem before. This was the cause, this was the fix, and these are the checks worth performing first.”&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Memory Is Evidence, Not Truth
&lt;/h2&gt;

&lt;p&gt;There is an important limitation.&lt;/p&gt;

&lt;p&gt;A similar error does not guarantee an identical root cause.&lt;/p&gt;

&lt;p&gt;Two timeout exceptions can come from completely different components.&lt;/p&gt;

&lt;p&gt;Two database connection failures can have different causes.&lt;/p&gt;

&lt;p&gt;Two model inference errors can look nearly identical while originating at different stages of the preprocessing pipeline.&lt;/p&gt;

&lt;p&gt;For that reason, RecallOps does not treat retrieved memory as an automatic command.&lt;/p&gt;

&lt;p&gt;It presents it as historical evidence.&lt;/p&gt;

&lt;p&gt;The interface should distinguish between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Similar Past Incident
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Confirmed Current Root Cause
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Likewise, there is a difference between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Suggested Resolution
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Engineer Confirmed Resolution
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This distinction matters because automatically applying an old fix to a new production failure could create an even larger problem.&lt;/p&gt;

&lt;p&gt;The agent helps narrow the investigation.&lt;/p&gt;

&lt;p&gt;The engineer still makes the final decision.&lt;/p&gt;




&lt;h2&gt;
  
  
  From a Stateless Agent to a Learning Workflow
&lt;/h2&gt;

&lt;p&gt;Most simple AI applications follow this pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;INPUT
  ↓
MODEL
  ↓
OUTPUT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each request is mostly independent.&lt;/p&gt;

&lt;p&gt;RecallOps behaves differently:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 ┌──────── MEMORY ────────┐
                 │                         │
                 ▼                         │
INCIDENT → RECALL → REASON → RESOLVE → RETAIN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That final retain step changes what the next execution knows.&lt;/p&gt;

&lt;p&gt;Consider a sequence of incidents:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;INCIDENT 01
Image preprocessing failure
        ↓
Resolution retained


INCIDENT 07
Similar model failure
        ↓
Incident 01 recalled
        ↓
New resolution retained


INCIDENT 18
Authentication configuration failure
        ↓
Resolution retained


INCIDENT 42
New production failure
        ↓
Multiple relevant engineering memories available
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model itself has not necessarily changed.&lt;/p&gt;

&lt;p&gt;The context surrounding it has.&lt;/p&gt;

&lt;p&gt;That is what makes long-term agent memory interesting.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Learned Building RecallOps
&lt;/h2&gt;

&lt;p&gt;The first lesson was that &lt;strong&gt;memory becomes much more useful when connected to a specific workflow&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;“An AI agent with memory” is a technical capability.&lt;/p&gt;

&lt;p&gt;“An incident agent that remembers how we fixed previous failures” is a concrete use case.&lt;/p&gt;

&lt;p&gt;The second lesson was that &lt;strong&gt;remembering everything is not the goal&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Logs already remember almost everything.&lt;/p&gt;

&lt;p&gt;What is frequently missing is the relationship between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SYMPTOM
   +
ROOT CAUSE
   +
CONFIRMED FIX
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the part worth retaining.&lt;/p&gt;

&lt;p&gt;The third lesson was &lt;strong&gt;recall before reasoning&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The agent should see relevant historical evidence before generating a theory.&lt;/p&gt;

&lt;p&gt;Otherwise it can generate a theory first and use later information simply to support what it already decided.&lt;/p&gt;

&lt;p&gt;The fourth lesson was that &lt;strong&gt;human confirmation matters&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Past solutions are useful evidence, but production systems change.&lt;/p&gt;

&lt;p&gt;Dependencies change.&lt;/p&gt;

&lt;p&gt;Configuration changes.&lt;/p&gt;

&lt;p&gt;Infrastructure changes.&lt;/p&gt;

&lt;p&gt;A previous fix should guide an investigation rather than automatically control it.&lt;/p&gt;

&lt;p&gt;Finally, I learned that persistent memory changes the way I think about AI applications.&lt;/p&gt;

&lt;p&gt;The most interesting improvement does not always come from using a larger model.&lt;/p&gt;

&lt;p&gt;Sometimes it comes from giving the existing model access to better context.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Goal Isn't to Replace Debugging
&lt;/h2&gt;

&lt;p&gt;RecallOps is not trying to remove engineers from incident response.&lt;/p&gt;

&lt;p&gt;Debugging still requires understanding the current system.&lt;/p&gt;

&lt;p&gt;Engineers still need to inspect logs, reproduce problems, verify assumptions and validate fixes.&lt;/p&gt;

&lt;p&gt;The problem RecallOps addresses is more specific:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;valuable debugging knowledge disappears too easily.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We solve a production issue.&lt;/p&gt;

&lt;p&gt;We close the incident.&lt;/p&gt;

&lt;p&gt;We continue building.&lt;/p&gt;

&lt;p&gt;Several months later, somebody encounters almost the same problem and starts the investigation again from the beginning.&lt;/p&gt;

&lt;p&gt;With persistent memory, that cycle can change.&lt;/p&gt;

&lt;p&gt;A resolved incident becomes context for a future incident.&lt;/p&gt;

&lt;p&gt;A previous root cause becomes a clue.&lt;/p&gt;

&lt;p&gt;A successful resolution becomes a debugging path worth checking.&lt;/p&gt;

&lt;p&gt;And every confirmed fix makes the system's engineering memory slightly more useful.&lt;/p&gt;

&lt;p&gt;The idea behind RecallOps can therefore be summarized in one question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before spending an hour solving this problem, have we already solved something like it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;With Hindsight providing persistent agent memory, RecallOps can actually ask that question every time.&lt;/p&gt;

&lt;p&gt;The new error may be unfamiliar.&lt;/p&gt;

&lt;p&gt;The fix does not have to be.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>security</category>
      <category>automation</category>
    </item>
  </channel>
</rss>
