DEV Community

Kalcion Duke
Kalcion Duke

Posted on

What Changes When Your Incident Agent Remembers Using Hindsight




 What Changes When Your Incident Agent Remembers Using Hindsight

The Same Alert, Two Different Responses

Here is what actually happened when we ran the same checkout-api latency incident through two versions of the agent—one without memory, one with Hindsight memory enabled.

Without memory: "Check application logs, restart the affected pods, and increase the timeout if the issue persists."

With memory: "Pause the retry amplification, then compare connection-pool saturation against the remembered deploy pattern for checkout-api."

Those are not variations on a theme. They represent different hypotheses about what is wrong and different first actions. The second one happened to be correct, because the agent recalled a specific lesson from a prior incident on that service. This is what observable behavior change looks like when agent memory is implemented well.

The Core Story: Why the Behavior Change Is Real

Before we added Hindsight, the agent was a stateless LLM call. Each new incident arrived with no context about the service's history. The checkout-api service had a well-documented connection pool exhaustion pattern—it had appeared twice in three months—but the agent treated every new latency spike as if it were the first time checkout-api had ever experienced problems.

The first attempt to fix this was caching recent incident summaries and passing them as context in every prompt. This worked until the context window filled up, at which point the most relevant historical incidents were competing with less relevant recent ones for attention. The agent started mentioning irrelevant patterns just because they were recent.

The right approach was structured, scoped, similarity-ranked retrieval—which is exactly what Hindsight's agent memory model provides. Instead of dumping history into every prompt, we retain specific lessons at resolution time and recall only the most semantically similar ones at intake time. The service scope key ensures that payments-worker memories don't pollute checkout-api recalls.

The memory store defines what Hindsight works with:

// App.tsx — Memory type: what is retained per resolved incident

type Memory = {

id: string;

title: string;

service: string; // recall scope — this is the scope key for all recalls

severity: Severity;

summary: string; // what happened

lesson: string; // what to do next time

hits: number; // recall frequency counter

};

When a new incident arrives for checkout-api, the recall returns memories scoped to checkout-api, ranked by semantic similarity to the current signal. The agent analysis then references those results directly.

How the Behavior Manifests

The comparison view in IncidentIQ was built specifically to demonstrate this behavioral difference. You select a service, enter an incident description, and run both paths side by side. The left panel shows the without-memory response; the right shows the memory-augmented response. A quality score bar at the bottom quantifies how specifically each response addresses the service's known failure patterns.

The initial memories in the system already demonstrate the expected behavior. For checkout-api, memory mem-1 captures:

// initialMemories from App.tsx — real memory data for checkout-api

{

id: 'mem-1',

title: 'Checkout latency after deploy',

service: 'checkout-api',

severity: 'P2',

summary: 'Connection pool exhaustion caused a slow cascade across regional checkout workers.',

lesson: 'Check pool saturation before increasing application timeouts.',

hits: 14,

}

The hits: 14 tells you this memory has been recalled 14 times. That's 14 incidents where the agent surfaced the pool saturation hypothesis before the engineer started guessing.

The memory gallery also shows what payments-worker has taught the agent:

// initialMemories — payments-worker lesson

{

id: 'mem-2',

title: 'Worker queue saturation',

service: 'payments-worker',

severity: 'P1',

summary: 'A retry storm amplified a provider timeout into a full queue backlog.',

lesson: 'Pause retries and protect the queue before restarting workers.',

hits: 9,

}

When a payments-worker incident arrives with symptoms matching a queue buildup, this lesson surfaces first. The agent's analysis tells the engineer to protect the retry path before attempting a restart—not because it reasoned that from general principles, but because the team already learned it and Hindsight retained it.

Before vs. After

BEFORE (without memory, for checkout-api latency spike):

The agent receives the incident description and returns: "Check application logs, restart the affected pods, and increase the timeout if the issue persists." The response is generic. It treats this as a novel event with no service-specific context. An engineer following this advice would spend the first thirty minutes on generic triage before discovering the pool saturation angle independently.

AFTER (with Hindsight memory, same signal):

Hindsight recalls the mem-1 lesson. The agent responds: "Pause the retry amplification, then compare connection-pool saturation against the remembered deploy pattern for checkout-api." The engineer has a specific testable hypothesis from the first minute. The behavioral difference is the lesson that somebody paid for three months ago, stored and retrieved without anyone manually surfacing it.

What I Learned

The behavioral change only appears once the memory store has real data. The comparison view shows a dramatic quality difference partly because mem-1 and mem-2 contain genuinely specific lessons. With empty or generic lessons, the recalled output would look more like the no-memory baseline. Quality in, quality out—more acutely than in most systems.

Context window injection is a poor substitute. Passing incident history directly into every prompt worked until it didn't. The fundamental problem is that recency and relevance are not the same thing. The checkout-api connection pool memory from eight weeks ago is more relevant than the edge-gateway certificate renewal from last week, but naive context injection would prioritize the recent one.

Recall frequency is a quality proxy, not a guarantee. A memory with 14 hits has survived 14 invocations where an engineer could have marked the recall unhelpful. That's weak validation, not strong. But it's better than nothing, and it lets engineers identify which patterns recur most often.

The lesson field requires active prompting. Engineers write root causes naturally; they write lessons only when explicitly prompted. The four-step modal prompt ("What should the agent remember?") changed what engineers wrote in that field. Lesson quality is a UI problem as much as a data problem.

Conclusion

Behavioral change is the only real test of an agent memory system. Generic responses becoming specific, guesses becoming grounded hypotheses, and time-to-first-hypothesis dropping from thirty minutes to thirty seconds—those are observable outcomes, not architectural claims. Hindsight makes the retain/recall infrastructure straightforward to integrate. The Hindsight documentation is the right starting point. The engineering work is designing what you retain, how you scope it, and when you recall it.

Top comments (0)