DEV Community

Cover image for I Stored Lessons, Not Incidents, in Hindsight
Rakshith Reddy
Rakshith Reddy

Posted on AI-assisted

I Stored Lessons, Not Incidents, in Hindsight


Enter fullscreen mode Exit fullscreen mode

Our worst Friday deploys all had the same shape, and nobody noticed until we made a memory system say it out loud. The postmortems were fine. The problem was that nobody rereads a postmortem before the next risky change.

So I built Preflight, a service that sits in front of risky changes (deploys, migrations, architecture changes) and checks the plan against what the organization has already learned the hard way. The memory layer is Hindsight, an open-source agent memory system. This post is about the one design decision that shaped everything else: the unit of memory is the lesson extracted from an experience, not the raw incident.

What Preflight does

You describe a plan in plain text: "Friday deploy of Billing API + 3 dependency updates." Preflight then does the following:

  1. Extracts two to six short trigger conditions from the plan (Friday deploy, dependency updates, bundled changes).
  2. Builds three or four recall queries from the plan and those conditions, and runs them against Hindsight.
  3. Merges and de-dupes what comes back.
  4. Asks an LLM to judge the plan using only the recalled memories, and to return a structured verdict with cited event IDs.

The backend is a small Node/TypeScript Express app. It holds every Hindsight and Groq call, and the frontend only talks to my own /api/* routes, so no key ever reaches a browser. There is no database. Hindsight is the only place anything is remembered, and the Hindsight docs describe the three operations I lean on: retain to write, recall to search, and reflect to reason across what's stored.

Alongside the preflight check, the app has an Ask History tab that reconstructs what happened in past events, a Patterns tab that surfaces recurring failures, and a form for adding new experiences.

The through-line: what is a memory?

The first version of my design retained postmortems verbatim. It failed quickly, and the reason is instructive.

A postmortem is written for people who were there. It has a timeline, quoted chat messages, and three paragraphs of context that only make sense if you remember the week. If you retain that as-is and later ask whether a Friday deploy with dependency bumps is risky, retrieval has to find the one sentence that matters inside a wall of narrative. The memory that does surface is whichever document happened to use similar words, not whichever one carries the relevant insight.

The useful part of an incident is what the organization learned. So I extract that first, in a fixed shape, and retain that:

export const lessonRecord = (x: Ev) =>
`LESSON RECORD ${x.id} | Type: ${x.type} | Date: ${x.date}
Situation: ${x.situation} | Believed: ${x.what_people_believed} | Decision: ${x.decision} | What happened: ${x.what_happened} |
Root cause: ${x.root_cause} | Recovery: ${x.recovery} | LESSON: ${x.lesson} | Trigger conditions: ${x.trigger_conditions.join("; ")}`;
Enter fullscreen mode Exit fullscreen mode

Two fields carry most of the weight.

Believed records what people thought at the time. For the dependency incident it reads roughly "minor version bumps are safe; one release is cheaper than four." That belief is what a new plan is most likely to repeat, so it deserves to be recalled on its own terms, separate from the outcome.

Trigger conditions are short phrases: Friday deploy, dependency updates, batch window 01:00-04:00, shared IOPS. At check time, I extract the same kind of phrases from the incoming plan and use them as recall queries. The plan and the memory end up described in the same vocabulary, which does more for retrieval quality than any amount of prompt tuning on the judge.

Counter-examples are memories too

If you only retain failures, the judge can only say "this resembles a past failure." That yields a system that refuses to let anyone do anything on a Friday, and people stop reading it within a week.

So I also retain the cases where the scary-looking thing went fine. One record is a Friday deploy behind a feature flag with no dependency changes. Another is a Saturday migration that ran outside the batch window after a load test. They use the same lesson format, with Type: COUNTER-EXAMPLE.

The payoff is in the recommendations. For the Friday plan above, the verdict cites two independent failures (E002 and E003) and points at E010, the flagged deploy that went fine. The advice becomes "separate the dependency update and put the change behind a flag," not "don't deploy on Fridays."

Recall: several small queries, merged

One query is rarely enough to cover a plan that touches several risk factors, so each request builds a handful of targeted queries and merges the results:

export async function recallMany(queries: string[]) {
  const t0 = Date.now();
  const per = await Promise.all(queries.map(async (q) => ({ query: q, results: await recall(q) })));
  const seen = new Set<string>(); const merged: string[] = [];
  for (const p of per) for (const t of p.results) if (!seen.has(t)) { seen.add(t); merged.push(t); }
  return { queries: per, merged, latencyMs: Date.now() - t0 };
}
Enter fullscreen mode Exit fullscreen mode

The raw queries, the raw recalled text and the latency all show up in a permanent Memory panel next to every verdict. When the judge says something surprising, I can open the panel and see whether the memory was missing or the reasoning was wrong. Those are very different bugs, and without the panel they look identical.

Retain is asynchronous on the server side, so newly written records may not be searchable immediately. After loading history, the UI polls recall for a known seeded fact every five seconds, up to a configured timeout, and shows a progress bar. I would rather show a loading state than let someone run a check against a half-indexed bank and conclude the system is broken.

Keeping the judge honest

The judge prompt is where the discipline lives. These are the rules that matter:

Verdicts: RED = plan matches >=2 conditions of a past failure, or >=2 independent failures share one condition.
YELLOW = partial match or mixed outcomes.
GREEN = nothing relevant recalled (state that absence of memory is not proof of safety).
Always look for and explain counter-evidence (similar situations that went fine).
Treat "PREFLIGHT OUTCOME" records as evidence...
Enter fullscreen mode Exit fullscreen mode

The model may use only what recall returned and must cite event IDs. GREEN is deliberately worded as "nothing relevant was recalled," never "this is safe." A memory system can only warn you about what it has seen, and the output should say so.

There is also a Memory OFF switch. With it off, the same plan goes to the LLM with no recall and comes back as a generic "looks reasonable" answer, labeled as using no organizational memory. That comparison is the fastest way I know to show a skeptical teammate what the memory is contributing.

Warnings that got ignored become evidence

The part I care about most is the feedback loop. Under every verdict there are buttons: "We changed the plan" or "We're proceeding anyway," and later "It went fine" or "It failed." Each click retains a record:

const rec = `PREFLIGHT OUTCOME | Plan: ${plan} | Verdict: ${verdict} | Cited: ${(cited ?? []).join(", ") || "none"} | Action: ${action ?? "not recorded"} | Result: ${result ?? "pending"} | Note: ${note ?? ""} | Date: ${new Date().toISOString().slice(0, 10)}`;
await hs.retain(rec, `outcome-${Date.now()}`);
Enter fullscreen mode Exit fullscreen mode

Because these land in the same bank as the lessons, later recalls can surface them, and the judge prompt treats them as evidence. If someone overrode a RED and the deploy failed, the next similar plan gets a higher-confidence warning that cites that outcome. If a warning was overridden and things went fine, that shows up too, and it should temper the next verdict.

I picked this design so the memory bank wouldn't stay frozen at whatever I seeded on day one. The system's own track record goes into the same store as everything else, and I didn't have to build a second store for it.

Patterns: asking the memory instead of counting

The Patterns tab asks one question: what mistakes are we at risk of repeating? I send that to Hindsight's reflect, then convert the result into structured patterns with a title, supporting event IDs, a failure count and a preventive rule.

If reflect errors or returns nothing useful, the server falls back to broad recall plus LLM clustering by shared trigger condition, and the UI states which path ran. I didn't want a silent fallback to make one path look like the other.

On our history, three clusters emerge: Friday deploys that change dependencies, datastore work without rehearsal or a load test, and architecture changes that shipped without documenting an operational constraint such as ordering, a shared transaction or an eviction policy. The last one now has a concrete rule attached: write an ADR before the change, not after the revert.

What I'd tell someone building this

Decide what a memory is before you write to the store. Everything downstream (recall quality, judge reasoning, the Patterns view) got easier once each record was a lesson with explicit trigger conditions.

Retain the boring cases and the near-misses. Counter-examples are what make warnings specific enough to act on.

Make memory visible. A panel showing raw queries, recalled text and latency turns "the AI said something odd" into a debuggable problem.

Let your own outputs become memories. Outcome records cost almost nothing to write and change how the next judgment goes.

Refuse to fake the memory layer.If a call fails, the UI shows the real error. There are no canned verdicts anywhere in the codebase. That rule caught integration problems early, because there was nothing to hide them.

If you're thinking about agent memory more broadly, Vectorize's overview of agent memory is a good place to see how retain, recall and reflect fit together. The Hindsight GitHub repository has the client library and self-hosting instructions, which I'd look at if you'd rather keep the memory bank in your own environment.

Top comments (0)