DEV Community

NIKHIL K
NIKHIL K

Posted on

How I Designed Evidence-Based Memory Retrieval With Hindsight

How I Designed Evidence-Based Memory Retrieval With Hindsight

Every process document I have read describes a company that does not exist. The real workflow lives in the cases: who approved what, in what order, and what went wrong afterward. I built ShadowOps to read those cases, compare them with the official process, and show the evidence behind every claim it makes.

What the system does

ShadowOps ingests an organization's history: approvals, security reviews, incidents, and outcomes. Each record becomes a memory with an ID, tags, and the step sequence it contributed to. From those memories the system derives patterns, workflows, and hidden dependencies, then lets you question them in plain English.

The app has six screens: a dashboard, an Ask interface (Intelligence), a Memory Explorer, discovered Workflows, a 3D Organization view, and a System Health page for the memory layer. It is a React 19 and TypeScript app built with Vite. Convex provides the database and authentication, and Three.js draws the organization view.

The rule I would not bend on is provenance. Every answer follows one chain:

Answer → Observed Pattern → Workflow → Historical Cases → Memories
Enter fullscreen mode Exit fullscreen mode

If a claim cannot walk that chain back to a memory, the system does not make it.

The through-line: drift between policy and practice

A vendor policy might say Procurement, then Manager approval. The cases may repeatedly show Security and Finance reviews between those steps, with IT at the end. Nobody changed the policy. The practice moved.

I keep four ideas separate in the data model:

  • Official process: the documented sequence.
  • Observed process: the sequence that repeats in cases.
  • Organizational memory: the decisions and outcomes that explain the sequence.
  • Hidden dependencies: step-to-step relationships that no document lists. The point is not to declare the policy wrong. It is to make the gap visible so an operator can decide whether to change the policy or the practice.

Why I put agent memory underneath it

A system like this fails in a specific way if its memory is a pile of text. It retrieves things that look similar and answers with confidence it has not earned. I wanted memory that keeps context across interactions and comes back only when it changes the answer.

That is the design idea behind Hindsight agent memory: retain what happened, recall what is relevant later. The Hindsight documentation covers the retain and recall model in detail, and Vectorize's write-up on what agent memory is explains why persistent context matters for agents that need to learn from outcomes.

In ShadowOps the split of responsibilities looks like this:

  • The memory layer retains cases and outcomes, and recalls the relevant ones for a question.
  • Pattern discovery summarizes repeated behavior across those memories.
  • The retrospective view asks what the outcomes reveal in hindsight.
  • Ask turns recalled context into an explanation. Keeping these apart means I can test and replace each layer without making one opaque agent responsible for everything.

Memory records that carry their own structure

Every ingested memory lands in one Convex table. The field that matters most is the optional step sequence:

ingestedMemories: defineTable({
  source: v.string(),      // "csv" | "json" | "manual"
  sourceName: v.string(),
  memoryId: v.string(),    // e.g. M-9101, assigned at ingest time
  kind: v.string(),        // episodic | decision | incident | outcome | procedural | workflow
  title: v.string(),
  summary: v.string(),
  date: v.string(),
  actors: v.array(v.string()),
  sequence: v.optional(v.array(v.string())),
  confidence: v.number(),
  createdAt: v.number(),
})
  .index("by_kind", ["kind"])
  .index("by_user", ["userId"]),
Enter fullscreen mode Exit fullscreen mode

A memory that carries ["Procurement", "Security", "Finance", "Manager", "IT"] in sequence is what makes workflow comparison possible. Without it, a memory is just text.

The ingest path

Records arrive as CSV, JSON, or manual entry. Each one is classified into a kind, given a confidence score, and stored under a new memory ID:

function classifyMemory(text: string): string {
  const t = text.toLowerCase();
  if (/incident|outage|sev|breach|blocked|failed|emergency/.test(t)) return "incident";
  if (/decision|approved|rejected|sign.?off|mandat|policy/.test(t)) return "decision";
  if (/outcome|result|completed|delivered|shipped|record/.test(t)) return "outcome";
  if (/procedure|runbook|process|sop|checklist/.test(t)) return "procedural";
  return "episodic";
}

function confidenceFor(kind: string, actors: number, seqLen: number): number {
  let c = 62 + actors * 5 + seqLen * 4;
  if (kind === "decision" || kind === "incident") c += 6;
  return Math.max(55, Math.min(97, c));
}
Enter fullscreen mode Exit fullscreen mode

This is deliberately simple. A record with more actors and a longer recorded sequence is more useful evidence than a one-line note, and the score reflects that. I would rather have a plain, readable heuristic I can argue with than an opaque classifier I cannot inspect. The clamp to 55-97 keeps any single record from looking either worthless or certain.

Observed workflows are computed, not drawn

The Ask pipeline reads the stored sequences and counts how often one step directly follows another:

function pairCount(from: string, to: string): number {
  const seqs = memories.filter((m) => m.sequence?.length).map((m) => m.sequence!);
  let n = 0;
  for (const s of seqs) {
    for (let i = 0; i < s.length - 1; i++) {
      if (s[i] === from && s[i + 1] === to) n++;
    }
  }
  return n;
}
Enter fullscreen mode Exit fullscreen mode

Adjacent-pair counts are the raw material for an observed workflow. If Security is followed by Finance in most recorded sequences and the official process never lists that edge, that gap is a hidden dependency.

Answers are structured, not paragraphs

An answer is a typed object. Each step points at the memory IDs behind it, and the caution travels with the answer:

export interface AgentStep {
  label: string;
  detail: string;
  memoryIds?: string[];
}

export interface AgentAnswer {
  intent: string;
  summary: string;
  kind: "workflow" | "explanation" | "steps" | "comparison" | "insufficient";
  stats: { label: string; value: string }[];
  steps: AgentStep[];
  memoryIds: string[];
  workflowMode?: "observed" | "official" | "compare";
  caution?: string;
  confidence: number;
}
Enter fullscreen mode Exit fullscreen mode

Notice the "insufficient" kind. When the recorded evidence does not support an answer, the system says so instead of producing prose. For the question "How do we approve a new vendor?", a supported answer separates four things:

  • Historical fact: Case M-1042 recorded Security approval on Jan 21, before Finance on Jan 28.
  • Observed pattern: the ordering repeats across cases.
  • Interpretation: Security sign-off likely gates Finance review for medium and high-risk data.
  • Recommendation: reflect Security before Finance in the documented workflow, or pilot parallel review for low-risk cases. A reader can see which part was recorded and which part was inferred. The recorded parts link to cases, and the inferred parts are labeled as inference.

Caveats live in the data

Operational records are observational. If cases that skipped a step have more incidents, that is a signal, not proof of cause. So every answer object carries a caution field:

caution: "Association ≠ causation. No controlled evidence establishes that Security causes Finance outcomes.",
Enter fullscreen mode Exit fullscreen mode

Other answers say "Observed ≠ intended. The documented flow may exist for reasons memory does not capture." I put the caveat in the data structure on purpose. A caveat in a UI template is easy to drop. A field on the answer is not.

Closing the loop with outcomes

A recommendation should not be the end of the story. In the design, an operator accepts, reviews, or dismisses a recommendation. After the workflow runs, the outcome is recorded as success, partial success, or failure, written back as a new memory, and pattern discovery runs again.

Recommendation
    → operator decision
    → workflow execution
    → recorded outcome
    → outcome memory
    → pattern rediscovery
Enter fullscreen mode Exit fullscreen mode

This is the part where retaining memory pays off most. The system learns from observed outcomes, not from its own generated text, and a human makes every decision explicitly.

What I would tell another engineer

  1. The documented process is one source of truth, not the only one. Comparing policy with case history teaches more than reading either.
  2. Make provenance a data structure. If each answer step stores its memory IDs, "show your evidence" is a lookup.
  3. Keep caveats next to claims. A caution field on the answer is harder to forget than a sentence in a template.
  4. Do not let similarity stand in for evidence. Retrieval should weigh more than semantic closeness, and every pattern should show the cases that contradict it as well as the cases that support it.
  5. Separate the layers. Memory, pattern discovery, retrospective analysis, and language generation solve different problems, and separate layers are easier to test. I do not want ShadowOps to replace an operator's judgment. I want it to make the organization's past easier to examine before the next decision gets made.

Top comments (0)