DEV Community

Sasuke
Sasuke

Posted on

The Hardest Part of Building an AI Memory System Is Making Its Failures Visible

I expected the difficult part of Waada to be retrieval.

It wasn't.

The harder part has been separating four things that easily collapse into one result:

  1. Was the right data retained?
  2. Was the right evidence recalled?
  3. Did the model interpret that evidence correctly?
  4. Was every external dependency healthy enough for the path to run?

To the user, all four look the same: an empty or wrong answer. That distinction has shaped how I'm building and evaluating the system.

Waada is a continuity system, not a chatbot

Waada is designed for B2B sales handoffs. When an account changes hands, the new rep inherits months of emails, Slack threads and call transcripts, and one sentence written weeks ago can decide whether the handoff goes well.

Waada imports account history from email, Slack, transcripts, audio and CRM data. It normalizes those sources, retains the interactions in an account-scoped Hindsight memory bank, and exposes several memory-aware operations:

  • Continuity brief: what the new owner needs to know
  • Commitment ledger: what was promised, by whom, and whether it's still open
  • Resolved-objection warnings ("landmines"): topics that were already settled
  • Question answering with citations
  • Context comparison against simpler baselines

The architecture:

imports
   |
   v
parseFiles()
   |
   v
Interaction[]
   |
   v
ingest()
   |
   +--------------------+
   |                    |
   v                    v
local JSON          Hindsight
                         |
                         v
                     recall
                         |
                         v
                    agent logic
                         |
                         v
                        LLM
Enter fullscreen mode Exit fullscreen mode

These layers are deliberately separate. That separation turned out to be essential once I started looking at failures.

The first reliability boundary is ingestion

Every source is converted into an Interaction:

type Interaction = {
  account: string;
  sourceId: string;
  type: "call" | "email" | "slack" | "meeting" | "note";
  date: string;
  title: string;
  participants: string[];
  content: string;
  source: "email" | "slack" | "transcript" | "audio" | "crm";
};
Enter fullscreen mode Exit fullscreen mode

The parser layer handles source-specific behavior. Email parsing uses message IDs (or derived IDs), strips quoted replies, and can fall back from HTML. Slack exports become channel/day interactions and resolve people through users.json. Transcripts support front matter and can use the LLM for metadata extraction. Audio goes through transcription first.

So a malformed or ambiguous source can fail before it becomes memory. I want that boundary. If the data is wrong at ingestion, blaming retrieval later is pointless.

Stable identity matters more than it looks

ingest() checks a local manifest of retained sourceId values before writing anything, and Hindsight also receives a document ID. The goal is idempotence:

import
  |
  v
sourceId already retained?
  |
  +-- yes --> skip
  |
  +-- no ---> retain
Enter fullscreen mode Exit fullscreen mode

Without this, uploading the same export twice creates duplicate evidence. That's especially dangerous in retrieval systems, because repetition can look like confidence. Repeated imports should be boring.

Hindsight introduces a new dependency boundary

Hindsight sits behind a HindsightMemory adapter. Each account maps to its own bank, derived from the account slug:

const bankId = bankIdFor(account);
Enter fullscreen mode Exit fullscreen mode

The adapter handles retain, recall, bank creation/update, document identity and retries. The Hindsight client import is confined to the memory package, so the agent only ever sees this:

agent
  |
  v
Memory.search()
  |
  v
Hindsight adapter
  |
  v
Hindsight
Enter fullscreen mode Exit fullscreen mode

The agent doesn't need to know whether memory is local or cloud-hosted. (See the Hindsight docs for how banks and recall work.)

A memory failure is not a retrieval failure

This became the central idea during evaluation:

  • Never retained: a perfect query still returns nothing. Ingestion failure.
  • Retained, but a poor query: recall returns the wrong evidence. Retrieval failure.
  • Right evidence, malformed output: generation failure.
  • Everything correct, but Hindsight Cloud returns an intermittent error: infrastructure failure.

All four can produce the same user-facing result. They should not be counted as the same problem. That's why the adapter retries transient network, 429 and 5xx errors, while the agent layer handles evidence and validates structured output.

Making failures visible: per-request traces

Handling failures isn't the same as seeing them. Retries and validation keep the system from breaking, but they also hide where things went wrong.

The next step I'm working toward is a trace per operation that records each boundary separately, something like:

{
  "operation": "ask",
  "account": "acme",
  "ingest": { "sourcesRetained": 42, "skippedDuplicates": 3 },
  "recall": { "hits": 11, "evidenceUsed": 6, "retries": 1 },
  "generation": { "validation": "repaired" },
  "infra": { "hindsightErrors": 1, "rateLimitWaits": 0 }
}
Enter fullscreen mode Exit fullscreen mode

With this, "the answer was wrong" becomes a specific question: was evidence missing, badly selected, or misread?

The commitment ledger exposes these boundaries

commitmentLedger() recalls promise-oriented evidence, deduplicates and chunks it, extracts structured commitments, merges duplicates, and sorts the result (simplified):

const hits = await recallPromiseEvidence(account);

const commitments = await extractCommitments(
  chunkAndDeduplicate(hits)
);

return sortLedger(commitments);
Enter fullscreen mode Exit fullscreen mode

Each step has its own failure mode. Empty recall means no usable evidence. If extraction fails, the system must not invent commitments. Duplicate evidence should merge, not multiply. Open, overdue commitments get priority.

Every ledger entry carries its evidence and source, not just a model assertion.

Landmines make false positives obvious

The landmine feature looks for objections that were resolved and agreements that were accepted. A false positive here is costly: if the system marks a topic "settled" when the conversation actually left it open, the new rep may avoid a discussion that still needs to happen.

objection
   |
   v
evidence of resolution?
   |
   v
landmine
Enter fullscreen mode Exit fullscreen mode

So the extractor works only from bounded, recalled evidence and validated structured output. The system should only surface a higher-level interpretation when there's supporting history.

Ask makes provenance visible

Ask is intentionally evidence-oriented. For a question like:

What changed since July?
Enter fullscreen mode Exit fullscreen mode

Waada recalls relevant history, caps the evidence, asks the model to answer only from that context, and returns citations (simplified):

const hits = await memory.search(account, question);
const evidence = capEvidence(hits);

return llm.chat({ question, evidence });
Enter fullscreen mode Exit fullscreen mode

In a live run on a synthetic test account ("Acme"), Ask correctly surfaced the go-live date moving from Q3 to Q4, with citations pointing back to the recalled contexts. The answer isn't unexplained model knowledge; it's traceable.

The LLM is treated as unreliable input

Structured output sounds safer than plain text, but it can still be malformed or semantically wrong. Waada validates generated objects with Zod:

structured output
      |
      v
Zod validation
      |
   success?
    /    \
  yes     no
  |        |
return   repair
           |
           v
       validate again
           |
        success?
        /     \
      yes      no
       |        |
     return    null
Enter fullscreen mode Exit fullscreen mode

The application fails safely instead of trusting a response because it happens to look like JSON.

Rate limits are part of correctness

The system has an explicit input budget of 5,000 tokens, estimated conservatively from character counts, plus rate-limit-aware retries. The continuity brief runs its major LLM calls sequentially because the provider enforces a shared rate window.

That isn't just a performance choice. A system that only works when calls happen to stay under a provider's limits isn't reliable enough to evaluate, so provider constraints are part of the design.

The evaluation record is intentionally mixed

The live evaluation log is one of the most useful artifacts in the repo, precisely because it doesn't turn every run into a success story. So far it includes:

  • Successful "What changed since July?" answers
  • Multiple runs that established the pricing landmine under some conditions
  • Structured-output variability
  • Rate-limit and prompt-size pressure
  • Intermittent Hindsight Cloud failures
  • A run invalidated by recall failures
  • One run where the summary-only baseline scored higher than Waada on its checks

That last one matters. When an account's full history fits comfortably inside the summary window, reading everything in order can beat selective recall. Memory earns its place when history outgrows what you can simply read, and proving that is part of what the evaluation still has to do.

The takeaway: the architecture and the observed model behavior are separate questions. Establishing a working memory path doesn't establish retrieval quality, and one good answer doesn't prove the system holds up under every dependency condition.

Comparison paths exist for the same reason

The comparison operation runs three contexts side by side:

CRM-only
summary-only
Waada memory-aware
Enter fullscreen mode Exit fullscreen mode

CRM-only reads just CRM fields. Summary-only reads normalized interactions chronologically, capped at 12,000 characters, without Hindsight. Waada uses recalled evidence.

This compares available context. It's not a benchmark of commercial CRM products, and I'm careful not to let one favorable result turn into a broad performance claim.

What Waada doesn't do yet

Waada is still in development. The repository does not yet include:

  • Production Gmail, Slack or HubSpot connector back ends
  • A Meet capture route or extension
  • An MCP server
  • The agent report() function
  • Encryption at rest
  • Production authentication and tenancy controls
  • Deletion/retention guarantees beyond bank deletion
  • Production-scale performance metrics

The integrations page is a product-facing status UI, not proof that every connector has a working backend.

I think this is worth stating plainly. AI applications become hard to trust when architecture diagrams quietly turn placeholders into capabilities.

What the failures have taught me so far

1. Separate memory from generation. First ask whether the right evidence was available. Then ask whether the model used it correctly.

2. Preserve provenance. A useful answer should trace back to recalled contexts and document IDs.

3. Validate every generated structure. JSON-shaped output is still external input.

4. Treat provider limits as system behavior. Rate and context limits shape architecture.

5. Keep degraded runs. A failed or invalid run is evidence about the system. Deleting it to improve the scorecard throws away information.

The broader lesson

Building Waada has changed what I consider a reliable AI system. Reliability doesn't start with choosing a stronger model. It starts with being able to answer, for any output:

Was the source ingested?
Was it retained?
Was it recalled?
Was the right evidence selected?
Did the model interpret it correctly?
Did the external service stay healthy?
Enter fullscreen mode Exit fullscreen mode

Hindsight provides the durable memory layer, and it's a major part of that chain. But memory doesn't remove the other engineering problems. It makes them more visible, and that's a good thing.

The most useful AI systems are the ones where you can tell not only what the model said, but why it had the information it had, and exactly where it can fail.


Waada is a work in progress. If you're building agent memory or evaluating retrieval systems, I'd love to hear how you separate these failure modes in your own work.

Top comments (0)