I expected the difficult part of Waada to be retrieval.
It wasn't.
The harder part has been separating four things that easily collapse into one result:
- Was the right data retained?
- Was the right evidence recalled?
- Did the model interpret that evidence correctly?
- Was every external dependency healthy enough for the path to run?
To the user, all four look the same: an empty or wrong answer. That distinction has shaped how I'm building and evaluating the system.
Waada is a continuity system, not a chatbot
Waada is designed for B2B sales handoffs. When an account changes hands, the new rep inherits months of emails, Slack threads and call transcripts, and one sentence written weeks ago can decide whether the handoff goes well.
Waada imports account history from email, Slack, transcripts, audio and CRM data. It normalizes those sources, retains the interactions in an account-scoped Hindsight memory bank, and exposes several memory-aware operations:
- Continuity brief: what the new owner needs to know
- Commitment ledger: what was promised, by whom, and whether it's still open
- Resolved-objection warnings ("landmines"): topics that were already settled
- Question answering with citations
- Context comparison against simpler baselines
The architecture:
imports
|
v
parseFiles()
|
v
Interaction[]
|
v
ingest()
|
+--------------------+
| |
v v
local JSON Hindsight
|
v
recall
|
v
agent logic
|
v
LLM
These layers are deliberately separate. That separation turned out to be essential once I started looking at failures.
The first reliability boundary is ingestion
Every source is converted into an Interaction:
type Interaction = {
account: string;
sourceId: string;
type: "call" | "email" | "slack" | "meeting" | "note";
date: string;
title: string;
participants: string[];
content: string;
source: "email" | "slack" | "transcript" | "audio" | "crm";
};
The parser layer handles source-specific behavior. Email parsing uses message IDs (or derived IDs), strips quoted replies, and can fall back from HTML. Slack exports become channel/day interactions and resolve people through users.json. Transcripts support front matter and can use the LLM for metadata extraction. Audio goes through transcription first.
So a malformed or ambiguous source can fail before it becomes memory. I want that boundary. If the data is wrong at ingestion, blaming retrieval later is pointless.
Stable identity matters more than it looks
ingest() checks a local manifest of retained sourceId values before writing anything, and Hindsight also receives a document ID. The goal is idempotence:
import
|
v
sourceId already retained?
|
+-- yes --> skip
|
+-- no ---> retain
Without this, uploading the same export twice creates duplicate evidence. That's especially dangerous in retrieval systems, because repetition can look like confidence. Repeated imports should be boring.
Hindsight introduces a new dependency boundary
Hindsight sits behind a HindsightMemory adapter. Each account maps to its own bank, derived from the account slug:
const bankId = bankIdFor(account);
The adapter handles retain, recall, bank creation/update, document identity and retries. The Hindsight client import is confined to the memory package, so the agent only ever sees this:
agent
|
v
Memory.search()
|
v
Hindsight adapter
|
v
Hindsight
The agent doesn't need to know whether memory is local or cloud-hosted. (See the Hindsight docs for how banks and recall work.)
A memory failure is not a retrieval failure
This became the central idea during evaluation:
- Never retained: a perfect query still returns nothing. Ingestion failure.
- Retained, but a poor query: recall returns the wrong evidence. Retrieval failure.
- Right evidence, malformed output: generation failure.
- Everything correct, but Hindsight Cloud returns an intermittent error: infrastructure failure.
All four can produce the same user-facing result. They should not be counted as the same problem. That's why the adapter retries transient network, 429 and 5xx errors, while the agent layer handles evidence and validates structured output.
Making failures visible: per-request traces
Handling failures isn't the same as seeing them. Retries and validation keep the system from breaking, but they also hide where things went wrong.
The next step I'm working toward is a trace per operation that records each boundary separately, something like:
{
"operation": "ask",
"account": "acme",
"ingest": { "sourcesRetained": 42, "skippedDuplicates": 3 },
"recall": { "hits": 11, "evidenceUsed": 6, "retries": 1 },
"generation": { "validation": "repaired" },
"infra": { "hindsightErrors": 1, "rateLimitWaits": 0 }
}
With this, "the answer was wrong" becomes a specific question: was evidence missing, badly selected, or misread?
The commitment ledger exposes these boundaries
commitmentLedger() recalls promise-oriented evidence, deduplicates and chunks it, extracts structured commitments, merges duplicates, and sorts the result (simplified):
const hits = await recallPromiseEvidence(account);
const commitments = await extractCommitments(
chunkAndDeduplicate(hits)
);
return sortLedger(commitments);
Each step has its own failure mode. Empty recall means no usable evidence. If extraction fails, the system must not invent commitments. Duplicate evidence should merge, not multiply. Open, overdue commitments get priority.
Every ledger entry carries its evidence and source, not just a model assertion.
Landmines make false positives obvious
The landmine feature looks for objections that were resolved and agreements that were accepted. A false positive here is costly: if the system marks a topic "settled" when the conversation actually left it open, the new rep may avoid a discussion that still needs to happen.
objection
|
v
evidence of resolution?
|
v
landmine
So the extractor works only from bounded, recalled evidence and validated structured output. The system should only surface a higher-level interpretation when there's supporting history.
Ask makes provenance visible
Ask is intentionally evidence-oriented. For a question like:
What changed since July?
Waada recalls relevant history, caps the evidence, asks the model to answer only from that context, and returns citations (simplified):
const hits = await memory.search(account, question);
const evidence = capEvidence(hits);
return llm.chat({ question, evidence });
In a live run on a synthetic test account ("Acme"), Ask correctly surfaced the go-live date moving from Q3 to Q4, with citations pointing back to the recalled contexts. The answer isn't unexplained model knowledge; it's traceable.
The LLM is treated as unreliable input
Structured output sounds safer than plain text, but it can still be malformed or semantically wrong. Waada validates generated objects with Zod:
structured output
|
v
Zod validation
|
success?
/ \
yes no
| |
return repair
|
v
validate again
|
success?
/ \
yes no
| |
return null
The application fails safely instead of trusting a response because it happens to look like JSON.
Rate limits are part of correctness
The system has an explicit input budget of 5,000 tokens, estimated conservatively from character counts, plus rate-limit-aware retries. The continuity brief runs its major LLM calls sequentially because the provider enforces a shared rate window.
That isn't just a performance choice. A system that only works when calls happen to stay under a provider's limits isn't reliable enough to evaluate, so provider constraints are part of the design.
The evaluation record is intentionally mixed
The live evaluation log is one of the most useful artifacts in the repo, precisely because it doesn't turn every run into a success story. So far it includes:
- Successful "What changed since July?" answers
- Multiple runs that established the pricing landmine under some conditions
- Structured-output variability
- Rate-limit and prompt-size pressure
- Intermittent Hindsight Cloud failures
- A run invalidated by recall failures
- One run where the summary-only baseline scored higher than Waada on its checks
That last one matters. When an account's full history fits comfortably inside the summary window, reading everything in order can beat selective recall. Memory earns its place when history outgrows what you can simply read, and proving that is part of what the evaluation still has to do.
The takeaway: the architecture and the observed model behavior are separate questions. Establishing a working memory path doesn't establish retrieval quality, and one good answer doesn't prove the system holds up under every dependency condition.
Comparison paths exist for the same reason
The comparison operation runs three contexts side by side:
CRM-only
summary-only
Waada memory-aware
CRM-only reads just CRM fields. Summary-only reads normalized interactions chronologically, capped at 12,000 characters, without Hindsight. Waada uses recalled evidence.
This compares available context. It's not a benchmark of commercial CRM products, and I'm careful not to let one favorable result turn into a broad performance claim.
What Waada doesn't do yet
Waada is still in development. The repository does not yet include:
- Production Gmail, Slack or HubSpot connector back ends
- A Meet capture route or extension
- An MCP server
- The agent
report()function - Encryption at rest
- Production authentication and tenancy controls
- Deletion/retention guarantees beyond bank deletion
- Production-scale performance metrics
The integrations page is a product-facing status UI, not proof that every connector has a working backend.
I think this is worth stating plainly. AI applications become hard to trust when architecture diagrams quietly turn placeholders into capabilities.
What the failures have taught me so far
1. Separate memory from generation. First ask whether the right evidence was available. Then ask whether the model used it correctly.
2. Preserve provenance. A useful answer should trace back to recalled contexts and document IDs.
3. Validate every generated structure. JSON-shaped output is still external input.
4. Treat provider limits as system behavior. Rate and context limits shape architecture.
5. Keep degraded runs. A failed or invalid run is evidence about the system. Deleting it to improve the scorecard throws away information.
The broader lesson
Building Waada has changed what I consider a reliable AI system. Reliability doesn't start with choosing a stronger model. It starts with being able to answer, for any output:
Was the source ingested?
Was it retained?
Was it recalled?
Was the right evidence selected?
Did the model interpret it correctly?
Did the external service stay healthy?
Hindsight provides the durable memory layer, and it's a major part of that chain. But memory doesn't remove the other engineering problems. It makes them more visible, and that's a good thing.
The most useful AI systems are the ones where you can tell not only what the model said, but why it had the information it had, and exactly where it can fail.
Waada is a work in progress. If you're building agent memory or evaluating retrieval systems, I'd love to hear how you separate these failure modes in your own work.
Top comments (0)