I started Waada with a simple observation: the most expensive piece of account history is often not a major decision. It is a small promise buried somewhere in the record.
A date someone committed to. A document someone said they would send. A pricing issue that was already resolved. A change in the go-live plan that never made it into the CRM.
I built the system around recovering that context without asking the incoming rep to reread an entire account.
Why promises disappear
Sales systems are optimized around records.
People are not.
An account might have CRM fields, dozens of emails, Slack conversations, meeting transcripts, and call notes. The CRM might say that an opportunity is in a particular stage. The messages contain the details that explain why.
That creates a strange information problem.
The data exists.
The problem is retrieval.
A new account owner doesn't usually ask:
Give me a complete summary of everything that happened.
They ask:
What are we still waiting on?
Or:
Did we already resolve pricing?
Or:
What changed after the July discussion?
Those questions require different pieces of history.
That is the problem Waada tries to solve.
The pipeline is intentionally explicit
The repository is a TypeScript pnpm workspace with packages/core and apps/web.
The core package owns the data model, parsers, ingestion, Hindsight adapter, LLM layer, and agent functions. The web application exposes import, briefing, commitments, Ask, comparison, settings, and pipeline views.
The ingestion path is:
files
-> parseFiles()
-> Interaction[]
-> ingest()
-> local persistence
-> Hindsight retain()
The normalized model is the key:
Interaction = {
account: string;
sourceId: string;
type: "call" | "email" | "slack" | "meeting" | "note";
date: string;
title: string;
participants: string[];
content: string;
source: "...";
}
I don't want the commitment logic to care whether the promise came from .eml, Slack JSON, or a transcript.
That knowledge belongs in the parser.
Once everything becomes an Interaction, the rest of the application can operate on one timeline.
I used Hindsight because the history needed to survive the current prompt
The memory layer is isolated under packages/core/src/memory/.
Each account gets a Hindsight bank derived from its slug:
const bankId = bankIdFor(account);
An imported interaction is retained with its content, date, context, document ID, and metadata.
The Hindsight adapter handles retain, recall, retries, and mapping external results into Waada's MemoryHit type.
The Hindsight GitHub repository and Hindsight documentation describe the underlying memory system.
For this application, the important distinction is between retention and retrieval.
Retention says:
This interaction belongs to this account's history.
Recall says:
Given what I'm trying to understand now, which pieces of that history should I see?
That second operation is where the commitment ledger starts to become interesting.
I made “open commitment” a retrieval problem
A naïve implementation would store commitments as they are discovered:
commitment -> status=open
That looks convenient.
It is also dangerous.
The source history may later contain the evidence that the commitment was fulfilled.
So I chose to derive commitment state from recalled evidence.
commitmentLedger() searches for promise-oriented evidence, deduplicates it, extracts structured commitments, and sorts them.
The conceptual flow is:
const evidence = await recallPromiseEvidence(account);
const uniqueEvidence = chunkAndDeduplicate(evidence);
const commitments = await extractCommitments(uniqueEvidence);
return sortLedger(commitments);
The sorting logic puts open items ahead of unclear and delivered items. Open items are then prioritized around deadlines.
That gives the incoming owner a practical view:
OPEN
overdue promise
upcoming promise
undated promise
UNCLEAR
...
DELIVERED
...
The important thing is that the list is reconstructed from history.
If the source history changes, the interpretation can change.
The September 2 promise is more useful than a generic summary
The Acme scenario gives the system a specific expected behavior: an outstanding September 2 SOC 2 commitment should lead the brief.
That is a useful test because it forces the system to preserve three things:
what was promised,
who was involved,
when it was due.
A generic summary could mention the SOC 2 discussion without making the outstanding obligation obvious.
The ledger turns it into an operational item.
This is why I think the phrase “AI summary” undersells what the system is doing.
The useful output is not simply shorter text.
It is structured interpretation of remembered evidence.
The opposite problem is the resolved objection
The second important object is a landmine.
I wanted to preserve decisions, not just unfinished work.
A resolved pricing objection can be more important than an open task if the incoming rep is about to reopen it without knowing what was already negotiated.
The landmine flow recalls objections, sensitive topics, and accepted agreements. The LLM extracts the topic, previous discussion, resolution, date, guidance, and source.
The result is a warning that says, in effect:
This topic has history. Read it before you reopen it.
The Acme scenario expects the resolved pricing issue to appear here.
That gives the handoff two kinds of continuity:
commitment -> remember what still needs doing
landmine -> remember what was already settled
I found that distinction more useful than trying to make everything fit into a generic “important facts” section.
The brief combines several memory questions
brief() runs the commitment ledger and landmine extraction, then recalls stakeholder and recent-change evidence before sending bounded context to the LLM.
The architecture is closer to a small query planner than to a single summarization call:
account memory
|
+---------------+---------------+
| | |
v v v
promises objections timeline
| | |
v v v
ledger landmines recent changes
\ | /
+--------------+--------------+
|
v
brief
This is also where Hindsight's role becomes concrete. The memory system is not responsible for writing the final brief. It supplies relevant historical evidence.
The LLM is responsible for turning that evidence into a readable result.
I kept those responsibilities separate.
Question answering is retrieval with a user-supplied query
The Ask path makes the design even simpler.
Suppose the rep asks:
What changed since July?
The system uses that question as the retrieval query, bounds the recalled excerpts, and gives the resulting evidence to the model.
const hits = await memory.search(account, question);
const evidence = capEvidence(hits);
return llm.chat({
question,
evidence,
});
The answer includes source contexts and document IDs.
A successful live evaluation recorded the expected Q3-to-Q4 go-live change for the Acme question.
That interaction is important because it demonstrates why I didn't want a single static summary.
The question itself changes the required context.
I had to make memory cheaper than “read everything”
Persistent memory creates a second problem: retrieval can return more evidence than the model should consume.
The LLM layer therefore uses a shared prompt budget of 5,000 tokens, represented through a conservative character limit. Evidence is bounded and chunked.
The practical flow becomes:
recall
-> deduplicate
-> select
-> chunk
-> budget
-> generate
I also made the brief legs sequential because the provider environment has shared rate limits.
That was painful at first. Parallel requests look cleaner in code.
But if several LLM calls share a tight token-per-minute limit, “more parallel” can simply mean “hit the limit faster.”
The implementation chooses controlled sequencing where necessary.
The LLM cannot be trusted with application state
I treat structured model output as external input.
The extraction layer attempts structured output, validates through Zod, retries with repair guidance, then attempts plain JSON parsing. Invalid results can safely become null.
const parsed = schema.safeParse(modelOutput);
if (parsed.success) {
return parsed.data;
}
const repaired = await repair(modelOutput);
return schema.safeParse(repaired).success
? schema.parse(repaired)
: null;
That matters because a model-generated commitment object eventually influences the user interface.
I don't want malformed JSON becoming a fake promise.
The same boundary discipline appears in the ingestion layer and server functions.
I kept comparison paths independent
Waada has a comparison view with CRM-only, raw-summary-only, and memory-aware outputs.
The CRM baseline reads only .waada/crm/.json.
The summary baseline reads normalized interactions chronologically and caps them at 12,000 characters. It does not call Hindsight.
The memory-aware path can add account-specific recalled evidence.
I intentionally made the baselines inspectable instead of presenting an opaque “AI versus old system” claim.
The evaluation record also contains degraded runs: structured-output variability, provider rate-limit pressure, intermittent Hindsight Cloud failures, and a run where the summary-only baseline scored higher on its checks.
That is exactly why I don't treat the comparison as a benchmark.
It is a context comparison.
The real lesson was about information shape
Building Waada changed how I think about sales history.
The problem isn't that there isn't enough data.
There is usually too much.
The hard part is shaping that history around the question the next person needs answered.
That led to several rules I would reuse elsewhere:
Normalize first
Every source should become one canonical interaction model before memory or reasoning starts.
Make memory boundaries explicit
An account should map to its own memory bank.
Preserve time
Sales questions are frequently temporal. “What changed?” is impossible to answer well without trustworthy dates.
Keep source history authoritative
Derived states such as commitment status should be reconstructable from evidence.
Budget retrieval
More memory is not automatically better context.
Treat models as untrusted
Validate every structured result before using it.
The architecture I would keep
The final shape is straightforward:
TanStack Start
|
v
server functions
|
v
@waada/core
| | |
v v v
ingest memory LLM
|
v
Hindsight
That simplicity is intentional.
The web layer doesn't know how Hindsight works.
The agent layer doesn't know how email parsing works.
The memory adapter doesn't know how the brief is rendered.
Each layer has one job.
And the result is a system that can preserve something a normal summary struggles with: not just what happened, but enough of what happened to answer the next question correctly.
That is the reason I built it this way.
The hardest part of a handoff isn't writing the summary.
It's finding the promise.
Top comments (0)