DEV Community

Dharmini
Dharmini

Posted on

I Didn’t Need a Better Sales Summary. I Needed Memory.

A sales handoff has a peculiar failure mode: everyone knows the account has history, but nobody can reliably find the one piece of history that matters right now.

That observation shaped Waada. Instead of compressing an account into one permanent summary, I built a system that normalizes sales history, stores it as account-scoped memory with Hindsight, and retrieves evidence according to the question being asked.
The problem with treating the handoff as a summary
When a salesperson leaves an account, the obvious solution is to create a handoff document.

I understand the appeal. A summary is easy to read, easy to store, and easy to pass around.
The problem is that sales history is not a static collection of facts.
It contains promises, objections, decisions, changes in dates, people involved in those decisions, and conversations that were resolved months ago. Those pieces of information have different importance depending on what the incoming owner is trying to do.

If the new owner asks, “What changed since July?”, the answer should come from timeline evidence.
If they ask, “Did we promise SOC 2 documentation?”, the relevant evidence is a commitment.
If they ask about pricing, the useful information may actually be a resolved objection that should not be reopened.
A single summary has to predict all of those future questions.
I decided not to make it do that.

Waada treats the account history as persistent memory instead.
The system starts by making every source look the same
The repository is a pnpm TypeScript workspace. packages/core contains the schemas, ingestion pipeline, Hindsight adapter, LLM layer, and agent functions. apps/web provides the TanStack Start application.
The first boundary is ingestion.
Waada accepts email exports, Slack day JSON, text/Markdown/VTT transcripts, audio, and CRM data. The individual parsers produce one normalized Interaction shape:

Interaction = {
account: string;
sourceId: string;
type: "call" | "email" | "slack" | "meeting" | "note";
date: string;
title: string;
participants: string[];
content: string;
source: "eml" | "slack_export" | "transcript" | "audio" |
"gmail" | "slack_api" | "hubspot" | "meet";
}

That decision sounds mundane, but it is one of the most important parts of the architecture.
Memory should not need to understand whether a fact came from an email header or a Slack export. Once the source has been parsed, the downstream system needs content, time, participants, source identity, and account identity.
I also use stable source IDs and a manifest so repeat imports can be skipped. The normalized interactions are persisted locally for baseline comparisons, while the same content is retained in Hindsight.

So the ingestion path becomes:
source files
-> parser
-> Interaction[]
-> validation
-> deduplication
-> Hindsight retain
That gives me one predictable boundary before any retrieval or generation starts.
Hindsight became the memory boundary
I kept Hindsight behind a dedicated adapter in
packages/core/src/memory/.
Each account maps to its own bank:
const bankId = bankIdFor(account);
// waada-acme
That account-level separation is deliberate. I don't want every downstream recall operation to depend on remembering a filter correctly. The storage boundary itself should express which account the memory belongs to.
The adapter handles the Hindsight operations used by the application: retain and recall, with reflect exposed as an interface even though the implemented MVP agent functions do not currently use it.
The Hindsight GitHub repository provides the underlying implementation, and the Hindsight documentation describes its memory model and APIs.

The important application-level flow is:
Interaction
|
v
Hindsight retain
|
v
account-scoped memory
|
v
query-specific recall
|
v
bounded evidence
|
v
LLM
That is fundamentally different from:
all history
|
v
one giant summary
|
v
LLM

The first approach lets the question determine the context.
I made the commitment ledger derive state from evidence
The most useful example is the commitment ledger.
A handoff should not just tell the incoming rep what happened. It should tell them what still needs attention.
commitmentLedger() performs several promise-oriented recall queries. It deduplicates the evidence, extracts structured commitments, and sorts them.
Open commitments come first. Within that group, overdue and upcoming deadlines receive priority.

The important design choice is that I do not treat the commitment status as permanent truth.

const hits = await recallPromiseEvidence(account);

const commitments = await extractCommitments(
chunkAndDeduplicate(hits)
);

return sortLedger(commitments);

Suppose someone promised a document on Monday. On Tuesday, another email says the document was delivered.
If I stored status = open as the authoritative state, I would need another synchronization mechanism to update it.
Instead, the ledger reconstructs the state from remembered evidence.
That makes the history the source and the ledger the current interpretation.
For the Acme scenario, the expected behavior is for the outstanding September 2 SOC 2 commitment to appear first in the brief.

That is a much more concrete outcome than “the AI understands the account.”
Then I added something a normal task list does not have
The second concept is a landmine.
A landmine is a resolved objection or sensitive topic that the incoming rep should not casually reopen.
This is different from an open commitment.

An open commitment means:
Someone still owes something.
A landmine means:
Something was already settled; preserve that decision.
The landmine implementation recalls objections, sensitive topics, and accepted agreements. Structured extraction produces the topic, what happened, resolution, date, guidance, and source.
That distinction changes what a handoff means.
A conventional task list preserves unfinished work.
A continuity system also preserves completed decisions.
For example, the Acme scenario expects the resolved pricing issue to appear as a landmine rather than as another unresolved task.
That is exactly the kind of information that gets lost when the entire account is compressed into a generic summary.

The brief is a composition of retrieval jobs
The final brief does not rely on one retrieval query.
brief() builds the commitment ledger and landmines, then recalls stakeholder and recent-change evidence. Those pieces are combined into bounded LLM context.

promise evidence -> commitment ledger
objection evidence -> landmines
stakeholder evidence -> people/context
recent evidence -> timeline changes
|
v
bounded prompt
|
v
brief
This also gives me a place to control prompt size.

The LLM layer uses a 5,000-token input budget and a conservative character approximation. Evidence is bounded and chunked rather than allowing every retrieved document into the prompt.
That was not an optimization I added at the end. It became part of the design.
A memory system can retrieve more information than a model should receive.
Those are different limits.
The Vectorize agent memory guide is useful background for thinking about this distinction: persistent memory is about making information available across interactions, not about stuffing the entire information store into every prompt.

Ask is where the architecture becomes obvious
The Ask page lets the incoming rep ask a question directly.

For example:
What changed since July?
ask() uses the question itself as the recall query.
const hits = await memory.search(account, question);
const evidence = capEvidence(hits);

return llm.chat({
question,
evidence,
});

The answer is generated from the recalled evidence and includes source contexts and document IDs as citations.
In the Acme scenario, a successful live evaluation found the Q3-to-Q4 go-live change for that question.
That behavior is exactly what I wanted.
The user did not need to know which email contained the answer. They supplied the question, and the memory layer supplied the relevant history.
I kept the comparison honest
Waada also has three context modes:
CRM-only
raw-summary-only
memory-aware Waada
The CRM baseline reads only imported CRM fields.

The summary baseline reads normalized interactions chronologically, without Hindsight, and caps the input at 12,000 characters.
The Waada path adds recalled account-specific evidence.
I deliberately kept those paths separate.
Otherwise it would be very easy to build a system where the “AI memory” version gets access to everything and the baseline gets an artificially weak input, then call the result a benchmark.
That is not what the repository establishes.
The evaluation record contains both successful behaviors and failures. Some runs found the expected pricing landmine or timeline change. Other runs documented structured-output variability, rate-limit pressure, prompt-size pressure, and intermittent Hindsight Cloud failures. One recorded run had the summary-only baseline score higher than Waada on its checks.
That does not invalidate the architecture.

It tells me that memory quality and model quality need to be evaluated independently.
The LLM is treated as an unreliable boundary
Structured extraction is validated through Zod.
The implementation attempts structured output, validates it, retries with repair guidance, and falls back to JSON parsing. If the result still does not validate, extraction can return null.
const parsed = schema.safeParse(modelOutput);

if (parsed.success) {
return parsed.data;
}

const repaired = await repair(modelOutput);

return schema.safeParse(repaired).success
? schema.parse(repaired)
: null;

That is a rule I use throughout the system:
Anything coming from an external system gets validated before it becomes application state.
The same principle applies to imported files, server-function inputs, dates, and external service failures.
What I learned
Memory should be queried by purpose
I don't want one retrieval query called “get everything about Acme.”
I want promise evidence, objection evidence, stakeholder evidence, recent-change evidence, or a question-specific query.
Narrow retrieval jobs are easier to reason about and easier to budget.
Account boundaries belong in storage
Per-account Hindsight banks make the intended isolation visible in the architecture.
That is easier to reason about than relying on every caller to remember a filter.
Derived state should remain derived
A commitment ledger is an interpretation of source history, not the source history itself.
That means a new email can change the current view without requiring a separate state synchronization pipeline.
Persistent memory does not solve context limits
Recall can produce useful evidence and still return too much of it.
I still need deduplication, selection, chunking, and prompt budgets.
A failed model response is not necessarily a failed architecture
The evaluation record showed model and service failures alongside successful retrieval behavior.
I found it more useful to preserve those distinctions than to turn the whole system into one “accuracy” number.
The conclusion I keep coming back to
The sales handoff problem looks like a summarization problem until you think about what the next person actually asks.
They don't need to know everything.
They need to know the right thing at the right moment.
That is why I built Waada around persistent, account-scoped memory instead of a single permanent summary.
The LLM is the final consumer of the context.
The more interesting engineering work happens before that: normalization, stable IDs, per-account memory, temporal metadata, retrieval strategy, evidence budgets, validation, retries, and failure handling.
For me, the central lesson is simple:
A summary tells me what happened. Memory gives me a way to find what matters now.

Top comments (0)