From Email Export to Agent Memory: The Pipeline I Actually Wanted
The interesting part of an AI application is often not the model call. It is the pipeline that decides what the model gets to remember.
I built Waada as a deal-continuity system, and the engineering problem that kept surfacing was straightforward: how do I turn messy account history into durable, queryable memory without making every downstream feature understand every source format?
The answer was a strict pipeline: normalize first, retain second, recall by purpose, and generate only from bounded evidence.
The architecture starts before the LLM
Waada accepts exported sales history from several sources:
.eml email
Slack day JSON
text, Markdown, and VTT transcripts
audio
CRM JSON
The application is a TypeScript pnpm workspace. packages/core owns the data model, parsers, ingestion, memory adapter, LLM layer, and agent functions. apps/web is a TanStack Start application exposing those capabilities through server functions.
The architecture is:
Uploaded files
|
v
parseFiles()
|
v
Interaction[]
|
v
ingest()
|
+-------------------+
| |
v v
local normalized data Hindsight
|
v
Memory.search()
|
v
bounded evidence
|
v
LLM
I like this shape because it makes the LLM the last stage rather than the center of the application.
Step one: normalize everything
The core model is a Zod-backed Interaction:
Interaction = {
account: string;
sourceId: string;
type: "call" | "email" | "slack" | "meeting" | "note";
date: string;
title: string;
participants: string[];
content: string;
source: "eml" | "slack_export" | "transcript" | "audio" |
"gmail" | "slack_api" | "hubspot" | "meet";
}
Once a file becomes an Interaction, downstream code doesn't care about its original syntax.
That matters because source formats are full of special cases.
Email parsing has headers, quoted replies, HTML fallbacks, and message IDs.
Slack exports have channel/day structure and optional user metadata.
Transcripts may have front matter or require metadata extraction.
Audio needs transcription before it can enter the same pipeline.
I wanted all of those problems to end at the parser boundary.
Step two: make ingestion idempotent
ingest() validates and sorts interactions by date, checks .waada/manifest.json, ensures the account bank, retains new material, and persists normalized interactions for the comparison baseline.
The manifest tracks retained sourceId values.
That gives me:
same import twice
|
v
same sourceId
|
v
manifest match
|
v
skip duplicate retain
The Hindsight document ID provides another layer of stable identity.
This matters more than it initially looks.
Sales exports are likely to be re-exported. Users may upload the same data again. If ingestion is not idempotent, retrieval quality can degrade through duplicated evidence.
The system therefore keeps the import path deterministic wherever possible.
Step three: retain into account-scoped memory
I use Hindsight as the durable memory layer.
The Hindsight adapter is isolated in packages/core/src/memory/, and every account maps to a bank derived from its slug:
const bankId = bankIdFor(account);
For each interaction, the adapter retains:
content
interaction date
context
document ID
string metadata
The Hindsight GitHub repository is the implementation reference, and the Hindsight documentation covers the memory API.
For Waada, Hindsight's job is not to decide what a commitment means. It provides the durable account history from which the agent can retrieve evidence.
That separation is important.
Memory storage and reasoning are different responsibilities.
Step four: recall according to the operation
Once the data is retained, I don't have one generic getAccountHistory() function.
The agent has different retrieval jobs.
The commitment ledger searches for promise-related evidence.
The landmine extractor searches for objections and resolved agreements.
The brief recalls stakeholder and recent-change evidence.
The Ask function uses the user's exact question.
This is the pattern:
agent operation
|
v
purpose-specific query
|
v
Hindsight recall
|
v
MemoryHit[]
|
v
evidence selection
This is where persistent memory differs from a static summary.
A summary has already decided what deserves attention.
A memory system postpones that decision until the question is known.
The commitment ledger is a retrieval pipeline of its own
commitmentLedger() performs multiple promise-oriented queries, deduplicates evidence, chunks it, runs structured extraction, merges duplicates, and sorts the result.
The resulting Commitment contains the deliverable, parties, date, optional due date, status, evidence, and source.
The sorting policy is straightforward:
open before unclear
unclear before delivered
overdue open items first
then upcoming deadlines
then newer/undated items
The important part is that the status is derived.
I don't want a permanently stored open value to become stale when a later interaction says the work was delivered.
The source history remains the evidence.
The landmine pipeline solves the opposite problem
A landmine is a resolved objection that should not be casually reopened.
The landmine function recalls objections, sensitive topics, and accepted agreements. It sends bounded evidence into structured extraction and returns an object containing the topic, what happened, resolution, date, guidance, and source.
That means the agent has two different operational memories:
commitment -> unfinished work
landmine -> settled history
This is a useful distinction for any system where continuity matters.
A handoff isn't only about preserving tasks.
It is also about preserving decisions.
Step five: budget before generation
Recall can return more information than an LLM should see.
Waada therefore uses evidence caps and a shared prompt budget. The configured maximum is 5,000 input tokens, approximated conservatively through character counts.
I intentionally don't call this an exact tokenizer budget.
The implementation uses it as a safety boundary.
The resulting pipeline is:
Hindsight recall
|
v
deduplicate
|
v
select relevant hits
|
v
chunk
|
v
prompt budget
|
v
LLM
This was one of the most useful design constraints in the project.
Without a budget, every retrieval function tends toward “send more context.”
With a budget, every query has to earn its place.
Why the brief is sequential
brief() runs the ledger and landmine stages sequentially. It separately recalls stakeholder and recent-change evidence before generating the final brief.
The reason is not conceptual purity.
It is provider capacity.
The live implementation uses a constrained provider environment, and parallel model calls can collide with a shared rate window. Sequential work makes the behavior more predictable.
It is a good example of a practical engineering trade-off: the theoretically faster architecture is not always the architecture that survives its provider limits.
Ask is the simplest complete path
The Ask operation makes the whole architecture easy to explain.
User:
What changed since July?
System:
question
-> Hindsight recall
-> bounded evidence
-> LLM
-> answer + citations
The implementation returns the recalled contexts and document IDs with the answer.
A successful live evaluation recorded the Q3-to-Q4 go-live change for the Acme scenario.
That is the kind of evidence I want the system to expose: not only a sentence, but the historical contexts that informed it.
The LLM layer has its own boundary
The LLM layer supports multiple providers through the AI SDK and has separate operations for chat, structured extraction, and transcription.
Structured extraction is treated as untrusted.
const parsed = schema.safeParse(modelOutput);
if (parsed.success) {
return parsed.data;
}
const repaired = await repair(modelOutput);
return schema.safeParse(repaired).success
? schema.parse(repaired)
: null;
The model can produce malformed output.
The application should survive that.
That sounds obvious, but it becomes important once generated data starts controlling downstream behavior.
The comparison path helped expose the real value
Waada compares three context sources:
CRM-only
raw chronological summary
memory-aware Waada
The CRM baseline reads only imported CRM data.
The summary baseline reads chronological normalized interactions without memory.
The Waada version can use Hindsight recall.
This isn't a benchmark against CRM vendors. It is a controlled comparison of context availability.
The live evaluation record is deliberately mixed. It contains successful timeline and landmine behaviors as well as structured-output variability, rate-limit pressure, intermittent memory failures, and a run where summary-only scored higher on its checks.
That distinction is important.
The architecture establishes a different context path.
It does not establish universal accuracy.
What I would reuse from this pipeline
End the messy-source problem at the parser boundary
Downstream code should consume one schema.
Make source identity explicit
Stable IDs and manifests are part of retrieval quality, not just ingestion bookkeeping.
Keep memory account-scoped
The account is the natural storage and retrieval boundary.
Preserve timestamps
Temporal questions are common in sales.
Separate retrieval from generation
Hindsight supplies evidence. The LLM interprets bounded evidence.
Validate generated structures
Model output is an external dependency.
Treat budgets as architecture
Prompt limits affect query design, chunking, sequencing, and error handling.
The pipeline is the product
When I look at Waada now, the most interesting component isn't the final brief.
It is the path that creates the brief:
messy exports
-> canonical interactions
-> idempotent ingest
-> account memory
-> purpose-specific recall
-> bounded evidence
-> validated generation
The model is only useful because the pipeline gives it the right history.
That is the engineering lesson I would carry into another memory-heavy application: don't start by asking which model should summarize the data.
Start by deciding what information must survive, how it will be identified, how it will be retrieved, and how much of it the model is allowed to see.
Then call the model.
Top comments (0)