DEV Community

Ranga sai
Ranga sai

Posted on

Context Has a Cost: What Building a Memory-Aware Agent Taught Me

The easiest way to make an LLM application look smart is to give the model more context.
The harder problem is deciding what context deserves to be there.
While building Waada, I found that persistent memory solved only half of the problem. Hindsight could give me a durable account history and retrieve relevant evidence, but I still had to control what reached the model.
That turned prompt budgeting into an architectural concern rather than a tuning detail.
More history is not the same as more useful context
Waada handles sales handovers.
The source material can include emails, Slack conversations, call and meeting transcripts, audio, and CRM information.
A naïve architecture would concatenate the account history and send it to the LLM.
That creates several problems immediately:
duplicated messages
irrelevant conversations
token limits
provider rate limits
expensive prompts
harder-to-explain answers
I wanted a different flow:
large account history
|
v
persistent memory
|
v
question-specific recall
|
v
bounded evidence
|
v
LLM
The important word is bounded.
The source data gets normalized first
The repository is a TypeScript pnpm workspace.
packages/core contains schemas, parsers, storage, Hindsight integration, the LLM layer, and agent functions. apps/web is the TanStack Start application.
Every imported source becomes an Interaction:
Interaction = {
account: string;
sourceId: string;
type: "call" | "email" | "slack" | "meeting" | "note";
date: string;
title: string;
participants: string[];
content: string;
source: "...";
}
This is important for context control because retrieval can't be intelligently bounded if every source uses a different representation.
The parser layer handles source-specific complexity.
The memory layer gets one canonical structure.
Hindsight gives me a durable source for retrieval
The Hindsight integration is isolated in packages/core/src/memory/.
Each account gets a bank based on its slug:
const bankId = bankIdFor(account);
The adapter retains the interaction content, date, context, document ID, and metadata.
The Hindsight GitHub repository and Hindsight documentation cover the memory layer itself.
For Waada, the important distinction is that I don't treat Hindsight as a second LLM.
It is the memory layer.
The application asks it for evidence. The LLM reasons over the evidence.
That separation lets me control the size of the final prompt.
The Vectorize agent memory guide is a useful way to frame the same distinction: memory gives an agent access to information across interactions; it doesn't mean every stored item belongs in every context window.
The commitment ledger became my first context filter
The commitment ledger is a good example of why retrieval has to be task-specific.
Instead of recalling everything, commitmentLedger() performs promise-oriented searches.
The evidence is deduplicated and chunked before structured extraction.
const hits = await recallPromiseEvidence(account);

const commitments = await extractCommitments(
chunkAndDeduplicate(hits)
);

return sortLedger(commitments);
That gives me a much smaller reasoning problem.
The model doesn't need the entire account to answer:
Which promises are still open?
It needs evidence about promises.
That is the first level of context budgeting: ask a narrower question.
Landmines are another context filter
Landmines search for a different class of history: objections, sensitive subjects, and accepted agreements.
The system wants to answer:
What should the new owner know before reopening a difficult topic?
That query is different from the commitment query.
This is why I avoided a single universal “retrieve account context” function.
Different operations need different evidence.
The resulting structure is:
account memory
|
+---- promise queries ------> ledger
|
+---- objection queries ----> landmines
|
+---- question query -------> answer
|
+---- timeline queries -----> recent changes
The memory store is shared.
The retrieval intent is not.
The 5,000-token budget changed how I designed the agent
The LLM layer uses a 5,000-token input budget, represented conservatively as approximately 12,500 characters.
It is not a precise tokenizer measurement.
It is a safety boundary.
That means every retrieval path eventually has to make a choice:
Hindsight returns N hits
|
v
Which hits matter?
|
v
How much text from each?
|
v
What fits the budget?
This is why the agent code contains evidence chunking, deduplication, and logging around dropped hits.
I would rather drop low-value context deliberately than let the prompt grow until the provider rejects it.
Sequential calls became a feature, not just a compromise
The brief builds several components:
commitment ledger
landmines
stakeholder evidence
recent changes
It would be natural to run every LLM operation concurrently.
The live implementation does not.
The brief runs the ledger and landmine work sequentially to fit the provider's shared rate window.
At first that felt like giving up performance.
But if a provider limits tokens per minute, parallelism can simply concentrate the load into a failure spike.
Sequential execution makes the dependency visible and predictable.
This is one of the less glamorous lessons from the project:
Provider limits are part of application architecture.
They affect control flow.
Ask demonstrates the smallest useful context
The Ask operation is almost a pure retrieval experiment.
A user asks:
What changed since July?
The application uses that exact question for recall, caps the returned evidence, and asks the model to answer only from the recalled contexts.
const hits = await memory.search(account, question);
const evidence = capEvidence(hits);

return llm.chat({
question,
evidence,
});
The response includes source contexts and document IDs.
A successful live evaluation recorded the expected Q3-to-Q4 go-live change for the Acme scenario.
The important architectural property is not the sentence itself.
It is that the question determines the evidence.
The system didn't have to include every July-to-September message in every prompt.
The comparison baseline exposed the context difference
Waada includes a comparison between:
CRM-only
raw-summary-only
memory-aware
The CRM path uses imported CRM fields.
The summary path reads chronological normalized interactions and caps them at 12,000 characters.
The memory-aware path uses Hindsight recall.
This gives me three different context strategies.
It is not a benchmark of commercial CRM systems, and the repository does not establish a universal accuracy advantage.
The live evaluation record actually contains mixed results. It documents successful retrieval behaviors alongside structured-output variability, rate-limit pressure, prompt-size problems, and intermittent Hindsight failures. One run recorded the summary-only baseline scoring higher on its checks.
That is exactly why I think the comparison is useful.
It shows that context construction is a variable worth testing.
Structured extraction adds another boundary
Once evidence has been selected, the LLM extracts commitments and landmines.
I don't trust that output blindly.
const parsed = schema.safeParse(modelOutput);

if (parsed.success) {
return parsed.data;
}

const repaired = await repair(modelOutput);

return schema.safeParse(repaired).success
? schema.parse(repaired)
: null;
The system tries structured output, validates it with Zod, retries with repair guidance, then falls back to JSON parsing.
Malformed results can safely fail instead of becoming application state.
This matters because context budgeting is pointless if the last stage can turn bad output into false commitments.
I learned to think of context as a pipeline
The useful mental model became:
retain
-> recall
-> deduplicate
-> select
-> chunk
-> budget
-> generate
-> validate
Each stage has a different job.
Hindsight handles durable memory and recall.
The agent decides what kind of evidence it needs.
The evidence layer controls quantity.
The LLM generates.
Zod validates.
Once I thought about the application this way, several design decisions became easier.
Five rules I would reuse

  1. Retrieval should have a purpose Don't retrieve “the account.” Retrieve promises, objections, changes, stakeholders, or the answer to a specific question.
  2. Budget after retrieval, not before it First find candidate evidence. Then decide what can fit.
  3. Preserve timestamps A bounded context without dates can still be misleading.
  4. Don't let the prompt become your database The memory store should remain the source of historical evidence. The prompt is a temporary reasoning window.
  5. Test degraded conditions Rate limits and external memory failures are part of the real system. They should appear in evaluation records instead of disappearing behind successful-looking UI. The conclusion The biggest lesson from Waada wasn't a new prompting trick. It was learning to treat context as a scarce engineering resource. Hindsight gave me a durable account memory. The retrieval layer gave me task-specific evidence. The budget prevented that evidence from turning into another giant prompt. And validation kept model output from becoming accidental truth. The model is still important. But the quality of the model call is constrained by the quality and quantity of the history I put in front of it. In a memory-aware application, context isn't just input. Context is architecture.

Top comments (0)