<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ranga sai</title>
    <description>The latest articles on DEV Community by Ranga sai (@ranga_sai).</description>
    <link>https://dev.to/ranga_sai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149147%2Fdedfaadf-1b76-42aa-acc3-25f1bd70c132.png</url>
      <title>DEV Community: Ranga sai</title>
      <link>https://dev.to/ranga_sai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ranga_sai"/>
    <language>en</language>
    <item>
      <title>Context Has a Cost: What Building a Memory-Aware Agent Taught Me</title>
      <dc:creator>Ranga sai</dc:creator>
      <pubDate>Tue, 29 Sep 2026 11:29:16 +0000</pubDate>
      <link>https://dev.to/ranga_sai/context-has-a-cost-what-building-a-memory-aware-agent-taught-me-578j</link>
      <guid>https://dev.to/ranga_sai/context-has-a-cost-what-building-a-memory-aware-agent-taught-me-578j</guid>
      <description>&lt;p&gt;The easiest way to make an LLM application look smart is to give the model more context.&lt;br&gt;
The harder problem is deciding what context deserves to be there.&lt;br&gt;
While building Waada, I found that persistent memory solved only half of the problem. Hindsight could give me a durable account history and retrieve relevant evidence, but I still had to control what reached the model.&lt;br&gt;
That turned prompt budgeting into an architectural concern rather than a tuning detail.&lt;br&gt;
More history is not the same as more useful context&lt;br&gt;
Waada handles sales handovers.&lt;br&gt;
The source material can include emails, Slack conversations, call and meeting transcripts, audio, and CRM information.&lt;br&gt;
A naïve architecture would concatenate the account history and send it to the LLM.&lt;br&gt;
That creates several problems immediately:&lt;br&gt;
duplicated messages&lt;br&gt;
irrelevant conversations&lt;br&gt;
token limits&lt;br&gt;
provider rate limits&lt;br&gt;
expensive prompts&lt;br&gt;
harder-to-explain answers&lt;br&gt;
I wanted a different flow:&lt;br&gt;
large account history&lt;br&gt;
        |&lt;br&gt;
        v&lt;br&gt;
persistent memory&lt;br&gt;
        |&lt;br&gt;
        v&lt;br&gt;
question-specific recall&lt;br&gt;
        |&lt;br&gt;
        v&lt;br&gt;
bounded evidence&lt;br&gt;
        |&lt;br&gt;
        v&lt;br&gt;
LLM&lt;br&gt;
The important word is bounded.&lt;br&gt;
The source data gets normalized first&lt;br&gt;
The repository is a TypeScript pnpm workspace.&lt;br&gt;
packages/core contains schemas, parsers, storage, Hindsight integration, the LLM layer, and agent functions. apps/web is the TanStack Start application.&lt;br&gt;
Every imported source becomes an Interaction:&lt;br&gt;
Interaction = {&lt;br&gt;
  account: string;&lt;br&gt;
  sourceId: string;&lt;br&gt;
  type: "call" | "email" | "slack" | "meeting" | "note";&lt;br&gt;
  date: string;&lt;br&gt;
  title: string;&lt;br&gt;
  participants: string[];&lt;br&gt;
  content: string;&lt;br&gt;
  source: "...";&lt;br&gt;
}&lt;br&gt;
This is important for context control because retrieval can't be intelligently bounded if every source uses a different representation.&lt;br&gt;
The parser layer handles source-specific complexity.&lt;br&gt;
The memory layer gets one canonical structure.&lt;br&gt;
Hindsight gives me a durable source for retrieval&lt;br&gt;
The Hindsight integration is isolated in packages/core/src/memory/.&lt;br&gt;
Each account gets a bank based on its slug:&lt;br&gt;
const bankId = bankIdFor(account);&lt;br&gt;
The adapter retains the interaction content, date, context, document ID, and metadata.&lt;br&gt;
The Hindsight GitHub repository and Hindsight documentation cover the memory layer itself.&lt;br&gt;
For Waada, the important distinction is that I don't treat Hindsight as a second LLM.&lt;br&gt;
It is the memory layer.&lt;br&gt;
The application asks it for evidence. The LLM reasons over the evidence.&lt;br&gt;
That separation lets me control the size of the final prompt.&lt;br&gt;
The Vectorize agent memory guide is a useful way to frame the same distinction: memory gives an agent access to information across interactions; it doesn't mean every stored item belongs in every context window.&lt;br&gt;
The commitment ledger became my first context filter&lt;br&gt;
The commitment ledger is a good example of why retrieval has to be task-specific.&lt;br&gt;
Instead of recalling everything, commitmentLedger() performs promise-oriented searches.&lt;br&gt;
The evidence is deduplicated and chunked before structured extraction.&lt;br&gt;
const hits = await recallPromiseEvidence(account);&lt;/p&gt;

&lt;p&gt;const commitments = await extractCommitments(&lt;br&gt;
  chunkAndDeduplicate(hits)&lt;br&gt;
);&lt;/p&gt;

&lt;p&gt;return sortLedger(commitments);&lt;br&gt;
That gives me a much smaller reasoning problem.&lt;br&gt;
The model doesn't need the entire account to answer:&lt;br&gt;
Which promises are still open?&lt;br&gt;
It needs evidence about promises.&lt;br&gt;
That is the first level of context budgeting: ask a narrower question.&lt;br&gt;
Landmines are another context filter&lt;br&gt;
Landmines search for a different class of history: objections, sensitive subjects, and accepted agreements.&lt;br&gt;
The system wants to answer:&lt;br&gt;
What should the new owner know before reopening a difficult topic?&lt;br&gt;
That query is different from the commitment query.&lt;br&gt;
This is why I avoided a single universal “retrieve account context” function.&lt;br&gt;
Different operations need different evidence.&lt;br&gt;
The resulting structure is:&lt;br&gt;
account memory&lt;br&gt;
      |&lt;br&gt;
      +---- promise queries ------&amp;gt; ledger&lt;br&gt;
      |&lt;br&gt;
      +---- objection queries ----&amp;gt; landmines&lt;br&gt;
      |&lt;br&gt;
      +---- question query -------&amp;gt; answer&lt;br&gt;
      |&lt;br&gt;
      +---- timeline queries -----&amp;gt; recent changes&lt;br&gt;
The memory store is shared.&lt;br&gt;
The retrieval intent is not.&lt;br&gt;
The 5,000-token budget changed how I designed the agent&lt;br&gt;
The LLM layer uses a 5,000-token input budget, represented conservatively as approximately 12,500 characters.&lt;br&gt;
It is not a precise tokenizer measurement.&lt;br&gt;
It is a safety boundary.&lt;br&gt;
That means every retrieval path eventually has to make a choice:&lt;br&gt;
Hindsight returns N hits&lt;br&gt;
        |&lt;br&gt;
        v&lt;br&gt;
Which hits matter?&lt;br&gt;
        |&lt;br&gt;
        v&lt;br&gt;
How much text from each?&lt;br&gt;
        |&lt;br&gt;
        v&lt;br&gt;
What fits the budget?&lt;br&gt;
This is why the agent code contains evidence chunking, deduplication, and logging around dropped hits.&lt;br&gt;
I would rather drop low-value context deliberately than let the prompt grow until the provider rejects it.&lt;br&gt;
Sequential calls became a feature, not just a compromise&lt;br&gt;
The brief builds several components:&lt;br&gt;
commitment ledger&lt;br&gt;
landmines&lt;br&gt;
stakeholder evidence&lt;br&gt;
recent changes&lt;br&gt;
It would be natural to run every LLM operation concurrently.&lt;br&gt;
The live implementation does not.&lt;br&gt;
The brief runs the ledger and landmine work sequentially to fit the provider's shared rate window.&lt;br&gt;
At first that felt like giving up performance.&lt;br&gt;
But if a provider limits tokens per minute, parallelism can simply concentrate the load into a failure spike.&lt;br&gt;
Sequential execution makes the dependency visible and predictable.&lt;br&gt;
This is one of the less glamorous lessons from the project:&lt;br&gt;
Provider limits are part of application architecture.&lt;br&gt;
They affect control flow.&lt;br&gt;
Ask demonstrates the smallest useful context&lt;br&gt;
The Ask operation is almost a pure retrieval experiment.&lt;br&gt;
A user asks:&lt;br&gt;
What changed since July?&lt;br&gt;
The application uses that exact question for recall, caps the returned evidence, and asks the model to answer only from the recalled contexts.&lt;br&gt;
const hits = await memory.search(account, question);&lt;br&gt;
const evidence = capEvidence(hits);&lt;/p&gt;

&lt;p&gt;return llm.chat({&lt;br&gt;
  question,&lt;br&gt;
  evidence,&lt;br&gt;
});&lt;br&gt;
The response includes source contexts and document IDs.&lt;br&gt;
A successful live evaluation recorded the expected Q3-to-Q4 go-live change for the Acme scenario.&lt;br&gt;
The important architectural property is not the sentence itself.&lt;br&gt;
It is that the question determines the evidence.&lt;br&gt;
The system didn't have to include every July-to-September message in every prompt.&lt;br&gt;
The comparison baseline exposed the context difference&lt;br&gt;
Waada includes a comparison between:&lt;br&gt;
CRM-only&lt;br&gt;
raw-summary-only&lt;br&gt;
memory-aware&lt;br&gt;
The CRM path uses imported CRM fields.&lt;br&gt;
The summary path reads chronological normalized interactions and caps them at 12,000 characters.&lt;br&gt;
The memory-aware path uses Hindsight recall.&lt;br&gt;
This gives me three different context strategies.&lt;br&gt;
It is not a benchmark of commercial CRM systems, and the repository does not establish a universal accuracy advantage.&lt;br&gt;
The live evaluation record actually contains mixed results. It documents successful retrieval behaviors alongside structured-output variability, rate-limit pressure, prompt-size problems, and intermittent Hindsight failures. One run recorded the summary-only baseline scoring higher on its checks.&lt;br&gt;
That is exactly why I think the comparison is useful.&lt;br&gt;
It shows that context construction is a variable worth testing.&lt;br&gt;
Structured extraction adds another boundary&lt;br&gt;
Once evidence has been selected, the LLM extracts commitments and landmines.&lt;br&gt;
I don't trust that output blindly.&lt;br&gt;
const parsed = schema.safeParse(modelOutput);&lt;/p&gt;

&lt;p&gt;if (parsed.success) {&lt;br&gt;
  return parsed.data;&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;const repaired = await repair(modelOutput);&lt;/p&gt;

&lt;p&gt;return schema.safeParse(repaired).success&lt;br&gt;
  ? schema.parse(repaired)&lt;br&gt;
  : null;&lt;br&gt;
The system tries structured output, validates it with Zod, retries with repair guidance, then falls back to JSON parsing.&lt;br&gt;
Malformed results can safely fail instead of becoming application state.&lt;br&gt;
This matters because context budgeting is pointless if the last stage can turn bad output into false commitments.&lt;br&gt;
I learned to think of context as a pipeline&lt;br&gt;
The useful mental model became:&lt;br&gt;
retain&lt;br&gt;
  -&amp;gt; recall&lt;br&gt;
  -&amp;gt; deduplicate&lt;br&gt;
  -&amp;gt; select&lt;br&gt;
  -&amp;gt; chunk&lt;br&gt;
  -&amp;gt; budget&lt;br&gt;
  -&amp;gt; generate&lt;br&gt;
  -&amp;gt; validate&lt;br&gt;
Each stage has a different job.&lt;br&gt;
Hindsight handles durable memory and recall.&lt;br&gt;
The agent decides what kind of evidence it needs.&lt;br&gt;
The evidence layer controls quantity.&lt;br&gt;
The LLM generates.&lt;br&gt;
Zod validates.&lt;br&gt;
Once I thought about the application this way, several design decisions became easier.&lt;br&gt;
Five rules I would reuse&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Retrieval should have a purpose
Don't retrieve “the account.”
Retrieve promises, objections, changes, stakeholders, or the answer to a specific question.&lt;/li&gt;
&lt;li&gt;Budget after retrieval, not before it
First find candidate evidence.
Then decide what can fit.&lt;/li&gt;
&lt;li&gt;Preserve timestamps
A bounded context without dates can still be misleading.&lt;/li&gt;
&lt;li&gt;Don't let the prompt become your database
The memory store should remain the source of historical evidence.
The prompt is a temporary reasoning window.&lt;/li&gt;
&lt;li&gt;Test degraded conditions
Rate limits and external memory failures are part of the real system.
They should appear in evaluation records instead of disappearing behind successful-looking UI.
The conclusion
The biggest lesson from Waada wasn't a new prompting trick.
It was learning to treat context as a scarce engineering resource.
Hindsight gave me a durable account memory.
The retrieval layer gave me task-specific evidence.
The budget prevented that evidence from turning into another giant prompt.
And validation kept model output from becoming accidental truth.
The model is still important.
But the quality of the model call is constrained by the quality and quantity of the history I put in front of it.
In a memory-aware application, context isn't just input.
Context is architecture.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
