<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aneesh Tummmala</title>
    <description>The latest articles on DEV Community by Aneesh Tummmala (@aneesh_tummmala_5205fa1af).</description>
    <link>https://dev.to/aneesh_tummmala_5205fa1af</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150701%2Fcf2acfbc-6e56-44db-98db-ac3628468d3e.jpg</url>
      <title>DEV Community: Aneesh Tummmala</title>
      <link>https://dev.to/aneesh_tummmala_5205fa1af</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aneesh_tummmala_5205fa1af"/>
    <language>en</language>
    <item>
      <title>Scaling Agent Context Across Multiple Meetings with Hindsight</title>
      <dc:creator>Aneesh Tummmala</dc:creator>
      <pubDate>Tue, 29 Sep 2026 17:59:46 +0000</pubDate>
      <link>https://dev.to/aneesh_tummmala_5205fa1af/scaling-agent-context-across-multiple-meetings-with-hindsight-2ohd</link>
      <guid>https://dev.to/aneesh_tummmala_5205fa1af/scaling-agent-context-across-multiple-meetings-with-hindsight-2ohd</guid>
      <description>&lt;p&gt;Most LLM-powered tools treat every request as a blank slate. You send a prompt, you get a response, and the model forgets everything the moment the connection closes. That works fine for one-shot tasks. It breaks completely when your application needs to reason across months of accumulated history.&lt;br&gt;
We were building a meeting intelligence system—something that could analyze a new executive meeting invite against every prior decision, open action item, and recorded conflict from the last six months. The naive approach is obvious and wrong: concatenate all your historical text into one giant prompt and send it to the model. You hit the context window ceiling by week three. You run up token costs that kill the unit economics. And the model starts hallucinating about events from two months ago because everything is weighted equally regardless of recency.&lt;/p&gt;
&lt;h2&gt;
  
  
  The real problem is not context size. It is context selection. This is what I spent most of my time building, and this is where Vectorize Hindsight fundamentally changed how I thought about agent memory.
&lt;/h2&gt;

&lt;p&gt;Why Naive Context Accumulation Fails&lt;br&gt;
When you are processing meeting number ten in a six-month history, the naive pipeline looks like this:&lt;br&gt;
Concatenate meetings 1 through 9 into a single string.&lt;br&gt;
Prepend to the current prompt.&lt;br&gt;
Send the whole block to Gemini.&lt;br&gt;
Meeting transcripts in our system average around 1,200 characters each. By meeting ten, you are injecting 10,800 characters of historical context before you even get to the actual prompt. By meeting thirty, you are well over the practical reasoning threshold where the model starts losing track of content in the middle of the context window.&lt;br&gt;
More critically, you are forcing the model to reason over everything with equal weight. The architecture decision from month one is treated with the same importance as yesterday's urgent security policy update. That is not how human memory works, and it is not how useful agent reasoning works either.&lt;br&gt;
What you actually want is semantic retrieval: given the topic of today's meeting, surface only the prior context that is most relevant to it.&lt;/p&gt;
&lt;h2&gt;
  
  
  &lt;img alt="The three-stage pipeline — each stage pushes transcripts into the memory bank and pulls relevant historical context before Gemini reasoning"&gt;
&lt;/h2&gt;

&lt;p&gt;The Memory Pipeline Architecture&lt;br&gt;
We broke the pipeline into three distinct phases that mirror how the system processes meetings over time. Each phase has a different relationship to the memory bank.&lt;br&gt;
Phase 1 — Retain: After every meeting, push the minutes into Hindsight. This is a fire-and-forget &lt;code&gt;POST /memories&lt;/code&gt; call. No embedding logic on our side, no schema to maintain.&lt;br&gt;
Phase 2 — Recall: Before analyzing a new meeting, fire a semantic recall query using the new meeting's subject line and agenda as the search string. Hindsight returns the most contextually relevant prior memories, not the most recent ones.&lt;br&gt;
Phase 3 — Reason: Pass the recalled context plus the new meeting to Gemini. Because the retrieved context is pre-filtered to be semantically relevant, the prompt stays concise and the model's reasoning stays sharp.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// backend/src/routes/auditRoute.ts (simplified)&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;processAudit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;invite&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// Step 1: recall relevant past context using the invite as the query&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;historicalContext&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;invite&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// Step 2: build the prompt with pre-filtered context&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;buildAuditPrompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;historicalContext&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// Step 3: reason over the current meeting against retrieved history&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;insights&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generateInsights&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// Step 4: retain this meeting for future recall&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;insights&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Notice the order: we recall before we reason, and we retain after we reason. This ensures that every meeting's analysis is grounded in relevant history while the new meeting is immediately available to inform future queries.
&lt;/h2&gt;

&lt;p&gt;What "Relevant" Actually Means at Scale&lt;br&gt;
The thing that surprised me about Vectorize agent memory is how well the semantic retrieval handles thematic relevance without any fine-tuning from our side.&lt;br&gt;
In our six-month corporate governance dataset, meeting five is about a staging environment migration test and MDM endpoint security. Meeting ten is about a vendor risk management audit that catches an unauthorized marketing analytics tool — a Shadow IT violation.&lt;br&gt;
When the system processes meeting ten's recall query, Hindsight surfaces meeting five's MDM and procurement policy text, not the database architecture discussion from meeting two. The semantic distance between "vendor risk audit" and "MDM compliance / software procurement policy" is small enough that the retrieval correctly identifies the relevant history.&lt;/p&gt;
&lt;h2&gt;
  
  
  If I had just dumped all prior meetings into the prompt, meeting two's database schema discussion would be there too, adding noise and consuming tokens. With semantic retrieval, the prompt for meeting ten contains approximately 2,400 characters of highly targeted historical context instead of 10,800 characters of indiscriminate history.
&lt;/h2&gt;

&lt;p&gt;Background Injection for Demo Realism&lt;br&gt;
One problem we had to solve was simulating the passage of time in a system that needed to demonstrate the accumulation of institutional memory over months.&lt;br&gt;
We could not make users sit and manually process ten meetings in sequence. We built a &lt;code&gt;POST /api/demo/advance-stage&lt;/code&gt; endpoint that takes a stage number and directly injects a pre-written array of past meeting transcripts into the Hindsight memory bank in a tight async loop. These injections bypass Gemini entirely — they are raw meeting text pushed straight to the &lt;code&gt;/memories&lt;/code&gt; endpoint.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// backend/src/routes/demoRoute.ts&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;stage2Background&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;stage3Background&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;../data/timeLapseData&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nx"&gt;router&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/advance-stage&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;stage&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;meetings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;stage&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;stage2Background&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;stage3Background&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// Inject all background meetings directly into the memory bank in parallel&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;meetings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)));&lt;/span&gt;

  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;injected&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;meetings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;stage&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  By the time a user reaches stage three in the live demo, Hindsight contains the memory of six months of corporate governance decisions. The next recall query draws on that full history. This pattern — direct memory injection, bypassing the reasoning step — is the equivalent of a database seed script for an agent's long-term memory.
&lt;/h2&gt;

&lt;p&gt;Lessons Learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Semantic recall is not a luxury, it is a necessity.
At ten meetings, brute-force context concatenation barely works. At thirty, it collapses. Build semantic retrieval into your architecture from day one, not as a later optimization.&lt;/li&gt;
&lt;li&gt;Retain immediately, recall selectively.
Every meeting should enter the memory bank unconditionally. Recall should be selective and query-driven. These are separate operations with different semantics — keep them that way in your code.&lt;/li&gt;
&lt;li&gt;Bypass the LLM for bulk historical injection.
When seeding a memory bank with historical data, you do not need the reasoning step. Raw text pushed directly to the memory API is orders of magnitude faster and cheaper than routing every document through an LLM first. Save Gemini calls for when you actually need to reason.&lt;/li&gt;
&lt;li&gt;Your recall query quality determines your context quality.
The meeting invite subject line turned out to be an excellent recall query. It is concise, topically precise, and naturally captures the semantic focus of the upcoming discussion. Do not overthink the query construction — start with what is already available in your existing data.&lt;/li&gt;
&lt;li&gt;Token efficiency compounds across a long context window.
Each selective recall that trims 8,000 characters of irrelevant history saves token costs on every single meeting processed going forward. At scale, the economics of semantic retrieval versus brute-force concatenation are not even close.
---
Scaling agent context across a six-month meeting history with Hindsight turned out to be less about building clever retrieval infrastructure and more about respecting what the model is actually good at. Models reason well over focused, relevant context. They struggle with long, indiscriminate dumps of historical text. Hindsight handles the selection problem so you can focus on building the reasoning logic that matters.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
