Most LLM-powered tools treat every request as a blank slate. You send a prompt, you get a response, and the model forgets everything the moment the connection closes. That works fine for one-shot tasks. It breaks completely when your application needs to reason across months of accumulated history.
We were building a meeting intelligence system—something that could analyze a new executive meeting invite against every prior decision, open action item, and recorded conflict from the last six months. The naive approach is obvious and wrong: concatenate all your historical text into one giant prompt and send it to the model. You hit the context window ceiling by week three. You run up token costs that kill the unit economics. And the model starts hallucinating about events from two months ago because everything is weighted equally regardless of recency.
The real problem is not context size. It is context selection. This is what I spent most of my time building, and this is where Vectorize Hindsight fundamentally changed how I thought about agent memory.
Why Naive Context Accumulation Fails
When you are processing meeting number ten in a six-month history, the naive pipeline looks like this:
Concatenate meetings 1 through 9 into a single string.
Prepend to the current prompt.
Send the whole block to Gemini.
Meeting transcripts in our system average around 1,200 characters each. By meeting ten, you are injecting 10,800 characters of historical context before you even get to the actual prompt. By meeting thirty, you are well over the practical reasoning threshold where the model starts losing track of content in the middle of the context window.
More critically, you are forcing the model to reason over everything with equal weight. The architecture decision from month one is treated with the same importance as yesterday's urgent security policy update. That is not how human memory works, and it is not how useful agent reasoning works either.
What you actually want is semantic retrieval: given the topic of today's meeting, surface only the prior context that is most relevant to it.
The Memory Pipeline Architecture
We broke the pipeline into three distinct phases that mirror how the system processes meetings over time. Each phase has a different relationship to the memory bank.
Phase 1 — Retain: After every meeting, push the minutes into Hindsight. This is a fire-and-forget POST /memories call. No embedding logic on our side, no schema to maintain.
Phase 2 — Recall: Before analyzing a new meeting, fire a semantic recall query using the new meeting's subject line and agenda as the search string. Hindsight returns the most contextually relevant prior memories, not the most recent ones.
Phase 3 — Reason: Pass the recalled context plus the new meeting to Gemini. Because the retrieved context is pre-filtered to be semantically relevant, the prompt stays concise and the model's reasoning stays sharp.
// backend/src/routes/auditRoute.ts (simplified)
export async function processAudit(req: Request, res: Response) {
const { transcript, invite } = req.body;
// Step 1: recall relevant past context using the invite as the query
const historicalContext = await recall(invite);
// Step 2: build the prompt with pre-filtered context
const prompt = buildAuditPrompt(transcript, historicalContext);
// Step 3: reason over the current meeting against retrieved history
const insights = await generateInsights(prompt);
// Step 4: retain this meeting for future recall
await retain(transcript);
res.json(insights);
}
Notice the order: we recall before we reason, and we retain after we reason. This ensures that every meeting's analysis is grounded in relevant history while the new meeting is immediately available to inform future queries.
What "Relevant" Actually Means at Scale
The thing that surprised me about Vectorize agent memory is how well the semantic retrieval handles thematic relevance without any fine-tuning from our side.
In our six-month corporate governance dataset, meeting five is about a staging environment migration test and MDM endpoint security. Meeting ten is about a vendor risk management audit that catches an unauthorized marketing analytics tool — a Shadow IT violation.
When the system processes meeting ten's recall query, Hindsight surfaces meeting five's MDM and procurement policy text, not the database architecture discussion from meeting two. The semantic distance between "vendor risk audit" and "MDM compliance / software procurement policy" is small enough that the retrieval correctly identifies the relevant history.
If I had just dumped all prior meetings into the prompt, meeting two's database schema discussion would be there too, adding noise and consuming tokens. With semantic retrieval, the prompt for meeting ten contains approximately 2,400 characters of highly targeted historical context instead of 10,800 characters of indiscriminate history.
Background Injection for Demo Realism
One problem we had to solve was simulating the passage of time in a system that needed to demonstrate the accumulation of institutional memory over months.
We could not make users sit and manually process ten meetings in sequence. We built a POST /api/demo/advance-stage endpoint that takes a stage number and directly injects a pre-written array of past meeting transcripts into the Hindsight memory bank in a tight async loop. These injections bypass Gemini entirely — they are raw meeting text pushed straight to the /memories endpoint.
// backend/src/routes/demoRoute.ts
import { stage2Background, stage3Background } from '../data/timeLapseData';
router.post('/advance-stage', async (req, res) => {
const { stage } = req.body;
const meetings = stage === 2 ? stage2Background : stage3Background;
// Inject all background meetings directly into the memory bank in parallel
await Promise.all(meetings.map((text) => retain(text)));
res.json({ injected: meetings.length, stage });
});
By the time a user reaches stage three in the live demo, Hindsight contains the memory of six months of corporate governance decisions. The next recall query draws on that full history. This pattern — direct memory injection, bypassing the reasoning step — is the equivalent of a database seed script for an agent's long-term memory.
Lessons Learned
- Semantic recall is not a luxury, it is a necessity. At ten meetings, brute-force context concatenation barely works. At thirty, it collapses. Build semantic retrieval into your architecture from day one, not as a later optimization.
- Retain immediately, recall selectively. Every meeting should enter the memory bank unconditionally. Recall should be selective and query-driven. These are separate operations with different semantics — keep them that way in your code.
- Bypass the LLM for bulk historical injection. When seeding a memory bank with historical data, you do not need the reasoning step. Raw text pushed directly to the memory API is orders of magnitude faster and cheaper than routing every document through an LLM first. Save Gemini calls for when you actually need to reason.
- Your recall query quality determines your context quality. The meeting invite subject line turned out to be an excellent recall query. It is concise, topically precise, and naturally captures the semantic focus of the upcoming discussion. Do not overthink the query construction — start with what is already available in your existing data.
- Token efficiency compounds across a long context window. Each selective recall that trims 8,000 characters of irrelevant history saves token costs on every single meeting processed going forward. At scale, the economics of semantic retrieval versus brute-force concatenation are not even close. --- Scaling agent context across a six-month meeting history with Hindsight turned out to be less about building clever retrieval infrastructure and more about respecting what the model is actually good at. Models reason well over focused, relevant context. They struggle with long, indiscriminate dumps of historical text. Hindsight handles the selection problem so you can focus on building the reasoning logic that matters.
Top comments (0)