The first version of my meeting agent could write a perfectly reasonable meeting brief. It just had no reliable memory of why the next meeting mattered.
That distinction turned out to be the whole problem. A language model can summarize the notes I give it, but a meeting assistant has to do something harder: recover the right pieces of history for one person, at the right time, and then turn those pieces into something I can act on. I built the system around Hindsight because I wanted the memory layer to be a real system boundary instead of a giant pile of prompts, transcript files, and application-side filtering.
The system I built
The application has a deliberately small architecture. A FastAPI service owns the application flow. A thin Hindsight wrapper handles long-term memory. Gemini handles language generation and multimodal transcription. The frontend is plain HTML, CSS, and JavaScript rather than a large client framework.
The interesting part is not the number of components. It is where I draw the boundaries between them.
For each contact, I create a distinct Hindsight memory bank. Meeting transcripts and notes are retained into that bank. When I need a brief, I ask Hindsight for memories relevant to that specific meeting context, then give only those recalled memories to Gemini. User formatting preferences live in a separate memory bank, so a preference like “keep the TL;DR to one bullet” is not mixed into Priya’s account history.
The backend flow is roughly:
Calendar / contact context
|
v
select contact
|
v
Hindsight recall
|
v
relevant contact memory
|
+------> user preference memory
|
v
Gemini brief prompt
|
v
structured meeting brief
The application entry point keeps that flow visible by importing the memory operations directly:
from hindsight_wrapper import recall_for_contact, reflect_for_contact, retain_meeting
from llm import build_brief, transcribe_media, draft_followup_email
I like this separation because it keeps the LLM layer from becoming the database layer. Gemini does not decide how I store or retrieve history. Hindsight does not decide what the final user-facing brief should look like. Each service has one job.
The problem was not summarization
The trap with meeting software is to define the problem as “summarize the meeting.” That is useful, but it is the easy part.
My real problem was continuity.
Imagine four conversations with the same contact. In the first, she describes the operational problem. In the second, she reacts positively to a product feature but rejects a workflow. In the third, she pushes on price and asks a compliance question. In the fourth, she is frustrated because two things promised in the previous conversation still have not happened.
A stateless summarizer can describe meeting four. It cannot reliably tell me that the frustration is connected to commitments made in meeting three unless I give it the right history.
That is where Hindsight became the core of the design rather than an add-on.
I retain each meeting against the contact’s bank:
async def retain_meeting(bank_id: str, content: str, document_id: str):
return await get_client().aretain(
bank_id=bank_id,
content=content,
document_id=document_id
)
The application does not have to invent its own memory schema for every sentence in a transcript. It sends the meeting content to the memory service and later asks for the context it needs.
The corresponding read path is intentionally simple:
recalled = await recall_for_contact(
bank_id,
f"Prep me for a meeting with {contact['name']}"
)
memory_text = "\n".join(item.text for item in recalled)
That query looks almost too small. That is the point. I do not want application logic that finds six transcripts, regexes for “promised,” filters by date, and merges commitments. That becomes a second retrieval system hiding inside the product.
The Hindsight documentation describes the distinction between retaining information, recalling relevant memories, and reflecting over memory. That separation maps well to what I needed: write history as conversations happen, retrieve relevant facts when building a brief, and use deeper reflection when I want patterns rather than raw facts.
Recall for facts, reflection for patterns
One design decision mattered more than I expected: I did not use one memory operation for everything.
The application exposes a memory view that performs both recall and reflection:
facts = await recall_for_contact(
bank_id,
"everything known about this contact"
)
models = await reflect_for_contact(
bank_id,
"what patterns have you noticed about this contact"
)
Those calls serve different purposes.
If I am trying to answer “What did Priya ask about compliance?” I want factual retrieval. If I am trying to answer “What tends to matter to Priya when we talk?” I want synthesis across multiple memories.
Hindsight's own guidance makes the same distinction: recall is for retrieving relevant memories, while reflect is for reasoning across them. In practice, that prevented me from turning every UI request into a heavier reasoning operation.
That also gave me a useful mental model for the rest of the application. I treat memory like a data access layer with different query semantics, not like a generic “AI context” bucket.
The memory boundary is also a hallucination boundary
I was deliberate about what crosses from Hindsight into Gemini.
The brief prompt says:
BRIEF_PROMPT = """You are an executive assistant preparing someone for a meeting with {contact}.
Use ONLY the memory below. Do not invent facts that are not in it.
Memory:
{memory}
That line is doing more work than it looks like.
I wanted the model to have enough freedom to turn history into a useful brief, but not enough freedom to silently fill gaps with plausible business details. So I made the memory retrieval step explicit and then constrained the generation step around that retrieved evidence.
The output is also structured. The model is asked for a fixed JSON shape containing a TL;DR, missed follow-ups, open or resolved follow-ups, contact quirks, and an icebreaker. That makes the result a contract between the LLM layer and the UI instead of a blob of prose that the frontend has to parse heuristically.
There is a subtle but important difference between “the model has memory” and “the model is given memory.” I built the latter.
The model remains stateless between requests. Hindsight owns continuity. That makes the architecture easier to reason about, test, and replace.
For a deeper background on why this separation matters, Vectorize's explanation of agent memory is a useful reference. My implementation follows the same basic idea: memory is an external capability that an agent can write to and query, not something I expect the model context window to magically preserve forever.
The most useful example was a missed commitment
The best behavior in the system is not a clever icebreaker. It is a boring, operationally useful reminder.
In my sample history, Priya Desai is evaluating cold-chain tracking. Early conversations establish the operational problem. Later, she focuses on cost and asks whether the platform is ready for FSMA 204 compliance. We offer an 8 percent loyalty discount contingent on signing by the end of the month and promise to confirm the compliance question.
In the next conversation, the contract still does not contain the discount, and the compliance question is still unanswered.
The current meeting brief is designed to surface exactly that kind of continuity. The prompt distinguishes “promises we previously made but have no record of completing” from ordinary open threads. That is an important product distinction because a missed commitment is not just another bullet point. It changes how I should enter the meeting.
The code reflects that distinction directly:
- missed_follow_ups: array of promises we previously made but have NO record of completing
- follow_ups: array of normal open threads, upcoming actions, or resolved items
I prefer this to a generic “summary of previous meetings” because summaries optimize for compression. Meeting preparation has to optimize for relevance and accountability.
Preferences are memory too
I also wanted the system to remember something other than customer facts: how I want the assistant to work.
That is why preferences use a separate Hindsight bank:
await retain_meeting(
bank_id="user_preferences",
content=req.preference,
document_id=str(uuid.uuid4())
)
Before generating a brief, I recall the formatting preference for the current user and inject it into the prompt alongside the contact memory.
This changed how I think about personalization. I do not need hard-coded settings for every possible request. Some preferences are naturally expressed in language, and a memory system is a reasonable place to preserve them.
The frontend makes that explicit by asking, “How would you like future briefs formatted?” rather than presenting a settings screen with an endless collection of toggles.
Meetings also become inputs to memory
The other half of long-term memory is writing good inputs.
The application can accept a pasted transcript or meeting notes, and it can capture a meeting recording, send it through Gemini for transcription, and retain the resulting transcript against the correct contact bank.
The browser-to-memory path is deliberately straightforward:
const res = await fetch(/api/ingest/${currentBank}, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ text: input.value })
});
On the backend, ingestion writes to Hindsight and also keeps a transcript log for the application UI.
That second storage path is intentional. I do not treat the memory index as a replacement for an auditable source record. The transcript log answers “what raw text did we receive?” Hindsight answers “what relevant knowledge can we retrieve from the history?” Those are different requirements.
In the production version, I also keep calendar integration behind the application API rather than coupling the UI directly to a provider. That lets the contact list represent the meetings I actually need to prepare for while Hindsight remains responsible for the historical context.
What I learned about building agent memory
- Start with the retrieval question, not the storage format
I initially thought more memory meant better memory. It does not.
The important question is: what do I need to know immediately before this meeting? That produces a better retrieval query and a better memory boundary than simply dumping every transcript into the model.
- Keep raw records and derived memory separate
I want the original transcript available for auditability and replay. I also want a memory layer optimized for retrieval. Those two representations can coexist without pretending they are the same thing.
- Use different operations for lookup and synthesis
Recall is appropriate when I need evidence. Reflection is appropriate when I need a synthesized pattern. Using the heavier operation everywhere is unnecessary, while using raw retrieval everywhere throws reasoning work back onto the application.
- Per-contact isolation matters
Giving each contact its own memory bank makes the security and relevance boundary obvious. Priya's history should not become context for Marcus just because both happen to discuss similar products.
- Memory needs product semantics, not just infrastructure
The most valuable fields in the brief are not “summary” and “topics.” They are things like missed commitments, open threads, communication quirks, and a grounded opening question. Those concepts tell the memory query what matters.
The bigger lesson
I used to think the hard part of an AI meeting assistant was getting the model to produce a good summary. After building this, I think the harder and more reusable engineering problem is deciding what the model should be allowed to remember and what it should be allowed to see for a particular task.
Hindsight gave me a clean place to put that decision.
The result is not a model with an enormous context window pretending to have a long-term relationship with a user. It is an application with an explicit memory layer, explicit retrieval queries, explicit generation constraints, and a source record I can inspect when something looks wrong.
That architecture is less magical, which is exactly why I like it.
A meeting agent does not need to remember everything. It needs to remember the right things when they become relevant.

Top comments (0)