My first attempt at agent memory was one giant text blob per brand. It looked simple, but recall returned a mixture of facts, old observations, and unrelated details. The problem was not only retrieval. I had not decided what each memory actually represented.
The fix was to classify memories before storing them.
I built the memory layer for a content-planning agent using Hindsight. The agent works with knowledge collected from brand information, content performance, experiments, audience feedback, and strategic observations. Instead of treating all of that as one large document, I separated it into five memory types and built a small service around Hindsight.
The result was a memory layer that could retain specific facts, recall relevant evidence, and reflect on that evidence to produce a higher-level view of a brand.
The problem: a brand knows more than its prompt
A content-planning agent needs more than the current conversation.
A marketing team can accumulate useful knowledge across spreadsheets, analytics exports, old discussions, published posts, experiments, and audience feedback. If that information is repeatedly pasted into prompts, the system becomes difficult to maintain and the agent still lacks a structured way to use what it learned previously.
I wanted the agent to answer questions such as:
- What is this brand's communication style?
- Which topics consistently perform well?
- What has the audience asked for?
- Which strategies have failed?
- Which areas of content are under-covered?
I owned the memory layer that answers those questions. I exposed it through hindsightService.ts, which is imported by both the API and the agent.
The important design decision was to give different kinds of information different meanings before sending them to Hindsight.
Five memory types
I split a brand's knowledge into five types:
| Memory type | Example |
|---|---|
| Brand | Technical, clear tone; avoid jargon and clickbait |
| Content | LinkedIn carousel published Aug 12 with high saves and low clicks |
| Experiment | Technical posts generated more saves than promotional posts |
| Feedback | Audience asked about implementation and cost |
| Strategic | Increasing technical content led to higher engagement |
Every record carries a brandId, which lets recall stay scoped to the correct brand.
For example, a content memory has a concrete structure:
export interface ContentMemory {
brandId: string;
postId: string;
title: string;
platform: 'linkedin' | 'x' | 'blog';
format: string;
topic: string;
publishedAt: string;
metrics: { saves: number; comments: number; clicks: number };
}
The interface itself is not the memory. It is the structure used to turn application data into a useful memory.
That distinction became important.
Retain: turn structured data into facts
One of the biggest improvements came from changing what I sent to Hindsight.
Instead of retaining an entire database record as a large object or text dump, I converted the important information into a short factual statement.
import { HindsightClient } from '@vectorize-io/hindsight-client';
const client = new HindsightClient({
baseUrl: process.env.HINDSIGHT_URL!
});
export async function retainPost(p: ContentMemory) {
const text =
`On ${p.publishedAt}, the ${p.format} post "${p.title}" about ${p.topic} ` +
`was published on ${p.platform}. It received ${p.metrics.saves} saves, ` +
`${p.metrics.comments} comments and ${p.metrics.clicks} clicks.`;
return withTimeout(
client.retain(p.brandId, text, { context: 'content-history' })
);
}
This wrapper gives the memory a clear context and keeps the retained information factual.
The lesson was straightforward: noisy input produced noisy facts. Short, specific sentences produced cleaner memories and sharper recall without changing the retrieval logic.
Hindsight provides the underlying memory system; my service decides what application information should become a memory and how the agent should access it. I used the Hindsight GitHub repository and Hindsight documentation while building these wrappers.
Recall: ask memory the same way the agent asks
The agent does not need to know how memories are stored. It needs an answer to a question.
I kept the recall interface small:
export async function recallForQuestion(
brandId: string,
question: string
) {
const res = await withTimeout(
client.recall(brandId, question)
);
return {
memories: res.results.map(r => r.text),
count: res.results.length
};
}
The brandId scopes the question to one brand, while the question determines what information is relevant.
I also return the number of memories retrieved.
That count is used by the UI to display the "N memories used" indicator. Returning it from the same call means the UI and the agent are working from the same retrieval result rather than maintaining separate counts.
Reflect: move from evidence to a mental model
Recall gives the agent evidence.
But sometimes the agent needs a higher-level understanding of that evidence.
For this, I use Hindsight's reflect capability. I defined recurring questions that represent the mental models I want the agent to maintain about a brand:
- What is this brand's voice?
- Which topics consistently perform well?
- What audience preferences have emerged?
- Which strategies have failed?
- Which topics are under-covered?
The application exposes those questions through a small wrapper:
const MODELS = {
voice: 'What is this brand\'s voice and communication style?',
topTopics: 'What content topics consistently perform well for this brand?',
gaps: 'What topics are currently under-covered for this brand?',
};
export const reflectOn = (
brandId: string,
k: keyof typeof MODELS
) =>
withTimeout(
client.reflect(brandId, MODELS[k])
);
The important distinction is between recall and reflection.
Recall gives me relevant pieces of evidence. Reflection helps turn accumulated evidence into a reusable summary or mental model.
That makes the memory layer more useful than simply keeping a growing collection of raw text.
For the API details and current behavior, I referred to the Hindsight documentation when implementing the wrappers.
The behavior change that mattered
The most useful test was simple:
retain one new piece of audience feedback, then ask the same question again.
Before retaining the feedback:
“We need more practical examples.”
the agent's recommendation was a conceptual technical explainer.
After retaining that feedback, the recommendation changed toward a technical carousel containing a real implementation example.
The prompt did not change.
The question did not change.
The new information came from memory.
That is the behavior I was looking for. The agent was no longer responding only to the information in its current context. It could use previously retained information when producing a new recommendation.
This is also why I think of the system as more than a longer prompt. The Vectorize explanation of agent memory describes the broader idea of giving agents a way to retain and use information across interactions.
What I learned
1. Decide the memory type first
Typed schemas forced me to answer a basic question before writing a retain function:
What exactly am I storing?
That made the memory layer easier to reason about.
2. Write memories as facts
A database record can contain many fields that are useful to an application but unnecessary for memory retrieval. Converting the important information into a short factual statement made the retained information easier to work with.
3. Keep retrieval and observability together
Returning the memory count alongside the results gave the UI one source of truth for displaying how many memories were actually retrieved.
4. Test memory independently
I tested the retain, recall, and reflect flow in isolation before connecting it to the LLM and UI.
That made it easier to determine whether a problem came from the memory layer or from another part of the agent.
What I have not solved
There is still a trade-off between freshness and cost.
Reflection is slower than recall, so refreshing all five mental models after every new memory would be wasteful. My current approach refreshes them on a schedule and after bulk loads.
That means a mental model can temporarily lag behind a newly retained memory.
The next improvement I would explore is event-driven refresh: when a new memory arrives, determine which mental models it could affect and refresh only those models.
For me, the biggest lesson was not simply that an agent needs memory. It was that memory needs structure.
Once I stopped treating everything as one large text blob and started deciding what each piece of information represented, the memory layer became easier to test, easier to explain, and more useful to the agent.


Top comments (0)