The hardest part of building a long-running AI agent is not picking the right model or writing a clever prompt. It is answering this question: what does the agent know, from when, and how confident should it be in that knowledge?
We were building an Institutional Memory Radar — a system that processes executive meeting transcripts over months and surfaces when new decisions contradict old ones, when action items go unresolved, and when organizational "fault lines" are forming below the surface. We needed the system to not just remember past meetings, but to understand the temporal relationships between them.
A decision made in January sets a constraint. A test run in March validates or invalidates it. A policy violation in May breaks it. These are not independent data points — they form a causal chain across time. Getting the agent to understand that chain is what I mean by a temporal memory graph, and it is why I chose Hindsight as the foundation for our memory layer.
The Difference Between Memory and Temporal Memory
Standard vector database retrieval answers the question: "What text in my corpus is most semantically similar to this query?"
That is useful, but it loses two critical pieces of information:
- When the memory was formed (recency)
- What prior memories that memory was building on (causality)
A purely semantic retrieval system will correctly surface the MDM security policy document when you ask about "endpoint security compliance." But it will not tell you that the policy was justified by a gap analysis from a month earlier, enforced by an HR mandate two weeks later, and violated by the marketing team three months after that. That chain is the temporal memory graph.
How We Represent Temporal Context
Our memory entries are not just raw transcripts. Every piece of text we push to Hindsight via Vectorize agent memory is a structured summary that embeds temporal markers directly in the content itself.
After Gemini processes a meeting, it returns a JSON payload with decisions, action items, and unresolved issues. We flatten this into a human-readable text block that includes explicit date references and forward-pointing pointers before retaining it.
// backend/src/routes/auditRoute.ts
function buildRetainableMemory(
transcript: string,
insights: AuditInsights,
meetingDate: string
): string {
const lines: string[] = [
`Meeting Date: ${meetingDate}`,
'',
'== SUMMARY ==',
insights.meetingSummary,
'',
];
if (insights.decisions.length > 0) {
lines.push('== DECISIONS (BINDING) ==');
insights.decisions.forEach((d) => lines.push(`- ${d}`));
lines.push('');
}
if (insights.actionItems.length > 0) {
lines.push('== OPEN ACTION ITEMS ==');
insights.actionItems.forEach((a) => lines.push(`- ${a}`));
lines.push('');
}
if (insights.unresolvedItems.length > 0) {
lines.push('== UNRESOLVED / DEFERRED ==');
insights.unresolvedItems.forEach((u) => lines.push(`- ${u}`));
lines.push('');
}
return lines.join('\n');
}
The result is a memory entry that any downstream recall can parse as structured text. When Hindsight returns five retrieved memories, Gemini can read "DECISIONS (BINDING): Approved Shared Database/Separate Schemas architecture" from January, and "OPEN ACTION ITEMS: Optimize Postgres RLS indexes" from February, and understand both the decision and its pending validation — without us having to maintain an explicit graph database.
Querying the Graph
The temporal graph emerges from the combination of structured memory entries and Hindsight's semantic retrieval. When a new meeting invite arrives, we use its subject and attendee list as the recall query:
// backend/src/adapters/hindsightAdapter.ts
export async function recall(query: string): Promise<string> {
const url = `${HINDSIGHT_URL}/banks/${process.env.HINDSIGHT_BANK_ID}/memories/recall`;
const response = await fetch(url, {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.HINDSIGHT_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({ query }),
});
const data = await response.json();
return data.results?.map((r: any) => r.content).join('\n\n---\n\n') ?? '';
}
The recall result surfaces multiple memory entries from different points in time. Because each entry contains the meeting date and structured decision/action labels, Gemini can reconstruct the causal chain in its reasoning step without us having to explicitly encode graph edges.
The Shadow IT Incident: A Temporal Chain in Practice
The clearest demonstration of temporal memory working as intended was when the system caught a Shadow IT violation five months into our six-month observation period.
Here is the causal chain that the agent successfully reconstructed:
| Date | Event | Memory Entry Type |
|---|---|---|
| Jan 29 | Architecture decision: Shared Database/Separate Schemas approved | DECISION |
| Feb 12 | SOC 2 gap analysis: missing centralized software procurement policy | UNRESOLVED |
| Feb 26 | Policy ratified: centralized software procurement process mandatory | DECISION |
| May 7 | Action item assigned: Vendor Risk Management audit required | ACTION ITEM |
| May 21 | Shadow IT found: marketing team using unauthorized analytics tool | VIOLATION |
When the system processed the May 21 meeting invite ("Vendor Risk Management Audit results"), the recall query surfaced the February 26 procurement policy and the February 12 gap analysis. Gemini read both and immediately flagged that the marketing tool violated a binding policy established three months prior. It also flagged that the audit itself was a direct result of an open action item from May 7 — agenda debt carried forward.
None of this required us to explicitly link these records. The temporal context embedded in the memory entries — the meeting dates, the BINDING/UNRESOLVED labels, the structured decision text — gave Gemini enough signal to reconstruct the chain on its own.
What Hindsight Gets Right About Temporal Memory
There are three specific design decisions in Hindsight that made this pattern work cleanly, rather than requiring us to build them ourselves.
Continuous memory, not snapshots. Hindsight is designed as a streaming memory bank, not a static document store. Every POST /memories call adds to a living context rather than versioning a point-in-time snapshot. This matches how institutional knowledge actually accumulates.
Retrieval by semantic intent, not by timestamp. The recall endpoint returns results ranked by semantic similarity to the query, not by recency. This is the correct default for a reasoning agent — you want relevant history, not recent history. Recency can be imposed at the prompt level if needed, but for most decisions, topical relevance matters more.
No schema enforcement. Our structured memory format (== DECISIONS (BINDING) ==, etc.) is just a text convention. Hindsight stores and retrieves plain text. This means we can evolve our memory schema without migration scripts or index rebuilds. The cost is that the consumer (Gemini) has to parse it — but LLMs are excellent at extracting structure from consistently formatted text.
Lessons Learned
1. Embed temporal markers in the memory content, not just metadata.
If your meeting date and decision labels live only in database metadata fields, your LLM cannot access them during reasoning. Put the date and structure inside the text itself.
2. Let the LLM reconstruct the graph, not your code.
We initially considered building an explicit graph database with typed edges (DECISION → validates → ACTION_ITEM). It was enormous overhead. Structured text in a semantic memory bank is sufficient for an LLM to reason about causal chains, and it is far simpler to maintain.
3. Labeling matters more than formatting.
The difference between "The team decided to use PostgreSQL" and "DECISION (BINDING): Approved PostgreSQL Shared Database/Separate Schemas architecture" is enormous for downstream reasoning accuracy. Invest in clear labeling conventions early.
4. The recall query should match the agent's actual uncertainty.
The most effective recall query is whatever question the agent is trying to answer, not a reformulation of the document it is analyzing. "What prior decisions constrain our database choice?" retrieves better context than "executive steering committee meeting agenda March 2026."
5. Temporal memory is a product decision, not just a technical one.
How far back the agent looks, what types of prior content it prioritizes, when a decision expires and becomes stale — none of these have technically correct answers. They are product decisions that determine how the agent behaves in practice. Define them explicitly before you ship.
Building a temporal memory graph on top of Hindsight did not require a graph database, a custom embedding pipeline, or a complex retrieval system. It required structuring what we stored, being deliberate about what we retrieved, and trusting the LLM to reason over the result. The infrastructure was simpler than I expected. The hard work was in deciding what institutional memory means for the specific reasoning task we were trying to automate.
Top comments (0)