Every time an engineer started a new session with our project assistant, the agent suffered from total amnesia. It could analyze architecture, propose deployment manifests, and draft pull requests, yet it routinely asked what cloud platform we were deploying to, which authentication pattern we preferred, and why we made key trade-offs last week.
Dumping thousands of lines of raw conversational history into prompt windows bloated token usage and diluted attention. Standard vector RAG on raw transcripts did not help either; querying "What is our deployment pattern?" fetched disjointed chat fragments rather than definitive engineering decisions.
We designed ProjectRecall, an agent workflow built around Hindsight and Microsoft Agent Framework to establish durable, session-independent context. By treating memory as an explicit lifecycle loopโretaining decisions, recalling structured facts, and reflecting on outcomesโwe transformed an ephemeral chat interface into a continuous engineering teammate.
1. The Core Technical Problem: Transcripts vs. Durable State
Most agent stacks treat memory as an unbroken text file or an uncurated vector database of raw user turns. This design creates three immediate operational failures:
- Context Bloat: Replaying raw transcripts rapidly exhausts context limits while increasing latency and inference cost.
- Context Drift: Irrelevant conversational tangents drown out past decisions and architectural constraints.
- No Learning Loop: When an agent executes a workflow (like deploying an infrastructure change or updating a schema), the eventual success or failure is forgotten once the run completes.
We realized we did not need to remember every conversational filler phrase. We needed to capture durable facts, engineering decisions, and execution outcomes. This insight led us to deploy Vectorize agent memory via Hindsight, isolating context into discrete, project-specific memory banks.
2. Architecture: The Memory Loop
ProjectRecall wraps the Microsoft Agent Framework lifecycle in a two-stage interception loop:
-
before_runhook: Intercepts the incoming user input, queries Hindsight for memories scoped strictly to the project'sbank_id, and injects only relevant context into the execution prompt. -
after_runhook: Analyzes the final interaction and tool outputs, identifies durable decisions or configuration choices, and calls Hindsight's retain endpoint to persist state.
text
+-------------------------------------------------------------------+
| ProjectRecall UI |
+-------------------------------------------------------------------+
|
v
+-------------------------------------------------------------------+
| Agent Orchestrator Loop |
| 1. before_run: recall(query, bank_id) -------------------------->|---+
| 2. Execute LLM Reasoning & Action Tools | |
| 3. after_run: retain(decision/outcome, bank_id) ---------------->|--+|
+-------------------------------------------------------------------+ ||
||
[Hindsight Memory Bank API] <-----+
- retain (Store durable context)
- recall (Retrieve task context)
- reflect (Synthesize insights)
import os
import httpx
from typing import List, Dict, Any
class HindsightMemoryProvider:
def __init__(self, api_base_url: str, bank_id: str):
self.api_base_url = api_base_url.rstrip("/")
self.bank_id = bank_id
self.client = httpx.Client(timeout=10.0)
def recall(self, query: str, limit: int = 5) -> List[Dict[str, Any]]:
"""Retrieve relevant durable memories prior to task execution."""
response = self.client.post(
f"{self.api_base_url}/banks/{self.bank_id}/recall",
json={"query": query, "limit": limit}
)
response.raise_for_status()
return response.json().get("memories", [])
def retain(self, content: str, metadata: Dict[str, Any] = None) -> None:
"""Store durable decisions, operational preferences, and outcomes."""
payload = {
"content": content,
"metadata": metadata or {}
}
response = self.client.post(
f"{self.api_base_url}/banks/{self.bank_id}/retain",
json=payload
)
response.raise_for_status()
Top comments (0)