Powering Long-Term Memory in OpsSentry Backend
While building OpsSentry, one of our primary technical challenges was ensuring that the AI agent could dynamically retain and recall user context across multiple sessions without bloating the primary prompt window.
To solve this, we implemented a dual-action memory workflow in our FastAPI backend powered by Hindsight.
1. Memory Recall Before Generation
Before generating any response with our LLM, our backend queries Hindsight using the incoming user message to retrieve relevant context from previous conversations. This gives the model access to earlier interactions without requiring us to manually paste the entire conversation into every prompt.
Here is how the recall integration is implemented in our backend service:
# Query Hindsight memory bank for context relevant to the current query
memory_context = hindsight.recall(
bank_id=HINDSIGHT_BANK_ID,
query=user_message,
limit=3
)
# Inject retrieved context directly into system prompt
system_prompt = f"""
You are OpsSentry, an AI assistant.
Relevant Past Context:
{memory_context}
"""
### 2. Retaining New Information
Recall is only useful if the system also stores useful information for future interactions.
After generating the response, our backend retains the interaction in Hindsight:
python
hindsight.retain(
bank_id=HINDSIGHT_BANK_ID,
content=f"User: {message} | AI: {ai_response}"
)
"""
This creates a simple memory loop:
- User asks a question
- Backend recalls relevant memories
- LLM generates response with context
- Backend retains interaction back to Hindsight
Key Takeaways & Architecture
Decoupled Memory: Keeping short-term context window management handled dynamically via Hindsight keeps latency minimal and responses highly accurate.
Database Sync: Metadata and chat logs are safely stored in Supabase, while vectorized long-term memory indexes live in Hindsight.
Seamless API: FastAPI handles asynchronous request routing between Groq, Supabase, and Hindsight seamlessly.
Top comments (0)