Stateless AI agents fail in complex operations workflows. When a user reports a recurring server fault or infrastructure issue, traditional LLM setups force them to re-explain the entire context from scratch.
To overcome this limitation, I built the asynchronous backend pipeline for OpsSentry AI. Our system integrates Vectorize’s Hindsight Memory SDK alongside FastAPI and Groq to create an agent that retains memory, recalls past context on demand, and adapts across sessions.
## The Backend Pipeline Strategy
Instead of relying on ephemeral chat histories or manually passing massive prompt histories on every request, our FastAPI service handles a three-stage memory lifecycle:
Context Retrieval (Recall): Before generating a response, the backend queries Hindsight to retrieve relevant past facts and troubleshooting history linked to the user.
Contextual Inference (LLM Completion): The recalled memory is injected directly into the system prompt before calling Groq (qwen/qwen3-32b).
Memory Persistence (Retain): Once the completion is generated, the interaction is written back into Hindsight for future recall.
Incoming Request (user_id + message)
│
▼
[ FastAPI Server ]
│
├──► 1. Query Hindsight Cloud ──► Fetch user-specific context
│
├──► 2. Call Groq API ─────────► Generate answer with context
│
└──► 3. Write to Hindsight ────► Store interaction into memory
**
Implementation Highlights
**
Below is an abstract of the backend logic running inside our FastAPI service:
import os
import logging
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from groq import Groq
from hindsight_client import Hindsight
app = FastAPI(title="OpsSentry AI Core Service")
# Initialize Clients
hindsight = Hindsight(api_key=os.getenv("HINDSIGHT_API_KEY"))
groq_client = Groq(api_key=os.getenv("GROQ_API_KEY"))
HINDSIGHT_BANK_ID = os.getenv("HINDSIGHT_BANK_ID", "customer_bank")
GROQ_MODEL = os.getenv("GROQ_MODEL", "qwen/qwen3-32b")
class ChatPayload(BaseModel):
user_id: str
message: str
@app.post("/api/chat")
async def process_chat(payload: ChatPayload):
user_id = payload.user_id.strip()
message = payload.message.strip()
# Stage 1: Recall context from Hindsight
recalled_context = []
if hindsight:
try:
search_query = f"User {user_id}: {message}"
recall_res = hindsight.recall(bank_id=HINDSIGHT_BANK_ID, query=search_query)
recalled_context = recall_res.results if hasattr(recall_res, "results") else recall_res
except Exception as err:
logging.error("Memory retrieval error: %s", err)
# Prepare system prompt with recalled memory
memory_str = "\n".join([f"• {m}" for m in recalled_context]) if recalled_context else "None"
system_instruction = (
f"You are OpsSentry AI assisting user '{user_id}'.\n"
f"Recalled User Context:\n{memory_str}\n\n"
f"Incorporate past facts naturally when answering."
)
# Stage 2: Inference via Groq
try:
response = groq_client.chat.completions.create(
model=GROQ_MODEL,
messages=[
{"role": "system", "content": system_instruction},
{"role": "user", "content": message}
],
temperature=0.7
)
ai_output = response.choices[0].message.content or ""
except Exception as err:
raise HTTPException(status_code=500, detail=f"Inference error: {err}")
# Stage 3: Store memory back into Hindsight
if hindsight:
try:
hindsight.retain(
bank_id=HINDSIGHT_BANK_ID,
content=f"User {user_id}: {message} | AI: {ai_output}"
)
except Exception as err:
logging.error("Memory persistence error: %s", err)
return {
"response": ai_output,
"recalled_memories": recalled_context
}
Building OpsSentry Backend: Integrating Long-Term AI Memory with Hindsight & FastAPI
ai
python
fastapi
backend
Powering Long-Term Memory in OpsSentry Backend
While building OpsSentry, one of our primary technical challenges was ensuring that the AI agent could dynamically retain and recall user context across multiple sessions without bloating the primary prompt window.
To solve this, we implemented a dual-action memory workflow in our FastAPI backend powered by Hindsight.
- Memory Recall Before Generation Before generating any response with our LLM, our backend queries Hindsight using the incoming user message to retrieve relevant context from previous conversations. This gives the model access to earlier interactions without requiring us to manually paste the entire conversation into every prompt.
Here is how the recall integration is implemented in our backend service:
# Query Hindsight memory bank for context relevant to the current query
memory_context = hindsight.recall(
bank_id=HINDSIGHT_BANK_ID,
query=user_message,
limit=3
)
# Inject retrieved context directly into system prompt
system_prompt = f"""
You are OpsSentry, an AI assistant.
Relevant Past Context:
{memory_context}
"""
### 2. Retaining New Information
Recall is only useful if the system also stores useful information for future interactions.
After generating the response, our backend retains the interaction in Hindsight:
python
hindsight.retain(
bank_id=HINDSIGHT_BANK_ID,
content=f"User: {message} | AI: {ai_response}"
)
"""
This creates a simple memory loop:
- User asks a question
- Backend recalls relevant memories
- LLM generates response with context
- Backend retains interaction back to Hindsight
**
Key Takeaways & Architecture
**
Decoupled Memory: Keeping short-term context window management handled dynamically via Hindsight keeps latency minimal and responses highly accurate.
Database Sync: Metadata and chat logs are safely stored in Supabase, while vectorized long-term memory indexes live in Hindsight.
Seamless API: FastAPI handles asynchronous request routing between Groq, Supabase, and Hindsight seamlessly
Top comments (0)