The most useful thing I did for my support agent was make its request handler boring. Three steps, in a fixed order, with all the hard memory work pushed into Hindsight, so the handler could stay small enough to read in one sitting.
What the system does
MemoryDesk is a customer support service. A customer sends a message, the service answers it, and the next time that same customer writes, the service remembers what happened before. If Rahul's payment failed last week and the fix was paying by card instead of UPI, the agent knows that the next time he says his order won't go through. It also doesn't tell Priya about Rahul's card.
The stack is deliberately small:
- FastAPI for the HTTP layer
-
Groq for inference, running
openai/gpt-oss-120b - Hindsight for long-term memory, through the Python client
There are three endpoints: a health check, POST /chat, and GET /memories/{customer_id}, which I built for debugging and which turned out to be useful well beyond that. The whole service is one file of about a hundred lines. That's on purpose. The interesting part is not the amount of code but where the state lives, and the answer is: not in my code.
The through-line: a stateless handler with a memory layer beside it
Most "agent with memory" designs I've seen try to be clever inside the request path. They keep a rolling summary in a database column, run a second LLM call to compress history, or maintain a hand-built vector index with a pile of metadata filters. I've written versions of all three. They all work, and they all become a second system that I have to operate, debug, and explain.
The alternative I settled on is to treat the handler as stateless and put a real memory service next to it. Every request runs the same loop:
- Recall what we know about this customer that's relevant to this message.
- Prompt the model with that context and the new message.
- Retain the exchange so the next request can recall it.
That's the whole architecture. What I want to go through is why each step is shaped the way it is, because the details are where this either works or quietly leaks data.
Step one: recall, scoped by tag
Here is the recall call from /chat:
Three decisions are packed in here.
One bank, many customers, isolation by tag. I use a single memory bank and tag every memory with user:<customer_id>. The alternative is one bank per customer. I considered it, but a bank-per-customer design pushes lifecycle management (creation, naming, cleanup) into my service. With tags, the customer boundary is a property of each memory, and the handler doesn't need to know whether a customer is new.
any_strict matters more than it looks. Tag matching modes differ in how they treat memories that carry no tags at all. The non-strict modes can include untagged memories in the results. The strict variant excludes them. In a multi-tenant support system that difference is the entire privacy story. If a memory ever gets written without a tag, through a bug, a script, or a future code path I haven't thought of, I want it invisible to customer-scoped recall rather than visible to everyone. I chose the mode that fails closed.
The query is the customer's message. I don't rewrite it, expand it, or add keywords. The customer says "my payment isn't working" and that string is what gets matched against what we know. It's the simplest possible retrieval query, and for support conversations, where people describe their problem in the same vocabulary they used last time, it's been a good starting point. I'd rather see where it breaks than tune it in advance.
I cap what goes into the prompt at the top five results. That's a budget decision: recalled memory competes with the system prompt and the current message for context, and I would rather give the model five relevant facts than fifteen loosely related ones.
Step two: prompt, and tell the model not to make things up
The recalled text goes into the system prompt:
response = groq_client.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[
{
"role": "system",
"content": f"""
You are MemoryDesk, a helpful and professional customer support agent.
Use previous customer memories when they are relevant.
Do not invent memories.
Previous relevant memories:
{memory_text}
"""
},
{"role": "user", "content": request.message}
]
)
There are two instructions there and they pull in opposite directions on purpose. "Use previous customer memories when they are relevant" gives the model permission to personalize. "Do not invent memories" stops it from pretending to recall things that aren't in the block above.
The second line is the one that saves you. A support agent that says "as we discussed last time, I've refunded the difference" when no such conversation exists is worse than one with no memory at all. Customers notice, and they stop trusting the agent.
I'll be honest about what this line is: a prompt, not a guarantee. It reduces the problem; it doesn't eliminate it. What makes it workable is that the memory block is plain text the model can see, so when the block is empty, there's nothing to draw on and the honest answer is the easy one. The response also returns memories_used, a count of what recall returned, so I can look at any reply and tell immediately whether the agent was answering from history or from nothing.
The model choice is intentionally decoupled from all of this. Memory lives in Hindsight, not in the model's context or in some model-specific format, so swapping the Groq model for another means changing one string. I've done it, and nothing about the customer's history had to move.
Step three: retain the exchange
After the answer is generated, the handler writes the interaction back:
This is the step I underestimated. My instinct was to store something curated: a summary, an extracted fact, a "resolution" field. Instead I hand Hindsight the raw exchange and let it decide what's worth keeping. That's the part of agent memory I didn't want to own. Deciding which facts to extract, how to relate them to each other, and how to surface them at recall time is a different job from handling support requests, and it's the job the memory layer is for.
Two details worth pointing out.
The tag is on the write, not just the read. Scoping recall does nothing if writes aren't scoped too. The same customer_tag variable is used on both sides, computed once per request, so the two can't drift apart.
The customer ID is in the content and in the tag, and only one of those is a security boundary. I put Customer ID: ... inside the text because it makes the stored memory self-describing when I read it back. But the text is just text. It's the tag that enforces isolation. I mention this because it's the kind of thing that's easy to get backwards: a stored ID that looks like an access control but isn't one.
What it looks like in practice
The scenario I used to validate the design is small on purpose. First, Rahul reports a payment failure while placing an order, and the resolution is to pay by card instead of UPI. That conversation gets retained under user:rahul.
Later, he writes:
My payment isn't going through again.
The recall step searches within his tag only. What comes back is the earlier failure and its resolution, and the model sees it in the system prompt. A reasonable reply now starts from what's known ("last time this happened, paying by card worked") instead of running through the generic checklist of retrying, checking the network, and contacting the bank.
If Priya sends the same message, her tag scope contains nothing about UPI, so memories_used is zero and the agent answers like a fresh support agent. That behavior is the important part. Same words, different customers, different context, and no path by which Rahul's history reaches Priya.
The debugging endpoint makes this checkable:
@app.get("/memories/{customer_id}")
def get_memories(customer_id: str):
customer_tag = f"user:{customer_id}"
memories = hindsight.recall(
bank_id=BANK_ID,
query="customer history previous issues preferences resolutions",
tags=[customer_tag],
tags_match="any_strict"
)
return {
"customer_id": customer_id,
"memories": [m.text for m in memories.results[:10]],
"memory_count": len(memories.results)
}
Hitting /memories/rahul shows what the agent knows about Rahul, and /memories/priya shows what it knows about Priya. It's the same recall path the chat handler uses, with a generic query and a higher cap. When someone asks "why did the agent say that?", the first thing I do is look here. It's also the natural place to point a support lead who wants to audit what the system has retained about a customer.
What went wrong, and what was painful
Retaining the agent's own replies is a double-edged decision. Storing Agent response alongside the customer's message means the memory contains what we told the customer, which is useful because it captures what was promised. It also means that if the agent said something wrong, that wrong answer is now part of the record. I haven't found a clean answer. Today I treat memory as a log of what happened rather than a source of truth, and the prompt reflects that: "relevant memories," not "facts."
A single bank is a choice you have to keep defending. Everything depends on the tag being present and correct on every write and every read. That's one string format (user:<id>) repeated in three places in the code. I'd like it to live in exactly one helper. It's the first refactor I would do, and I'd recommend making that helper before you write the second endpoint, not after.
Synchronous calls in the request path add up. A chat request now makes a recall, a model call, and a retain, in sequence. The retain doesn't need to finish before the customer sees a reply, so it's the obvious candidate to move off the critical path. I've left it inline because it keeps the ordering guarantee simple: when a response returns, that exchange is already retained.
Lessons learned
1. Make the handler stateless and let the memory layer be a separate thing. The moment your request code owns summarization, deduplication, and indexing, you have two products to maintain. A fixed recall, prompt, retain loop is easy to reason about and easy to test.
2. Choose the tag-matching mode that fails closed. In anything multi-tenant, the question isn't "does scoping work when everything is right?" but "what happens when a write is missing its tag?" any_strict answers that the way I want.
3. Scope both sides of the loop with the same value. A tag on recall without a tag on retain is a false sense of security. Compute it once, use it twice.
4. Give the model permission to use memory and an instruction not to invent it. You need both. Then expose a signal, like memories_used, so you can tell which mode a given answer came from.
5. Build the inspection endpoint early. GET /memories/{customer_id} cost almost nothing to write and answered more debugging questions than any log line. If you can't see what the agent remembers, you can't tell whether a bad answer came from bad memory or bad reasoning.
The whole thing is still a hundred-odd lines of Python, which is what I wanted. If you want to try the same approach, the Hindsight repository and documentation are the best places to start, and the Vectorize overview of agent memory is a good primer on why keeping memory outside your handler is worth the extra moving part.






Top comments (0)