DEV Community

Cover image for Swapping Hardcoded Customer Context for Hindsight Recall
Sreeja
Sreeja

Posted on

Swapping Hardcoded Customer Context for Hindsight Recall

Swapping Hardcoded Customer Context for Hindsight Recall

The first version of my support agent had a great memory, as long as I typed it in by hand. Every request carried a customer_context field, and every test passed because I was the one writing the context. Then I asked the obvious question: who fills that field in production?

This is the story of replacing that field with real memory, using Hindsight, and of the design decisions that mattered more than the swap itself.

What the system does

The project is a customer support agent. A customer sends a message, and the agent works out what kind of problem it is, decides whether a human needs to get involved, pulls in what it knows about that customer, and produces a response. The stack is deliberately plain: Python, FastAPI, and a local Llama 3.2 model served by Ollama. There is no hosted LLM API and no API key to manage. Everything runs on hardware I control, which matters when the input is support tickets full of account details.

Every message moves through a fixed pipeline:

Customer Message
        ↓
Intent Detection
        ↓
Escalation Check
        ↓
Customer Context
        ↓
Ollama AI Agent
        ↓
Support Response
Enter fullscreen mode Exit fullscreen mode

Intent detection sorts messages into six categories: PAYMENT, LOGIN, ACCOUNT, SUBSCRIPTION, TECHNICAL, and OTHER. The escalation check decides whether the model should be involved at all. Only after those two steps does the customer context get assembled and handed to Llama 3.2.

Here is the same flow with Hindsight in place. Recall sits in the request path, retain sits at the end, and both talk to a per-customer bank:

Architecture diagram

The model side is unremarkable, and that's the point. Llama 3.2 runs locally through Ollama, and a quick ollama run llama3.2 is all it takes to confirm it's answering before I involve any of my own code:

Ollama Running in terminal

The API surface is a single endpoint:

POST /chat

{
    "customer_id": "C001",
    "message": "My payment failed",
    "customer_context": {}
}
Enter fullscreen mode Exit fullscreen mode

That customer_context field is the subject of this article.

The problem with context you pass in

In the early version, the caller supplied the context:

{
    "previous_issue": "Payment failed before",
    "previous_solution": "Retrying the payment resolved the issue"
}
Enter fullscreen mode Exit fullscreen mode

This was useful for proving the prompt worked. With that context present, the agent stopped suggesting generic fixes and started referring to what had already been tried. The tests I wrote (payment issues, login issues, escalation, customer context) all passed.

But the design had a flaw that no test could catch. The frontend, or whatever sits in front of the agent, would have to know what is relevant about a customer, fetch it, format it, and send it on every request. That pushes the hardest part of the problem onto the caller. It also means the agent itself has no memory. It only has whatever the caller remembered to send.

I wanted the agent to own this. The request should need only a customer_id and a message.

Why I used a memory layer instead of a table

My first instinct was a Postgres table: customer_id, previous_issue, previous_solution. That works until you notice what the fields imply. Support history isn't a fixed schema. One customer's history is "card declined, retry fixed it." Another's is "locked out after a phone change, verified by email, asked twice about downgrading." Forcing that into two columns means deciding up front what matters, and I wouldn't know until I saw real tickets.

What I actually needed was to store conversation outcomes as they happened, and later retrieve the ones relevant to the current message. That is what agent memory is for, and it's why I looked at Hindsight instead of building retrieval on top of a database myself. It's open source, so I could read how it works and run it next to the rest of the stack. The Hindsight documentation covers the model well, and there is a Python client, hindsight-client, which is already in my requirements.txt alongside fastapi, uvicorn, python-dotenv, and requests.

The through-line: one memory bank per customer

The decision that shaped everything else was scoping. Hindsight organizes memory into banks. I create one bank per customer, keyed by the same customer_id the API already receives.

from hindsight_client import Hindsight

client = Hindsight(base_url=HINDSIGHT_URL)

def bank_for(customer_id: str) -> str:
    return f"customer-{customer_id}"
Enter fullscreen mode Exit fullscreen mode

This gives me isolation for free. When the agent recalls memory for C001, it queries only C001's bank. There is no filtering step that could be forgotten, and no shared index where one customer's payment failure could surface in another customer's reply. For a support system, a cross-customer leak is the worst class of bug, and I preferred to make it structurally impossible rather than something I had to test for.

Retrieval replaces the field that used to arrive in the request body:

def get_customer_context(customer_id: str, message: str) -> str:
    result = client.recall(
        bank_id=bank_for(customer_id),
        query=message,
    )
    return format_memories(result)
Enter fullscreen mode Exit fullscreen mode

The query is the customer's current message. If someone writes "my payment failed again," the recall is driven by that text, so payment-related history ranks above an old login problem. That is the behavior I was faking by hand before, except now it selects from real history instead of a string I typed.

Where recall sits in the pipeline

Placement mattered. Recall happens after the escalation check, not before.

The reasoning: escalation is a decision about the message and the category, and it should not depend on a network call to a memory service. If a message needs a human, I want that path to be fast and to work even if the memory service is down. Recall only runs on the path where the model will actually use it. That keeps the failure modes separate. A memory outage degrades the agent into a competent but generic support bot. It doesn't take escalation with it.

The empty case needed the same care. A new customer has an empty bank, and recall returns nothing. That is a normal condition, not an error. The context block becomes a plain statement that there is no prior history, and the prompt is written so the model doesn't invent any. This was the first thing I tested after the swap, because a model handed an empty "history" section will sometimes fill the gap with confident nonsense.

Writing memory back

Recall is only half of it. After the agent replies and the outcome is known, I write it back:

def remember_outcome(customer_id: str, message: str, resolution: str) -> None:
    client.retain(
        bank_id=bank_for(customer_id),
        content=(
            f"Customer reported: {message}. "
            f"Outcome: {resolution}."
        ),
    )
Enter fullscreen mode Exit fullscreen mode

I write the outcome, not the raw transcript. The useful sentence is "payment failed, retrying resolved it," not forty lines of greetings. That is the same information as the old hardcoded previous_issue and previous_solution fields, except it now accumulates on its own, and I don't have to decide the schema ahead of time.

There's a judgment call here about escalated tickets. When a case goes to a human, the agent doesn't know how it ended. I record that it was escalated and why, so the next conversation can acknowledge that a person is already handling it rather than starting over.

Behavior in practice

The clearest way I found to see what memory changes is to send nearly the same message twice. These two captures are from the FastAPI /docs page, from the stage where I still passed the context by hand. That is exactly the payload recall now produces on its own, so they show the effect cleanly.

First, a customer with history. The context says a previous payment failure was fixed by retrying, and the message is "My payment failed again":

FastAPI docs showing a POST /chat request with previous_issue and previous_solution in customer_context, and a response that references the earlier failure and the retry that fixed it

The response acknowledges that this has happened before, mentions that retrying resolved it last time, and then asks what the customer has already tried and whether their payment method or account details changed. It reads like someone who has seen this customer's file.

Now a customer with nothing to recall. Same endpoint, same customer ID, empty context, message "My payment failed":

FastAPI docs showing a POST /chat request with an empty customer_context and a generic response asking which payment method was used and why the payment might have failed

The intent is still classified as PAYMENT and the action is still ASSIST, but the reply is a generic intake: which payment method, was the card expired, were there insufficient funds. Nothing in it claims a history that doesn't exist.

Same model, same prompt template, same intent. The only difference is the recalled context. With Hindsight in the loop, the request body shrinks to customer_id and message, and the first response is what the agent produces when the bank has a relevant memory. The second is what it produces when the bank is empty.

What I'd tell someone doing the same thing

Move context ownership into the agent. If the caller has to assemble memory, you've built a prompt template, not an agent. The request should carry identity and intent; everything else is the agent's job.

Scope memory by the identifier you already have. One bank per customer_id meant isolation was a property of the data layout rather than a filter I had to get right on every query.

Decide where a dependency sits in the failure path. Putting recall after escalation means a memory outage can only degrade answer quality. It can't block the paths that need to work no matter what.

Treat empty memory as a first-class case. The new-customer path is the one most likely to produce fabricated history. Write the prompt for it explicitly and test it before anything else.

Store outcomes, not transcripts. What a future conversation needs is what happened and what fixed it. Short, factual entries recall better than raw chat logs, and they read cleanly when injected into a prompt.

Where it ended up

The pipeline diagram barely changed. It gained one arrow: between the escalation check and the model, customer context now comes from Hindsight recall instead of the request body. The customer_context field went from load-bearing to unnecessary, and the frontend got simpler because it stopped needing to know anything about customer history.

If you're building a similar agent, the Hindsight repository is the place to start, and the docs explain banks, retain, and recall in more detail than I can here. The most useful thing I got out of this wasn't a feature. It was the ability to stop hand-writing what my agent was supposed to remember.

Top comments (0)