DEV Community

Misbah
Misbah

Posted on

How Hindsight Turned My Chatbot Into an Agent That Remembers

What Happens When Your Support Agent Finally Remembers You

Every support chatbot I'd used before had the same problem: it forgot me the second I closed the tab. Tell it about a damaged order today, come back tomorrow, and you're explaining it from scratch again. So my team of five set out to fix that one specific annoyance, using Hindsight as the memory layer.

We built ResolveIQ Lite, a small customer-support agent that remembers a customer's past incidents and uses them the next time that customer writes in, even in a brand-new session. It's a prototype with fictional customers and orders, not a real support system, but the memory behavior is real.

The problem in one sentence

A support agent without memory treats every message as the first message. Ask it "I still haven't received my replacement" with no context, and it has nothing to work with. That's the gap we wanted to close.

How it's wired together

The flow is short on purpose:

  1. The customer types a message in a Streamlit screen.
  2. We recall that customer's history from Hindsight.
  3. An LLM (via Groq) writes a reply using the current message plus whatever was recalled.
  4. We retain the new exchange back into Hindsight.
def _bank(customer_id):
    return "resolveiq-" + customer_id.lower()

def retain_incident(customer_id, text):
    client.retain(
        bank_id=_bank(customer_id),
        content=text,
        context="customer support interaction",
    )

def recall_incidents(customer_id, query):
    try:
        results = client.recall(bank_id=_bank(customer_id), query=query)
        return [r.text for r in results.results]
    except Exception:
        return []
Enter fullscreen mode Exit fullscreen mode

The one design decision I'd point to is the bank naming: one Hindsight memory bank per customer (resolveiq-c101, resolveiq-c102, and so on). It's a small line of code, but it's what guarantees customer isolation — C102 can never see C101's damaged order, because they're not even querying the same bank.

The part that made memory feel real

The moment that convinced me this wasn't just a chat log was the recalled-memory panel. It doesn't show the raw conversation history — it shows what Hindsight actually surfaced as relevant:

"The customer has not received their replacement item. | When: 2026-09-28 | Involving: customer"

"Customer's order A104 arrived damaged, they requested a replacement, and as of September 28, 2026, they have not yet received it."

Those are synthesized facts, not copy-pasted messages. That's the difference between memory and a transcript.

Before and after, same message

We sent the exact same line, "I still haven't received my replacement," with memory off and memory on.

Memory OFF:

Memory ON:

"I see you're still waiting for the replacement item you mentioned on 2026-09-28. Could you please confirm the order number and the shipping address it should go to? I'll pass this information to our support team for review, and they'll follow up with the next steps."

The second reply references a real date from the customer's own history instead of starting cold. That's agent memory doing its job.

What surprised me

Honestly, less than I expected. The core loop — recall, generate, retain — worked close to how we designed it on paper. If anything, that was the surprising part: memory retrieval that "just works" from a documented API is not something I'd experienced before this project.

An honest limitation

The reply above still asks the customer to "confirm the order number," even though an earlier message in the same customer's history mentions order A104 by name. Hindsight recalled several memory lines, but the most specific one (the one naming A104) didn't rank first for this particular query. It's not a failure of memory, it's a retrieval-tuning problem: worth exploring further, and worth admitting rather than papering over. For a prototype with fictional data and a small seeded history, it's a reasonable rough edge, not a dealbreaker.

Lessons I'd pass on

  1. Isolate memory by key early. Deciding "one bank per customer" on day one saved us from ever worrying about data leaking between customers.
  2. Show, don't just build, the recall. Putting "recalled memory" in its own visible panel made the value of memory obvious to anyone watching the demo, not just to us.
  3. Coordinating five people on one Hindsight instance was the actual hard part. Not the code — waiting on teammates to push their piece before you could test yours. Small, frequent commits made this bearable.
  4. A generic fallback message can hide a real bug. Ours briefly failed silently behind a "please try again" message before we found the underlying model issue. Log the real error, always.
  5. Memory only feels valuable when the "before" is genuinely broken. The clearest way to prove memory matters is a clean before/after with the identical input.

If you're building anything where a user comes back more than once, plain chat history isn't enough. Structured, queryable memory is a different thing, and it changes what the agent is capable of saying.

Top comments (0)