DEV Community

member_b47cac70
member_b47cac70

Posted on

Hindsight Memory Made My Support Agent Stop Saying "Restart Your Router"

Ravi has had slow internet three times this year. The first time, switching his laptop to the 5 GHz network fixed it. The second time, it was a weak signal upstairs. The third time, a power cut had reset his router. Each time he contacted support, he got the same five steps: run a speed test, restart the router, check the cables, count your devices, move the router.

He was annoyed by the second time. By the third, he was ready to switch providers.

That is the problem I wanted to fix. Not "make the chatbot smarter", but something much more specific: a support agent should never make a customer repeat themselves. So I built SupportMind, a support agent for an internet provider that remembers every customer, every fix that worked, and every promise the company made. The memory layer is Hindsight, an open-source agent memory system, and it ended up shaping the whole design.

Same message with memory off, showing generic steps and an empty memory panel

What the system does

SupportMind is a small Python app with three parts:

  • A Streamlit chat screen. On the left, a customer chats with the agent. On the right, a panel shows exactly which memories the agent recalled for that answer.
  • An LLM on Groq (openai/gpt-oss-120b, with qwen/qwen3-32b as a fallback) that writes the replies.
  • Hindsight as the memory. Every customer message triggers a recall() before the LLM sees it, and every exchange is stored with retain() after.

The data is a fictional fiber provider, SwiftNet, with 15 customers across Andhra Pradesh and Telangana and 27 historical tickets: routers resetting after power cuts, duplicate OTT charges, gamers with ping spikes, a retired teacher who just wants her Sunday video call with her son to work.

There is also a Memory switch in the sidebar. With it off, the agent only sees the account basics: name, plan, router model. That is roughly what most support bots see today.

The core decision: two kinds of memory

My first version had one memory bank for everything. It recalled the wrong customers' problems, and it was hard to reason about. I rewrote it around a simple split:

  1. One bank per customer. Their history, preferences, moods, and promises made to them. Nothing leaks between customers.
  2. One shared playbook bank. Fixes that actually worked, tagged with the router model. This is how a fix found for one customer helps the next.

In code, the whole memory layer is a small file. Recall looks like this:

def recall_customer(customer_id: str, query: str, limit: int = 8) -> list[dict]:
    """What do we remember about THIS customer that is relevant to their message?"""
    response = get_client().recall(
        bank_id=customer_bank(customer_id),
        query=query,
        budget="mid",
        max_tokens=2048,
    )
    return _to_list(response, limit)
Enter fullscreen mode Exit fullscreen mode

And the "learning" part is just two retain() calls when an agent clicks Mark issue as resolved:

def retain_resolution(customer_id, name, router, issue, fix):
    now = datetime.now()
    get_client().retain(
        bank_id=customer_bank(customer_id),
        content=f"Issue for {name} was resolved. Problem: {issue}. Fix that worked: {fix}.",
        context="ticket resolved", timestamp=now, retain_async=True,
    )
    get_client().retain(
        bank_id=PLAYBOOK_BANK,
        content=f"On router {router}, the problem '{issue}' was fixed by: {fix}.",
        context="confirmed fix", timestamp=now, retain_async=True,
    )
Enter fullscreen mode Exit fullscreen mode

The fix itself is summarized from the chat by the LLM, so the human agent does not have to type anything.

Notice what I am not doing: I am not writing extraction logic, embedding pipelines, or deduplication. I pass Hindsight raw text like "Ravi reported slow speed after a power cut, restarting and reconnecting to 5 GHz fixed it in 5 minutes", and it extracts the facts, entities, and dates itself. That is the part I would have spent a week on otherwise.

Before and after, on the same message

Here is Ravi typing My internet is slow again.

Memory off:

Hello Ravi, I'm sorry you're experiencing slow speeds. Let's try a few quick checks: 1. Speed test... 2. Restart the router... 3. Check cables... 4. Device load... 5. Location...

Five generic steps. Technically correct and completely useless to someone who has already done all of them.

Memory on:

Hi Ravi, I'm sorry you're seeing slow speeds again. This often happens when the router falls back to the 2.4 GHz band or resets its channel after a power cut. Let's try the same fix that worked before: power-cycle the TP-Link Archer C6, then connect to the 5 GHz Wi-Fi "SwiftNet_5G"... If speeds are still low, we can discuss a mesh extender for the upstairs area.

Same model, same prompt template. The only difference is what Hindsight recalled: the June 5 GHz fix, the September power-cut reset, and the August conversation about a mesh extender for his first-floor room. Ravi also prefers short instructions, so the answer is three steps instead of five.

SupportMind with memory on, showing past fixes and recalled memories

The other customers show different sides of the same idea:

  • Priya was double-charged in July and promised a 10% loyalty discount that never appeared. Ask "Why is my bill still wrong?" without memory and you get a generic billing FAQ. With memory, the agent apologizes for the missed discount first, because that is what she is actually angry about.
  • Lakshmi is 67, not technical, and uses the internet mainly for Sunday calls with her son. With memory, the agent slows down and drops the jargon.

The part that surprised me: observations

I expected Hindsight to give me back the facts I put in. What I did not expect was the observation memories.

After seeding Ravi's three tickets, the recall panel showed a memory I never wrote:

Ravi Kumar experienced internet connectivity issues, including slow speeds and Wi-Fi drops, which caused him frustration and annoyance; on September 10, 2026, a power cut reset his router to default channel settings...

Hindsight had consolidated three separate tickets into one belief about Ravi, with the dates intact. Then, after a live chat, the same observation grew a new clause: "on September 28, 2026, he again reported slow speeds and was instructed to power-cycle his router." I did not write any code for that. The chat was retained with retain_async=True, and a few seconds later the customer's history had updated itself.

This is the difference between memory and a transcript log. A log grows forever and you have to search it. Agent memory summarizes, merges, and keeps the timeline straight.

Watching it learn across customers

The demo I care most about goes like this:

  1. Karthik, on a TP-Link Archer C6, says his uploads are stuck. In chat, changing DNS to 8.8.8.8 fixes it. The agent clicks Mark issue as resolved.
  2. A minute later, Ravi, on the same router model, says "My uploads are stuck."
  3. The agent suggests a firmware update (from Karthik's June ticket) and the DNS change (from the chat that just happened).

Ravi never had an upload problem. The agent is using what the team learned from somebody else. That is the thing a human support team does naturally in a shared Slack channel, and what a stateless bot can never do.

Lessons learned

1. Split memory by who owns it. Per-customer banks for personal history, a shared bank for team knowledge. It made recall far more accurate, and it gives you privacy boundaries for free.

2. Tell the model where each memory came from. My first version of the prompt put customer history and playbook fixes in the same context. The result: the agent told Ravi "let's apply the fix that cleared this before" for a problem he had never had. It had confused another customer's fix with his own history. The fix was two lines in the system prompt:

"- The first list is THIS customer's own history. The second list is from OTHER customers.\n"
"- Only say 'last time' or 'before' if the problem appears in this customer's own history.\n"
Enter fullscreen mode Exit fullscreen mode

Memory makes an agent more confident. That means provenance matters more, not less.

3. Make memory visible. The right-hand panel was originally a debugging aid. It turned out to be the most convincing part of the product. When the agent says "same fix as before," you can see the June 14 ticket it is referring to. Support leads trust what they can audit.

4. Store raw events, not your own summaries. I retain full chat turns and ticket text, and let Hindsight extract facts and build observations. Every time I tried to pre-summarize, I threw away something that later turned out to matter, like a customer's mood.

5. Never block the customer on a write. retain_async=True keeps replies fast. Recall is on the critical path; retain is not.

Try it

The code is on GitHub:

https://github.com/lokeshgavara1/SupportMind

You need a Hindsight Cloud account or a self-hosted instance and a Groq key. Run python seed_memory.py to load the 15 customers, then streamlit run app.py.

The next thing I want to add is a Hindsight mental model per customer, a live "customer summary" a human agent can read in five seconds before picking up a call. The raw material is already in memory. It just needs to be surfaced.

If you have ever had to explain your problem to support for the third time, you already know why this matters.

Top comments (0)