DEV Community

Misbah
Misbah

Posted on

Building ResolveIQ Lite: A Customer Support Agent That Remembers, with Hindsight

Most support chatbots forget you the moment you close the tab. Report a damaged order today, come back tomorrow, and you're explaining it all over again.

Our team of five built ResolveIQ Lite to fix that one annoyance. It's a small customer-support agent that remembers each customer's past incidents and uses them in later sessions, with Hindsight as the memory layer. It's a prototype with fictional customers and orders, but the memory behavior is real.

The problem

A customer reports: "My order A104 arrived damaged. I want a replacement."

Later they write: "I still haven't received my replacement."

Without memory, the agent has no idea which order or replacement this means. With memory, it can connect the message to A104, but it must not invent details like shipping dates or tracking numbers that were never provided.

Architecture

Customer message
      ↓
Streamlit UI
      ↓
Hindsight recall
      ↓
LLM (Groq) generates reply
      ↓
Hindsight retain
Enter fullscreen mode Exit fullscreen mode

The order matters: recall, then generate, then retain. If you store the current message before recalling, the system can treat the current request as old history.

Team and responsibilities

  • Memory layer (Hindsight): Shaik Afreen Firdose
  • LLM agent and prompt design: Anjali Jakka
  • Streamlit UI: Heena Meheraj
  • Data and QA: Muhammad Shazia
  • [Repo, Docs & Video Lead]: Misbah

Memory layer: one bank per customer

The key design decision was one Hindsight memory bank per customer instead of one shared bank:

def _bank(customer_id):
    return "resolveiq-" + customer_id.lower()

def retain_incident(customer_id, text):
    client.retain(
        bank_id=_bank(customer_id),
        content=text,
        context="customer support interaction",
    )

def recall_incidents(customer_id, query):
    try:
        results = client.recall(bank_id=_bank(customer_id), query=query)
        return [r.text for r in results.results]
    except Exception:
        return []
Enter fullscreen mode Exit fullscreen mode

C101 maps to resolveiq-c101, C102 to resolveiq-c102, and so on. Because customers never query the same bank, one customer's history can't leak into another's. Recall failures return an empty list so a missing memory doesn't break the whole support flow.

LLM layer: memory needs boundaries

The recalled memories go into the prompt under a PAST HISTORY section, kept separate from the current message. The prompt rules tell the agent to:

  • use past history when relevant
  • never pretend to remember something that isn't available
  • never invent tracking numbers, dates, or policies
  • never claim a refund or replacement was completed without evidence
  • ask for missing information when needed
  • keep replies short, warm, and specific

One failure shaped the prompt. The agent replied "Thank you for confirming it's order A104", but the customer had never confirmed anything. The model had retrieved A104 from memory and then treated it as something the customer just said. Retrieval wasn't the issue; the prompt needed clear rules separating what is known, what is remembered, and what the customer actually confirmed.

We also added model fallback: the agent tries the configured models in sequence, and <think>...</think> sections are stripped before the reply reaches the customer.

UI: making memory visible

The Streamlit app has a sidebar with a customer selector, a Hindsight memory ON/OFF toggle, and a "Start new session" button. The main area shows three panels: the current query, the generated response, and the recalled memory.

Showing recalled memory in its own panel turned out to be the most useful design choice. It makes clear why the agent answered the way it did.

Before and after, same message

We sent "I still haven't received my replacement." with memory off and on.

Memory OFF: the agent asks generically for the order number and item details.

Memory ON: the recalled-memory panel shows facts Hindsight synthesized, such as the customer's order A104 arriving damaged and the replacement not yet received. The reply builds on that history instead of starting cold.

Testing

We seeded fictional customers (C101 to C106), orders, and past incidents, then ran five scenarios: initial complaint, recall in a new session, cross-customer isolation, a new customer with no history, and an ambiguous follow-up like "Any update?"

The most important one was isolation. C102 sent the same message as C101, and Hindsight returned an empty list for C102. The agent asked for the order ID instead of mentioning A104.

Scenario Customer Expected Result
Report damaged order A104 C101 Log incident Pass
"Still haven't received replacement" C101 Recall A104 Pass
Same message C102 No memories returned Pass
"Where is my order?" C106 Ask for order ID, no invention Pass
"Any update?" C101 Link to pending replacement Pass

Problems we hit

  • Configuration: the Hindsight base URL and environment variables (HINDSIGHT_BASE_URL, HINDSIGHT_API_KEY, GROQ_API_KEY) were initially wrong or empty.
  • 402 Payment Required: memory operations stopped until account credits were available.
  • 404 on missing banks: recalling a bank that doesn't exist yet throws an error, so we handle it.
  • Interface mismatch: the UI's stand-in retain_incident() had a different signature from the real one in memory.py.
  • Silent failures: a generic "please try again" message briefly hid a real model error. Log the actual exception.

Limitations

  • Retrieval ranking: in one run, the reply still asked the customer to confirm the order number even though A104 was in their history. Hindsight recalled several memories, but the one naming A104 didn't rank first. That's a retrieval-tuning problem worth exploring.
  • Indexing delay: in early tests, recalling within a few seconds of a retain occasionally returned nothing.
  • Very short queries like "Status?" retrieved less reliably than queries with a bit of context.
  • Synthetic data: everything is fictional and pre-seeded. A real system would need integrations with platforms like Shopify or Zendesk.

What we learned

  1. Memory needs boundaries. More context doesn't automatically mean more accuracy.
  2. Isolate memory by key from day one.
  3. Always test with and without memory using the identical input.
  4. Test failure cases: missing tracking numbers, dates, and policies reveal invention fast.
  5. Prompt engineering is part of system design, because it defines what the agent can claim.
  6. Make recall visible in the UI so people can see why the agent answered as it did.
  7. Coordinating five people on one integration was harder than the code. Small, frequent commits helped.

Links

Top comments (0)