DEV Community

TejaAndhoju
TejaAndhoju

Posted on AI-assisted

Hindsight Recall Was Easy. Deciding What to Retain Wasn't.

My sales agent told a stalled CFO deal to try milestone-based payments. It got that idea from a different company's closed deal, weeks earlier, which nobody had mentioned in the conversation. Recall worked on the first try. What took me a lot longer was deciding what the agent should be allowed to remember.

This is a write-up of how I wired Hindsight agent memory into a FastAPI sales assistant, and the design mistake I only noticed when I read my own save endpoint out loud.

What the system does

DealMind is a small assistant for B2B sales reps. A rep opens a deal that has stalled, asks "what should we do?", and gets a strategy back. The pieces:

  • SQLite holds the CRM: 20 sample deals with stage, value, contact role, blocker, and a success_reason for the ones that closed.
  • FastAPI exposes /api/chat as a server-sent-events stream, plus endpoints to save a strategy and update a deal's outcome.
  • Hindsight is the memory layer. I use its Python client to recall past strategies and retain new ones.
  • An LLM on Groq (openai/gpt-oss-120b) writes the final answer from the CRM context plus whatever memory came back.

The CRM answers "what is this deal?" Hindsight answers "have we solved this before?" Those are different questions, and keeping them in different systems turned out to matter.

Why not just put the closed deals in the prompt?

I could have pasted every success_reason into the system prompt. With 20 deals that works. With 2,000 it doesn't, and even at 20 it hides the real problem: the model gets no signal about which past win is relevant to this blocker.

Hindsight does the matching. I hand it a query and it returns the memories that are semantically closest. I never write the "if blocker mentions payment terms, look up the payment playbook" logic myself, which is the part I would have gotten wrong.

Recall: the blocker is the query

The first design decision was what to search on. The user's question ("what should we do?") is nearly useless as a query. The useful text is the deal's blocker, so I append it:

search_q = query
if active_deal and active_deal['blocker']:
    search_q += f" {active_deal['blocker']}"

memories = await search_hindsight(search_q, limit=3)
Enter fullscreen mode Exit fullscreen mode

For Stark Industries the blocker in the CRM is "Pushing back hard on upfront costs. Wants Net-90." The call underneath is short:

api = get_hindsight_api()
req = RecallRequest(query=query)
res = await api.recall_memories(
    bank_id=BANK_ID,
    recall_request=req,
    authorization=f"Bearer {HINDSIGHT_API_KEY}",
)
Enter fullscreen mode Exit fullscreen mode

The memory bank contains a strategy written in completely different words, from a company that isn't Stark:

When a prospect pushes back hard on upfront costs and wants Net-90 terms (like Oscorp), structure a milestone-based payment plan tied to deployment phases.

No keyword overlap on the company name, no shared deal ID. It matched on meaning. The recalled text goes into the system prompt under a "what worked in the past" heading, and the model tailors it to the current deal.

I also stream a visible "memory match" message to the UI before the LLM starts, so the rep can see the recalled playbook itself and not only the model's rewrite of it. In a sales tool, being able to check the source matters more than a polished paragraph.

The before and after

Without memory, the prompt falls back to rule 2 in my system prompt: "generate a fresh, highly tactical market-based sales solution." You get the advice any LLM gives about discounts and negotiation. It's fine and interchangeable.

With memory, the model has a concrete precedent, including the mechanism (payments tied to deployment phases) and why it worked (it removed the upfront risk). The output stops being generic and starts referencing something the company has actually done.

I haven't run a controlled comparison, so I'm not going to put a number on "better." What I can say is that the two answers are visibly different, and the second one is traceable to a specific memory.

Retain: where I got it wrong

Retaining is symmetrical to recalling, and that symmetry made it look easy:

content_text = f"Novel solution successfully applied for {deal_company}: {solution}"
item = MemoryItem(content=Content(actual_instance=content_text))
req = RetainRequest(items=[item])
await api.retain_memories(bank_id=BANK_ID, retain_request=req, authorization=...)
Enter fullscreen mode Exit fullscreen mode

The UI shows an "Accept Strategy & Save to Hindsight" button after the model proposes a fix for a stalled deal. Clicking it retains the text and moves the deal to a "Strategy Applied" stage:

conn.execute(
    "UPDATE deals SET stage = 'Strategy Applied', success_reason = ? WHERE company = ?",
    (req.solution_text, req.deal_company),
)
Enter fullscreen mode Exit fullscreen mode

Read that back. The memory says "successfully applied." The moment of saving is when a rep agrees with the advice, not when the deal closes. I'm storing a hypothesis under a label that says it's a result.

That is a real bug in how the memory bank behaves over time. If a rep accepts three plausible-sounding strategies and two of the deals die anyway, the next rep who asks a similar question gets those strategies recalled with the same weight as the ones that actually closed. Recall doesn't know the difference, because I never told the bank.

There's already an /api/update_outcome endpoint that marks a deal Closed Won or Closed Lost. It updates SQLite and never touches Hindsight. The fix I'm working toward is to split retention into two events:

  1. On acceptance, retain the strategy as proposed, clearly labeled.
  2. On outcome, retain a second memory that links the result to the strategy: worked, or didn't, and for which deal.

Then a recall for a payment-terms blocker can return the playbook along with its track record.

Cold start

An empty memory bank gives you nothing to recall, and a brand-new deployment is empty. I wrote a small script that retains four playbooks drawn from the sample deals that had closed, covering upfront costs, migration downtime, vendor lock-in, and on-prem requirements. It's honest seeding: those strategies came from the CRM data, not from thin air. But anyone reading this should know that the first recall in a fresh install works because I planted the memory, not because the agent learned it. Real learning starts after that first batch of accepted, resolved deals.

Things that bit me

  • Unclosed HTTP sessions. The generated client uses aiohttp, and I was creating a new client per request without closing it. Sockets leaked until I added an explicit await api.api_client.close() on both the success and error paths. If you build on any async SDK, check this first.
  • Silent failures hide missing memory. My first integration swallowed an import error, so the agent was answering without memory and looked fine. Now search_hindsight logs a warning when the key is missing, and the UI shows whether memories were found. An agent that quietly loses its memory is the worst failure mode because the output still sounds confident.
  • Sample data resets on restart. My database initializer reloads the CRM every time the app starts, so local stage changes disappear. Hindsight persists, the SQLite copy doesn't. That inconsistency is next on my list.

Lessons

  1. Query with the problem, not the question. The user's words were the weakest search text I had. The blocker from the CRM was the strongest.
  2. Show the recalled memory, not just the answer. Reps trust a strategy more when they can read the precedent behind it.
  3. Retain outcomes, not agreement. "The user liked this advice" and "this advice worked" are different facts. Store them separately.
  4. Be upfront about seeding. A pre-loaded bank is fine for cold start, but say it, and don't present it as learning.
  5. Make missing memory loud. Log it and surface it, so you notice when the agent is running blind.

The code is on GitHub. If you're deciding whether to add memory to an agent, start with Hindsight and read the documentation. Then spend more time than you expect on what gets written into the bank, because the retrieval side will mostly take care of itself.

Top comments (0)