The first version of my sales-deal assistant got better every time I pasted more deal notes into the prompt. It also got slower, more expensive, and worse at knowing which of those notes mattered. After the third rewrite of a prompt template that was mostly deal history, I stopped and asked a different question: why was I the one deciding what the model should see?
This is the story of replacing that prompt-stuffing with a real memory layer, using Hindsight, and what broke along the way.
What the system does
The tool answers one kind of question for a sales team: "We're in a deal with company X and they're worried about Y. What has worked before?"
It's a Streamlit app with three moving parts:
Hindsight holds everything the team has learned: deal registrations, stakeholders, objections, competitors, what worked, and outcomes.
Groq serves the LLM (an OpenAI-compatible endpoint, called through plain httpx).
Streamlit is the UI: a deal selector in the sidebar, a query box, and two buttons, Answer WITH Memory and Answer WITHOUT Memory.
That last pair of buttons is the most useful piece of the whole app. Every skeptical engineer I showed it to did the same thing: ask the same question both ways and compare. If memory doesn't visibly change the answer, it isn't doing its job, and I want that to be obvious rather than something I have to argue.
After a call, the rep types a free-text outcome into a "Log Call Outcome" box. That text is turned into structured fields by one LLM call, then retained into Hindsight. The next question about a similar objection can recall it.
The code is small on purpose. There are two modules that matter: llm_utils.py for the model and hindsight_utils.py for memory. The app file just wires them to widgets.
The through-line: memory is a write path, not a bigger prompt
My first design treated memory as a read problem. Collect all deal notes, rank them somehow, cram the top N into the prompt. That gave me three problems at once: token cost that scaled with history, a ranking function I had to maintain, and no clean answer to "how does new knowledge get in?"
Once I treated memory as a write path with its own recall semantics, the design got smaller. The app does two things with Hindsight: retain when something worth remembering happens, and recall when a question needs it. Everything else, including relevance, is the memory system's job.
The thin wrapper looks like this:
python
def retain_memory(bank_id: str, content: str):
"""Retain a memory unit in the specified Hindsight bank."""
if client is None:
print("Error: Hindsight client is not initialized.")
return None
try:
return client.retain(bank_id=bank_id, content=content)
except Exception as e:
print(f"Error retaining memory: {e}")
return None
That's it. No embedding code, no vector store to size, no chunking strategy. If you want the background on why an agent needs this layer in the first place, Vectorize's overview of agent memory is the clearest explanation I've found, and the Hindsight docs cover the retain/recall model in detail.
Decision 1: retain prose that reads like a fact, not a database row
Hindsight extracts facts from what you give it, so what you write matters. When a deal is added, I don't serialize a dict. I write a small document whose sentences are individually meaningful:
python
```content = (
f"Company Name: {company}\n"
f"Deal Registration: {company} is registered as an active deal in the portfolio.\n"
f"Stakeholders: {deal_data.get('stakeholders', '')}\n"
f"Objections: {deal_data.get('objections', '')}\n"
f"Competitors: {deal_data.get('competitors', '')}\n"
f"What Worked: {deal_data.get('what_worked', '')}\n"
f"Outcome: {deal_data.get('outcome', '')}"
)
Every line names the company. That's deliberate. A recalled memory gets pulled out of its surrounding document, so a line like Objections: migration downtime is useless on its own. Acme Corp ... Objections: concerns about migration downtime and SOC2 compliance is something the model can actually use, and something I can attribute to a deal when I show it in the UI.
I learned this the annoying way, by recalling fragments that had lost their subject and watching the answers get vague.
Decision 2: structure the outcomes at write time
Free-text call notes are how reps actually talk. "They pushed back on price again, CFO wants milestone billing, following up Thursday." I don't want to force a form on that, but I also don't want to leave every downstream feature to re-parse prose.
So the outcome goes through one extraction call before it's retained:
python
prompt = (
"Extract structured information from this call outcome into a JSON object.\n"
"Return ONLY a valid JSON object with these exact keys:\n"
"- sentiment: (e.g. Positive, Neutral, Negative)\n"
"- objection: (key objection raised, or 'None')\n"
"- decision_maker: (stakeholder or title involved, or 'Unspecified')\n"
"- next_step: (action item or next milestone, or 'None')\n\n"
f"Call Outcome: {outcome_text}\n\n"
"JSON:"
)
The parsing around it is defensive because models wrap JSON in code fences, or add a sentence before it, whenever they feel like it:
python
if "" in cleaned:
match = re.search(r"
```(?:json)?\s*([\s\S]*?)\s*```
", cleaned)
if match:
cleaned = match.group(1).strip()
json_match = re.search(r"\{[\s\S]*\}", cleaned)
if json_match:
cleaned = json_match.group(0)
If extraction fails, the function returns None and the raw text still gets retained. Structure is a bonus, never a gate. A failed parse should not mean a lost call note.
Decision 3: dedupe on recall, because memory returns near-duplicates
This one surprised me. When you retain the same deal in slightly different phrasings (a registration, then an outcome, then a follow-up), recall can hand back several memories that say nearly the same thing. Put those in a prompt and the model treats repetition as importance.
The fix is small and lives in the read path:
python
seen_texts = set()
deduped = []
for m in raw_results:
text = getattr(m, "text", None)
if text is None and isinstance(m, dict):
text = m.get("text")
if text:
text = text.split("|")[0].strip()
normalized = text.strip() if text else ""
if normalized and normalized not in seen_texts:
seen_texts.add(normalized)
deduped.append(m)
The split("|") matters. Recalled text can carry trailing metadata after a pipe (a "When: ..." suffix, for example), and two otherwise identical memories with different timestamps would slip past a naive comparison. Stripping the suffix before comparing is what made the dedupe actually work.
The part that hurt: using memory as a registry
Here's where I got clever and paid for it.
The sidebar needs a list of deals. I didn't want a second datastore for something so small, so I retained a "registry" memory (Deal Metadata Registry: Active Deals: Acme Corp, TechFlow Systems, ...) and recalled it to populate the dropdown, parsing the names back out with regexes.
It works. I would not do it again without a better reason.
python
registry_content = f"Deal Metadata Registry: Active Deals: {', '.join(current_deals)}"
retain_memory(bank_id, registry_content)
The trouble is that recall is a relevance operation, not a key-value lookup. I'm asking a system built to answer "what's related to this?" to behave like SELECT name FROM deals. The regexes that parse the answer back into names are the smell: two of them, one for the registry sentence and one for individual registration facts, because either might be what comes back.
I kept it because it made the system self-contained, and because the registration fact acts as a fallback if the registry memory doesn't surface. But the lesson is clear: use memory for what's true and relevant, and use boring storage for what's enumerable. If you need to list things, list them.
Aggregating across deals: the objection playbook
The feature that convinced me memory was worth it isn't a chat answer at all. It's the objection playbook, which recalls across the whole bank and groups what it finds:
python
CATEGORIES = {
"Pricing & Budget": ["price", "pricing", "cost", "budget", "upfront", "discount", "expensive"],
"Migration & Downtime": ["downtime", "migration", "cutover", "zero-downtime"],
"Security & Compliance": ["compliance", "soc2", "hipaa", "residency", "security", "vault", "baa"],
...
}
Each recalled memory is matched to a known deal, split into an objection and, where the text supports it, what resolved it, then bucketed by category. The output is a list like: Security & Compliance: Apex Healthcare, Acme Corp; what worked: shared HIPAA BAA certification, automated SOC2 audit reports.
The keyword categories are crude, and I say so. But they're crude in a place where crude is fine: a human reads the result. The value is that nobody hand-maintains this document. It falls out of what the team has already retained.
What it looks like in use
Take a new conversation with a prospect worried about downtime during migration. Asked without memory, the model returns competent, generic advice: run a pilot, plan a phased cutover, communicate clearly.
Asked with memory, the recalled context includes a past deal where the customer had the same concern together with SOC2 questions, and what worked was a two-week proof of concept showing zero-downtime cutover plus automated audit reports. The answer now names that approach, ties it to the compliance angle that came up in that account, and can suggest the same stakeholders to involve. The UI shows the recalled memories next to the answer, so a rep can see exactly what the advice rests on.
That last part matters more than it sounds. An answer you can trace to a specific past deal is one you can trust or challenge. An answer that's merely fluent is neither.
Then the loop closes. The rep logs how the call went, the outcome is structured and retained, and the next question can recall it. Nobody edits a prompt.
What I learned
Stop hand-curating context. If you're deciding, per query, which notes go into the prompt, you've built a bad retrieval system by hand. Give the job to something designed for it, like Hindsight, and spend your time on what you retain.
Write memories to survive being alone. A recalled fact is stripped of its document. Put the subject in every sentence.
Make the with/without comparison a first-class feature. A/B buttons cost almost nothing and turn "I think memory helps" into something you and your users can see.
Dedupe the read path, and normalize before you compare. Near-duplicate memories are normal. Metadata suffixes will defeat exact-match dedupe unless you strip them first.
Don't make memory a database. Enumerating records through relevance search works until it doesn't. The regexes I wrote to parse deal names back out were the warning sign I ignored longest.
I also made two smaller choices I'd repeat: model name comes from an environment variable, because the model I originally picked disappeared from the provider and returned a 404 one afternoon; and every external call returns a readable error string instead of raising into the UI.
The bigger shift was conceptual. I used to think of the agent's quality as a function of how much context I could fit. Now I think of it as a function of what the system has been allowed to learn, and whether I can see it doing so.
Top comments (0)