The worst call my voice agent ever made went like this: "Hi Imran, Aisha from TableTap. Our billing app starts at just nine hundred ninety nine rupees a month." Imran runs a cloud kitchen. It was one o'clock, he had forty orders on screen, and he hung up after two sentences with a clear instruction: don't call during lunch rush.
A human sales rep would remember that forever. My agent, like almost every AI calling agent I've seen, forgot it the moment the call ended. The next cloud kitchen owner got the same price-first pitch at the same bad hour.
So I gave it a memory. Now every call, good or bad, becomes something the agent reads before it dials again.
What the system does
Aisha is an outbound voice sales rep for TableTap, a billing app for Indian restaurants, cafes and cloud kitchens. Her only job on a call is to book a twenty-minute demo.
The stack is small:
- Voice: LiveKit Agents with AssemblyAI for speech-to-text, Gemma 4 31B as the LLM, the LiveKit turn detector, and Sarvam's Bulbul v3 for an Indian English voice.
- Memory: Hindsight, an open-source agent memory system from Vectorize, holding every past call.
- Frontend: a Next.js dashboard that shows what the agent remembers about the prospect before the call, the live transcript, and the outcome after.
The flow has three steps. Before the call, recall fetches what this person told us and reflect works out what works with people like them. During the call, the LiveKit pipeline talks and a record_call_outcome tool labels the result. After the call, retain writes the transcript and outcome back.
The core idea: remember the person, learn the persona
When I started, I thought "agent memory" meant one thing: remember what this user said last time. For a sales agent that is only half the value.
Take Rahul, who owns a cafe in Bengaluru. On the first call he told Aisha his billing software crashes every weekend, he's opening a second outlet in November, his staff "are not tech people", and a competitor quoted him fifteen hundred rupees a month. He asked for a demo video on WhatsApp and a weekday callback after the twenty-fifth. That's person memory, and it makes the callback sound like a follow-up instead of a cold call.
Now take Sana, who runs a cloud kitchen in Hyderabad. Aisha has never spoken to her, so there is no person memory at all. But Aisha has spoken to Imran, Neha, Abdul and Pooja, all cloud kitchen owners. Imran hung up on a price-first pitch. Abdul hung up on a feature tour ("My customers are Swiggy's customers, not mine. What loyalty program?"). Neha and Pooja booked demos after Aisha asked about missed Swiggy and Zomato orders at rush hour. That's persona memory, and it's what lets a first call go well.
I did not want two storage systems for this. The trick that made it simple was one Hindsight bank and tags.
One bank, scoped by tags
Every call is saved with three tags:
def prospect_tags(prospect: dict) -> list[str]:
return [
f"prospect:{prospect['id']}",
f"persona:{prospect['persona']}",
f"industry:{prospect['industry']}",
]
Before a call, the agent asks two different questions of the same bank. For the person, it uses Hindsight's recall, filtered strictly to that prospect's tag:
response = await client.arecall(
bank_id=BANK_ID,
query=(
f"What happened in past calls with {prospect['name']} at "
f"{prospect['company']}: pains, objections, competitors, promises "
"and agreed follow-ups?"
),
tags=[f"prospect:{prospect['id']}"],
tags_match="any_strict",
max_tokens=1500,
types=["world", "experience", "observation"],
prefer_observations=True,
)
For the persona, it uses reflect, filtered to the persona tag. reflect doesn't just return matching facts; it reasons over them and writes an answer:
response = await client.areflect(
bank_id=BANK_ID,
query=(
f"Across our past calls with a {label}, which opening angle and value "
"points led to booked meetings, which objections came up and what "
"answer worked, and what must the rep avoid? Reply as at most six "
"short, direct coaching bullets."
),
tags=[f"persona:{persona}"],
tags_match="any_strict",
budget="low",
)
any_strict matters. Without it, untagged memories can leak into the result, and Rahul's cafe history would show up in a cloud kitchen playbook. With it, one bank behaves like many narrow ones, and I can add a new axis (city, deal size) by adding a tag instead of a new bank.
The bank is configured with missions that tell Hindsight what to extract from each transcript and how to reason. My favourite line in the codebase is the reflect mission: "You are a blunt sales coach reviewing past calls. Base every point on evidence from calls, prefer what led to booked meetings, and call out approaches that got prospects to hang up." The Hindsight docs cover these bank-level missions, and they did more for output quality than any prompt tweak I made on the agent side.
Memory must never block the phone
The rule I care most about: memory makes a call better, but it must never stop a call from happening. Recall and reflect run in parallel, each with a timeout, and any failure falls back to an empty brief:
async def guarded(coro, default):
try:
return await asyncio.wait_for(coro, timeout)
except Exception:
logger.exception("memory lookup failed, continuing without it")
return default
history, playbook = await asyncio.gather(
guarded(recall_prospect_history(client, prospect), []),
guarded(reflect_persona_playbook(client, prospect["persona"]), ""),
)
The brief is appended to the system prompt as two sections, "What you remember about Rahul" and "Playbook learned from past calls with a Cafe Owner". The agent is told never to reveal it has a memory, because "According to my notes…" is a terrible thing to hear on a sales call.
A small script, show_brief.py, prints exactly what the agent reads before it dials. This is Rahul's brief, straight from Hindsight:
Closing the loop: retain after every call
During the call, the LLM calls a record_call_outcome tool as soon as the result is clear: meeting_booked, follow_up, callback_requested or not_interested, plus its opening angle and the objections it heard. When the call ends, a shutdown hook saves the transcript and outcome into Hindsight with retain, under the same three tags.
That tool is the quiet hero. A raw transcript tells Hindsight what was said. The outcome label tells it whether it worked. Without the label, reflect can describe what cloud kitchen owners talk about; with it, reflect can say which openers get meetings and which get hang-ups.
Before and after
Here is the same callback to Rahul with memory off and on. The dashboard has a switch for exactly this.
Memory off. Aisha opens with the generic pitch about faster billing and daily sales on WhatsApp, then asks how he bills today. Rahul: "You called me three weeks ago. I told you all this already." She apologises and starts discovery again. A real owner hangs up here.
Memory on. The pre-call panel shows everything Hindsight recalled from the first call. Aisha calls on a weekday, asks whether the video helped, and brings up the new outlet. When Rahul raises his staff again, she offers on-site setup and ten-minute training, the answer that won over another cafe owner who had the same worry. He books a demo.
For Sana, a first call, the cloud kitchen playbook says: open with missed Swiggy and Zomato orders at rush hour, never lead with price or features, never call during lunch rush, and answer "Swiggy's tablet is free" with the cost of juggling two tablets. Each point traces back to a specific earlier call. Nobody wrote that script. It was learned from what worked and what didn't.
Lessons learned
1. Split memory into "this person" and "people like this". Per-user recall is the obvious feature. Per-segment reflection is where the agent actually improves. Tags in one bank gave me both.
2. Label outcomes at the source. Have the agent record the result with a tool call while the call is still happening. Inferring "did this work?" from transcripts later is much weaker than an explicit label.
3. Treat memory as optional at runtime. Timeouts plus empty defaults mean a slow memory lookup costs you personalisation, not a call. In a voice product, a phone that doesn't ring is worse than a generic pitch.
4. Expect duplicates and clean them. The same fact ("don't call on weekends") can come back from more than one call. prefer_observations=True lets Hindsight return its consolidated observations instead of every raw fact, and I still dedupe the text before it goes into the prompt.
5. Small samples produce confident advice. This is the honest limitation. With a handful of calls per persona, reflect will turn two data points into a rule. The evidence-first mission helps, but the playbook is only as good as the calls behind it, and a brand-new persona gets no playbook at all.
Where this goes next
The part that surprised me most is how little sales logic I wrote. I wrote down what happened on each call, labelled whether it worked, and Hindsight's agent memory turned that into two kinds of knowledge: what this person told us, and what works with people like them. Call one is generic. Call five is personal. By call fifty, its playbook comes from its own market instead of the script I started with.
If you're building any agent that talks to the same people more than once, or to many people of the same kind, try giving it memory before you give it a longer prompt. The Hindsight repository on GitHub is a good place to start.
The full code for this agent, including the Hindsight memory module and the dashboard, is on GitHub: tabletap-sdr-hindsight.





Top comments (0)