<!-- Target query: do AI voice agents remember callers between calls --> <!-- dev.to title: Do AI Voice Agents Remember Callers Between Calls? No, Unless You Wire the Memory Yourself --> <!-- vilix.ai-blog title: Your Voice Agent Treats Every Caller Like a First-Time Caller. Cross-Call Memory Is the Fix --> <!-- dev.to slug (base): do-ai-voice-agents-remember-callers-between-calls --> <!-- vilix.ai-blog slug (base): voice-agent-treats-every-caller-like-first-time-caller -->
Do AI Voice Agents Remember Callers Between Calls? No, Unless You Wire the Memory Yourself
A customer calls your Vapi agent on Monday about order #4821. On Wednesday she calls again. The agent asks for her name, her order number, and what the problem is, like they have never spoken. She has spoken to it. It just does not remember.
That is the default. AI voice agents do not remember callers between calls unless you build the memory layer yourself. Here is how the platforms actually behave, and the pattern operators use to fix it.
What Vapi and Retell actually persist
Vapi keeps state inside a call. Its Workflows product stores variables extracted during the conversation and injects them back into prompts, so a 20-turn call holds together without losing the thread. Retell similarly retains context within a call. The moment the call ends, that working state is gone as far as the next call is concerned. Neither platform keeps a caller profile across calls out of the box.
Both give you the hooks to build it. Vapi fires server URL events: assistant-request when a call starts, where you can override the assistant config including the system prompt, and end-of-call-report when it finishes. Retell has webhooks for call start and call end. Bland and Pipecat-based stacks expose the same two moments. The pattern is identical on every stack:
- Identify the caller from call metadata, usually the phone number, sometimes an account ID from a CRM lookup.
- Recall at call start: fetch that caller's memory and inject the relevant slice into the prompt before the first word is spoken.
- Retain at call end: persist what happened, keyed to that caller.
- Repeat for every caller, in strict isolation.
The two-event memory pattern
The cleanest implementation is two events and one endpoint. On the call-start event, you look up the caller and inject a compact briefing into the system prompt. On Vapi this goes through assistantOverrides. This happens once, before the call connects, so recall does not add per-turn latency once the conversation is live. A recall budget in the low hundreds of milliseconds is comfortable here, because the caller is still hearing the ring.
On the end-of-call event, you take the transcript and retain it: a short summary plus the durable facts. Names, order numbers, stated preferences, open issues, promises the agent made. Fire and forget. If retention fails, you lose one transcript, not the call. A failing memory call should never break a live call.
Do not store raw transcripts as the memory itself. Store the summary and the facts, and keep the transcript archived separately if you need an audit trail. Retrieval at call start should pull "what happened with this caller" in a few hundred tokens, not replay the entire last call into a prompt that is already expensive.
Per-caller isolation is not optional
The moment a voice agent serves more than one caller, memory becomes a data-model decision, not a feature toggle. Every caller needs their own isolated memory, and those memories must never leak across callers. Key the store by phone number or account ID, and treat a lookup miss as a first-time caller, not an error. If your operation has any deletion obligation, keying per caller makes "delete everything about this person" a single operation instead of a scavenger hunt through a shared table.
What to store, and what not to
Store the caller's identity facts, open issues and their status, preferences they stated, promises the agent made, and a rolling summary of the last few interactions. Update facts when they change instead of appending contradictions. Last write wins keeps the memory honest, and retrieval should be recency-aware so the newest version is what the agent sees.
Do not store payment details spoken on the call, anything the caller asked you to forget, or internal debugging notes that would confuse the agent if injected into a prompt. Voice calls collect sensitive data casually, mid-sentence, without anyone typing it into a labeled field. Decide the redaction policy before the first production call, not after.
DIY or a hosted memory layer
You can build the two-event pattern yourself: a webhook handler, a Postgres table or a vector store, retrieval code, an isolation scheme, and a deletion path. It works. It is also another service to operate, and retrieval quality is the part most teams underestimate. Keyword matching finds the order number. Semantic search finds "the thing she complained about last time" when she phrases it differently this time. You want both in the same lookup, and that is real engineering.
The alternative is a hosted memory layer: one place where call memories live, with semantic plus keyword retrieval, per-caller isolation, and export and delete built in. Vilix AI is built for exactly this shape. It is cloud-hosted, so there is no infrastructure to run. It stores full conversation history, not just extracted facts, so the real exchange is there when you need it. Retrieval is semantic plus keyword, so it finds what the caller meant and the exact strings they used. Memory is isolated per account, and you can export everything in a portable format or delete individual memories or wipe the account instantly, anytime. The free plan is free forever, and the 7-day Pro trial needs no credit card. Because it is one memory over MCP, the same caller context your voice agent reads is visible to every other AI tool you connect, so the follow-up email your automation drafts and the note your support chat writes all start from the same memory instead of from zero.
Learn more at https://vilix.ai?utm_source=devto&utm_medium=article&utm_campaign=do-ai-voice-agents-remember-callers-between-calls
The bottom line
In text, a forgetful agent is an annoyance. In voice, it is rude. Callers expect to be known, and they judge the second call against the first one. The platforms give you the hooks. The two-event pattern gives you the shape. The only question is whether you operate the memory yourself or hand it to a layer built for it. Either way, stop letting your voice agent meet every caller for the first time.
Top comments (0)