agent memory / hindsight / payecho
How Hindsight Taught My AI Agent to Remember What Actually Worked
The first time I tested it, the agent told an already-annoyed customer to check their email — for the third time. That's when I realized the problem was never the prompt.
A build log on giving a revenue-recovery agent real memory, instead of a database it had to guess about.
“Send an email reminder.” That was the agent's answer the first time I tested it — for a customer who had ignored three email reminders in a row and only ever responded to WhatsApp. I stared at that output for a minute, annoyed, because technically it wasn't wrong. It just wasn't useful. The agent had no idea it was repeating a mistake, because it had no idea there'd been a mistake at all. That's the moment the whole project stopped being about prompting and became about memory.
I was working on the memory layer for PayEcho, an agent that recommends what to do about overdue invoices — which channel to use for a reminder, when to follow up, and, for repeat cases, whether a customer's payment history should weigh into a new credit decision. It doesn't approve or deny anything itself; it recommends, with its reasoning attached, and a person decides. The rest of the team built the recovery engine, the dashboard, and the action triggers. My job was making sure that by the time a recommendation reached the screen, it was actually grounded in what had happened with that specific customer before, instead of a templated guess dressed up as AI.
The plain table wasn't the answer
My first instinct was the obvious one, and it turned out to be wrong. I logged every event to a plain table — customer ID, channel, response time, outcome — and before making a recommendation, queried it, filtered by customer, and dumped the rows into a prompt. It fell apart almost immediately, because a row like {channel: "whatsapp", outcome: "paid", days: 3} has no story in it. Handed a handful of rows like that, the model had to reconstruct what actually happened every single time, and it did that inconsistently — sometimes it caught the pattern, sometimes it just defaulted to “send an email,” which is exactly the output that started this whole thing.
That's when I actually looked into Hindsight properly, instead of treating it as a database with extra branding. It's a memory layer built specifically for agents, structured around two operations — retain and recall — rather than a table you're expected to query yourself, and the difference between those two operations turned out to be the whole story.
Retain: writing a note, not a row
Retain is where the shift happened. Instead of writing structured rows, I started writing memories the way a person would actually write a note:
def log_recovery_outcome(customer_id, invoice_id, channel, days_to_respond, outcome):
memory.retain(
agent_id="payecho-recovery",
user_id=customer_id,
content=(
f"Sent a payment reminder for invoice {invoice_id} via {channel}. "
f"Customer responded after {days_to_respond} days. Outcome: {outcome}."
),
metadata={
"channel": channel,
"invoice_id": invoice_id,
"outcome": outcome,
},
)
The content field is the part that changed everything — it reads like something a collections analyst would actually jot down, not a log line. Hindsight ties it to user_id (the customer) and agent_id (which agent wrote it), so memories don't leak across customers or across the other agents running in the system. The metadata is still useful for filtering and reporting, but it's the narrative content that gets recalled and reasoned over later, and getting that phrasing right mattered more than I expected going in.
Here is the text after the image.
def get_recommendation(customer_id, invoice_amount):
memories = memory.recall(
agent_id="payecho-recovery",
user_id=customer_id,
query="past recovery attempts, channel responses, and payment outcomes",
)
prompt = build_recommendation_prompt(
customer_id=customer_id,
invoice_amount=invoice_amount,
past_experiences=memories,
)
return llm.generate(prompt)
With that in place, the model sees a short, relevant slice of history instead of a wall of records — and that's the actual reason the output ends up specific instead of generic. The clearest proof of that wasn't a metric, it was watching identical code produce two different answers depending on what memory existed.
A fresh customer with no history still gets “Send an email reminder for the overdue invoice.” The same customer later — once Hindsight has retained an ignored email, a WhatsApp response, and a payment three days after a follow-up — gets something else entirely.
Nothing in the recommendation logic changed between those two calls — the only difference is what the agent could recall. That's the entire thesis of agent memory as a design pattern: an agent acting on retained experience is doing something structurally different from an agent reasoning from a stateless prompt, not just a better-worded version of the same thing.
Expected output: the PayEcho dashboard showing AI Customer Memory recall and an AI Recovery Assistant recommendation grounded in a customer's missed promise history
Expected output — recall surfaces the customer's history, and the recommendation is built directly on top of it, reasoning and tone included.
Expected output — recall surfaces the customer's history, and the recommendation is built directly on top of it, reasoning and tone included.
Two histories, kept apart on their own
What actually caught me off guard was a case I hadn't specifically coded for. A customer had two completely different histories on two different invoices — one paid fast after a phone call, another dragged on for weeks despite WhatsApp reminders. I expected the agent to average those out into one generic “this customer is medium-risk” recommendation. It didn't. Recall pulled back both experiences, and the recommendation actually kept them apart: it suggested the phone-call approach for the new invoice, while flagging the WhatsApp-dragging pattern as a reason to loop in someone before extending further credit.
I hadn't written any logic to keep those two threads separate — it fell out naturally from having two distinct memories to reason over instead of one blended score. The same idea carries into credit decisions more broadly: when a customer with a pattern of late payments requests new credit, PayEcho surfaces that history to whoever's approving it, instead of letting the request get evaluated as if it were the first interaction ever. The agent doesn't decide — it just makes sure the person deciding isn't working from a blank slate.
How it fits together
Retain and recall don't live in isolation — they sit inside the same loop as everything else PayEcho does: an interaction happens, it gets retained, the next decision recalls it, a recommendation goes out, and the outcome feeds the next recall.
What I'd tell myself on day one
A few things stood out looking back on how this came together.
Fewer, fuller memories beat more, thinner ones. I retained too much, too early — my first pass logged a separate memory for every micro-event (sent, opened, viewed), and recall came back cluttered enough that the recommendations went vague again, because the model was piecing together ten fragments instead of reading one clear note. Collapsing related events into a single memory per completed recovery attempt fixed that.
Query the decision, not a keyword. Early on I searched for literal terms like exact channel names and missed memories phrased slightly differently. Writing the query around the decision itself — “past recovery attempts and outcomes” — got noticeably better retrieval.
An empty recall needs an honest answer. Left unhandled, the agent's behavior on an empty recall was undefined, and once it produced a plausible-sounding “history” that didn't actually exist. Now an empty recall means the agent says so, out loud, and falls back to a general first-contact strategy.
Metadata is for machines; content is for the model. The rule that held up best throughout: treat metadata as data for filtering and reporting, and treat the retained content as narrative — the field that actually gets reasoned over needs to read like something a person wrote, not something a person would query.
The last thing worth saying is that a recommendation only earns trust if whoever's reading it can see what it's based on — putting the recalled history right next to the recommendation, instead of hiding it behind the output, mattered as much as getting recall itself right. If you're building anything where the same entity — a customer, a user, an account — should get different treatment depending on what's actually happened with them before, Hindsight's retain/recall model is a smaller, more precise tool than it looks at first. It's not about storing more. It's about deciding what's worth remembering, writing it so a model can reason over it, and pulling back only what's relevant to the decision in front of you.
resources
Hindsight on GitHub →
Hindsight Documentation →
Agent Memory — Vectorize →



Top comments (0)