I wanted to ask one question: if an agent gets the same overdue invoice and uses exactly the same language model, how much can its suggestion change when the only new thing is customer memory?
So I created PayRecall, a test for B2B accounts receivable. It does not send emails make calls or link to an accounting system. Given an invoice it suggests the next step to take: **who to talk to which way to communicate what tone to use, when to act and how to start the conversation.
The part that is interesting is the controlled test. The memory-off path sees a simple customer profile. The memory-on path sees the invoice and model but also gets information about the customer from Hindsight.
*Why the Baseline Keeps Making the Same Mistake
*
An overdue invoice has details like the amount, due date, customer and contacts. It does not have the history that shows whether a collection strategy will work.
For example previous emails to a shared accounts inbox may not have been read calls to an AP decision-maker may have led to a promise to pay or a strong message may have started a disagreement. The person who was contacted before might have left the company.
A normal AI does not know any of this unless that information is in the prompt. Of adding historical messages each time PayRecall uses Hindsight as the memory part that keeps finds, puts together and thinks about customer history.
What Hindsight Brings
For the memory-on path PayRecall finds customer-specific memories using tags. It also uses a customer model and looks at the same information again.
The design keeps memory separate from the decision part. A special memory part handles Hindsight. The decision agent uses the information that is found.
Customer tags also help keep things separate: a question about one customer should not accidentally find another customers history.
When the Suggestion Changes
The example uses Meridian Retail. An overdue invoice, INV-2041.
Without memory the model can suggest an action: contact the accounts-payable contact by email use a polite tone and ask for a payment date.
With memory the situation changes. Previous records show that emails to the shared accounts inbox were not read calls to the AP decision-maker were successful a strong message led to a dispute and the old contact left. A new AP contact also has a time when they can be reached.
The memory-based suggestion can then be a call to the contact during their available time using a friendly rather than strict way.
The main point is not that calling is always better than emailing. The point is that the agent now has proof that this customer might need a way.
From Past to a Customer Plan
The data has 35 stories covering several months. They look like collection notes: emails, call summaries WhatsApp attempts, promises to pay, arguments and changes in contacts.
Of turning every note into a fixed format PayRecall keeps the stories and lets Hindsight find useful facts.
Repeated events can then become higher-level trends. For example one event says an email was sent. A customer-level trend might show that the shared accounts inbox often does not get a response.
This information helps build the customers model, which is updated when new results come in.
The Agent Can Learn From New Results
After a suggestion PayRecall lets a person record what happened. The system saves the agents action and the customers response separately.
Hindsight can then find facts from both update the customer model and affect the suggestion.
One test taught a lesson: I said the customer preferred email. Later suggestions started to use that preference. The system was not going wrongβit had learned what I told it.
This means memory tests must be treated like database tests. Test data can become part of the agents actions and must be found kept separate and cleaned up.
What the Test Shows
The strongest result is not one suggestion. It is the difference between two tests:
Same invoice. Same model. Same process. Different facts available.
Without memory the model suggests an action. With customer memory it can use results changes in contact timing preferences and old strategies that did not work.
That is the role I wanted Hindsight to have in PayRecall: not another prompt with old text but a long-term part, between what happened before and what the agent decides next.
PayRecall is still a small model using fake data and human-entered results.. The test shows the main idea clearly: memory can change an agents actions without changing the model itself.
Hindsight GitHub: https://github.com/vectorize-io/hindsight documentation: https://hindsight.vectorize.io/ Vectorize agent memory: https://vectorize.io/what-is-agent-memory




Top comments (0)