For five months, a logistics company added a separate shipping charge to every bill it sent us, and for five months someone on the finance team approved it. Then on August 1 the contract changed: shipping was now supposed to be included in the price. The interesting question wasn't whether my AI agent could remember five approvals. It was whether it could notice that one newer fact made all five of them wrong.
This is the story of building an accounts payable exception agent on Hindsight agent memory, and what I learned about designing memory that has a sense of time.
Invoice checking in 60 seconds
If you've never worked near a finance team, here's all the background you need.
When a company buys something from a supplier, it first creates a purchase order: a document saying what it agreed to buy, how many, and at what price. Later the supplier sends an invoice, which is the bill. Before anyone pays, the accounts payable team (the people who pay the company's bills) checks that the invoice matches the purchase order.
Think of ordering furniture. You agree on a sofa for $800, delivery included. If the bill arrives saying $800 plus $90 delivery, something doesn't match. In finance, that mismatch is called an exception, and a person has to decide what to do: pay it, send the bill back, or ask someone senior.
Here's the part that makes it a memory problem: most exceptions aren't new. The same supplier adds the same shipping charge every month. Another supplier writes its own order number on the bill instead of yours. A third prints the wrong payment deadline. Experienced reviewers know which of these are harmless habits and which are real problems. That knowledge lives in their heads, and it leaves when they do.
What the agent does
The agent handles each invoice in five steps:
- Detect mismatches with plain Python, comparing each invoice line, price and tax amount against the purchase order.
- Remember by asking Hindsight how similar mismatches from this supplier were handled before.
- Decide with an LLM (running on Groq), which turns the mismatches plus the remembered history into a structured recommendation.
- Guardrail: the agent may only approve on its own when it has remembered evidence, high confidence and at least two past cases to point to.
- Learn: when a reviewer approves, rejects or escalates an invoice, that outcome goes back into memory.
Why an agent without memory fails here
Give an LLM a bill from our logistics supplier, Meridian, with a $450 shipping charge that isn't on the purchase order, and no history. It does the only sensible thing: it flags the charge and asks a human. That's safe, but a simple rule could do the same.
What an experienced reviewer adds is context. In June they'd say: "Meridian always does this, it's fine." In September they'd say: "Meridian used to do this, but their new contract says shipping is included, so this charge is wrong." Without memory, the agent can't tell those two situations apart.
The memory design
I made three decisions early that shaped everything else.
One memory bank, organized with tags. In Hindsight, a bank is a separate store of memories, and a search only looks inside one bank. One bank per supplier would have been tidy, but it would have hidden company-wide rules that apply to everyone. So there's one bank for the whole finance team, and every memory is labeled with tags like vendor:meridian, the type of mismatch, and the outcome.
Memories are sentences, not raw data. Each past invoice decision becomes one plain-English memory, dated with the day it was decided:
memories.append({
"doc_id": f"res_{inv['invoice_number']}_{inv['invoice_date']}",
"content": header + body + verdict,
"context": "AP exception resolution",
"timestamp": datetime.fromisoformat(res["resolved_date"]),
"tags": [f"vendor:{inv['vendor_id']}",
f"exception:{res['exception_type'] or 'none'}",
f"decision:{res['decision']}"],
})
A typical memory reads: "Invoice MER-2205 from Meridian Logistics, dated 2026-07-13, total $3,885.94. Exception: shipping charge of $469.00 not on the purchase order. Reviewer Priya Nair approved it on 2026-07-15." Hindsight pulls out the names, dates and facts by itself.
Setting the real date on each memory turned out to be essential. Without it, every memory would look like it happened on the day I loaded them, and a question like "has anything changed recently?" would be meaningless.
The bank has a mission. Hindsight lets you describe what a memory bank is for. Mine says to pay attention to past decisions and the reasons behind them, and that a newer contract or policy always overrides older habits.
Recall and reflect do different jobs
Hindsight offers two ways to use memory, and understanding the difference was the most useful thing I learned.
tags = [f"vendor:{inv['vendor_id']}", "policy"]
recalled = self.memory.recall(**kwargs, tags=tags, tags_match="any")
# ...
answer = self.memory.reflect(bank_id=BANK_ID, query=question, budget="mid",
tags=tags, tags_match="any")
recall is like searching your notes: it returns the individual past invoices and facts that match. I show these on screen so a reviewer can check the agent's homework. reflect is like asking a colleague who has read all the notes: you ask a question and get a reasoned answer. My question is "what do past decisions say about this mismatch, and has anything changed that overrides them?" That answer is what the agent's final decision relies on.
The contract change test
Before building the full agent, I wrote a small test. It stored four memories (three approved shipping charges and the August contract change), then asked reflect about a new September bill with a shipping charge. I expected to need extra rules to make it work. I didn't. Reflect said the bill should not be approved, listed the three earlier approvals, and explained that the August contract made them no longer apply, because the new bill came after the change.
In the full agent, with 66 past memories loaded, the decision for the September Meridian bill read:
The August 1 2026 contract mandates freight be included in the PO unit price, so a separate $450 freight line violates current terms; recent invoices show compliance with this rule.
(Freight is the finance word for shipping, and PO means purchase order.) It pointed to the contract change plus the two bills since August that followed the new rule. The same model with the same instructions but no memory said only: "No prior history for Meridian Logistics and the freight charge is not on the PO, requiring human review." The right outcome, but no understanding of why.
Keep the LLM away from arithmetic
Finding mismatches is deliberately plain Python. Comparing amounts is not a job for a language model, and doing it in code gives the model a clean, correct list of what's wrong.
Even so, testing caught a bug. My first version flagged the Meridian bill twice: once for the shipping charge and once for a $36 tax difference. But the tax only differed because tax was being charged on the extra shipping line. Reporting it separately was noise, and a "$36 tax difference" sitting next to a company rule that allows "up to 5 cents of tax rounding" is exactly the kind of thing that confuses a model. The fix was one condition:
# Only a tax issue if the lines match; otherwise the tax difference just follows the line differences.
if abs(inv["subtotal"] - po["subtotal"]) < 0.001 and abs(inv["tax"] - po["tax"]) > 0.001:
Results
On eight test invoices, memory raised the number of correct decisions from 5 to 7, and the number handled without a human from 2 to 5. Decisions with memory take about 10 seconds, against about 1 second without, and nearly all of that time is the reflect step. Loading 66 memories and running all my tests cost about ten cents in Hindsight credits.
All the data is made up for testing, generated the same way every time so results can be reproduced, and eight invoices is a small test. I'd treat these numbers as a demonstration of behavior, not a benchmark.
Lessons
- Dates are part of the memory, not an extra. An agent can only notice that something changed if its memories know when things happened. Load history with the real dates.
- Write memories the way a person would explain them. Plain sentences that include the reason ("approved because...") gave much better answers than structured fields would.
- Use recall to show your work and reflect to make the call. Showing the retrieved facts builds trust; relying on reflect's reasoning produces better decisions.
- Let code find the problem and memory explain it. Plain rules plus remembered context is more reliable than asking a model to do both.
- Test the hardest case first. My first test checked the contract change. Knowing the riskiest part worked made every later decision easier.
If you want to try this pattern, the Hindsight documentation covers memory banks, tags and reflect, and Vectorize has a clear overview of what agent memory is and why it matters. Full Story on github:)

Top comments (0)