The most important line of code in my accounts payable agent doesn't touch memory or a language model. It's a plain if that forces an invoice into the human review queue when the bank account changes, and I've stopped trusting any version of the system that doesn't have it.
Everything else in the agent is about learning what normal looks like for each vendor. That part runs on Hindsight, an open-source agent memory system. This post is about how the two fit together: what I let memory decide, what I refuse to let it decide, and what went wrong when I got that split wrong.
What the system does
VendorSense is an accounts payable agent. An invoice comes in as a PDF, PNG, JPG, or JPEG. The agent extracts the fields that matter, recalls what it knows about that vendor, and decides one of two things: AUTO_PROCESS (route it through normal AP handling) or EXCEPTION (send it to a person).
The pipeline is deliberately linear:
-
Extraction. PDFs go through
pypdfand a set of regexes for vendor, invoice ID, PO, payment terms, currency, bank last four, and total. Images go to a Groq vision model with a prompt that says, in so many words, extract fields and make no decisions. Both paths produce the same dictionary. - Archive. The original file is stored under a name that includes the vendor, invoice ID, and a SHA-256 prefix, so the record and the file can never drift apart.
- Recall. I query Hindsight for what the agent has learned about this vendor.
-
Decision.
openai/gpt-oss-120bon Groq gets the invoice plus the recalled memory and returns JSON. - Human review. For exceptions, a reviewer approves or rejects and writes a note. Both go back into Hindsight.
Here is how the pieces connect, and where Hindsight sits:
Invoice (PDF / PNG / JPG)
│
▼
Extract fields (pypdf + regex, or vision model)
│
▼
LLM decision (gpt-oss-120b on Groq) ◄── recall ─┐
│ │
▼ ┌────────┴─────────┐
Guard rules in code │ Hindsight │
(bank change → EXCEPTION, │ one memory bank │
any error → EXCEPTION) │ per user │
│ └────────▲─────────┘
▼ │
AUTO_PROCESS or EXCEPTION │ retain
│ │
▼ │
Human review ────────────────────┘
(approve / reject + note)
Each user gets a private workspace: their own invoice database, their own file archive, and their own Hindsight memory bank. More on that later, because it turned out to matter more than I expected.
I ask the model for HIGH, MEDIUM, or LOW confidence rather than a float. Early on I asked for a number between 0 and 1 and got 0.85 and 0.9 back with no way to tell what separated them. A coarse label at least doesn't pretend to be calibrated.
Why memory instead of a rules table
"Normal" is per vendor. Take Apex Industrial Supplies, who send invoices for industrial pump components somewhere in the ₹40,000–₹60,000 range, on Net 30 terms, with purchase orders in an AIS- format and a standard footer. Another vendor's normal looks nothing like that.
The obvious approach is a rules table: vendor, expected amount range, expected terms, verified bank account. It works until someone has to maintain it. The rules go stale, nobody remembers why a threshold was set, and every exception a reviewer resolves teaches the system nothing.
What I wanted instead is what a good AP clerk has: a memory of what happened last time and why. That is a different data shape from a rules row. It's an experience with a decision and a reason attached, and it's exactly what agent memory is meant to hold. So I stopped writing rules and started writing down what reviewers do.
What goes in: retaining the reviewer's judgment
When a reviewer approves or rejects an exception, the app writes the whole event to Hindsight:
experience = f"""
VendorSense AP learning event.
Vendor: {invoice.get('vendor')}
Invoice ID: {invoice.get('invoice_id')}
Amount: ₹{invoice.get('amount')}
Purchase Order: {invoice.get('purchase_order')}
Payment Terms: {invoice.get('payment_terms')}
Bank ending: {invoice.get('bank_account_last4')}
Human decision: {decision}
Human review note:
{reason}
"""
hindsight.retain(bank_id=BANK_ID, content=experience)
Two choices here are deliberate.
First, I store the reviewer's note, not just the verdict. "Approved" tells the agent nothing it can reason about. "Invoice was verified against the purchase order; bank account ending 7821 was verified; AIS standard footer confirmed as legitimate" tells it which properties a human actually checked. Rejections matter just as much as approvals: a rejected invoice with a note is the closest thing the system has to a record of what suspicious looks like for that vendor.
Second, the content is plain text. I'm not building a schema for memory; I'm writing down what a person concluded and letting retrieval handle the rest. That kept the write path to about ten lines, and it means the Hindsight documentation is the only thing I had to read to get it working.
What comes out: recall as a question
Recall is where I spent the most time, because the quality of the decision depends on the quality of what comes back.
result = hindsight.recall(
bank_id=BANK_ID,
query=query,
max_tokens=2500,
budget="mid",
)
The query is not a keyword. It's a description of the situation followed by a checklist: previous approvals, previous rejections, typical amounts, payment terms, PO patterns, verified bank information, previous exceptions. Writing it as a list of things I want to know, rather than the vendor name alone, is what got me back the reviewer's reasoning and not just a matching invoice.
The system prompt then frames whatever comes back with one sentence I repeat in several places: Hindsight memory is evidence, not absolute truth. That matters because memory can be incomplete or out of date. A vendor legitimately changes banks. A reviewer approves something they shouldn't have. I don't want the model to treat a recalled approval as a permission slip, and I don't want it declaring an invoice fraudulent on thin evidence either. The prompt says a changed bank account "does NOT automatically mean fraud" and asks for "human verification is required" wording instead.
One thing I had to design for explicitly: empty recall. For a new vendor, Hindsight has nothing to say, and my flattening code turns that into the string "No relevant previous vendor experience found." The prompt treats that as a real state with a real rule: a vendor with no relevant experience should normally go to a human. That review then becomes the first memory. The system starts at zero and learns by being corrected, which is the whole point.
The through-line: memory informs, code decides
Here is the story I promised. Memory and the model are good at fuzzy judgment: does this amount look right for this vendor, does this PO match the pattern, does this footer look familiar. They are the wrong tool for the one rule that must never fail.
Bank account changes are the classic payment-redirection attack, and a model that reads a friendly recalled approval next to an invoice that looks right can talk itself into a yes. I tell the model the rule in the prompt:
4. A bank account change must trigger EXCEPTION.
And then I don't rely on it:
result = json.loads(content)
if result.get("bank_change_detected") is True:
result["decision"] = "EXCEPTION"
The same principle applies to failure. If the Groq call raises, or the JSON doesn't parse, the function returns an EXCEPTION with LOW confidence and the reason "Automated analysis failed safely." An unrecognized decision value is normalized to EXCEPTION too.
Look at the shape of those overrides. Every one of them moves the outcome in one direction: toward a human. The code can turn an AUTO_PROCESS into an EXCEPTION, but nothing in the code path can turn an EXCEPTION into an AUTO_PROCESS. That asymmetry is the design. The model and Hindsight get to be helpful and occasionally wrong, and being wrong can only cost me a reviewer's few minutes, never a payment to the wrong account.
I'll admit the first version got this wrong in a more basic way. The early scripts compared the invoice's last four digits against a hardcoded constant for the one vendor I was testing with. That was fine for one vendor and useless for two. Now the verified account is stored per vendor as structured data and compared in code, and the model's job is to explain the evidence. Hindsight's job is to hold the reason the account was verified in the first place. Three components, three different jobs, and the one that must be deterministic is.
Per-user memory banks
The other decision I'd defend is isolating memory at the bank level. Two companies using this shouldn't share what "Apex Industrial Supplies" normally looks like, and one company's rejection shouldn't leak into another's recall. So every account gets its own Hindsight bank:
def user_slug(username):
return hashlib.sha256(
username.strip().lower().encode("utf-8")
).hexdigest()[:16]
def user_bank_id(username):
return f"vendorsense-user-{user_slug(username)}"
The bank is created idempotently on login, with a mission that states the boundary in words: private AP memory for this user only, storing that user's vendor experiences and human-confirmed decisions. The slug is deterministic, so the same username always resolves to the same bank, and it's filesystem-safe, so the same slug names the user's database file and invoice archive too. One identifier, three stores.
Two things were annoying. Accounts created before isolation existed had to be migrated: the original single bank became the first account's bank, and later accounts got fresh ones. And create_bank is wrapped in a bare except because calling it on an existing bank is harmless. I don't love that, because it also swallows real failures. They surface later, at the first recall or retain, which is at least loud, but I'd rather they surfaced at creation.
What it does in practice
The clearest example is the pair of Apex invoices I use as a regression check.
A reviewer approves invoice INV-AIS-1047: ₹48,620, PO AIS-2417, Net 30, bank ending 7821. Their note records that they matched it to the purchase order and verified the account. That goes into the bank.
The next invoice, INV-AIS-1048, arrives with the same amount, the next PO in the same AIS- format, the same terms, the same footer, and a bank account ending 9143. Recall returns the earlier approval and its note. Everything about the invoice looks routine except the one thing that matters. The result is an EXCEPTION, the reason reads as "human verification is required" rather than an accusation, and the evidence list includes the mismatch between 9143 and the previously verified account.
Notice that the reviewer's original note is what makes this useful. Because it recorded that 7821 was verified, the evidence can say what the account was compared against and why the earlier one was trusted.
The other case is a vendor the agent has never seen. Recall comes back empty, the decision is an EXCEPTION, and the reviewer's note becomes that vendor's first memory. I don't have a benchmark for how many invoices it takes before a vendor starts auto-processing, and I'm not going to invent one. What I can say is that the mechanism is visible: each decision either adds to the evidence or doesn't.
AUTO_PROCESS here means the invoice is routed through AP handling. The agent's prompt explicitly forbids claiming that money moved, and I've kept that wording deliberately.
Lessons
Store decisions with their reasons. A verdict is a label. A verdict plus what the human checked is something a later invocation can reason over. If you're writing to agent memory, write the note.
Write recall queries as questions with a checklist. A vendor name retrieves the vendor. A description of the situation plus a list of things you want to know retrieves the reasoning. Budget the tokens for it.
Tell the model memory is evidence, then enforce the important parts anyway. Prompt language is a request. For any rule where a miss is expensive, put the consequence in code.
Make overrides one-directional. Every deterministic guard I wrote can only move a decision toward human review. That makes the guards safe to add freely, because they can never create a bad approval.
Design the empty state and the failure state first. New vendor and broken call both resolve to a human. That single decision removed a whole category of edge cases.
Isolate memory at the boundary, early. Retrofitting per-user banks onto a system with one shared bank is doable but tedious. If more than one party will use the system, give each its own bank from the first commit.
Memory made the agent noticeably better at recognizing routine invoices. The if statement made it safe to let it try.
Top comments (0)