I Built an Agent That Improves Every Session With Hindsight
The first time my accounts payable agent saw an invoice from Apex Industrial Supplies, it flagged it as an exception. That was the right call. It had no idea what a normal Apex invoice looked like.
What interested me was everything after that: what happens when a human says "this one is fine," and whether the agent can carry that judgment into the next invoice.
What VendorSense does
VendorSense is an accounts payable agent built around one idea: it learns what normal looks like for every vendor. You upload an invoice (PDF, PNG, or JPG), and the agent decides whether it can be processed automatically or needs a person to look at it.

The overview tracks how many invoices were processed automatically, how many went to a human, and how many became learning events.
Every invoice moves through the same six-step pipeline:
- Invoice captured: the file is stored and its fields are extracted
- Hindsight recall: the agent asks memory what it knows about this vendor
- Groq reasoning: an LLM compares the invoice against what was recalled
- Decision: process automatically, or raise an exception with a confidence level
- Human review, if needed: a person approves or rejects, with a note
-
Hindsight retain: that decision becomes memory
The agent itself is stateless. Everything it knows about vendors lives in Hindsight, an open-source agent memory system running as a separate service. Each user gets a private workspace with isolated data, so one team's vendor history never leaks into another's.
Why fixed rules don't work for invoices
Classic accounts payable automation is a pile of thresholds: flag anything over a set amount, flag anything without a PO. Those rules have the wrong shape for the problem.
"Normal" is a property of the vendor, not of the company. One vendor bills a nearly identical amount on Net 30 terms every month, and a 4% difference is worth a look. Another vendor's amounts swing widely by design. A single global threshold either floods reviewers with false alarms or lets the strange invoice through.
I wanted the agent to build that per-vendor picture from the people who actually review the invoices, and to keep refining it over time. That is a memory problem, not a rules problem.
Why Hindsight
Hindsight gives me three operations, and I only needed them to map onto two moments in the pipeline:
- Retain stores what happened
- Recall retrieves what's relevant before a decision
- Reflect reasons over stored memories to surface patterns
Recall uses several retrieval strategies in parallel (semantic, keyword, graph, and temporal) rather than one similarity search. For invoices that matters. "What do we know about Apex Industrial Supplies, and how have their amounts changed?" is a question about an entity over time, and a plain embedding lookup handles it poorly. If you want the background on why this matters for agents, Vectorize's explainer on agent memory is a good primer.
The core loop
The pattern is recall before deciding, retain after a human weighs in. The wiring lives in [FILE PATH, e.g. agent/memory.py], and it looks roughly like this:
from hindsight_client import Hindsight
client = Hindsight(base_url="http://localhost:8888")
def decide(invoice, bank_id):
# 1. Ask memory what we know about this vendor BEFORE reasoning
recalled = client.recall(
bank_id=bank_id,
query=f"Previous invoices and decisions for {invoice.vendor}",
)
experience = "\n".join(f"- {r.text}" for r in recalled.results)
# 2. Reason over the invoice plus recalled experience
return reason_with_llm(invoice, experience)
def learn_from_review(invoice, decision, reviewer_note, bank_id):
# 3. After a human approves or rejects, retain the decision AND the why
client.retain(
bank_id=bank_id,
content=(
f"Vendor {invoice.vendor}, invoice {invoice.invoice_id}, "
f"amount {invoice.amount}: human {decision}. "
f"Reviewer note: {reviewer_note}"
),
context="human invoice review",
)
(Swap in the real functions from your repo. These show the shape using the Hindsight client's recall and retain calls.)
The cold start: an exception is the right answer
On a vendor the agent knows nothing about, the interface says so plainly. Hindsight returns nothing relevant, and the agent raises an exception with a confidence level and a reason.

I like that the panel doesn't pretend. "No relevant previous vendor experience found" is followed by a line saying this is a new learning situation and that a human decision can become future experience. The empty memory is a state the system handles honestly, not an error.
The evidence trail
When the agent does have history, it shows its work. In the review screen, it lists the specific evidence behind its decision: a prior approved invoice from the same vendor on Net 30 terms, a current amount that is close to the earlier one but not identical, and identifiers that don't look like the vendor's earlier structured IDs.

Reviewers shouldn't have to trust a bare "exception." Citing the prior invoice, the amount comparison, and the malformed fields lets a person confirm or overrule the agent in seconds.
Human review is the teaching step
The review form has a note field ("Explain what you verified...") and two buttons: Approve & Teach Agent and Reject & Teach Agent. The wording is deliberate. A click isn't just closing a ticket; it's writing to memory.
The reviewer's note is retained alongside the decision. So the memory isn't a pile of yes/no labels. It contains the reasoning of the person who checked, and that's what later recalls can lean on.

A "Learning Status" counter shows how many human decisions have been retained, so it's visible when the system has actually learned something rather than just processed something.
What it looks like across sessions
First invoice from a vendor
Agent: Exception. No prior experience with this vendor. Confidence: medium.
Reviewer: Approves, with a note on what they verified.
Later invoice from the same vendor
Agent: [YOUR REAL EXAMPLE: e.g. "Matches prior approved Apex invoice pattern. Auto-processed."]
Nothing in the prompt changed between the two. The difference is what came back from recall.
What was painful
- Extraction errors look like vendor anomalies. In one run, the extracted invoice ID came back as the generic word "Invoice" and the purchase order as a garbled fragment. The agent correctly treated that as suspicious, but the vendor wasn't the problem; the extraction was. Memory can't fix bad input, and I had to be careful not to teach the agent that a vendor's format had changed when really the OCR had failed.
- Deciding what to retain. [YOUR REAL EXPERIENCE, e.g. "storing raw extracted text made recall noisy; storing the decision and the reviewer's reason worked better."]
- Cold start. Every new vendor starts as an exception. That's the correct behavior, but it means the first weeks of use are review-heavy.
Lessons learned
- Treat "normal" as per-entity state. Global thresholds fight the shape of the problem. Memory keyed to the vendor matches it.
- Keep the agent stateless. Pushing memory into Hindsight made the agent easier to restart, test, and reason about.
- Retain the reason, not just the verdict. A reviewer's note is worth more than an approve/reject flag.
- Make empty memory a first-class state. "I don't know this vendor yet" should be shown, not hidden.
- Show the evidence. People will trust an agent that cites the prior invoice and the specific mismatch far sooner than one that just says "exception."
Try it yourself
If you're building anything where decisions should improve as humans correct them, start with the Hindsight docs and the open-source repo. Retain and recall can be working in an afternoon; the interesting part is when your agent starts using a reviewer's judgment you'd forgotten you gave it.
Top comments (0)