The first time an invoice agent makes a decision, it has almost no reason to trust its own context.
The interesting part starts with the second invoice from the same vendor.
A human accounts-payable reviewer does not treat that second invoice as a completely new problem. They remember the vendor’s usual purchase-order format, payment terms, typical amounts, previous exceptions, and—most importantly—what was actually confirmed by a human before. I wanted VendorSense to have that kind of continuity without turning old decisions into blind trust.
VendorSense is an Accounts Payable agent built around that idea: extract the current invoice, retrieve relevant vendor experience with Hindsight, reason over the current evidence and historical context, route uncertain cases to a human, and retain confirmed outcomes for the next invoice.
The workflow is deliberately small
I kept the core loop narrow:
Invoice
↓
Extract structured fields
↓
Hindsight Recall
↓
LLM reasoning
↓
AUTO_PROCESS or EXCEPTION
↓
Human review when required
↓
Hindsight Retain
↓
Future invoices use the experience
The VendorSense dashboard exposes the invoice pipeline and learning activity.
The important boundary is that AUTO_PROCESS is an application routing decision in the prototype—it does not execute a real payment.
Why memory comes before reasoning
My first instinct with an invoice agent was to focus on extraction and prompting. But extraction is only the beginning.
Suppose the current invoice says:
Vendor: Apex Industrial Supplies
Invoice: INV-AIS-1049
Amount: ₹52,100
Purchase Order: AIS-2419
Payment Terms: Net 30
Bank ending: 7821
Those fields describe what is happening now. They do not tell the agent whether this is normal for the vendor.
That is where Hindsight changes the input to the reasoning step.
For each invoice, VendorSense constructs a query from the current vendor and invoice context and asks Hindsight for relevant experience: previous approvals and rejections, amount patterns, payment terms, purchase-order patterns, verified bank information, previous exceptions, human decisions, and other learned vendor behavior.
The core recall looks like this:
result = hindsight.recall(
bank_id=BANK_ID,
query=query,
max_tokens=2500,
budget="mid",
)
I then put the current invoice and the retrieved memory into the reasoning prompt:
user_prompt = f"""
CURRENT INVOICE:
{json.dumps(invoice, indent=2)}
HINDSIGHT MEMORY:
{memory_text}
Evaluate this invoice.
"""
That ordering is intentional.
Without the second column, every invoice starts from a clean slate.
With it, the reasoning layer can ask a more useful question: Does this invoice fit the experience we already have with this vendor?
I used Hindsight on GitHub, the Hindsight documentation, and Vectorize’s agent memory overview while building this layer.
Memory is evidence, not permission
There is a trap in giving an agent memory: a familiar vendor can start looking safe simply because it has been seen many times.
I explicitly did not want that.
The reasoning instructions treat Hindsight memory as evidence rather than absolute truth. The agent is expected to use relevant history, but it still has to evaluate the current invoice.
That distinction matters most when something important changes.
Imagine that several Apex Industrial Supplies invoices have matched previous patterns and have been confirmed by humans. A new invoice arrives with a different bank account.
A naive memory system could reason:
I know this vendor.
Previous invoices were approved.
Therefore, this invoice is probably fine.
That is exactly the behavior I wanted to avoid.
The intended safety path is:
Previous verified bank: 7821
Current bank: 9143
↓
Bank changed
↓
EXCEPTION
↓
Human review
The agent output includes a bank_change_detected field, and the application can force the exception path when that condition is present:
result = json.loads(content)
if result.get("bank_change_detected") is True:
result["decision"] = "EXCEPTION"
return result
The principle is simple: historical familiarity should help the agent recognize normal behavior, not erase sensitivity to meaningful changes.
One implementation detail I would strengthen further before production is making the bank comparison independently deterministic from the stored verified vendor data, rather than relying on the model to identify every change. Critical financial controls should not depend solely on an LLM-generated flag.
The part that actually teaches the system
I did not want the agent to learn from its own previous recommendations.
If the model says “approve,” that is a recommendation. It should not automatically become a trusted fact about the vendor.
So VendorSense puts a human between recommendation and learning.
When an invoice requires review, the reviewer can approve or reject it and add a note. The application turns that outcome into an explicit learning event:
experience = f"""
VendorSense AP learning event.
Vendor: {invoice.get('vendor')}
Invoice ID: {invoice.get('invoice_id')}
Amount: ₹{invoice.get('amount')}
Purchase Order: {invoice.get('purchase_order')}
Payment Terms: {invoice.get('payment_terms')}
Bank ending: {invoice.get('bank_account_last4')}
Human decision: {decision}
Human review note:
{reason}
This is a human-confirmed AP experience.
Use this experience when evaluating future invoices
from this vendor.
"""
Then it is retained:
hindsight.retain(
bank_id=BANK_ID,
content=experience,
)
That creates a learning boundary:
Agent recommendation
↓
Human confirmation
↓
Confirmed experience
↓
Hindsight Retain
↓
Future Hindsight Recall
I prefer this over self-reinforcement. The system can accumulate experience, but the strongest learning signal comes from an actual reviewed outcome.
The before-and-after behavior
The clearest way to understand the value of memory is to follow the same vendor twice.
First interaction
An invoice arrives from a vendor with little or no relevant history.
VendorSense can still extract the vendor, invoice number, amount, purchase order, payment terms, bank information, and description, but it has less vendor-specific evidence available.
The system can therefore route the case conservatively when the evidence is insufficient.
A human reviews it.
Suppose the reviewer approves the invoice and records why. That confirmed experience is retained in Hindsight.
Later interaction
Another invoice from the same vendor arrives.
Now the agent can retrieve the previous experience before reasoning about the new invoice.
The behavioral difference is:
FIRST INVOICE
Little vendor-specific context
↓
Conservative evaluation
↓
Human decision
↓
Memory created
LATER INVOICE
Relevant vendor experience available
↓
Historical context informs evaluation
↓
Routine patterns are easier to recognize
↓
Meaningful deviations can stand out
The goal is not to make the agent blindly approve more invoices. The goal is to make the reasoning context better.
Keeping memory inspectable
Another design choice I made was to keep the memory lifecycle visible.
The application keeps processed invoices separate from learned Hindsight experiences, so operational records and agent memory remain distinct.
The interface can expose the same lifecycle:
Current invoice
↓
Memory retrieved
↓
Reasoning
↓
Recommendation
↓
Human confirmation
↓
Memory updated
That makes surprising decisions easier to debug: I can inspect what was extracted, what memory was recalled, what the model concluded, what the human decided, and what was retained for next time.
The human-review step exposes the evidence before a confirmed decision becomes learning data.
What I learned building it
- Retrieval only matters when it changes the reasoning context
It is easy to call a memory API and display a list of old records.
- Human confirmation is a useful learning boundary
I do not want the system to teach itself that its own recommendation was correct.
A human-confirmed approval or rejection is a much clearer signal for future evaluations.
- Familiarity should never become permanent permission
A vendor having a long history of approved invoices does not make every future invoice safe.
Bank changes and other meaningful deviations still need explicit attention.
- Critical rules belong outside the model
LLMs are useful for contextual reasoning, but some financial controls should be deterministic.
The more consequential the condition, the less I want its enforcement to depend on whether a model noticed it.
The pattern I would reuse
The invoice parser and dashboard are useful pieces, but the part of VendorSense I would reuse in another agent is the memory loop:
Recall the past.
↓
Reason about the present.
↓
Ask a human when the evidence is insufficient.
↓
Retain the confirmed outcome.
↓
Use it the next time.
That is the difference between an agent that merely has access to information and one that can accumulate useful experience over time.
For me, the most interesting result was not that an LLM could read an invoice. It was seeing how a confirmed human decision could become part of the context for the next decision—while still keeping the current invoice, historical memory, model reasoning, and human judgment as separate parts of the system.
That separation is what makes the architecture useful beyond invoices.


Top comments (3)
The decision to make bank account changes trigger an exception outside the prompt is the sharpest architectural call here. In agent pipelines touching money or credentials, letting an LLM decide whether an anomaly matters is where silent drift happens.
The second half that usually bites in AP workflows is vendor aliasing. A vendor's legal entity on the invoice might differ slightly from the registered entity, or a subsidiary shares a tax ID with a different remittance address. When memory retrieval keys on surface vendor names, you get split histories. Having an explicit human review step that anchors the memory to an immutable vendor ID keeps the recall layer from drifting across invoice cycles.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.