An Accounts Payable Agent Needs More Than Invoice Extraction
An invoice can be perfectly readable and still be the wrong invoice to approve.
That became one of the most important engineering lessons for me while working on our Accounts Payable Agent.
At first, accounts payable automation looks like a document-processing problem.
Read the invoice.
Extract the vendor.
Find the invoice number.
Find the amount.
Find the purchase order.
But extraction only answers one question:
What information is present in the document?
It does not answer:
Does this information make sense in context?
That is where validation, decision-making, and agent memory become important.
From invoice to decision
I designed the workflow around several stages:
Invoice
↓
Extraction
↓
Normalization
↓
Validation
↓
Historical context
↓
Decision
Each stage has a different responsibility.
Extraction identifies information.
Validation checks whether the information satisfies the required conditions.
Memory provides relevant historical context.
The decision layer determines what should happen next.
Keeping these responsibilities separate makes the workflow easier to understand.
Why extraction is not enough
Imagine an invoice contains:
Vendor: ABC Supplies
Invoice Number: INV-2048
PO: PO-8831
Total: ₹72,000
The extraction system may correctly identify every field.
But suppose the purchase order expects ₹60,000.
The invoice is readable.
The extracted information is correct.
Yet something still requires investigation.
The agent needs context before deciding what to do.
Validation with historical context
This is where Hindsight becomes useful.
Suppose the same vendor previously submitted an invoice with a similar discrepancy.
The previous case may contain:
Issue:
Invoice amount differed from PO.
Action:
Escalated for review.
Resolution:
Additional goods were confirmed.
Final result:
Invoice approved.
When another similar invoice arrives, the agent can retrieve this context.
The workflow becomes:
invoice = extract_invoice(document)
history = memory.search(
vendor=invoice.vendor,
invoice_context=invoice
)
result = validate(
invoice=invoice,
history=history
)
Again, the exact code should correspond to the implementation in the project.
The important point is that the current invoice and historical context are considered together.
A before-and-after workflow
Without memory
Invoice
↓
Extract
↓
Validate
↓
Apply rules
↓
Decision
With memory
Invoice
↓
Extract
↓
Retrieve relevant history
↓
Validate current invoice
↓
Decision
↓
Save outcome
The second workflow does not eliminate rules.
It gives the rules and agent more context.
Why relevant retrieval matters
More memory does not necessarily mean better decisions.
Suppose the system retrieves ten old invoices when only one is relevant.
The agent now has more information but potentially more noise.
For an AP workflow, useful retrieval can focus on:
Same vendor
Similar invoice issue
Previous validation failure
Previous escalation
Previous human resolution
Similar payment conditions
The objective should be relevant context, not maximum context.
Memory as evidence
I found it useful to think about memory as evidence.
For example:
Current invoice:
₹48,500
Historical context:
A similar invoice from the same vendor had a
PO mismatch and was manually verified.
Current validation:
PO mismatch exists again.
The memory doesn't tell the agent:
Approve this invoice.
Instead, it tells the agent:
This situation has happened before, and here is what happened then.
The agent can then evaluate the current case.
That distinction prevents old decisions from becoming automatic rules.
Human review is still part of the architecture
One of the most important parts of an automated financial workflow is knowing when to stop.
A simple decision structure is:
Information sufficient?
│
┌───┴───┐
Yes No
│ │
Validate Review
│
Decision
If the information is incomplete or conflicting, the agent can escalate.
That is not a failure of automation.
It is a deliberate boundary.
Why the decision should be explainable
An AP system should not only produce:
{
"status": "approved"
}
It is more useful if the system preserves a reason:
{
"status": "approved",
"reason": "Validated against the available invoice and purchase-order context."
}
The exact fields should follow the actual application.
The broader idea is to preserve enough information to understand the decision later.
That also makes the resulting memory more useful.
Hindsight as part of the architecture
Hindsight provides the persistent memory layer around the agent.
The overall architecture can be represented as:
┌──────────────┐
│ Invoice │
└──────┬───────┘
↓
┌──────────────┐
│ Extraction │
└──────┬───────┘
↓
┌──────────────┐
│ Validation │
└──────┬───────┘
↓
┌────────────────────┐
│ Hindsight Memory │
│ Relevant history │
└─────────┬──────────┘
↓
┌──────────────┐
│ AP Decision │
└──────┬───────┘
↓
┌──────────────┐
│ Outcome │
└──────┬───────┘
↓
Store useful
memory
I used the Hindsight GitHub repository� and its documentation� as the basis for the agent-memory integration.
The practical lesson
The biggest lesson from this part of the project was that agentic automation is not simply:
LLM + Invoice
A useful workflow is closer to:
Current data
+
Validation rules
+
Historical context
+
Decision boundaries
+
Human escalation
Each part solves a different problem.
Extraction tells us what the document contains.
Validation checks the current information.
Memory tells us whether relevant situations have occurred before.
Decision-making determines the next action.
Escalation handles uncertainty.
Three reusable lessons
Separate extraction from decision-making
Correctly extracting a value does not mean the value should automatically be trusted.
Treat memory as context
Previous interactions should inform current decisions rather than blindly determine them.
Design escalation from the beginning
An agent should have a clear path when information is incomplete or contradictory.
Final thought
The hardest part of an AP Agent isn't teaching it to read invoices.
It is teaching the system to understand when an invoice is routine, when it resembles something that happened before, and when it needs additional review.
That's where validation and memory become valuable.
The result is not simply an invoice-processing script.
It is a workflow that can use previous context while still evaluating every new invoice on its own terms.
Top comments (0)