How We Built an Invoice Exception Agent That Cites Its Sources
The hardest constraint we set for ourselves was this: the agent is not allowed to make a recommendation unless it can point to a specific past decision that justifies it. No citation, no recommendation. Escalate instead.
That sounds simple. It turned out to shape every architectural decision we made.
What the System Does
MemoryOps is an accounts-payable agent that processes invoice exceptions — mismatches between what a vendor invoiced, what was ordered on the purchase order, and what was actually received. Five exception types: price variance, quantity mismatch, tax error, duplicate invoice, missing PO.
For each exception, the agent does three things:
Recalls up to five relevant precedents from Hindsight — a vector memory layer that stores every past human decision
Passes those precedents to a Groq LLM, which picks an action (approve / reject / adjust / escalate) and must cite which precedent IDs it used
Validates the citation before trusting the output
If the LLM cites a precedent that wasn't in the recalled set, or recommends an action without citing anything, the system overrides it and escalates. Every time.
How It Hangs Together
The system has three clearly separated layers.
Detection is fully deterministic. detector/match.py compares invoice lines against the PO and goods receipts using decimal.Decimal with ROUND_HALF_UP throughout — no floats anywhere near money. A 3% price variance is a 3% price variance, not 2.9999999%.
Memory lives in Hindsight, a vector memory bank. Every human decision — approver name, role, reason, exception type, vendor, size — is retained as a structured text record with tags. The memory bank persists independently of the application database. Wipe SQLite and the memory survives.
Recommendation combines recalled precedents with a constrained LLM call. The LLM does not see raw invoice data. It sees the exception type, the pre-computed options (calculated deterministically from the PO), and the recalled precedents. It picks from a fixed menu.
The Core Technical Story: Constrained LLM Output
The earliest version of this system let the LLM recommend any amount it wanted. That lasted about ten minutes before it recommended approving an invoice for a figure that appeared nowhere in the PO, the invoice, or any precedent. We killed that immediately.
The current design pre-computes all valid options before the LLM is involved:
# detector/options.py — options are computed from the PO, not invented
def build_options(exc_type, details, invoice, po_lines, received, tax_rate):
if exc_type == "price_variance":
return [
{"action": "approve", "amount": str(invoice["subtotal"]), "label": "Approve as billed"},
{"action": "adjust", "amount": str(expected_subtotal), "label": "Short-pay to PO prices"},
{"action": "reject", "amount": "0", "label": "Reject invoice and return to vendor"},
{"action": "escalate","amount": None, "label": "Escalate to a senior approver"},
]
The LLM receives this list and picks one. It cannot invent amounts. It cannot recommend an action that isn't on the list.
Then validation runs in code, not in the prompt:
# agent/recommender.py
cited = [c for c in dict.fromkeys(cited) if c in relevant_ids]
if action != "escalate" and not cited:
return result("escalate", 0.0,
"The chooser did not cite any recalled precedent, so its choice cannot be trusted. Escalating.",
[], recalled, "guard")
If the LLM returns cited IDs that don't match what was actually recalled, they're stripped. If nothing remains, escalate. This runs on every response, every time, before anything reaches the human reviewer.
There's one more guard: the approval ceiling. If the LLM recommends approving an invoice that's larger than any previously approved exception of the same type for the same vendor, it's overridden:
if action == "approve" and magnitude is not None:
approved = [float(r["fields"]["magnitude"]) for r in decisions
if r["fields"].get("decision") == "approved"]
if approved and magnitude > max(approved) + tolerance:
return result("escalate", 0.3,
rationale + " [Escalated by guard: larger than any approved precedent.]",
cited, recalled, engine + "+guard")
Results
Starting from an empty memory bank, every exception in the first batch escalated — 17 out of 17. That's correct behaviour. The agent had nothing to draw from.
After replaying decisions from four senior approvers into Hindsight, the second batch showed a different picture: 1 escalation out of 15 exceptions. The 14 resolved exceptions all carried cited precedent IDs. Human reviewers agreed with the agent's recommendation on 79% of graded decisions.
The third batch demonstrated something more interesting. A vendor that had previously been approved started getting rejected because the two most recent decisions were overrides. The agent detected the pattern shift from the precedent history alone — no rule was written for this. It came from the recalled records.
Lessons Learned
Separation of concerns is what made this debuggable. When a recommendation looks wrong, there are exactly three places to look: the detection result, the recalled precedents, or the LLM output. Because each layer has a single responsibility, we've never had to guess which one caused a problem.
The citation requirement changed how we thought about memory. We originally stored decisions as flat text blobs. When we added the requirement that cited IDs must be traceable to recalled records, we had to redesign the memory records to carry explicit IDs that survive the recall round-trip. Hindsight's document_id and metadata fields made this possible without building a lookup layer ourselves.
Hard rules in code beat hard rules in prompts. We tried encoding the citation requirement in the system prompt. The LLM followed it most of the time. "Most of the time" isn't acceptable for finance decisions. Moving the validation into Python code — run after every LLM response — eliminated the failure mode entirely.
An agent that escalates honestly is more useful than one that guesses confidently. When the memory bank is empty, the right answer is "I don't know, escalate." The temptation to make the agent look capable by having it reason from general principles is real. We resisted it. The result is a system that finance teams can actually trust because its citations are verifiable.
For anyone building something similar, agent memory is the part worth investing in early. It's what separates an agent that degrades gracefully from one that hallucinates quietly.
Top comments (0)