I've watched LLMs confidently recommend approving an invoice for an amount that doesn't appear in the purchase order, the invoice, or any past decision. They do it with high confidence and well-structured reasoning. The reasoning sounds right. The number is invented.
For accounts-payable decisions, that's not an acceptable failure mode. I spent more time on the prevention mechanism than on any other part of MemoryOps.
The Root Cause
LLMs hallucinate in recommendation tasks for a specific reason: they're asked to reason about a space of possible answers, and they have no reliable way to distinguish "I recall this from the context" from "I generated this from training." When the context is thin — few precedents, novel vendor — the model fills the gap with plausible-sounding output.
The standard mitigation is prompt engineering. "Only use information from the provided context." "Do not invent amounts." This works most of the time. Most of the time isn't good enough for finance.
The approach in MemoryOps is different: structurally prevent the model from outputting anything that isn't traceable to real data.
Prevention Layer 1: Pre-computed Options
The LLM never computes an amount. All valid options — with their exact amounts — are calculated deterministically from the purchase order before the LLM is involved:
# detector/options.py
def build_options(exc_type, details, invoice, po_lines, received, tax_rate):
if exc_type == "price_variance":
inv_sub = money(invoice["subtotal"])
expected = sum(money(l["qty"] * po_price[l["sku"]]) for l in invoice["lines"]
if l["sku"] in po_price)
return [
{"action": "approve", "amount": str(inv_sub), "label": "Approve as billed"},
{"action": "adjust", "amount": str(expected), "label": "Short-pay to PO prices"},
{"action": "reject", "amount": "0", "label": "Reject and return to vendor"},
{"action": "escalate","amount": None, "label": "Escalate to senior approver"},
]
` ``
{% endraw %}
{% raw %}`
The LLM receives this list and picks one. It cannot output a number that isn't on the list. The amounts come from `decimal.Decimal` arithmetic on the actual PO data — no floats, no rounding surprises.
## Prevention Layer 2: Citation Validation
The LLM is required to cite which precedent IDs it used from [Hindsight](https://hindsight.vectorize.io/). After every response, before anything reaches the human reviewer, this runs:
` `` `python
# agent/recommender.py
cited = [c for c in dict.fromkeys(cited) if c in relevant_ids]
if action != "escalate" and not cited:
return result(
"escalate", 0.0,
"The chooser did not cite any recalled precedent, so its choice cannot be trusted. Escalating.",
[], recalled, "guard"
)
` `` `
`relevant_ids` is the set of record IDs actually returned by the [Hindsight](https://github.com/vectorize-io/hindsight) recall call for this exception. If the LLM cites an ID that wasn't in that set — because it made it up — it gets stripped. If nothing remains, the entire recommendation is discarded and replaced with an escalation.
The LLM cannot cite a record it didn't receive. If it tries, the validation catches it.
## Prevention Layer 3: The Approval Ceiling
Even if the LLM correctly cites a real precedent, there's a third check: the recommended action cannot approve an invoice larger than the largest previously approved exception of the same type for the same vendor.
` `` `python
# agent/recommender.py
if action == "approve" and magnitude is not None:
approved = [
float(r["fields"]["magnitude"]) for r in decisions
if r["fields"].get("decision") == "approved"
and _is_num(r["fields"].get("magnitude"))
]
if approved:
amax = max(approved)
if magnitude > amax + offline._range_tol(exc_type, amax):
return result(
"escalate", 0.3,
rationale + " [Escalated by guard: larger than any approved precedent.]",
cited, recalled, engine + "+guard"
)
` `` `
LLMs tend to over-generalise from precedents. If 3% variances were approved, the model may reason that a 15% variance should also be approved. The ceiling prevents that class of error without requiring the model to reason about it correctly.
## Prevention Layer 4: The Fallback Chain
If both the primary and fallback Groq models fail, the system doesn't try harder — it escalates:
` `` `python
# agent/recommender.py
if choice is None:
return result(
"escalate", 0.0,
"The language model did not return a valid recommendation after retries "
"(primary and fallback). Escalating; no precedent was used.",
[], recalled, "guard", self.llm.last_errors[-3:]
)
` `` `
No valid LLM output → escalate. Always. The agent does not invent a recommendation when it can't produce a real one.
## What the Guard Layer Looks Like in Practice
Every recommendation carries a label showing which engine produced it:
- `groq:openai/gpt-oss-120b` — primary model, valid citation
- `groq:openai/gpt-oss-120b+guard` — guard modified the output (ceiling exceeded)
- `guard` — recommendation overridden entirely; escalation from validation layer
When a recommendation shows `+guard`, the rationale includes the guard's note. When it shows just `guard`, the agent is being transparent that it produced nothing trustworthy.
## The Tradeoff
This architecture makes the agent less capable on paper. It can't approve something outside precedent bounds, even if that approval would be correct. It escalates aggressively when uncertain.
That's the design. An agent that escalates when uncertain is one a finance team will trust. An agent that occasionally invents a justification for a large approval creates a liability.
The hard rules are a deliberate choice about where human judgment belongs — it belongs on novel cases, not on routine ones with clear precedent.
[Agent memory](https://vectorize.io/what-is-agent-memory) via [Hindsight](https://hindsight.vectorize.io/) made the citation requirement practical. Without a memory layer that returns structured records with stable IDs, there's nothing to validate citations against. Hindsight's `document_id` and `metadata` fields make that check reliable.
Top comments (0)