DEV Community

Lokesh Reddy
Lokesh Reddy

Posted on

My Agent Never Hallucinates Invoice Decisions — Here's the Hard Rule That Prevents It

I've watched LLMs confidently recommend approving an invoice for an amount that doesn't appear in the purchase order, the invoice, or any past decision. They do it with high confidence and well-structured reasoning. The reasoning sounds right. The number is invented.

For accounts-payable decisions, that's not an acceptable failure mode. I spent more time on the prevention mechanism than on any other part of MemoryOps.

The Root Cause

LLMs hallucinate in recommendation tasks for a specific reason: they're asked to reason about a space of possible answers, and they have no reliable way to distinguish "I recall this from the context" from "I generated this from training." When the context is thin — few precedents, novel vendor — the model fills the gap with plausible-sounding output.

The standard mitigation is prompt engineering. "Only use information from the provided context." "Do not invent amounts." This works most of the time. Most of the time isn't good enough for finance.

The approach in MemoryOps is different: structurally prevent the model from outputting anything that isn't traceable to real data.

Prevention Layer 1: Pre-computed Options

The LLM never computes an amount. All valid options — with their exact amounts — are calculated deterministically from the purchase order before the LLM is involved:

# detector/options.py
def build_options(exc_type, details, invoice, po_lines, received, tax_rate):
    if exc_type == "price_variance":
        inv_sub = money(invoice["subtotal"])
        expected = sum(money(l["qty"] * po_price[l["sku"]]) for l in invoice["lines"]
                       if l["sku"] in po_price)
        return [
            {"action": "approve", "amount": str(inv_sub),    "label": "Approve as billed"},
            {"action": "adjust",  "amount": str(expected),   "label": "Short-pay to PO prices"},
            {"action": "reject",  "amount": "0",             "label": "Reject and return to vendor"},
            {"action": "escalate","amount": None,            "label": "Escalate to senior approver"},
        ]
` ``
{% endraw %}
 {% raw %}`

The LLM receives this list and picks one. It cannot output a number that isn't on the list. The amounts come from `decimal.Decimal` arithmetic on the actual PO data — no floats, no rounding surprises.

## Prevention Layer 2: Citation Validation

The LLM is required to cite which precedent IDs it used from [Hindsight](https://hindsight.vectorize.io/). After every response, before anything reaches the human reviewer, this runs:

` `` `python
# agent/recommender.py
cited = [c for c in dict.fromkeys(cited) if c in relevant_ids]
if action != "escalate" and not cited:
    return result(
        "escalate", 0.0,
        "The chooser did not cite any recalled precedent, so its choice cannot be trusted. Escalating.",
        [], recalled, "guard"
    )
` `` `

`relevant_ids` is the set of record IDs actually returned by the [Hindsight](https://github.com/vectorize-io/hindsight) recall call for this exception. If the LLM cites an ID that wasn't in that set — because it made it up — it gets stripped. If nothing remains, the entire recommendation is discarded and replaced with an escalation.

The LLM cannot cite a record it didn't receive. If it tries, the validation catches it.

## Prevention Layer 3: The Approval Ceiling

Even if the LLM correctly cites a real precedent, there's a third check: the recommended action cannot approve an invoice larger than the largest previously approved exception of the same type for the same vendor.

` `` `python
# agent/recommender.py
if action == "approve" and magnitude is not None:
    approved = [
        float(r["fields"]["magnitude"]) for r in decisions
        if r["fields"].get("decision") == "approved"
        and _is_num(r["fields"].get("magnitude"))
    ]
    if approved:
        amax = max(approved)
        if magnitude > amax + offline._range_tol(exc_type, amax):
            return result(
                "escalate", 0.3,
                rationale + " [Escalated by guard: larger than any approved precedent.]",
                cited, recalled, engine + "+guard"
            )
` `` `

LLMs tend to over-generalise from precedents. If 3% variances were approved, the model may reason that a 15% variance should also be approved. The ceiling prevents that class of error without requiring the model to reason about it correctly.

## Prevention Layer 4: The Fallback Chain

If both the primary and fallback Groq models fail, the system doesn't try harder — it escalates:

` `` `python
# agent/recommender.py
if choice is None:
    return result(
        "escalate", 0.0,
        "The language model did not return a valid recommendation after retries "
        "(primary and fallback). Escalating; no precedent was used.",
        [], recalled, "guard", self.llm.last_errors[-3:]
    )
` `` `

No valid LLM output → escalate. Always. The agent does not invent a recommendation when it can't produce a real one.

## What the Guard Layer Looks Like in Practice

Every recommendation carries a label showing which engine produced it:

- `groq:openai/gpt-oss-120b` — primary model, valid citation
- `groq:openai/gpt-oss-120b+guard` — guard modified the output (ceiling exceeded)
- `guard` — recommendation overridden entirely; escalation from validation layer

When a recommendation shows `+guard`, the rationale includes the guard's note. When it shows just `guard`, the agent is being transparent that it produced nothing trustworthy.

## The Tradeoff

This architecture makes the agent less capable on paper. It can't approve something outside precedent bounds, even if that approval would be correct. It escalates aggressively when uncertain.

That's the design. An agent that escalates when uncertain is one a finance team will trust. An agent that occasionally invents a justification for a large approval creates a liability.

The hard rules are a deliberate choice about where human judgment belongs — it belongs on novel cases, not on routine ones with clear precedent.

[Agent memory](https://vectorize.io/what-is-agent-memory) via [Hindsight](https://hindsight.vectorize.io/) made the citation requirement practical. Without a memory layer that returns structured records with stable IDs, there's nothing to validate citations against. Hindsight's `document_id` and `metadata` fields make that check reliable.
Enter fullscreen mode Exit fullscreen mode

Top comments (0)