DEV Community

Tata Venkata Karthik
Tata Venkata Karthik

Posted on

My Invoice Agent Went From 100% Escalation to 7% — Here's the Memory Layer That Did It

The first time I ran the agent against a batch of invoices, it escalated every single one. 17 exceptions, 17 escalations, 17 "no precedent found." I had built a very expensive escalation machine.

Three batches later, the escalation rate was 7%. The agent was resolving 93% of exceptions autonomously, with citations. Here's what changed.

The Starting Point

MemoryOps processes accounts-payable invoice exceptions. When a vendor invoices an amount that doesn't match the purchase order or the goods receipt, the system flags it and the agent recommends what to do: approve it, reject it, adjust it, or escalate it to a human.

The agent uses Hindsight as its memory layer. Every time a human makes a decision — approved, rejected, adjusted, with a reason — that decision is stored in a Hindsight memory bank as a structured record. The next time a similar exception comes in, the agent recalls relevant precedents and uses them to make a recommendation.

The key word is "next time." The first time a given vendor and exception type appears, there's nothing to recall. The agent escalates honestly. That's the 100%.

What Gets Stored

When a human records a decision, services.py writes three types of records to Hindsight:

PREC-* — one per decision. Exact details: vendor, invoice number, exception type, size, approver name, role, reason, date, and whether it overrode the agent's previous recommendation.

VF-* — vendor-level pattern facts. Rebuilt deterministically after every decision. Summarises the full history for a vendor+type combination: how many times each outcome occurred, the size range, the most recent decision and reason.

AP-* — approver preference facts. If one person has made every decision for a particular vendor and exception type, a record is written saying "route this to them."

# services.py — the vendor fact record, rebuilt after every decision
text = (f"Vendor pattern for {v.name} ({v.code}), {label}: {len(rows)} past decisions: "
        f"{'; '.join(parts)}. Most recent: {last.decision} on {last.decided_on} "
        f"by {last.approver_name} - \"{last.reason}\".")
Enter fullscreen mode Exit fullscreen mode

This means the agent doesn't just recall individual decisions — it recalls patterns. By the time 5 decisions exist for a vendor, the VF record summarises all of them in one chunk.

The Recall in Numbers

Batch 1 — memory empty:

  • 17 exceptions
  • 17 escalations (100%)
  • 0 citations
  • Memory bank: empty

After batch 1, 17 decisions from senior approvers were retained into Hindsight. 17 PREC records, plus vendor facts and approver preference records for each vendor+type combination that had 2+ decisions.

Batch 2 — memory populated:

  • 15 exceptions
  • 1 escalation (7%)
  • 14 resolved with citations
  • Human agreement rate: 79%

The single escalation in batch 2 was correct — it was a new vendor with no precedent.

Batch 3 — memory includes overrides:

  • 16 exceptions
  • 3 escalations (19%)

The 3 escalations in batch 3 were all correct: one new vendor, one case with directly conflicting precedents (53% approve vs 47% reject across recalled records), one invoice significantly larger than any previously approved amount for that vendor.

Why the Numbers Look the Way They Do

The 100% → 7% drop between batch 1 and batch 2 is entirely explained by memory. The agent didn't change. The prompts didn't change. The only difference was that Hindsight now had 17 precedents to draw from.

The 19% in batch 3 is higher than 7%, and that's intentional. After batch 2 included a human override — a case where the agent recommended approval but the human rejected it — the agent picked up on the pattern shift in batch 3. It saw that recent decisions for that vendor were rejections, weighted them more heavily than older approvals, and recommended rejection. The override taught the agent something.

# agent/recommender.py — the guard that prevents approvals above the precedent ceiling
if action == "approve" and magnitude is not None:
    approved = [float(r["fields"]["magnitude"]) for r in decisions
                if r["fields"].get("decision") == "approved"]
    if approved:
        amax = max(approved)
        if magnitude > amax + offline._range_tol(exc_type, amax):
            return result("escalate", 0.3,
                rationale + " [Escalated by guard: larger than any approved precedent.]",
                cited, recalled, engine + "+guard")
Enter fullscreen mode Exit fullscreen mode

One of the batch 3 escalations was triggered by this guard. The invoice was 40% larger than any previously approved price variance for that vendor. The agent recommended approval but the guard caught it and escalated instead.

What the 7% Number Actually Means

The 7% figure is not the goal. The goal is that the right cases escalate.

An agent that escalates 0% would be concerning — it would mean it's resolving cases it shouldn't be resolving autonomously. An agent that escalates 100% is useless. The right escalation rate depends on how novel the incoming invoices are relative to the precedent bank.

What the numbers show is that agent memory works as a mechanism for reducing unnecessary escalations over time. Cases that have been resolved before don't need a human every time. Cases that are genuinely new, conflicted, or outside precedent bounds still get escalated.

Lessons

The empty-memory baseline is important. Running batch 1 with no precedents gave us a clear before-state. Without that, we'd have no way to measure what the memory layer was actually contributing.

Agreement rate matters more than escalation rate. 79% human agreement on non-escalated cases means the agent was wrong about 21% of the cases it resolved. That's the number to improve, not the escalation rate.

Pattern facts compound. Each new decision doesn't just add one precedent — it rebuilds the vendor fact record. After 5 decisions, the agent has a synthesised summary of all 5, not just 5 individual records to reason over. That's what makes the pattern detection work.

The memory layer is the whole thing. Without Hindsight, this system is a three-way match detector that escalates everything. With it, the agent learns from every decision and gets more useful over time.

Top comments (0)