DEV Community

Cover image for Without Hindsight Memory, My AP Agent Was Safe but Useless
sp-182
sp-182

Posted on

Without Hindsight Memory, My AP Agent Was Safe but Useless

The first time we scored our agent, the version without memory got 5 out of 8 invoices right. That number was misleading, and figuring out why taught me more about evaluating agents than the version that actually worked.

We built an accounts payable exception agent: it compares invoices to purchase orders, looks up how similar exceptions were resolved before using Hindsight agent memory, and recommends what to do. My part was the question every agent project eventually has to answer: how do you prove memory is doing anything?

Setting up a fair comparison

The comparison only means something if memory is the only variable. So every demo invoice is processed twice, with the same model (openai/gpt-oss-120b on Groq), the same system prompt and the same exception detection. The only difference is the evidence section of the prompt. With memory on, it contains Hindsight's reflect answer and recalled facts. With memory off, it says:

Memory evidence:
NONE. No memory is available for this decision.
Enter fullscreen mode Exit fullscreen mode

The eight test invoices each have an expected outcome: two clean invoices, three routine exceptions a senior reviewer would approve without thinking (tax rounding of a few cents, a vendor that quotes its own sales order number, a vendor that prints the wrong payment terms), one exact duplicate, one freight line that a new contract made invalid, and one new vendor with no history at all.

The guardrail came first

Before any scoring, we decided what the agent must never do: auto-approve an invoice without evidence. The rule lives in code, not in the prompt:

if d["action"] == "auto_approve":
    reasons = []
    if not memory_on:
        reasons.append("no memory evidence")
    if d["confidence"] < AUTO_APPROVE_MIN_CONFIDENCE:
        reasons.append(f"confidence {d['confidence']:.2f} below {AUTO_APPROVE_MIN_CONFIDENCE}")
    if len(d["cited_evidence"]) < AUTO_APPROVE_MIN_EVIDENCE:
        reasons.append(f"fewer than {AUTO_APPROVE_MIN_EVIDENCE} cited precedents")
    if reasons:
        d["action"] = "recommend_approve"
Enter fullscreen mode Exit fullscreen mode

Reject and escalate are always recommendations a person confirms. A wrongly rejected invoice damages a vendor relationship, so the agent only gets autonomy for the cheapest mistake to reverse.

Why the first score was wrong

The first run gave memory off 5 out of 8. Looking at the individual answers, two of those "passes" were rejections where we expected escalation. My scorer accepted either, reasoning that both mean "don't pay this."

In accounts payable they're different outcomes. Reject sends the invoice back to the vendor. Escalate sends it to a senior person to investigate. The memory-off agent had rejected an invoice from a brand-new staffing vendor over eight extra hours, without any investigation. That's poor practice, and my scorer was calling it correct.

After making the scoring strict and adding a prompt rule ("choose escalate, not reject, when the vendor has no relevant history"), the memory-off agent behaved sensibly: it escalated everything it couldn't verify. Its accuracy stayed at 5 of 8, and that's when the real finding appeared.

Accuracy was the wrong metric

A cautious agent without memory can score well by escalating anything unfamiliar. It's rarely wrong, and it saves nobody any work. So we added a second metric that matters to an AP team: invoices resolved without a human.

Memory OFF Memory ON
Correct outcome 5 of 8 7 of 8
Resolved without a human 2 of 8 5 of 8

Memory off only resolved the two invoices that matched their PO exactly, which needs no intelligence at all. Memory on resolved all three routine exceptions itself and still escalated the new vendor.

The Northwind tax case shows the difference in one line each:

  • Memory off: "No prior history for Northwind Office Supply, so we cannot determine if the tax difference is acceptable."
  • Memory on: "The $0.03 tax difference is within the $0.05 rounding tolerance defined in the policy effective 2026-03-01 and matches multiple prior approvals."

I also audited every citation against the source data. Across all runs, the agent never invented an invoice number. When it cited Northwind precedents, they were real Northwind invoices with tax differences.

The Kestrel lesson: models need a threshold

One routine exception kept failing even with memory. Kestrel quotes its own sales order number on every invoice, and all seven past cases were matched and approved. The agent recognized the pattern, cited the precedents at 0.95 confidence, and still only recommended approval.

The model wasn't wrong so much as uninstructed. Nothing told it when precedent is strong enough to act. We added one general rule: with at least three consistent past resolutions of the same exception for this vendor, nothing newer overriding them, and facts that fit the pattern, auto-approve. Kestrel passed on the next run. The rule applies to every vendor, which matters: a fix written for one test case would have been overfitting.

Temperature 0 is not deterministic

The same Meridian invoice, with identical inputs, came back as "escalate" in one run and "reject" in the next. Both were defensible given the contract change, and we updated our expected answers to accept either for that invoice. But it's a real risk for anyone planning a live demo: Hindsight's reflect answer is phrased slightly differently each time, and small wording changes can tip the decision. If you're evaluating agents with memory, run each case more than once.

Cost and latency

Memory-on decisions took about 10 seconds against roughly 1 second without, and nearly all of that is the reflect call. For invoice processing that's acceptable, since a human takes minutes. Seeding 66 memories and running several full evaluations cost about ten cents in Hindsight credits, so cost was never the constraint. Reflect was the priciest operation per call at about five cents.

Limitations

Eight synthetic invoices is a small test set, and the guardrail blocks auto-approval without memory by design, so part of the workload gap is structural. We report it anyway because it reflects how the system is meant to behave, but these numbers demonstrate behavior; they're not a benchmark.

Lessons

  1. Measure the work saved, not just correctness. A memoryless agent can look accurate by escalating everything.
  2. Make your scorer as strict as the domain. "Close enough" outcomes hid the most important difference in our results.
  3. Put guardrails in code. Prompts are suggestions; an if-statement is a guarantee.
  4. Give the model an explicit threshold for acting on precedent. Otherwise it recognizes patterns and still hesitates.
  5. Run evaluations more than once. Identical inputs don't guarantee identical decisions.

The Hindsight docs explain recall, reflect and bank configuration, and Vectorize's explainer on agent memory is a good starting point if you're new to the idea. Our full code and evaluation script are on GitHub: https://github.com/satya-harika08/ap-exception-agent

Top comments (0)