DEV Community

Cover image for Showing Reviewers What Hindsight Remembered, Not Just the Answer
Sathwika Kancherla
Sathwika Kancherla

Posted on

Showing Reviewers What Hindsight Remembered, Not Just the Answer

An accounts payable reviewer doesn't need an agent that says "approve." They need one that says "approve, because of these three invoices," and then lets them check.

I built the interface for our AP exception agent, and almost every design decision came down to one question:

Can the reviewer see why?

The agent itself compares invoices to purchase orders, asks Hindsight agent memory how similar exceptions were resolved before, and recommends an action.

This article is about the part people actually look at.

Who Uses This and What They Need

The user is an AP clerk working through a queue of invoices that didn't match their POs.

They know the vendors better than any model does. What they lack is time and a perfect memory of what their colleagues decided three months ago.

So the interface has three jobs:

Show what's wrong. Show what the agent recommends. Show what it remembered that led there.

We built it in Streamlit. The layout is an invoice queue on the left and the selected invoice on the right, with invoice lines next to PO lines, then the detected exceptions, then the agent's decision.

Show the Problem Before the AI Says Anything

The first thing a reviewer sees is not the agent's opinion.

It's the invoice and the PO side by side, with every line that isn't on the PO highlighted in red. A tax difference shows in red, and a payment terms mismatch gets its own red warning line.

Below that, the exceptions detected by plain Python are listed in words.

This matters for trust.

If the first thing on screen is an AI verdict, people either rubber-stamp it or dismiss it. If they see the problem first, the agent's recommendation reads as a second opinion on something they already understand.

The most useful screen we built runs the same invoice twice, once without memory and once with, and shows both decisions next to each other.

For the Meridian Logistics invoice with a separate freight line, the left column says the agent has no history and escalates.

The right column explains that an August 1 contract put freight into the unit price, cites the contract change and the two compliant invoices since then, and recommends against approval.

We built it for the demo, but it turned out to be the fastest way to explain the product to anyone.

Nobody needs a definition of agent memory after seeing those two columns.

*Showing Memory, Not Just Using It *

Every decision card lists the evidence the agent cited, and an expander shows exactly what Hindsight returned.

There are two kinds of content in there, and separating them was important.

Hindsight's reflect gives a synthesized answer: what past precedent says and whether anything changed.

recall gives the underlying facts.

Our first version dumped all recalled items into one list, and it looked repetitive, because recall returns both raw facts (world type) and consolidated observation entries about the same events.

So the expander shows three sections:

obs = [f for f in mem["recalled"] if f["type"] == "observation"] facts = [f for f in mem["recalled"] if f["type"] != "observation"]

First the reflect synthesis, then the consolidated observations, then the raw facts.

A reviewer who trusts the agent reads the first paragraph. A skeptical one can drill down to individual invoices with dates and reviewer names.

We also checked the citations against the source data.

The agent never cited an invoice that didn't exist, which is the property that makes showing citations worth doing at all.

The Dollar Sign Bug

A small one, but it cost me time—and it will cost you time too.

Our totals line rendered as broken green monospace text.

Streamlit's markdown treats text between two dollar signs as a math formula, and an invoice line like:

Subtotal $742.38 | Tax $59.39

has two dollar signs.

Every piece of text the agent generates is full of amounts, so every string gets escaped before display:

def esc(text): """Streamlit reads $...$ as math, so escape every dollar sign in displayed text.""" return str(text).replace("$", "\$")

Ten Seconds Is a Design Problem

Without memory, a decision takes about a second.

With memory, about ten, almost all of it in Hindsight's reflect call.

Ten seconds of a frozen screen feels broken.

Ten seconds with a spinner that says:

"Consulting Hindsight memory..."

feels like the agent doing research—which is what it's doing.

Streamlit also reruns the whole script on every click, so without caching, clicking between invoices would repeat every API call.

Results are cached per invoice and mode:

def run(inv, memory_on): key = (inv_key(inv), memory_on) if key not in state.results: msg = "Consulting Hindsight memory..." if memory_on else "Deciding without memory..." with st.spinner(msg): state.results[key] = agent.process(inv, memory_on=memory_on) return state.results[key]

Buttons That Teach the Agent

Below each decision are Approve, Reject, and Escalate buttons.

Clicking one retains the outcome into Hindsight as a new memory, so the next similar invoice benefits.

That closes the loop between the reviewer and the agent.

It also created a problem we didn't anticipate:

Rehearsal clicks are permanent.

If you click Approve on a test invoice while practicing a demo, the agent now "remembers" an approval that never happened.

We added a warning in the sidebar and planned to reseed a clean memory bank before recording, which costs pennies.

Any agent that learns from its UI needs a story for test data.

Lessons

  1. Show the problem before the recommendation.
    Reviewers trust an agent more when they already understand what it's judging.

  2. Make memory inspectable at two levels.
    A synthesized summary for speed, raw facts for verification.

  3. Citations are only worth showing if they're real.
    Audit them against your data before you put them in front of users.

  4. Design for latency.
    A spinner that names what the agent is doing turns a wait into a feature.

  5. If the UI teaches the agent, plan for rehearsals.
    Keep a way to reset memory to a clean state.

The Hindsight documentation covers recall result types and reflect, and Vectorize has a clear explainer on what agent memory is.

Our full code, including the Streamlit app, is on GitHub:

https://github.com/satya-harika08/ap-exception-agent

Top comments (0)