DEV Community

Bhavi0816
Bhavi0816

Posted on

I Built an Agent That Improves Every Session With Hindsight

I Built an Agent That Improves Every Session With Hindsight

The first time my accounts payable agent saw an invoice from Apex Industrial Supplies, it flagged it as an exception. That was the right call. It had no idea what a normal Apex invoice looked like.

What interested me was everything after that: what happens when a human says "this one is fine," and whether the agent can carry that judgment into the next invoice.

What VendorSense does

VendorSense is an accounts payable agent built around one idea: it learns what normal looks like for every vendor. You upload an invoice (PDF, PNG, or JPG), and the agent decides whether it can be processed automatically or needs a person to look at it.

The overview tracks how many invoices were processed automatically, how many went to a human, and how many became learning events.

Every invoice moves through the same six-step pipeline:

  1. Invoice captured: the file is stored and its fields are extracted
  2. Hindsight recall: the agent asks memory what it knows about this vendor
  3. Groq reasoning: an LLM compares the invoice against what was recalled
  4. Decision: process automatically, or raise an exception with a confidence level
  5. Human review, if needed: a person approves or rejects, with a note
  6. Hindsight retain: that decision becomes memory

The agent itself is stateless. Everything it knows about vendors lives in Hindsight, an open-source agent memory system running as a separate service. Each user gets a private workspace with isolated data, so one team's vendor history never leaks into another's.

Why fixed rules don't work for invoices

Classic accounts payable automation is a pile of thresholds: flag anything over a set amount, flag anything without a PO. Those rules have the wrong shape for the problem.

"Normal" is a property of the vendor, not of the company. One vendor bills a nearly identical amount on Net 30 terms every month, and a 4% difference is worth a look. Another vendor's amounts swing widely by design. A single global threshold either floods reviewers with false alarms or lets the strange invoice through.

I wanted the agent to build that per-vendor picture from the people who actually review the invoices, and to keep refining it over time. That is a memory problem, not a rules problem.

Why Hindsight

Hindsight gives me three operations, and I only needed them to map onto two moments in the pipeline:

  • Retain stores what happened
  • Recall retrieves what's relevant before a decision
  • Reflect reasons over stored memories to surface patterns

Recall uses several retrieval strategies in parallel (semantic, keyword, graph, and temporal) rather than one similarity search. For invoices that matters. "What do we know about Apex Industrial Supplies, and how have their amounts changed?" is a question about an entity over time, and a plain embedding lookup handles it poorly. If you want the background on why this matters for agents, Vectorize's explainer on agent memory is a good primer.

The core loop

The pattern is recall before deciding, retain after a human weighs in. The wiring lives in [FILE PATH, e.g. agent/memory.py], and it looks roughly like this:

from hindsight_client import Hindsight

client = Hindsight(base_url="http://localhost:8888")

def decide(invoice, bank_id):
    # 1. Ask memory what we know about this vendor BEFORE reasoning
    recalled = client.recall(
        bank_id=bank_id,
        query=f"Previous invoices and decisions for {invoice.vendor}",
    )
    experience = "\n".join(f"- {r.text}" for r in recalled.results)

    # 2. Reason over the invoice plus recalled experience
    return reason_with_llm(invoice, experience)
Enter fullscreen mode Exit fullscreen mode
def learn_from_review(invoice, decision, reviewer_note, bank_id):
    # 3. After a human approves or rejects, retain the decision AND the why
    client.retain(
        bank_id=bank_id,
        content=(
            f"Vendor {invoice.vendor}, invoice {invoice.invoice_id}, "
            f"amount {invoice.amount}: human {decision}. "
            f"Reviewer note: {reviewer_note}"
        ),
        context="human invoice review",
    )
Enter fullscreen mode Exit fullscreen mode

(Swap in the real functions from your repo. These show the shape using the Hindsight client's recall and retain calls.)

The cold start: an exception is the right answer

On a vendor the agent knows nothing about, the interface says so plainly. Hindsight returns nothing relevant, and the agent raises an exception with a confidence level and a reason.

I like that the panel doesn't pretend. "No relevant previous vendor experience found" is followed by a line saying this is a new learning situation and that a human decision can become future experience. The empty memory is a state the system handles honestly, not an error.

The evidence trail

When the agent does have history, it shows its work. In the review screen, it lists the specific evidence behind its decision: a prior approved invoice from the same vendor on Net 30 terms, a current amount that is close to the earlier one but not identical, and identifiers that don't look like the vendor's earlier structured IDs.

Reviewers shouldn't have to trust a bare "exception." Citing the prior invoice, the amount comparison, and the malformed fields lets a person confirm or overrule the agent in seconds.

Human review is the teaching step

The review form has a note field ("Explain what you verified...") and two buttons: Approve & Teach Agent and Reject & Teach Agent. The wording is deliberate. A click isn't just closing a ticket; it's writing to memory.

The reviewer's note is retained alongside the decision. So the memory isn't a pile of yes/no labels. It contains the reasoning of the person who checked, and that's what later recalls can lean on.

A "Learning Status" counter shows how many human decisions have been retained, so it's visible when the system has actually learned something rather than just processed something.

What it looks like across sessions

First invoice from a vendor

Agent: Exception. No prior experience with this vendor. Confidence: medium.
Reviewer: Approves, with a note on what they verified.

Later invoice from the same vendor

Agent: [YOUR REAL EXAMPLE: e.g. "Matches prior approved Apex invoice pattern. Auto-processed."]

Nothing in the prompt changed between the two. The difference is what came back from recall.

What was painful

  • Extraction errors look like vendor anomalies. In one run, the extracted invoice ID came back as the generic word "Invoice" and the purchase order as a garbled fragment. The agent correctly treated that as suspicious, but the vendor wasn't the problem; the extraction was. Memory can't fix bad input, and I had to be careful not to teach the agent that a vendor's format had changed when really the OCR had failed.
  • Deciding what to retain. [YOUR REAL EXPERIENCE, e.g. "storing raw extracted text made recall noisy; storing the decision and the reviewer's reason worked better."]
  • Cold start. Every new vendor starts as an exception. That's the correct behavior, but it means the first weeks of use are review-heavy.

Lessons learned

  1. Treat "normal" as per-entity state. Global thresholds fight the shape of the problem. Memory keyed to the vendor matches it.
  2. Keep the agent stateless. Pushing memory into Hindsight made the agent easier to restart, test, and reason about.
  3. Retain the reason, not just the verdict. A reviewer's note is worth more than an approve/reject flag.
  4. Make empty memory a first-class state. "I don't know this vendor yet" should be shown, not hidden.
  5. Show the evidence. People will trust an agent that cites the prior invoice and the specific mismatch far sooner than one that just says "exception."

Try it yourself

If you're building anything where decisions should improve as humans correct them, start with the Hindsight docs and the open-source repo. Retain and recall can be working in an afternoon; the interesting part is when your agent starts using a reviewer's judgment you'd forgotten you gave it.

Top comments (0)