DEV Community

Pulla Indira Keerthana
Pulla Indira Keerthana

Posted on

What happened when I tested my agent on messy invoices

The Reality Check: When Clean Architecture Meets Messy Enterprise Data
You spend days setting up a pristine agent architecture. Your 3-way matching logic is tight, your tool calls are well-defined, and your state management handles database lookups without breaking a sweat.

Then, you feed it a real batch of invoices from actual corporate vendors.

And your agent immediately starts face-planting.

When I ran my invoice-processing agent through a batch of genuinely messy, unstructured, and edge-case-heavy enterprise data, everything broke in ways the clean test scripts never warned me about. Here is what happened, why it broke, and how I had to fix it.

  1. The "Invisible Variance" Problem (OCR Noise) The first failure came from standard invoice data ingestion. In theory, a line item says ₹500,000. In practice, a PDF generated from a legacy vendor portal formats it as INR 5,00,000.00 (Incl. Freight), or worse, splits the tax across two unindexed lines.

What the agent did: Because the string didn't match the purchase order's exact numeric token format, the 3-way matching tool threw a catastrophic mismatch error on an invoice that was actually completely fine.

The fix: I had to implement a normalization layer before the agent ever touched the data. Instead of letting the LLM guess numbers from raw text blocks, I forced a strict parsing step that strips currency symbols, normalizes Indian numbering formats (lakhs/crores vs. millions), and separates tax metadata into distinct JSON fields.

  1. The Memory Trap (Blindly Trusting Past Resolutions) This was the most dangerous failure mode.

An invoice arrived from a recurring vendor (let's call them ACME Corp) with a ₹50,000 discrepancy. My agent's memory module kicked in, recalled that last month ACME had a ₹50,000 discrepancy, and cheerfully concluded: "Ah, this is just that approved contract amendment again. Auto-approving."

Except... it was a completely different purchase order, and the contract amendment had expired two weeks prior.

What the agent did: It over-indexed on vector similarity (vendor name match + exact variance amount) and ignored the underlying metadata context.

The fix: I had to strip the agent's autonomy to make final financial decisions. I introduced a strict Metadata Guardrail: memory can suggest a past resolution, but if a primary key like the PO reference or contract ID doesn't match the current document payload, the agent is hard-coded to trigger a conflict warning and halt. Memory became an investigative hint, not a shortcut.

  1. Tool-Calling Loops on Ambiguous Surcharges When dealing with messy logistics invoices, unexpected line items like "Emergency Fuel Surcharge" or "Handling Fee" pop up with zero prior documentation.

What the agent did: When it couldn't find a matching clause in the master contract, the agent panicked and entered an infinite tool-calling loop—querying the vendor history, re-fetching the goods receipt, and searching vector memory five times in a row for something that literally did not exist. It blew through tokens and timed out.

The fix: I introduced circuit breakers for tool execution limits. If an agent calls the same tool or fails to synthesize a hypothesis within 3 iteration steps, it must abort the loop, summarize what it did find, and escalate the case to a human reviewer with a clear note: "Missing documentation for fuel surcharge."

What I Learned From the Mess
Messy data exposes the difference between a toy agent demo and a production-grade workflow.

LLMs are terrible at raw arithmetic matching: Let your deterministic code handle math and comparison logic. Use the LLM strictly for contextual reasoning and synthesis.

State and memory need strict boundaries: If you let an agent search vector memory without verifying relational keys (like PO numbers and timestamps), your agent will hallucinate patterns out of coincidence.

Graceful failure is a feature: An agent that knows when to stop, raise its hands, and say "I need a human to look at this" is infinitely more valuable than one that confidently automates a financial mistake.

Have you ever deployed an agent into a domain with messy, real-world data? What broke first for you? Let’s discuss in the comments.

Top comments (0)