DEV Community

NihalMaredla
NihalMaredla

Posted on

Building Guardrails in FinOps Agents with Hindsight: Preventing Hallucinations

Building Guardrails in FinOps Agents with Hindsight: Preventing Hallucinations

When engineers talk about testing AI agents, they usually focus on prompt evaluations or synthetic benchmarks like MMLU. In enterprise finance operations, those benchmarks are useless. If an autonomous Accounts Payable agent misinterprets a policy and auto-approves a ₹12,000 unapproved supplier price hike, the engineering team has introduced an automated vulnerability directly into the company’s treasury.

As the test and integration engineer on LedgerMind, my job was to answer a difficult question: How do you write deterministic unit and integration tests for an agent whose decisions change as it learns?

Testing an agent with persistent memory is fundamentally different from testing stateless software. In a standard test suite, calling a function with the same input yields the same output. With memory, calling evaluate_invoice() today must yield a different—and better—result than calling it yesterday.

Here is how we built an automated verification pipeline to validate Vectorize Hindsight, test episodic memory recall, and enforce strict safety guardrails.


The Testing Dilemma: Stateful Agents vs Deterministic Tests

In our architecture, LedgerMind sits on top of enterprise ERP systems (like SAP S/4HANA) to intercept invoices flagged with discrepancies against Purchase Orders.

When we integrated Vectorize agent memory, our system had to handle three distinct lifecycle states for any given invoice:

  1. Unseen Exceptions (Cold Start): The agent must safely flag the invoice and refuse to guess.
  2. Precedent Recall (Warm State): When an approved policy rule exists, the agent must recall the human precedent, check mathematical bounds, and clear payment.
  3. Contract Violations (Hostile State): When a vendor attempts an unauthorized price increase, the agent must strictly reject auto-approval, even if other fees from that vendor were previously approved.

If your test suite does not explicitly test state isolation and boundary conditions, an agent with memory can easily become a dumb rubber stamp that blindly approves everything.


The Automated Verification Suite (test_flow.py)

To verify our memory loop before deploying to production, we built an end-to-end Python test suite that simulates time progression and validates memory boundaries:

[ STEP 1: Cold Start ] ──► Assert Decision == "FLAGGED" (Amnesia Baseline)
            │
            ▼
[ STEP 2: Retain Memory ] ──► Store Sarah Jenkins' Precedent in Hindsight
            │
            ▼
[ STEP 3: Warm Recall ] ──► Assert Decision == "AUTO_APPROVED" (₹3,500 <= ₹4,000)
            │
            ▼
[ STEP 4: Guardrail Test ] ──► Assert Decision == "REJECTED_AUDIT" (₹12,000 Unit Price Hike)
Enter fullscreen mode Exit fullscreen mode

Here is the exact implementation from our verification harness:

# test_flow.py - End-to-end verification harness for memory and guardrails
import os, json
from hindsight_service import hindsight_service
from agent import agent

def run_test_suite():
    # Step 0: Ensure deterministic state by purging previous test memories
    hindsight_service.reset_memories()

    with open("data/invoices.json", "r") as f:
        invoices = json.load(f)

    inv1 = invoices[0]  # Acme Industrial (₹3,500 freight - cold start)
    inv2 = invoices[1]  # Acme Industrial (₹3,500 freight - recurring)
    inv3 = invoices[2]  # NovaTech (₹12,000 unit price increase - hostile)

    # 1. Verify stateless baseline failure
    res1 = agent.evaluate_invoice(inv1, use_memory=False)
    assert res1["decision"] == "FLAGGED", "Failed: Agent should flag unseen discrepancy"

    # 2. Retain managerial precedent into Vectorize Hindsight
    mem = hindsight_service.retain_precedent(
        vendor_name=inv1["vendor_name"],
        invoice_id=inv1["id"],
        discrepancy_type=inv1["discrepancy_type"],
        discrepancy_amount=inv1["discrepancy_amount"],
        precedent_note="Authorized expedited air-freight surcharge up to ₹4,000 during Q3 warehouse relocation project.",
        approved_by="Sarah Jenkins (Finance Lead)"
    )
    assert len(hindsight_service.get_all_memories()) == 1, "Failed: Memory not persisted"

    # 3. Verify that warm recall auto-approves recurring invoice
    res2 = agent.evaluate_invoice(inv2, use_memory=True)
    assert res2["decision"] == "AUTO_APPROVED", "Failed: Agent should recall precedent"
    assert res2["confidence"] >= 0.90, "Failed: Confidence score too low"

    # 4. Critical Negative Test: Guardrail enforcement on price increase
    res3 = agent.evaluate_invoice(inv3, use_memory=True)
    assert res3["decision"] == "REJECTED_AUDIT", "Failed: Guardrail breached on price hike"
Enter fullscreen mode Exit fullscreen mode

Negative Testing: Engineering the Guardrail

The most critical test in our entire pipeline is Step 4: The Negative Guardrail Test.

When developing with agent memory, the biggest danger is over-generalization. If Sarah Jenkins authorizes an emergency freight surcharge for Acme Industrial, an unconstrained LLM might synthesize a rule like: "All extra charges from Acme are acceptable."

To protect against this, our agent logic in agent.py couples Hindsight’s semantic recall with explicit negative boundary filters:

# agent.py - Enforcing strict negative guardrails before memory execution
def evaluate_invoice(self, invoice, use_memory=True):
    disc_amt = invoice.get("discrepancy_amount", 0.0)
    disc_type = invoice.get("discrepancy_type", "")

    # Query Hindsight memory engine
    recall_result = hindsight_service.recall_precedents(invoice.get("vendor_name"), disc_type)

    if recall_result.get("matched"):
        mem = recall_result["memory"]
        # Mathematical boundary verification: Is surcharge within authorized limit?
        if disc_amt <= mem.get("discrepancy_amount", 0.0) + 500.0:
            return {
                "decision": "AUTO_APPROVED",
                "confidence": 0.95,
                "summary": f"Matches past approval by {mem.get('approved_by')}.",
                "memory_recalled": mem
            }

    # Strict Security Guardrail: Never auto-approve unannounced base price increases
    if "price increase" in disc_type.lower():
        return {
            "decision": "REJECTED_AUDIT",
            "confidence": 0.98,
            "summary": f"Unauthorized ₹{disc_amt:,.2f} price increase.",
            "reasoning": "The supplier raised unit prices without approval. No past rule permits base price hikes. Escalated to procurement."
        }
Enter fullscreen mode Exit fullscreen mode

By consulting Hindsight documentation, we designed our memory structures so that exceptions are strictly keyed to operational categories (freight, statutory taxes, handling), making it impossible for a freight approval to bleed into a unit price increase.


Verification Results: Real Test Output

When executing python test_flow.py, the test suite prints deterministic traces confirming that the memory loop and safety guardrails are functioning:

============================================================
  LEDGERMIND VERIFICATION SUITE: HINDSIGHT MEMORY LOOP (INR ₹)
============================================================

[STEP 1] Testing Stateless Baseline (Without Memory) on Invoice 1...
Decision: FLAGGED
Reasoning: STATELESS LLM: Detected unapproved surcharge of ₹3,500.00. Payment frozen.
>>> PASS: Stateless LLM flagged discrepancy due to amnesia.

[STEP 2] Simulating Human Approval & Retaining Precedent in Hindsight...
Memory Retained: ID=MEM-A48F9B12
Stored Rule: Acme Industrial: Exception authorized for 'Expedited Air Freight Surcharge' up to ₹4,000.00.
>>> PASS: Hindsight memory successfully retained.

[STEP 3] Testing LedgerMind WITH Hindsight Memory on Invoice 2...
Decision: AUTO_APPROVED
Confidence: 95.0%
Cited Memory: Approved by Sarah Jenkins on 2024-10-12. Surcharge within authorized tolerance.
>>> PASS: LedgerMind successfully recalled precedent and AUTO-APPROVED Invoice 2!

[STEP 4] Testing Safety Guardrail on Unauthorized Price Increase (Invoice 3)...
Decision: REJECTED_AUDIT
Reasoning: Unauthorized ₹12,000.00 price increase. No past rule permits price hikes.
>>> PASS: Agent strictly rejected unauthorized price increase without precedent.

============================================================
  ALL VERIFICATION TESTS PASSED SUCCESSFULLY!
============================================================
Enter fullscreen mode Exit fullscreen mode

Honest Dead Ends & Lessons Learned

  1. State Leaks Between Test Runs: During our first sprint, tests would randomly fail because previous runs left active memories in the local database. We had to implement an atomic reset_memories() fixture before every test execution.
  2. Fuzzy Matching Causes False Positives: We initially experimented with broad cosine similarity thresholds on invoice line items. This led to dangerous mistakes—like matching an "emergency transit fee" with a "routine delivery charge". We tuned Hindsight’s recall to require strict vendor isolation and tight semantic confidence scores ($\ge 0.90$).
  3. The Importance of Negative Tests: In AI engineering, 80% of your tests should be negative tests: verifying what the agent refuses to do. Proving that LedgerMind rejects unauthorized ₹12,000 price increases gave our team far more confidence than proving it could auto-approve freight.

Conclusion

Testing stateful, memory-augmented AI agents requires shifting from static input-output assertions to lifecycle verification: testing cold starts, memory retention, warm recall, and strict safety guardrails.

With Vectorize Hindsight, we were able to build an autonomous system that continuously adapts to managerial decisions while maintaining complete mathematical safety.

Our complete test scripts, verification suites, and API endpoints are available on our GitHub repository.

Top comments (0)