DEV Community

Cover image for The Bug I Found Wasn't in My Code — It Was in a Buyer's Promise
PAVAN SIVA KRISHNA GADDAM
PAVAN SIVA KRISHNA GADDAM

Posted on

The Bug I Found Wasn't in My Code — It Was in a Buyer's Promise

The Bug I Found Wasn't in My Code — It Was in a Buyer's Promise

The first time a buyer says "payment is stuck," it's a reason. The fifth time the same pattern shows up after a missed commitment, it's history — and that distinction is why I built Vasool around Hindsight instead of another stateless prompt.

What Vasool Does

Vasool is an AI collections assistant for B2B trade credit. It doesn't just generate reminders. It builds a daily collection plan from open invoices, understands how each buyer has actually behaved, predicts an expected payment date, and drafts a follow-up that reflects the buyer's prior communication and commitments.

The architecture splits into two responsibilities: Hindsight provides durable agent memory, and Groq provides fast LLM inference for prediction and message generation. A Streamlit console sits on top, exposing a prioritized daily plan, a buyer memory view, an interactive simulator, and a learning-curve backtest.

I don't treat the raw event log as the agent's memory just because the data exists on disk. events.json is the source ledger. Hindsight is the memory layer the agent actually queries when it needs context — and that split matters because the project needs three distinct operations: store an interaction, retrieve relevant history, and turn accumulated history into behavioral understanding. That maps directly onto Hindsight's retain, recall, and reflect primitives, described in the Hindsight docs as ingestion, retrieval, and synthesis over memory, respectively.

I think of Vasool as an accounting problem wrapped around a memory problem.

The Core Decision: A Reminder Is Not the Unit of Context

My biggest design decision was simple: store the interaction, not just the latest prompt.

Collections work is sequential — invoice, reminder, buyer reply, promise, missed promise, maybe another reply, maybe a payment. A stateless LLM sees each message in isolation unless I keep rebuilding that history into the prompt every time. That's the wrong abstraction.

Every useful interaction gets retained:

def retain_memory(content: str, context: str = "buyer_reply") -> bool:
    with get_hindsight() as client:
        client.retain(bank_id=BANK_ID, content=content, context=context)
Enter fullscreen mode Exit fullscreen mode

When a buyer replies "Client payment stuck, I will pay Friday," a generic agent just sends another reminder. Vasool first asks memory:

result = client.recall(
    bank_id=BANK_ID,
    query=f"What is the full history of {name}? Include all promises, reminders, replies, and payments.",
)
Enter fullscreen mode Exit fullscreen mode

Recall tells me what happened. It doesn't give me the higher-order pattern I actually need, so Vasool follows up with reflect:

response = client.reflect(
    bank_id=BANK_ID,
    query=f"What is {name}'s payment behavior? Do they keep promises? What tactics have worked?",
)
Enter fullscreen mode Exit fullscreen mode

That's the point where memory turns into operational context — a buyer profile, not a pile of retrieved messages. This is also why I built on Hindsight instead of a thin vector-search wrapper around a transcript store: the broader agent-memory model, as Vectorize describes it, is persistent accumulation of facts and beliefs over time, not just search over old text.

The Before/After That Changed the Design

A buyer writes: "Payment is delayed because our client payment is stuck. I'll clear it next Friday."

Before (stateless):

Namaste team, this is a reminder regarding the pending invoice.
Kindly confirm when payment will be made.
Enter fullscreen mode Exit fullscreen mode

Nothing's wrong with it. It's just context-free.

After (memory-backed): Vasool already knows this buyer has missed earlier promised dates.

Namaste team,

Aapne is invoice ke liye pehle bhi payment date confirm ki thi.
Aapne ab Friday ka commitment diya hai. Kindly confirm the
expected transfer date once the client payment is received.

Dhanyavaad.
Enter fullscreen mode Exit fullscreen mode

The difference isn't smarter phrasing. The message is conditioned on a remembered commitment — that's the behavior I wanted: the next action changes because the agent learned from what happened before.

Keeping Prediction Separate From Memory

Hindsight answers "what do we know about this buyer?" Groq answers "given that history, what date should I predict, and how should I phrase this?" Keeping those separate gives me a clean test surface: in the backtest, I only pass events that existed before the invoice being predicted, so the evaluation can't leak future information into the forecast. A model that's "accurate" because it already saw the answer isn't forecasting anything.

Results: What the Agent Sees Over Time

The repo includes a synthetic six-month event history and a backtest measuring prediction error against how much buyer history is available:

History size Mean absolute date error
8 events 19.4 days
13 events 10.0 days
20 events 15.1 days

I'm not claiming "more memory always improves accuracy" — it doesn't, and the largest bucket here regresses. I kept that in rather than cherry-picking a clean upward curve, because the honest result is more useful: context helps, but it isn't a substitute for a better forecasting model. One synthetic buyer sequence does go from a 22-day error on an early prediction to a 1-day error later, once more prior interactions are available — which is the shape I was hoping to see, just not a guarantee.

The operational output is a ranked collection plan: open invoices, overdue amount, days overdue, predicted date, priority, channel, tone, and an evidence trail pointing back to the memory that justified it. That's a more useful artifact than a chatbot saying "please send a reminder."

What I Learned

Memory should change decisions, not just prompts. The worst version of memory is invisible — stored somewhere, retrieved occasionally, barely affecting output. A remembered commitment should visibly move the forecast, the profile, or the next message.

Recall and reflect solve different problems. Recall is for evidence. Reflect is for synthesized understanding. Treating them as interchangeable makes the system harder to reason about.

Keep the source ledger separate from the memory layer. I still keep events.json as a deterministic source of truth for replay and debugging, even though Hindsight is what the agent actually queries.

Test for temporal leakage. "History" has to mean history — the backtest only passes events that existed before the date being predicted, which is much closer to the real decision boundary the agent would face live.

Personalization needs guardrails. Vasool's system rules say: never threaten, never invent facts, use only supplied memories, stay firm but respectful. Memory makes personalization possible, but it can just as easily amplify a bad assumption without rules constraining what the agent's allowed to say.

Honest Limitation

Vasool's results are currently demonstrated on synthetic data — six months of buyer behavior archetypes, not a production portfolio, and a backtest that's explicitly a synthetic evaluation rather than a real-world benchmark. The current build drafts WhatsApp-ready messages but doesn't yet run a live WhatsApp Business API loop, and invoice reconciliation is represented through the project's event data rather than a real accounting-system connector. I'm using these numbers to validate the architecture and the evaluation method, not to claim production-grade collection lift.

Closing

An LLM writing a polite collections message is the easy part. The interesting part is what happens when the agent remembers that this buyer already made — and missed — the same kind of promise before. Once that history is part of the agent's operating context, a reminder stops being a template and becomes the next step in an ongoing record. That's the practical definition of agent memory I ended up with: not remembering everything, but remembering enough to make the next decision different.




Top comments (1)

Collapse
 
jami_saidinesh_bb9ace5d2 profile image
Jami Sai Dinesh •

Good work 👏