DEV Community

uday kiran chellu
uday kiran chellu

Posted on

Why I Built a Strict A/B Test Into Our Sales Agent Before Writing Any Logic

Why I Built a Strict A/B Test Into Our Sales Agent Before Writing Any Logic

When engineers start building AI agents, they usually make a massive, unspoken assumption: more context is always better.

When I was building DealMind AI—a sales assistant designed to prepare reps for their next call—I almost fell into this trap. I knew that feeding the latest CRM notes into a large language model (LLM) wasn't enough. The agent needed historical memory to understand the full narrative of a deal. I decided to integrate Hindsight to retrieve relevant past interactions.

But how could I mathematically prove that this retrieved memory was actually helping the agent, rather than just confusing it with older, irrelevant noise?

Instead of just blindly wiring up the retrieval pipeline and hoping for the best, I built a strict A/B comparison routing architecture directly into our FastAPI backend. Here is the story of how I forced our AI to prove its worth, and why you should do the same.

The Problem with Untested Memory

If you dump a user's prompt, the current CRM state, and five retrieved historical memories into an LLM, it will confidently generate a response. But you have no baseline. If the agent recommends a "phased rollout", did it do so because the historical memory revealed a budget constraint, or did it just hallucinate a generic B2B sales tactic?

Without a baseline, you cannot tune your retrieval strategy. If your retrieval is too broad, you dilute the LLM's attention. I needed a way to isolate the exact impact of agent memory.

The Architecture: Dual-Routing the Agent

To solve this, I designed a /comparison endpoint. When the frontend requests call preparation, it doesn't just hit a single logic flow. It triggers a dual-execution path.

The key to a fair A/B test is isolating variables. Both paths must use the exact same user query and the exact same underlying LLM. The only variable we toggle is the contextual payload.

Here is what the routing logic looks like in our prepare.py file:


@router.post("/comparison")
def compare_preparation(request: PrepareRequest):
    # Fetch ONLY the most recent CRM note for the baseline
    latest_context = get_latest_crm_context(request.deal_id)
    without_memories = [latest_context] if latest_context else []

    response = {
        "status": "success",
        "deal_id": request.deal_id,
        "without_memory": None,
        "with_hindsight": None
    }

    # PATH A: The Baseline (Without Memory)
    if request.mode in ["both", "without_memory"]:
        try:
            # Analyze using only the latest snapshot
            response["without_memory"] = analyzer.analyze_deal(
                memories=without_memories, 
                query=request.query
            )
        except Exception as e:
            logger.error(f"Without Memory Analysis Error: {e}")

    # PATH B: The Enriched Route (With Hindsight)
    if request.mode in ["both", "with_hindsight"]:
        # ... [Hindsight retrieval logic executes here] ...
Enter fullscreen mode Exit fullscreen mode

By supporting a mode="both" flag, I could hit the endpoint via cURL during development and immediately see a side-by-side JSON comparison of how the agent's logic shifted when historical memory was injected.

The Results: Catching the Invisible Objections


The A/B test immediately paid off. By comparing the two outputs side-by-side, the value of historical retrieval became undeniably clear.

In the "Without Memory" baseline, the agent saw a prospect scrutinizing pricing in the latest CRM note. Predictably, it recommended generic defensive tactics, like offering a 10% discount.

But in the "With Hindsight" flow, the agent received the chronological history retrieved by Hindsight. It saw that three weeks prior, the prospect had explicitly stated budget was not a blocker. The agent's recommendation entirely flipped. Instead of offering a discount, it correctly diagnosed a structural shift in the prospect's stance and recommended proposing a phased rollout to mitigate the newfound risk.

I didn't have to guess if the memory integration worked; the A/B test proved it.

Lessons Learned

Building this A/B testing architecture taught me a few critical lessons about engineering LLM applications:

1. Establish a naive baseline early

Never build a complex RAG (Retrieval-Augmented Generation) pipeline before building a naive baseline. If your baseline (just the latest CRM note) gets you 80% of the way there, you might not need a complex memory layer. If it fails terribly, you now have a yardstick to measure your improvements against.

2. Fair comparisons require isolated variables

When testing your agent's memory, you cannot change the underlying model temperature, the system prompt, or the user query. The only thing that should change between your control and your experiment is the retrieved contextual payload.

3. Expose the A/B test to the user

We ended up keeping the A/B test in the final UI. Users can literally click a toggle to see what the agent would have recommended without historical memory. This contrast builds massive user trust, as they can visually see the exact value the memory layer provides.

If you are building an AI agent, stop assuming your context layer is working. Build a strict A/B test, isolate your baseline, and check out the Hindsight docs to see how surgical memory retrieval can completely change your agent's reasoning.

Shout-out to: Code.in.

Top comments (0)