DEV Community

Mohammed Musthafa
Mohammed Musthafa

Posted on

What a Sales Agent Learned After Watching 36 Deals Win and Lose

I didn't expect a side project to change how I think about "AI memory," but building a
memory-backed sales agent on top of Hindsight
did exactly that. The part that stuck with me wasn't the recall — it was watching the
agent turn a pile of old call notes into an actual recommendation, and then watching that
recommendation get better the moment I fed it one more real conversation.

This is the story of that part of the build.

What the system does

The project is a small internal tool for B2B sales reps: a Deal Intelligence Agent.
Every call a rep logs — an objection, a stakeholder, a discount offered — gets stored as a
memory in Hindsight, tagged by company. From there the agent does three things:

  1. Pre-call briefing — recall everything known about one deal and summarize it before the next call.
  2. Cross-deal pattern insight — recall across every deal and surface what actually correlates with winning.
  3. Live learning — a rep can log a brand-new call on the spot, and the agent immediately has it available for the next question.

The first one is useful. The second and third are the reason I'm writing this.

The story: teaching it to notice what wins

I generated a synthetic dataset of 36 B2B deals across a mix of industries, each with a
pricing objection handled one of two ways: the rep either quantified the ROI, or just
offered a discount. I planted the pattern on purpose — I wanted to see whether the agent
could find it from raw call notes, not just repeat it back to me.

Here's the recall step, scoped to search across every deal instead of one:

def get_pattern_insight():
    ctx = recall_context(
        "pricing objections, rep approach and deal outcomes across all deals",
        max_tokens=3500,
        budget="high",
    )
    return ask_llm(
        f"Notes from many past deals:\n\n{ctx}\n\n"
        "Identify which objection-handling approaches correlated with won "
        "vs lost/stuck deals. Cite specific deal names as evidence. "
        "Give one clear recommendation."
    )
Enter fullscreen mode Exit fullscreen mode

The first time I ran it, the output named actual companies from my dataset and drew the
line correctly: deals where the rep led with a quantified ROI closed, deals where the rep
only discounted mostly didn't. I checked the math myself afterward, computed directly from
the data with no LLM involved, and it held up — roughly 100% of the "ROI framing" deals in
my set closed won, against 0% of the "discount only" deals.

That's a synthetic, self-planted pattern, so I'm not claiming anything about real sales
data. What mattered to me was that the agent found it from unstructured notes, not from a
label I handed it.

The part that actually surprised me: live learning

The pattern insight is a nice trick on data I already had. The part that felt different was
adding a way to teach the agent something new, mid-session, and have it show up
immediately.

def log_call():
    ss = st.session_state
    company = ss.get("log_new_name", "").strip()
    note = ss.get("log_note", "").strip()
    content = (
        f"Deal: {company} (${int(ss.get('log_size', 0)):,}, "
        f"stage: {ss.get('log_stage')}). "
        f"Call logged with {ss.get('log_who', '').strip() or 'the prospect'}: "
        f"{note} [Objection: {ss.get('log_obj')}]"
    )
    hindsight.retain(
        bank_id=BANK,
        content=content,
        tags=[f"company:{company}", f"objection:{ss.get('log_obj')}"],
    )
Enter fullscreen mode Exit fullscreen mode

I tested this live: I typed in a company that didn't exist in my dataset — "Helix
Robotics" — with a call note saying the CFO thought the price was 20% over budget and was
comparing us to a competitor. I hit save, switched to the pre-call brief screen, and the
agent generated a full, specific brief for a deal that hadn't existed sixty seconds
earlier — objection, stakeholder, a suggested opening line, all pulled from the one note I'd
just given it.

That's the moment memory stopped feeling like a database lookup and started feeling like
the thing it's supposed to be: context that accumulates.

An honest limitation

I want to be upfront about what this isn't. Thirty-six deals is a small, synthetic sample,
and I built the win/loss pattern into the data on purpose to see if the agent would find
it — this isn't evidence about what wins real sales deals. The system also only handles
pricing objections well right now; competitor and timing objections need more varied
training notes before the advice gets specific rather than generic. And it's single-user —
there's no separation between reps, so two people using it would share one memory, which
isn't how a real team would want it.

What I'd build next

Separate memory per rep is the obvious next step, so the agent can learn each person's
style instead of pooling everyone's calls together. After that, a feedback loop — letting a
rep mark whether a piece of advice actually worked — would let the pattern insight get more
honest over time instead of relying on a single snapshot of historical data.

Takeaways

  • Recall isn't the interesting part — cross-record synthesis is. Pulling back one memory is a lookup. Pulling back thirty and asking "what's the pattern" is closer to actual reasoning.
  • Tag your memories, or you'll regret it. Early on I only used semantic search and it occasionally blended unrelated deals together. Tagging by company and filtering recall on that tag fixed it — Hindsight's docs cover this properly, and it's worth reading before you build anything you plan to demo.
  • Test the live-write path, not just the read path. It's easy to build a demo that only reads from a fixed dataset. Actually writing a new memory and immediately recalling it is a much better test of whether the system works the way you think it does.
  • Plant your dataset's pattern on purpose if you're testing recall, not generation. It let me verify the agent was actually finding the signal, rather than me just trusting that it was.

If you're curious about agent memory as a
concept, or want to see the full project, the code is here:
github.com/ghaneeshkumar83-lab/deal-intelligence-agent.

Top comments (0)