DEV Community

Monky
Monky

Posted on

I Didn't Want an AI That Remembered. I Wanted One I Could Catch Lying

Building a sales agent with Hindsight taught me that memory isn't the hard part. Knowing when to trust the memory is.

The first version looked impressive.

Give the agent a deal.

Let it recall similar historical deals.

Ask it what usually happened.

It produced convincing answers.

Then I checked the numbers.

That's when the interesting part started.

The agent said a particular objection had caused most deals to fail.

The actual data didn't support that conclusion.

The model hadn't necessarily hallucinated a deal.

It had done something more subtle:

It had reasoned from the memories it happened to retrieve.

That isn't the same as reasoning from the complete dataset.

So the architecture changed.

The model writes the words. Code does the math.

That became the rule for the entire system.

Hindsight, Vectorize's open-source agent memory system, handles the memory layer.

The application handles deterministic aggregation.

The pipeline is:

retain → extract → recall → reflect → score → explain → draft

Hindsight's documentation describes retain, recall, and reflect as its core operations.

I use those operations to give the agent historical experience, but I don't let an LLM decide statistical facts about that experience.

The retrieval problem nobody notices at first

Suppose the team has 200 closed deals.

A new prospect mentions a budget freeze.

The agent recalls the most relevant historical deals.

Maybe it gets 10.

That's great for an LLM.

Ten relevant examples are plenty of context for generating a response.

But now ask:

«"How many budget-freeze deals at the Evaluation stage were lost?"»

Suddenly those ten memories aren't necessarily enough.

The answer might depend on 17 records.

Or 43.

Or 6.

Retrieval is optimized for relevance.

Statistics are optimized for completeness.

Those are different problems.

So I split them.

Recall is for context.

Complete retrieval is for counting.

The actual pattern comes from code

The historical deals are structured by objection, stage, outcome, and other attributes.

Then the application calculates the pattern:

scoped = [
d for d in deals
if d.objection == objection
and d.objection_stage == stage
]

lost = sum(d.outcome == "lost" for d in scoped)
won = len(scoped) - lost

Now the agent has something much safer to explain:

6 lost / 10 total

rather than an LLM-generated approximation of what happened.

Hindsight's reflection can still add useful context around the recalled experiences.

But if reflection says something different from the deterministic count, the count wins.

Then I found the second problem

What happens when there are zero wins?

One historical competitor had:

5 losses

0 wins

A naive loss ratio gives:

1.0

Then "logit(1.0)" goes to infinity.

Not exactly the kind of thing you want appearing in a live sales dashboard.

So I clamped the ratio before converting it to log-odds.

The exact clamp range is an engineering choice.

The important lesson was simpler:

Even a reasonable mathematical model can break at the edges of real data.

Real datasets contain zeroes.

They contain tiny samples.

They contain weird distributions.

They contain historical exceptions.

The agent has to survive those cases.

A live deal makes the difference obvious

Imagine an Evaluation-stage prospect.

First, a competitor is mentioned.

The agent recalls seven historical deals involving that competitor.

Six were lost.

The score moves slightly because a competitor mention alone isn't strong enough evidence.

Then the prospect says:

«"Finance has frozen new vendor spend."»

Now the relevant historical pattern is:

6 of 10 lost.

The system explains the change:

«"Dropped 15 pts: budget freeze mentioned, and 6 of 10 past budget_freeze deals at Evaluation were lost."»

The pattern crosses the drafting threshold.

The agent recalls the four historical deals that survived the same objection and uses those responses to draft a follow-up.

Later, the prospect responds positively.

The score recovers.

The historical risk doesn't disappear.

That's intentional.

A positive response doesn't erase what happened historically.

It just adds new evidence.

Memory changes behavior

This is where Hindsight becomes more than a retrieval layer.

Add another lost budget-freeze deal.

The historical pattern changes:

6 of 10 → 7 of 11

Nothing else changed.

No prompt.

No retraining.

No model update.

The memory changed.

The agent's behavior changed with it.

That means a newly closed deal isn't just another CRM record.

It's another piece of evidence available to the next call.

But memory can also become dangerous

A shared bank means one bad record can potentially influence future decisions.

So I added an integrity check.

def integrity_check(client, expected):
recalled = fetch_all_deals(
client,
bank_id="sales-team-shared"
)

return (
    "COUNTS VERIFIED"
    if tally(recalled) == expected
    else "COUNT MISMATCH"
)
Enter fullscreen mode Exit fullscreen mode

If the complete dataset can't be reconstructed, the system doesn't silently produce a statistic.

And if Hindsight isn't reachable, the application falls back to a local store and says so.

The system should be allowed to say:

"I don't have enough verified memory to make this claim."

That's much more useful than a confident number built from incomplete retrieval.

Five rules I ended up with

  1. Use memory for experience, not blind authority. A recalled example is evidence, not automatically the complete truth.

  2. Separate retrieval from aggregation. What an LLM needs for context isn't necessarily what code needs for statistics.

  3. Keep the numbers traceable. Every important count should lead back to actual records.

  4. Treat edge cases as part of the design. Zero wins, tiny samples, and contradictory evidence aren't unusual enough to ignore.

  5. Let the agent explain uncertainty. A system that knows when its evidence is incomplete is more useful than one that always produces an answer.

That's ultimately what I wanted from Hindsight.

Not a system that simply remembers everything.

A system that lets an agent use past experience without pretending that memory is automatically truth.

The model can write the explanation.

Hindsight can provide the experience.

The code can verify the numbers.

And the rep gets something much more useful than:

«"Based on my analysis, this deal looks risky."»

They get:

«"Here is what happeneI Didn't Want an AI That Remembered. I Wanted One I Could Catch Lying.

Building a sales agent with Hindsight taught me that memory isn't the hard part. Knowing when to trust the memory is.

The first version looked impressive.

Give the agent a deal.

Let it recall similar historical deals.

Ask it what usually happened.

It produced convincing answers.

Then I checked the numbers.

That's when the interesting part started.

The agent said a particular objection had caused most deals to fail.

The actual data didn't support that conclusion.

The model hadn't necessarily hallucinated a deal.

It had done something more subtle:

It had reasoned from the memories it happened to retrieve.

That isn't the same as reasoning from the complete dataset.

So the architecture changed.

The model writes the words. Code does the math.

That became the rule for the entire system.

Hindsight, Vectorize's open-source agent memory system, handles the memory layer.

The application handles deterministic aggregation.

The pipeline is:

retain → extract → recall → reflect → score → explain → draft

Hindsight's documentation describes retain, recall, and reflect as its core operations.

I use those operations to give the agent historical experience, but I don't let an LLM decide statistical facts about that experience.

The retrieval problem nobody notices at first

Suppose the team has 200 closed deals.

A new prospect mentions a budget freeze.

The agent recalls the most relevant historical deals.

Maybe it gets 10.

That's great for an LLM.

Ten relevant examples are plenty of context for generating a response.

But now ask:

«"How many budget-freeze deals at the Evaluation stage were lost?"»

Suddenly those ten memories aren't necessarily enough.

The answer might depend on 17 records.

Or 43.

Or 6.

Retrieval is optimized for relevance.

Statistics are optimized for completeness.

Those are different problems.

So I split them.

Recall is for context.

Complete retrieval is for counting.

The actual pattern comes from code

The historical deals are structured by objection, stage, outcome, and other attributes.

Then the application calculates the pattern:

scoped = [
d for d in deals
if d.objection == objection
and d.objection_stage == stage
]

lost = sum(d.outcome == "lost" for d in scoped)
won = len(scoped) - lost

Now the agent has something much safer to explain:

6 lost / 10 total

rather than an LLM-generated approximation of what happened.

Hindsight's reflection can still add useful context around the recalled experiences.

But if reflection says something different from the deterministic count, the count wins.

Then I found the second problem

What happens when there are zero wins?

One historical competitor had:

5 losses

0 wins

A naive loss ratio gives:

1.0

Then "logit(1.0)" goes to infinity.

Not exactly the kind of thing you want appearing in a live sales dashboard.

So I clamped the ratio before converting it to log-odds.

The exact clamp range is an engineering choice.

The important lesson was simpler:

Even a reasonable mathematical model can break at the edges of real data.

Real datasets contain zeroes.

They contain tiny samples.

They contain weird distributions.

They contain historical exceptions.

The agent has to survive those cases.

A live deal makes the difference obvious

Imagine an Evaluation-stage prospect.

First, a competitor is mentioned.

The agent recalls seven historical deals involving that competitor.

Six were lost.

The score moves slightly because a competitor mention alone isn't strong enough evidence.

Then the prospect says:

«"Finance has frozen new vendor spend."»

Now the relevant historical pattern is:

6 of 10 lost.

The system explains the change:

«"Dropped 15 pts: budget freeze mentioned, and 6 of 10 past budget_freeze deals at Evaluation were lost."»

The pattern crosses the drafting threshold.

The agent recalls the four historical deals that survived the same objection and uses those responses to draft a follow-up.

Later, the prospect responds positively.

The score recovers.

The historical risk doesn't disappear.

That's intentional.

A positive response doesn't erase what happened historically.

It just adds new evidence.

Memory changes behavior

This is where Hindsight becomes more than a retrieval layer.

Add another lost budget-freeze deal.

The historical pattern changes:

6 of 10 → 7 of 11

Nothing else changed.

No prompt.

No retraining.

No model update.

The memory changed.

The agent's behavior changed with it.

That means a newly closed deal isn't just another CRM record.

It's another piece of evidence available to the next call.

But memory can also become dangerous

A shared bank means one bad record can potentially influence future decisions.

So I added an integrity check.

def integrity_check(client, expected):
recalled = fetch_all_deals(
client,
bank_id="sales-team-shared"
)

return (
    "COUNTS VERIFIED"
    if tally(recalled) == expected
    else "COUNT MISMATCH"
)
Enter fullscreen mode Exit fullscreen mode

If the complete dataset can't be reconstructed, the system doesn't silently produce a statistic.

And if Hindsight isn't reachable, the application falls back to a local store and says so.

The system should be allowed to say:

"I don't have enough verified memory to make this claim."

That's much more useful than a confident number built from incomplete retrieval.

Five rules I ended up with

  1. Use memory for experience, not blind authority. A recalled example is evidence, not automatically the complete truth.

  2. Separate retrieval from aggregation. What an LLM needs for context isn't necessarily what code needs for statistics.

  3. Keep the numbers traceable. Every important count should lead back to actual records.

  4. Treat edge cases as part of the design. Zero wins, tiny samples, and contradictory evidence aren't unusual enough to ignore.

  5. Let the agent explain uncertainty. A system that knows when its evidence is incomplete is more useful than one that always produces an answer.

That's ultimately what I wanted from Hindsight.

Not a system that simply remembers everything.

A system that lets an agent use past experience without pretending that memory is automatically truth.

The model can write the explanation.

Hindsight can provide the experience.

The code can verify the numbers.

And the rep gets something much more useful than:

«"Based on my analysis, this deal looks risky."»

They get:

«"Here is what happened to ten deals like this one, here is what survived, and here is exactly how we know."»

Code: [repo link] · Demo: [demo link]d to ten deals like this one, here is what survived, and here is exactly how we know."»

Code: [repo link] · Demo: [demo link]I Didn't Want an AI That Remembered. I Wanted One I Could Catch Lying.

Building a sales agent with Hindsight taught me that memory isn't the hard part. Knowing when to trust the memory is.

The first version looked impressive.

Give the agent a deal.

Let it recall similar historical deals.

Ask it what usually happened.

It produced convincing answers.

Then I checked the numbers.

That's when the interesting part started.

The agent said a particular objection had caused most deals to fail.

The actual data didn't support that conclusion.

The model hadn't necessarily hallucinated a deal.

It had done something more subtle:

It had reasoned from the memories it happened to retrieve.

That isn't the same as reasoning from the complete dataset.

So the architecture changed.

The model writes the words. Code does the math.

That became the rule for the entire system.

Hindsight, Vectorize's open-source agent memory system, handles the memory layer.

The application handles deterministic aggregation.

The pipeline is:

retain → extract → recall → reflect → score → explain → draft

Hindsight's documentation describes retain, recall, and reflect as its core operations.

I use those operations to give the agent historical experience, but I don't let an LLM decide statistical facts about that experience.

The retrieval problem nobody notices at first

Suppose the team has 200 closed deals.

A new prospect mentions a budget freeze.

The agent recalls the most relevant historical deals.

Maybe it gets 10.

That's great for an LLM.

Ten relevant examples are plenty of context for generating a response.

But now ask:

«"How many budget-freeze deals at the Evaluation stage were lost?"»

Suddenly those ten memories aren't necessarily enough.

The answer might depend on 17 records.

Or 43.

Or 6.

Retrieval is optimized for relevance.

Statistics are optimized for completeness.

Those are different problems.

So I split them.

Recall is for context.

Complete retrieval is for counting.

The actual pattern comes from code

The historical deals are structured by objection, stage, outcome, and other attributes.

Then the application calculates the pattern:

scoped = [
d for d in deals
if d.objection == objection
and d.objection_stage == stage
]

lost = sum(d.outcome == "lost" for d in scoped)
won = len(scoped) - lost

Now the agent has something much safer to explain:

6 lost / 10 total

rather than an LLM-generated approximation of what happened.

Hindsight's reflection can still add useful context around the recalled experiences.

But if reflection says something different from the deterministic count, the count wins.

Then I found the second problem

What happens when there are zero wins?

One historical competitor had:

5 losses

0 wins

A naive loss ratio gives:

1.0

Then "logit(1.0)" goes to infinity.

Not exactly the kind of thing you want appearing in a live sales dashboard.

So I clamped the ratio before converting it to log-odds.

The exact clamp range is an engineering choice.

The important lesson was simpler:

Even a reasonable mathematical model can break at the edges of real data.

Real datasets contain zeroes.

They contain tiny samples.

They contain weird distributions.

They contain historical exceptions.

The agent has to survive those cases.

A live deal makes the difference obvious

Imagine an Evaluation-stage prospect.

First, a competitor is mentioned.

The agent recalls seven historical deals involving that competitor.

Six were lost.

The score moves slightly because a competitor mention alone isn't strong enough evidence.

Then the prospect says:

«"Finance has frozen new vendor spend."»

Now the relevant historical pattern is:

6 of 10 lost.

The system explains the change:

«"Dropped 15 pts: budget freeze mentioned, and 6 of 10 past budget_freeze deals at Evaluation were lost."»

The pattern crosses the drafting threshold.

The agent recalls the four historical deals that survived the same objection and uses those responses to draft a follow-up.

Later, the prospect responds positively.

The score recovers.

The historical risk doesn't disappear.

That's intentional.

A positive response doesn't erase what happened historically.

It just adds new evidence.

Memory changes behavior

This is where Hindsight becomes more than a retrieval layer.

Add another lost budget-freeze deal.

The historical pattern changes:

6 of 10 → 7 of 11

Nothing else changed.

No prompt.

No retraining.

No model update.

The memory changed.

The agent's behavior changed with it.

That means a newly closed deal isn't just another CRM record.

It's another piece of evidence available to the next call.

But memory can also become dangerous

A shared bank means one bad record can potentially influence future decisions.

So I added an integrity check.

def integrity_check(client, expected):
recalled = fetch_all_deals(
client,
bank_id="sales-team-shared"
)

return (
    "COUNTS VERIFIED"
    if tally(recalled) == expected
    else "COUNT MISMATCH"
)
Enter fullscreen mode Exit fullscreen mode

If the complete dataset can't be reconstructed, the system doesn't silently produce a statistic.

And if Hindsight isn't reachable, the application falls back to a local store and says so.

The system should be allowed to say:

"I don't have enough verified memory to make this claim."

That's much more useful than a confident number built from incomplete retrieval.

Five rules I ended up with

  1. Use memory for experience, not blind authority. A recalled example is evidence, not automatically the complete truth.

  2. Separate retrieval from aggregation. What an LLM needs for context isn't necessarily what code needs for statistics.

  3. Keep the numbers traceable. Every important count should lead back to actual records.

  4. Treat edge cases as part of the design. Zero wins, tiny samples, and contradictory evidence aren't unusual enough to ignore.

  5. Let the agent explain uncertainty. A system that knows when its evidence is incomplete is more useful than one that always produces an answer.

That's ultimately what I wanted from Hindsight.

Not a system that simply remembers everything.

A system that lets an agent use past experience without pretending that memory is automatically truth.

The model can write the explanation.

Hindsight can provide the experience.

The code can verify the numbers.

And the rep gets something much more useful than:

«"Based on my analysis, this deal looks risky."»

They get:

«"Here is what happened to ten deals like this one, here is what survived, and here is exactly how we know."»

Code: [repo link] · Demo: [demo link]

Top comments (1)

Collapse
 
cailab profile image
CAI •

The split between "model writes words, code does the math" is the kind of architectural decision most people don't make until they've been burned by it. I think the same principle extends beyond claims into actions. An agent can initiate an action, but making sure that action is actually authorized through a channel the agent can't inject into is the action-side version of the same problem. Your integrity check pattern (COUNTS VERIFIED vs COUNT MISMATCH) is essentially forcing the system to acknowledge what it can and can't prove. That's harder than it sounds, and you've laid out the reasoning clearly enough that someone building their first sales agent should read this before wiring up their retrieval pipeline.