<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Monky</title>
    <description>The latest articles on DEV Community by Monky (@monky_234d98fa1dab436cfac).</description>
    <link>https://dev.to/monky_234d98fa1dab436cfac</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149001%2F7e98f330-ff24-4e90-b8bc-744ccacf02b2.png</url>
      <title>DEV Community: Monky</title>
      <link>https://dev.to/monky_234d98fa1dab436cfac</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/monky_234d98fa1dab436cfac"/>
    <language>en</language>
    <item>
      <title>I Didn't Want an AI That Remembered. I Wanted One I Could Catch Lying</title>
      <dc:creator>Monky</dc:creator>
      <pubDate>Tue, 29 Sep 2026 08:26:47 +0000</pubDate>
      <link>https://dev.to/monky_234d98fa1dab436cfac/i-didnt-want-an-ai-that-remembered-i-wanted-one-i-could-catch-lying-227</link>
      <guid>https://dev.to/monky_234d98fa1dab436cfac/i-didnt-want-an-ai-that-remembered-i-wanted-one-i-could-catch-lying-227</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F73awsyk6svygflcnimee.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F73awsyk6svygflcnimee.jpg" alt=" " width="800" height="731"&gt;&lt;/a&gt;&lt;a href="https://dev.tourl"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzveet8m22gslozijfaz2.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzveet8m22gslozijfaz2.jpg" alt=" " width="800" height="731"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0lf03ku98s09mubfbqx.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0lf03ku98s09mubfbqx.jpg" alt=" " width="800" height="383"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3xmlorx0bthkbhm8i27y.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3xmlorx0bthkbhm8i27y.jpg" alt=" " width="800" height="384"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh19mryzhr5bee6gp9i6m.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh19mryzhr5bee6gp9i6m.jpg" alt=" " width="800" height="386"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Building a sales agent with Hindsight taught me that memory isn't the hard part. Knowing when to trust the memory is.&lt;/p&gt;

&lt;p&gt;The first version looked impressive.&lt;/p&gt;

&lt;p&gt;Give the agent a deal.&lt;/p&gt;

&lt;p&gt;Let it recall similar historical deals.&lt;/p&gt;

&lt;p&gt;Ask it what usually happened.&lt;/p&gt;

&lt;p&gt;It produced convincing answers.&lt;/p&gt;

&lt;p&gt;Then I checked the numbers.&lt;/p&gt;

&lt;p&gt;That's when the interesting part started.&lt;/p&gt;

&lt;p&gt;The agent said a particular objection had caused most deals to fail.&lt;/p&gt;

&lt;p&gt;The actual data didn't support that conclusion.&lt;/p&gt;

&lt;p&gt;The model hadn't necessarily hallucinated a deal.&lt;/p&gt;

&lt;p&gt;It had done something more subtle:&lt;/p&gt;

&lt;p&gt;It had reasoned from the memories it happened to retrieve.&lt;/p&gt;

&lt;p&gt;That isn't the same as reasoning from the complete dataset.&lt;/p&gt;

&lt;p&gt;So the architecture changed.&lt;/p&gt;

&lt;p&gt;The model writes the words. Code does the math.&lt;/p&gt;

&lt;p&gt;That became the rule for the entire system.&lt;/p&gt;

&lt;p&gt;Hindsight, Vectorize's open-source agent memory system, handles the memory layer.&lt;/p&gt;

&lt;p&gt;The application handles deterministic aggregation.&lt;/p&gt;

&lt;p&gt;The pipeline is:&lt;/p&gt;

&lt;p&gt;retain → extract → recall → reflect → score → explain → draft&lt;/p&gt;

&lt;p&gt;Hindsight's documentation describes retain, recall, and reflect as its core operations.&lt;/p&gt;

&lt;p&gt;I use those operations to give the agent historical experience, but I don't let an LLM decide statistical facts about that experience.&lt;/p&gt;

&lt;p&gt;The retrieval problem nobody notices at first&lt;/p&gt;

&lt;p&gt;Suppose the team has 200 closed deals.&lt;/p&gt;

&lt;p&gt;A new prospect mentions a budget freeze.&lt;/p&gt;

&lt;p&gt;The agent recalls the most relevant historical deals.&lt;/p&gt;

&lt;p&gt;Maybe it gets 10.&lt;/p&gt;

&lt;p&gt;That's great for an LLM.&lt;/p&gt;

&lt;p&gt;Ten relevant examples are plenty of context for generating a response.&lt;/p&gt;

&lt;p&gt;But now ask:&lt;/p&gt;

&lt;p&gt;«"How many budget-freeze deals at the Evaluation stage were lost?"»&lt;/p&gt;

&lt;p&gt;Suddenly those ten memories aren't necessarily enough.&lt;/p&gt;

&lt;p&gt;The answer might depend on 17 records.&lt;/p&gt;

&lt;p&gt;Or 43.&lt;/p&gt;

&lt;p&gt;Or 6.&lt;/p&gt;

&lt;p&gt;Retrieval is optimized for relevance.&lt;/p&gt;

&lt;p&gt;Statistics are optimized for completeness.&lt;/p&gt;

&lt;p&gt;Those are different problems.&lt;/p&gt;

&lt;p&gt;So I split them.&lt;/p&gt;

&lt;p&gt;Recall is for context.&lt;/p&gt;

&lt;p&gt;Complete retrieval is for counting.&lt;/p&gt;

&lt;p&gt;The actual pattern comes from code&lt;/p&gt;

&lt;p&gt;The historical deals are structured by objection, stage, outcome, and other attributes.&lt;/p&gt;

&lt;p&gt;Then the application calculates the pattern:&lt;/p&gt;

&lt;p&gt;scoped = [&lt;br&gt;
    d for d in deals&lt;br&gt;
    if d.objection == objection&lt;br&gt;
    and d.objection_stage == stage&lt;br&gt;
]&lt;/p&gt;

&lt;p&gt;lost = sum(d.outcome == "lost" for d in scoped)&lt;br&gt;
won = len(scoped) - lost&lt;/p&gt;

&lt;p&gt;Now the agent has something much safer to explain:&lt;/p&gt;

&lt;p&gt;6 lost / 10 total&lt;/p&gt;

&lt;p&gt;rather than an LLM-generated approximation of what happened.&lt;/p&gt;

&lt;p&gt;Hindsight's reflection can still add useful context around the recalled experiences.&lt;/p&gt;

&lt;p&gt;But if reflection says something different from the deterministic count, the count wins.&lt;/p&gt;

&lt;p&gt;Then I found the second problem&lt;/p&gt;

&lt;p&gt;What happens when there are zero wins?&lt;/p&gt;

&lt;p&gt;One historical competitor had:&lt;/p&gt;

&lt;p&gt;5 losses&lt;/p&gt;

&lt;p&gt;0 wins&lt;/p&gt;

&lt;p&gt;A naive loss ratio gives:&lt;/p&gt;

&lt;p&gt;1.0&lt;/p&gt;

&lt;p&gt;Then "logit(1.0)" goes to infinity.&lt;/p&gt;

&lt;p&gt;Not exactly the kind of thing you want appearing in a live sales dashboard.&lt;/p&gt;

&lt;p&gt;So I clamped the ratio before converting it to log-odds.&lt;/p&gt;

&lt;p&gt;The exact clamp range is an engineering choice.&lt;/p&gt;

&lt;p&gt;The important lesson was simpler:&lt;/p&gt;

&lt;p&gt;Even a reasonable mathematical model can break at the edges of real data.&lt;/p&gt;

&lt;p&gt;Real datasets contain zeroes.&lt;/p&gt;

&lt;p&gt;They contain tiny samples.&lt;/p&gt;

&lt;p&gt;They contain weird distributions.&lt;/p&gt;

&lt;p&gt;They contain historical exceptions.&lt;/p&gt;

&lt;p&gt;The agent has to survive those cases.&lt;/p&gt;

&lt;p&gt;A live deal makes the difference obvious&lt;/p&gt;

&lt;p&gt;Imagine an Evaluation-stage prospect.&lt;/p&gt;

&lt;p&gt;First, a competitor is mentioned.&lt;/p&gt;

&lt;p&gt;The agent recalls seven historical deals involving that competitor.&lt;/p&gt;

&lt;p&gt;Six were lost.&lt;/p&gt;

&lt;p&gt;The score moves slightly because a competitor mention alone isn't strong enough evidence.&lt;/p&gt;

&lt;p&gt;Then the prospect says:&lt;/p&gt;

&lt;p&gt;«"Finance has frozen new vendor spend."»&lt;/p&gt;

&lt;p&gt;Now the relevant historical pattern is:&lt;/p&gt;

&lt;p&gt;6 of 10 lost.&lt;/p&gt;

&lt;p&gt;The system explains the change:&lt;/p&gt;

&lt;p&gt;«"Dropped 15 pts: budget freeze mentioned, and 6 of 10 past budget_freeze deals at Evaluation were lost."»&lt;/p&gt;

&lt;p&gt;The pattern crosses the drafting threshold.&lt;/p&gt;

&lt;p&gt;The agent recalls the four historical deals that survived the same objection and uses those responses to draft a follow-up.&lt;/p&gt;

&lt;p&gt;Later, the prospect responds positively.&lt;/p&gt;

&lt;p&gt;The score recovers.&lt;/p&gt;

&lt;p&gt;The historical risk doesn't disappear.&lt;/p&gt;

&lt;p&gt;That's intentional.&lt;/p&gt;

&lt;p&gt;A positive response doesn't erase what happened historically.&lt;/p&gt;

&lt;p&gt;It just adds new evidence.&lt;/p&gt;

&lt;p&gt;Memory changes behavior&lt;/p&gt;

&lt;p&gt;This is where Hindsight becomes more than a retrieval layer.&lt;/p&gt;

&lt;p&gt;Add another lost budget-freeze deal.&lt;/p&gt;

&lt;p&gt;The historical pattern changes:&lt;/p&gt;

&lt;p&gt;6 of 10 → 7 of 11&lt;/p&gt;

&lt;p&gt;Nothing else changed.&lt;/p&gt;

&lt;p&gt;No prompt.&lt;/p&gt;

&lt;p&gt;No retraining.&lt;/p&gt;

&lt;p&gt;No model update.&lt;/p&gt;

&lt;p&gt;The memory changed.&lt;/p&gt;

&lt;p&gt;The agent's behavior changed with it.&lt;/p&gt;

&lt;p&gt;That means a newly closed deal isn't just another CRM record.&lt;/p&gt;

&lt;p&gt;It's another piece of evidence available to the next call.&lt;/p&gt;

&lt;p&gt;But memory can also become dangerous&lt;/p&gt;

&lt;p&gt;A shared bank means one bad record can potentially influence future decisions.&lt;/p&gt;

&lt;p&gt;So I added an integrity check.&lt;/p&gt;

&lt;p&gt;def integrity_check(client, expected):&lt;br&gt;
    recalled = fetch_all_deals(&lt;br&gt;
        client,&lt;br&gt;
        bank_id="sales-team-shared"&lt;br&gt;
    )&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;return (
    "COUNTS VERIFIED"
    if tally(recalled) == expected
    else "COUNT MISMATCH"
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;If the complete dataset can't be reconstructed, the system doesn't silently produce a statistic.&lt;/p&gt;

&lt;p&gt;And if Hindsight isn't reachable, the application falls back to a local store and says so.&lt;/p&gt;

&lt;p&gt;The system should be allowed to say:&lt;/p&gt;

&lt;p&gt;"I don't have enough verified memory to make this claim."&lt;/p&gt;

&lt;p&gt;That's much more useful than a confident number built from incomplete retrieval.&lt;/p&gt;

&lt;p&gt;Five rules I ended up with&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Use memory for experience, not blind authority. A recalled example is evidence, not automatically the complete truth.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Separate retrieval from aggregation. What an LLM needs for context isn't necessarily what code needs for statistics.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Keep the numbers traceable. Every important count should lead back to actual records.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Treat edge cases as part of the design. Zero wins, tiny samples, and contradictory evidence aren't unusual enough to ignore.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Let the agent explain uncertainty. A system that knows when its evidence is incomplete is more useful than one that always produces an answer.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's ultimately what I wanted from Hindsight.&lt;/p&gt;

&lt;p&gt;Not a system that simply remembers everything.&lt;/p&gt;

&lt;p&gt;A system that lets an agent use past experience without pretending that memory is automatically truth.&lt;/p&gt;

&lt;p&gt;The model can write the explanation.&lt;/p&gt;

&lt;p&gt;Hindsight can provide the experience.&lt;/p&gt;

&lt;p&gt;The code can verify the numbers.&lt;/p&gt;

&lt;p&gt;And the rep gets something much more useful than:&lt;/p&gt;

&lt;p&gt;«"Based on my analysis, this deal looks risky."»&lt;/p&gt;

&lt;p&gt;They get:&lt;/p&gt;

&lt;p&gt;«"Here is what happeneI Didn't Want an AI That Remembered. I Wanted One I Could Catch Lying.&lt;/p&gt;

&lt;p&gt;Building a sales agent with Hindsight taught me that memory isn't the hard part. Knowing when to trust the memory is.&lt;/p&gt;

&lt;p&gt;The first version looked impressive.&lt;/p&gt;

&lt;p&gt;Give the agent a deal.&lt;/p&gt;

&lt;p&gt;Let it recall similar historical deals.&lt;/p&gt;

&lt;p&gt;Ask it what usually happened.&lt;/p&gt;

&lt;p&gt;It produced convincing answers.&lt;/p&gt;

&lt;p&gt;Then I checked the numbers.&lt;/p&gt;

&lt;p&gt;That's when the interesting part started.&lt;/p&gt;

&lt;p&gt;The agent said a particular objection had caused most deals to fail.&lt;/p&gt;

&lt;p&gt;The actual data didn't support that conclusion.&lt;/p&gt;

&lt;p&gt;The model hadn't necessarily hallucinated a deal.&lt;/p&gt;

&lt;p&gt;It had done something more subtle:&lt;/p&gt;

&lt;p&gt;It had reasoned from the memories it happened to retrieve.&lt;/p&gt;

&lt;p&gt;That isn't the same as reasoning from the complete dataset.&lt;/p&gt;

&lt;p&gt;So the architecture changed.&lt;/p&gt;

&lt;p&gt;The model writes the words. Code does the math.&lt;/p&gt;

&lt;p&gt;That became the rule for the entire system.&lt;/p&gt;

&lt;p&gt;Hindsight, Vectorize's open-source agent memory system, handles the memory layer.&lt;/p&gt;

&lt;p&gt;The application handles deterministic aggregation.&lt;/p&gt;

&lt;p&gt;The pipeline is:&lt;/p&gt;

&lt;p&gt;retain → extract → recall → reflect → score → explain → draft&lt;/p&gt;

&lt;p&gt;Hindsight's documentation describes retain, recall, and reflect as its core operations.&lt;/p&gt;

&lt;p&gt;I use those operations to give the agent historical experience, but I don't let an LLM decide statistical facts about that experience.&lt;/p&gt;

&lt;p&gt;The retrieval problem nobody notices at first&lt;/p&gt;

&lt;p&gt;Suppose the team has 200 closed deals.&lt;/p&gt;

&lt;p&gt;A new prospect mentions a budget freeze.&lt;/p&gt;

&lt;p&gt;The agent recalls the most relevant historical deals.&lt;/p&gt;

&lt;p&gt;Maybe it gets 10.&lt;/p&gt;

&lt;p&gt;That's great for an LLM.&lt;/p&gt;

&lt;p&gt;Ten relevant examples are plenty of context for generating a response.&lt;/p&gt;

&lt;p&gt;But now ask:&lt;/p&gt;

&lt;p&gt;«"How many budget-freeze deals at the Evaluation stage were lost?"»&lt;/p&gt;

&lt;p&gt;Suddenly those ten memories aren't necessarily enough.&lt;/p&gt;

&lt;p&gt;The answer might depend on 17 records.&lt;/p&gt;

&lt;p&gt;Or 43.&lt;/p&gt;

&lt;p&gt;Or 6.&lt;/p&gt;

&lt;p&gt;Retrieval is optimized for relevance.&lt;/p&gt;

&lt;p&gt;Statistics are optimized for completeness.&lt;/p&gt;

&lt;p&gt;Those are different problems.&lt;/p&gt;

&lt;p&gt;So I split them.&lt;/p&gt;

&lt;p&gt;Recall is for context.&lt;/p&gt;

&lt;p&gt;Complete retrieval is for counting.&lt;/p&gt;

&lt;p&gt;The actual pattern comes from code&lt;/p&gt;

&lt;p&gt;The historical deals are structured by objection, stage, outcome, and other attributes.&lt;/p&gt;

&lt;p&gt;Then the application calculates the pattern:&lt;/p&gt;

&lt;p&gt;scoped = [&lt;br&gt;
    d for d in deals&lt;br&gt;
    if d.objection == objection&lt;br&gt;
    and d.objection_stage == stage&lt;br&gt;
]&lt;/p&gt;

&lt;p&gt;lost = sum(d.outcome == "lost" for d in scoped)&lt;br&gt;
won = len(scoped) - lost&lt;/p&gt;

&lt;p&gt;Now the agent has something much safer to explain:&lt;/p&gt;

&lt;p&gt;6 lost / 10 total&lt;/p&gt;

&lt;p&gt;rather than an LLM-generated approximation of what happened.&lt;/p&gt;

&lt;p&gt;Hindsight's reflection can still add useful context around the recalled experiences.&lt;/p&gt;

&lt;p&gt;But if reflection says something different from the deterministic count, the count wins.&lt;/p&gt;

&lt;p&gt;Then I found the second problem&lt;/p&gt;

&lt;p&gt;What happens when there are zero wins?&lt;/p&gt;

&lt;p&gt;One historical competitor had:&lt;/p&gt;

&lt;p&gt;5 losses&lt;/p&gt;

&lt;p&gt;0 wins&lt;/p&gt;

&lt;p&gt;A naive loss ratio gives:&lt;/p&gt;

&lt;p&gt;1.0&lt;/p&gt;

&lt;p&gt;Then "logit(1.0)" goes to infinity.&lt;/p&gt;

&lt;p&gt;Not exactly the kind of thing you want appearing in a live sales dashboard.&lt;/p&gt;

&lt;p&gt;So I clamped the ratio before converting it to log-odds.&lt;/p&gt;

&lt;p&gt;The exact clamp range is an engineering choice.&lt;/p&gt;

&lt;p&gt;The important lesson was simpler:&lt;/p&gt;

&lt;p&gt;Even a reasonable mathematical model can break at the edges of real data.&lt;/p&gt;

&lt;p&gt;Real datasets contain zeroes.&lt;/p&gt;

&lt;p&gt;They contain tiny samples.&lt;/p&gt;

&lt;p&gt;They contain weird distributions.&lt;/p&gt;

&lt;p&gt;They contain historical exceptions.&lt;/p&gt;

&lt;p&gt;The agent has to survive those cases.&lt;/p&gt;

&lt;p&gt;A live deal makes the difference obvious&lt;/p&gt;

&lt;p&gt;Imagine an Evaluation-stage prospect.&lt;/p&gt;

&lt;p&gt;First, a competitor is mentioned.&lt;/p&gt;

&lt;p&gt;The agent recalls seven historical deals involving that competitor.&lt;/p&gt;

&lt;p&gt;Six were lost.&lt;/p&gt;

&lt;p&gt;The score moves slightly because a competitor mention alone isn't strong enough evidence.&lt;/p&gt;

&lt;p&gt;Then the prospect says:&lt;/p&gt;

&lt;p&gt;«"Finance has frozen new vendor spend."»&lt;/p&gt;

&lt;p&gt;Now the relevant historical pattern is:&lt;/p&gt;

&lt;p&gt;6 of 10 lost.&lt;/p&gt;

&lt;p&gt;The system explains the change:&lt;/p&gt;

&lt;p&gt;«"Dropped 15 pts: budget freeze mentioned, and 6 of 10 past budget_freeze deals at Evaluation were lost."»&lt;/p&gt;

&lt;p&gt;The pattern crosses the drafting threshold.&lt;/p&gt;

&lt;p&gt;The agent recalls the four historical deals that survived the same objection and uses those responses to draft a follow-up.&lt;/p&gt;

&lt;p&gt;Later, the prospect responds positively.&lt;/p&gt;

&lt;p&gt;The score recovers.&lt;/p&gt;

&lt;p&gt;The historical risk doesn't disappear.&lt;/p&gt;

&lt;p&gt;That's intentional.&lt;/p&gt;

&lt;p&gt;A positive response doesn't erase what happened historically.&lt;/p&gt;

&lt;p&gt;It just adds new evidence.&lt;/p&gt;

&lt;p&gt;Memory changes behavior&lt;/p&gt;

&lt;p&gt;This is where Hindsight becomes more than a retrieval layer.&lt;/p&gt;

&lt;p&gt;Add another lost budget-freeze deal.&lt;/p&gt;

&lt;p&gt;The historical pattern changes:&lt;/p&gt;

&lt;p&gt;6 of 10 → 7 of 11&lt;/p&gt;

&lt;p&gt;Nothing else changed.&lt;/p&gt;

&lt;p&gt;No prompt.&lt;/p&gt;

&lt;p&gt;No retraining.&lt;/p&gt;

&lt;p&gt;No model update.&lt;/p&gt;

&lt;p&gt;The memory changed.&lt;/p&gt;

&lt;p&gt;The agent's behavior changed with it.&lt;/p&gt;

&lt;p&gt;That means a newly closed deal isn't just another CRM record.&lt;/p&gt;

&lt;p&gt;It's another piece of evidence available to the next call.&lt;/p&gt;

&lt;p&gt;But memory can also become dangerous&lt;/p&gt;

&lt;p&gt;A shared bank means one bad record can potentially influence future decisions.&lt;/p&gt;

&lt;p&gt;So I added an integrity check.&lt;/p&gt;

&lt;p&gt;def integrity_check(client, expected):&lt;br&gt;
    recalled = fetch_all_deals(&lt;br&gt;
        client,&lt;br&gt;
        bank_id="sales-team-shared"&lt;br&gt;
    )&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;return (
    "COUNTS VERIFIED"
    if tally(recalled) == expected
    else "COUNT MISMATCH"
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;If the complete dataset can't be reconstructed, the system doesn't silently produce a statistic.&lt;/p&gt;

&lt;p&gt;And if Hindsight isn't reachable, the application falls back to a local store and says so.&lt;/p&gt;

&lt;p&gt;The system should be allowed to say:&lt;/p&gt;

&lt;p&gt;"I don't have enough verified memory to make this claim."&lt;/p&gt;

&lt;p&gt;That's much more useful than a confident number built from incomplete retrieval.&lt;/p&gt;

&lt;p&gt;Five rules I ended up with&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Use memory for experience, not blind authority. A recalled example is evidence, not automatically the complete truth.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Separate retrieval from aggregation. What an LLM needs for context isn't necessarily what code needs for statistics.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Keep the numbers traceable. Every important count should lead back to actual records.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Treat edge cases as part of the design. Zero wins, tiny samples, and contradictory evidence aren't unusual enough to ignore.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Let the agent explain uncertainty. A system that knows when its evidence is incomplete is more useful than one that always produces an answer.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's ultimately what I wanted from Hindsight.&lt;/p&gt;

&lt;p&gt;Not a system that simply remembers everything.&lt;/p&gt;

&lt;p&gt;A system that lets an agent use past experience without pretending that memory is automatically truth.&lt;/p&gt;

&lt;p&gt;The model can write the explanation.&lt;/p&gt;

&lt;p&gt;Hindsight can provide the experience.&lt;/p&gt;

&lt;p&gt;The code can verify the numbers.&lt;/p&gt;

&lt;p&gt;And the rep gets something much more useful than:&lt;/p&gt;

&lt;p&gt;«"Based on my analysis, this deal looks risky."»&lt;/p&gt;

&lt;p&gt;They get:&lt;/p&gt;

&lt;p&gt;«"Here is what happened to ten deals like this one, here is what survived, and here is exactly how we know."»&lt;/p&gt;

&lt;p&gt;Code: [repo link] · Demo: [demo link]d to ten deals like this one, here is what survived, and here is exactly how we know."»&lt;/p&gt;

&lt;p&gt;Code: [repo link] · Demo: [demo link]I Didn't Want an AI That Remembered. I Wanted One I Could Catch Lying.&lt;/p&gt;

&lt;p&gt;Building a sales agent with Hindsight taught me that memory isn't the hard part. Knowing when to trust the memory is.&lt;/p&gt;

&lt;p&gt;The first version looked impressive.&lt;/p&gt;

&lt;p&gt;Give the agent a deal.&lt;/p&gt;

&lt;p&gt;Let it recall similar historical deals.&lt;/p&gt;

&lt;p&gt;Ask it what usually happened.&lt;/p&gt;

&lt;p&gt;It produced convincing answers.&lt;/p&gt;

&lt;p&gt;Then I checked the numbers.&lt;/p&gt;

&lt;p&gt;That's when the interesting part started.&lt;/p&gt;

&lt;p&gt;The agent said a particular objection had caused most deals to fail.&lt;/p&gt;

&lt;p&gt;The actual data didn't support that conclusion.&lt;/p&gt;

&lt;p&gt;The model hadn't necessarily hallucinated a deal.&lt;/p&gt;

&lt;p&gt;It had done something more subtle:&lt;/p&gt;

&lt;p&gt;It had reasoned from the memories it happened to retrieve.&lt;/p&gt;

&lt;p&gt;That isn't the same as reasoning from the complete dataset.&lt;/p&gt;

&lt;p&gt;So the architecture changed.&lt;/p&gt;

&lt;p&gt;The model writes the words. Code does the math.&lt;/p&gt;

&lt;p&gt;That became the rule for the entire system.&lt;/p&gt;

&lt;p&gt;Hindsight, Vectorize's open-source agent memory system, handles the memory layer.&lt;/p&gt;

&lt;p&gt;The application handles deterministic aggregation.&lt;/p&gt;

&lt;p&gt;The pipeline is:&lt;/p&gt;

&lt;p&gt;retain → extract → recall → reflect → score → explain → draft&lt;/p&gt;

&lt;p&gt;Hindsight's documentation describes retain, recall, and reflect as its core operations.&lt;/p&gt;

&lt;p&gt;I use those operations to give the agent historical experience, but I don't let an LLM decide statistical facts about that experience.&lt;/p&gt;

&lt;p&gt;The retrieval problem nobody notices at first&lt;/p&gt;

&lt;p&gt;Suppose the team has 200 closed deals.&lt;/p&gt;

&lt;p&gt;A new prospect mentions a budget freeze.&lt;/p&gt;

&lt;p&gt;The agent recalls the most relevant historical deals.&lt;/p&gt;

&lt;p&gt;Maybe it gets 10.&lt;/p&gt;

&lt;p&gt;That's great for an LLM.&lt;/p&gt;

&lt;p&gt;Ten relevant examples are plenty of context for generating a response.&lt;/p&gt;

&lt;p&gt;But now ask:&lt;/p&gt;

&lt;p&gt;«"How many budget-freeze deals at the Evaluation stage were lost?"»&lt;/p&gt;

&lt;p&gt;Suddenly those ten memories aren't necessarily enough.&lt;/p&gt;

&lt;p&gt;The answer might depend on 17 records.&lt;/p&gt;

&lt;p&gt;Or 43.&lt;/p&gt;

&lt;p&gt;Or 6.&lt;/p&gt;

&lt;p&gt;Retrieval is optimized for relevance.&lt;/p&gt;

&lt;p&gt;Statistics are optimized for completeness.&lt;/p&gt;

&lt;p&gt;Those are different problems.&lt;/p&gt;

&lt;p&gt;So I split them.&lt;/p&gt;

&lt;p&gt;Recall is for context.&lt;/p&gt;

&lt;p&gt;Complete retrieval is for counting.&lt;/p&gt;

&lt;p&gt;The actual pattern comes from code&lt;/p&gt;

&lt;p&gt;The historical deals are structured by objection, stage, outcome, and other attributes.&lt;/p&gt;

&lt;p&gt;Then the application calculates the pattern:&lt;/p&gt;

&lt;p&gt;scoped = [&lt;br&gt;
    d for d in deals&lt;br&gt;
    if d.objection == objection&lt;br&gt;
    and d.objection_stage == stage&lt;br&gt;
]&lt;/p&gt;

&lt;p&gt;lost = sum(d.outcome == "lost" for d in scoped)&lt;br&gt;
won = len(scoped) - lost&lt;/p&gt;

&lt;p&gt;Now the agent has something much safer to explain:&lt;/p&gt;

&lt;p&gt;6 lost / 10 total&lt;/p&gt;

&lt;p&gt;rather than an LLM-generated approximation of what happened.&lt;/p&gt;

&lt;p&gt;Hindsight's reflection can still add useful context around the recalled experiences.&lt;/p&gt;

&lt;p&gt;But if reflection says something different from the deterministic count, the count wins.&lt;/p&gt;

&lt;p&gt;Then I found the second problem&lt;/p&gt;

&lt;p&gt;What happens when there are zero wins?&lt;/p&gt;

&lt;p&gt;One historical competitor had:&lt;/p&gt;

&lt;p&gt;5 losses&lt;/p&gt;

&lt;p&gt;0 wins&lt;/p&gt;

&lt;p&gt;A naive loss ratio gives:&lt;/p&gt;

&lt;p&gt;1.0&lt;/p&gt;

&lt;p&gt;Then "logit(1.0)" goes to infinity.&lt;/p&gt;

&lt;p&gt;Not exactly the kind of thing you want appearing in a live sales dashboard.&lt;/p&gt;

&lt;p&gt;So I clamped the ratio before converting it to log-odds.&lt;/p&gt;

&lt;p&gt;The exact clamp range is an engineering choice.&lt;/p&gt;

&lt;p&gt;The important lesson was simpler:&lt;/p&gt;

&lt;p&gt;Even a reasonable mathematical model can break at the edges of real data.&lt;/p&gt;

&lt;p&gt;Real datasets contain zeroes.&lt;/p&gt;

&lt;p&gt;They contain tiny samples.&lt;/p&gt;

&lt;p&gt;They contain weird distributions.&lt;/p&gt;

&lt;p&gt;They contain historical exceptions.&lt;/p&gt;

&lt;p&gt;The agent has to survive those cases.&lt;/p&gt;

&lt;p&gt;A live deal makes the difference obvious&lt;/p&gt;

&lt;p&gt;Imagine an Evaluation-stage prospect.&lt;/p&gt;

&lt;p&gt;First, a competitor is mentioned.&lt;/p&gt;

&lt;p&gt;The agent recalls seven historical deals involving that competitor.&lt;/p&gt;

&lt;p&gt;Six were lost.&lt;/p&gt;

&lt;p&gt;The score moves slightly because a competitor mention alone isn't strong enough evidence.&lt;/p&gt;

&lt;p&gt;Then the prospect says:&lt;/p&gt;

&lt;p&gt;«"Finance has frozen new vendor spend."»&lt;/p&gt;

&lt;p&gt;Now the relevant historical pattern is:&lt;/p&gt;

&lt;p&gt;6 of 10 lost.&lt;/p&gt;

&lt;p&gt;The system explains the change:&lt;/p&gt;

&lt;p&gt;«"Dropped 15 pts: budget freeze mentioned, and 6 of 10 past budget_freeze deals at Evaluation were lost."»&lt;/p&gt;

&lt;p&gt;The pattern crosses the drafting threshold.&lt;/p&gt;

&lt;p&gt;The agent recalls the four historical deals that survived the same objection and uses those responses to draft a follow-up.&lt;/p&gt;

&lt;p&gt;Later, the prospect responds positively.&lt;/p&gt;

&lt;p&gt;The score recovers.&lt;/p&gt;

&lt;p&gt;The historical risk doesn't disappear.&lt;/p&gt;

&lt;p&gt;That's intentional.&lt;/p&gt;

&lt;p&gt;A positive response doesn't erase what happened historically.&lt;/p&gt;

&lt;p&gt;It just adds new evidence.&lt;/p&gt;

&lt;p&gt;Memory changes behavior&lt;/p&gt;

&lt;p&gt;This is where Hindsight becomes more than a retrieval layer.&lt;/p&gt;

&lt;p&gt;Add another lost budget-freeze deal.&lt;/p&gt;

&lt;p&gt;The historical pattern changes:&lt;/p&gt;

&lt;p&gt;6 of 10 → 7 of 11&lt;/p&gt;

&lt;p&gt;Nothing else changed.&lt;/p&gt;

&lt;p&gt;No prompt.&lt;/p&gt;

&lt;p&gt;No retraining.&lt;/p&gt;

&lt;p&gt;No model update.&lt;/p&gt;

&lt;p&gt;The memory changed.&lt;/p&gt;

&lt;p&gt;The agent's behavior changed with it.&lt;/p&gt;

&lt;p&gt;That means a newly closed deal isn't just another CRM record.&lt;/p&gt;

&lt;p&gt;It's another piece of evidence available to the next call.&lt;/p&gt;

&lt;p&gt;But memory can also become dangerous&lt;/p&gt;

&lt;p&gt;A shared bank means one bad record can potentially influence future decisions.&lt;/p&gt;

&lt;p&gt;So I added an integrity check.&lt;/p&gt;

&lt;p&gt;def integrity_check(client, expected):&lt;br&gt;
    recalled = fetch_all_deals(&lt;br&gt;
        client,&lt;br&gt;
        bank_id="sales-team-shared"&lt;br&gt;
    )&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;return (
    "COUNTS VERIFIED"
    if tally(recalled) == expected
    else "COUNT MISMATCH"
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;If the complete dataset can't be reconstructed, the system doesn't silently produce a statistic.&lt;/p&gt;

&lt;p&gt;And if Hindsight isn't reachable, the application falls back to a local store and says so.&lt;/p&gt;

&lt;p&gt;The system should be allowed to say:&lt;/p&gt;

&lt;p&gt;"I don't have enough verified memory to make this claim."&lt;/p&gt;

&lt;p&gt;That's much more useful than a confident number built from incomplete retrieval.&lt;/p&gt;

&lt;p&gt;Five rules I ended up with&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Use memory for experience, not blind authority. A recalled example is evidence, not automatically the complete truth.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Separate retrieval from aggregation. What an LLM needs for context isn't necessarily what code needs for statistics.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Keep the numbers traceable. Every important count should lead back to actual records.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Treat edge cases as part of the design. Zero wins, tiny samples, and contradictory evidence aren't unusual enough to ignore.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Let the agent explain uncertainty. A system that knows when its evidence is incomplete is more useful than one that always produces an answer.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's ultimately what I wanted from Hindsight.&lt;/p&gt;

&lt;p&gt;Not a system that simply remembers everything.&lt;/p&gt;

&lt;p&gt;A system that lets an agent use past experience without pretending that memory is automatically truth.&lt;/p&gt;

&lt;p&gt;The model can write the explanation.&lt;/p&gt;

&lt;p&gt;Hindsight can provide the experience.&lt;/p&gt;

&lt;p&gt;The code can verify the numbers.&lt;/p&gt;

&lt;p&gt;And the rep gets something much more useful than:&lt;/p&gt;

&lt;p&gt;«"Based on my analysis, this deal looks risky."»&lt;/p&gt;

&lt;p&gt;They get:&lt;/p&gt;

&lt;p&gt;«"Here is what happened to ten deals like this one, here is what survived, and here is exactly how we know."»&lt;/p&gt;

&lt;p&gt;Code: [repo link] · Demo: [demo link]&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
