DEV Community

Labeequa Waheed
Labeequa Waheed

Posted on

I Planted Six Hidden Biases and Let Hindsight Find Them

I Planted Six Hidden Biases and Let Hindsight Find Them

Priya marks a deal at 90%. Her deals with one contact and no finance person close
less than half the time. Nobody told the agent that. It worked it out by
remembering, and then it told her manager to ask who controls the budget.

Every sales forecast I have seen is a sum of guesses, and every guesser is biased in
a personal way. One rep's "90%" is another rep's coin flip. A revenue plan built on
those numbers misses, and the misses are expensive.

So I built a forecast calibration agent, and the interesting part was not the
forecasting. It was working out how to prove the agent learned something real.

What the system does

The agent remembers every forecast each rep has made and what actually happened. It
learns each rep's personal bias, corrects new forecasts, and explains the correction
in plain words with the evidence attached.

The pieces:

  • Data. Deals, accounts, products, dates and won/lost outcomes come from a public CRM dataset (Maven Analytics' CRM Sales Opportunities: 8,800 opportunities for a fictional hardware company). It has no column for a rep's own forecast.
  • Forecast simulator. I simulate only that one column: each rep's stated probability. It is the true win rate of similar deals, plus noise, plus a hidden bias I write down in advance.
  • Replay engine. It plays history in date order, once with memory ON and once with memory OFF.
  • Calibrator agent. It uses Hindsight for memory and Groq-hosted models for reasoning and explanations.
  • Dashboard. FastAPI backend, React frontend. One chart shows what reps said, what the agent said, and what actually closed.

The through-line: planting the answer key

Most agent demos have an evaluation problem. The agent "learns", but you can't tell
whether it discovered something or whether you nudged it toward the answer you wanted.

So I wrote the bias into the data generator first, as plain data:

  • Priya is overconfident on deals with one contact and no finance person.
  • Arjun is a sandbagger: he says about 40% on deals that close about 63% of the time.
  • Meera is accurate, except on big-ticket products.
  • Rahul over-calls deals opened in the last two weeks of a quarter.
  • Sana is overconfident in Q1 and Q2, then accurate from Q3, because she improved.
  • Karan is hidden until Q3, as if he just joined.

The deals and outcomes are real Maven data, so I can't tune them. The only simulated
number is the rep's stated forecast, and its bias is on record before the agent sees
anything. That lets me automatically check whether the agent's beliefs match the
planted biases. The checker also counts false alarms: strong rules for biases that
don't exist.

How memory is used

The agent saves three kinds of facts, each with a real timestamp and strict tags:

def retain_forecast(deal: dict):
    client.retain(
        bank_id=BANK,
        content=(f"On {deal['forecast_date']}, {deal['rep_name']} forecast "
                 f"{deal['account']} ({deal['n_contacts']} contact(s), "
                 f"finance contact: {'yes' if deal['has_finance_contact'] else 'no'}) "
                 f"at {int(deal['stated_prob'] * 100)}%."),
        context="sales forecast entered by a rep",
        timestamp=deal["forecast_date"] + "T10:00:00Z",
        tags=[f"rep:{deal['rep_id']}", f"quarter:{deal['quarter']}", "kind:forecast"],
        document_id=f"forecast-{deal['deal_id']}",
    )
Enter fullscreen mode Exit fullscreen mode

Forecasts and outcomes are what the agent was told. The third kind is what it did:
a self-check memory. When a deal closes, the agent saves something like "I adjusted
this forecast from 90% to 50%. The deal was lost, so my correction was right." So it
tracks its own accuracy alongside the reps'.

Hindsight also forms beliefs, which it calls observations, from these memories. I
gave the bank a mission to focus them on stated versus actual confidence by deal type,
plus two directives that act as hard rules:

  • Never adjust a forecast based on a rep's personal traits. Use only deal evidence and track record.
  • Always state how many deals a belief is based on.

The first one matters. An agent that judges people is a liability. An agent that
judges track records is a tool.

Think once per rep per quarter

My first design called the model for every deal, which would burn through a free-tier
token budget quickly and make every run slow and unrepeatable. So I changed the
shape. At the start of each quarter, the agent calls Hindsight's reflect once per
rep and gets a structured calibration card:

def reflect_calibration_card(rep_id: str, rep_name: str):
    response = client.reflect(
        bank_id=BANK,
        query=(f"How does {rep_name}'s stated forecast confidence compare with "
               f"actual outcomes, by type of deal? Give adjustment rules with "
               f"evidence counts. Only use deals that have closed."),
        tags=[f"rep:{rep_id}"],
        tags_match="all_strict",
        budget="mid",
        response_schema=CARD_SCHEMA,
    )
    return response.structured_output, response.text
Enter fullscreen mode Exit fullscreen mode

Two details do the heavy lifting here. tags_match="all_strict" means one rep's
memories never leak into another's. And CARD_SCHEMA forces every rule to name a
trait from a fixed list (single_contact_no_finance, large_deal, end_of_quarter,
overall) and a direction (over, under, accurate). That is what lets the
checker compare the agent's rules to my planted biases without a human reading prose.

Per deal, the card is applied with plain Python: fast, free, repeatable. The model
only writes explanation text where a person will read it.

The no-peeking rule

A replay that lets the agent see the future proves nothing. The replay engine plays
events in date order, and a card used for a forecast may only be built from deals
that closed before that forecast's date. There is a test for exactly that.

The other timing trap is that Hindsight processes memories and forms beliefs in the
background. Saving and reflecting in the same instant is a bug waiting to happen, so
the engine pauses between quarters, and I measured how long beliefs take to appear
before choosing the pause.

What it looks like in use

A forecast comes in through the live form:

  • Rep: Priya. Account: Zenith Corp. Amount: ₹60L.
  • Contacts: 1. Finance contact: no. Stated probability: 90%.

The agent's response:

  • Corrected probability: 50%
  • Confidence: high
  • Why: Priya's deals with one contact and no finance person closed 4 of 9 times.
  • Evidence: the specific past deals, with stated probability and outcome.
  • Question to ask: Who controls the budget at Zenith Corp?

Before Hindsight, the agent could only echo "90%". Now it can say why not, and
show its work. For a new rep with no history it says so and falls back to team-wide
patterns. With fewer than three similar deals it makes only a small adjustment and
labels it low confidence.

The other behavior I like is the one that shows beliefs are alive. Sana's card
softens after Q2 because her recent forecasts are close to reality. The old belief
stays visible in her timeline, so you can see the agent change its mind and when.

Where it is honest about its limits

  • The forecasts are simulated. The deals and outcomes are real CRM data about a fictional company; the reps' stated probabilities are not. That is the trade-off that makes the answer key possible, and I say so plainly.
  • Win rates are flat. Maven's win rates are similar across agents and products (roughly 55–70%), so the signal is each rep's forecasting bias, not big product effects. That is the story, but it means results on real data may look different.
  • Beliefs lag. Memory processing is asynchronous, so the agent is always slightly behind the newest evidence.
  • Rate limits are real. Every model response is cached to disk by prompt, so re-runs are free, and the final demo run is saved as JSON so the deployed site needs no live model calls.

[YOUR RESULTS HERE: replace this line with your real numbers from the final run: the
hidden-bias checker headline (for example "found X of 6", false alarms), and the
Brier score for memory ON vs OFF vs "trust the rep" in Q4.]

Lessons learned

  1. Plant the answer key. If you can define the truth before the agent runs, you can score learning instead of eyeballing it.
  2. Reflect on a schedule, not per item. One structured reflection per rep per quarter beat calling the model on every deal, in cost and in repeatability.
  3. Constrain memory output to enums. Free-text beliefs are nice to read and impossible to test. A fixed trait list made evaluation automatic.
  4. Use strict tags for scoping. Per-rep tags with strict matching stopped cross-contamination between reps.
  5. Make the agent grade itself. Self-check memories give the agent a track record of its own corrections.
  6. Guard against time travel. Write the no-peeking test on day one.

If you want to try this pattern, start with the
Hindsight docs and the
Hindsight GitHub repo. The
Vectorize guide to agent memory explains
why retaining, recalling and reflecting is a different job from stuffing a context
window.

The whole point of memory is that an agent can be wrong on Monday and better on
Friday, and you can see exactly why.
![ ]
 (https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/pa5lesj46ovj3tehfptm.png)

Top comments (0)