DEV Community

SHAIK RAYHAN
SHAIK RAYHAN

Posted on

Why I Shrunk Hindsight’s Forecast Corrections

An agent that remembers every forecast can still become less accurate. We learned that the hard way: in one complete replay, the memory-enabled version scored worse than the sales reps whose forecasts it was meant to improve.

That result changed the question we were asking. It was no longer just “Can an agent remember what happened?” It became “Can it use the right memory, at the right time, without turning a noisy pattern into an overconfident rule?”

The problem: forecasts need history

Calibrate is a sales-forecast calibration system. A rep enters a win probability for a deal; later, the deal is won or lost. Over time, a manager wants to know whether that rep tends to overstate or understate the odds for certain kinds of deals.

A one-shot language model can summarize a rep’s history if the history is placed in its prompt. That approach becomes awkward once the history spans many deals, quarters, and deal traits. It also makes it easy to mix evidence from the wrong period into a current forecast.

We built a replay engine around historical Maven sales data so we could study this as a sequence of events. The Maven outcomes stay intact. The simulated rep probabilities add a controlled “gut feel” and planted rep biases, giving us something concrete to learn. That makes this a useful calibration experiment, not a claim that the system is already validated on live forecasts.

Keep forecast time separate from outcome time

The central design choice is chronological replay. A forecast exists on the date a rep enters it. Its outcome becomes available only when the deal closes. Those are different events, and treating them as one row at one point in time would let the agent learn from information it did not yet have.

The replay engine builds two event types and sorts them by date. It processes a forecast before an outcome if both occur on the same day. It also writes an audit date for each quarter so we can check that prior evidence really predates the next set of predictions.

That temporal boundary matters more than it first appears. A “Q2 deal” can mean a deal forecast in Q2 or a deal that closed in Q2. If those labels are used interchangeably, a memory lookup can return the wrong cohort. We now tag a memory with the forecast quarter as well as the close quarter, and ask the calibration card to use the previous forecast cohort.

Here is the shape of the event ordering:

forecast_day = pd.Timestamp(deal["forecast_date"])
forecast_q = f"{forecast_day.year}-Q{forecast_day.quarter}"
events[forecast_q].append((forecast_day, 0, "forecast", deal))

if deal.get("close_date"):
    outcome_day = pd.Timestamp(deal["close_date"])
    outcome_q = f"{outcome_day.year}-Q{outcome_day.quarter}"
    events[outcome_q].append((outcome_day, 1, "outcome", deal))
Enter fullscreen mode Exit fullscreen mode

The sequence is simple on purpose: the forecast goes into memory when it is made; the outcome and the agent’s self-check arrive later. Open deals stay pending until there is an actual close date.

Hindsight is the memory layer, not the scoring shortcut

We use Hindsight for agent memory to retain forecast facts, CRM outcomes, and the agent’s own correction checks. Each memory has a timestamp and tags for the rep, event type, and quarter. At the start of a later quarter, the system asks Hindsight to reflect on the prior forecast cohort and return a structured calibration card.

That card contains evidence such as the number of closed deals, average stated probability, observed win rate, and a proposed adjustment. The Python calibrator applies the card to a new probability. The replay engine remains independent of the Hindsight client through a small adapter, which lets us exercise the memory-off path without contacting the service.

We batch retained memories instead of saving one at a time:

for offset in range(0, len(memories), RETAIN_BATCH_SIZE):
    batch = memories[offset:offset + RETAIN_BATCH_SIZE]
    services.retain_many(batch)
Enter fullscreen mode Exit fullscreen mode

After a quarter’s batch is retained, the adapter waits 30 seconds for background processing before the next reflection. The engine also checkpoints output files and quarter progress, so an interrupted replay can resume from a completed quarter instead of starting over.

This separation gives us a useful comparison: replay the exact same historical events with memory off and memory on. We can then inspect the cards, beliefs, audit dates, and per-quarter results rather than relying on a persuasive summary.

The uncomfortable result

In the first full memory-on run, the agent did not beat the reps. Its Brier score was 0.1546, compared with 0.1467 for the reps; lower is better. At a 50% win/loss cutoff, it classified 80.8% of closed deals correctly, versus 82.5% for the reps.

The weakness was clearest in Q3. The memory-enabled Brier score was 0.1660, compared with 0.1477 for the reps. The saved cards showed why a single all-history correction is risky: a strong bias from an earlier quarter can remain in the card even after the rep’s forecasts move closer to actual outcomes.

Sana’s stated win rate was about 8.5 percentage points above her actual win rate in Q3 and about 4.4 points above it in Q4. The earlier memory card reflected a much larger gap. Multiplying a current probability by that older ratio could overcorrect the forecast.

That is not a result to hide behind a demo. It is a useful failure: memory was present, retrieval worked, and the agent still made worse probability estimates.

Smaller corrections and fresher evidence

We changed the calibration path in three ways. First, each card is scoped to the previous forecast quarter instead of mixing that cohort with older history. Second, the calibrator uses only a fraction of the card’s proposed adjustment: 25% at high confidence, 20% at medium, and 12.5% at low. Third, if a reflection call fails, replay can reuse the last valid card rather than silently acting as if no useful history exists.

The correction is still bounded to the valid probability range:

strength = STRENGTH.get(rule.get("confidence", "low"), 0.5)
factor = 1 + (float(rule["adjustment"]) - 1) * strength
corrected = min(PROB_MAX, max(PROB_MIN, stated * factor))
Enter fullscreen mode Exit fullscreen mode

This is a conservative design, not a measured victory. The saved benchmark predates these changes, and we have not run another full Hindsight replay to claim that the new policy performs better. The next evaluation should compare the same quarters and report both Brier score and classification accuracy. We should keep memory off as the cheap baseline, then spend on a memory-on run only when the local analysis gives us a reason to expect a useful result.

For implementation details, the Hindsight documentation explains the memory primitives; Vectorize’s overview of agent memory gives the broader context. Calibrate’s replay and data code is in the project repository.

What I’d carry into the next agent project

Time is part of the data model. Store when a claim was made and when it was confirmed. A memory system that ignores those timestamps can leak future evidence into an answer.

A memory card is an estimate. More retrieved context does not make a proposed adjustment correct. Evidence count, recency, and the size of the correction should all affect how much the agent changes a forecast.

Compare against a boring baseline. A memory-enabled agent should earn its complexity by beating the same forecasts with memory disabled. If it does not, inspect the intermediate cards and event chronology before spending more on a larger run.

Keep the failed run. The poor score is a more useful engineering artifact than a made-up success story. It told us to test recency and correction strength before scaling the memory workload.

That is the part of agent memory I find most interesting: the hard problem is not merely helping an agent remember. It is helping it remember what matters, when it mattered, and how uncertain it should remain.

Top comments (0)