DEV Community

Roshan
Roshan

Posted on

Hindsight Memory Caught a Pattern No Single Call Showed

"My salary came almost two weeks late last month. Madam said there was some problem with the bank." A month later, on a different call: "This month also the salary is late. I had to borrow money for my son's school fees."

Heard on their own, neither sentence sounds alarming. Heard together, they are exactly the kind of pattern a home-care agency should notice. I built the safety layer of TrustMemory, a voice agent for home-care agencies, to catch that pattern without ever turning it into an accusation. This is how it works, and why Hindsight, the open-source agent memory engine, ended up at the centre of it.

Safety across calls: two dated mentions, below the threshold for a check

Why a single-call classifier was the wrong tool

TrustMemory calls domestic helpers on behalf of their agency, in English, Hindi or Telugu. The coordinator sees the call live, and memory carries context from one conversation to the next.

Safety is a different problem from ordinary conversation memory. A helper rarely describes a serious problem in one sentence. She mentions late pay on one call, long hours on another, trouble leaving the house weeks later. Evaluate each call alone and you miss the pattern. Let a model draw a strong conclusion from two sentences and you create a different kind of harm.

So I split the problem in two: capture the helper's own words as dated signals, then evaluate the pattern across recent calls.

How it hangs together

Every call runs through the same path. The voice agent recalls what the agency already knows before it speaks. When the call ends, an extraction step pulls out the outcome and any safety concerns, using only the helper's words. The concerns go into a safety_signals table in SQLite, and the whole call is retained to Hindsight under her helper:<id> tag.

SQLite owns the decision: whether a flag is open, reviewed or closed. Hindsight owns the long memory: everything she has said across every call, which fills the flag with context a coordinator can actually use.

Where Hindsight sits in the safety path

The rules are deliberately conservative

The rules live in server/care.js:

const WINDOW_DAYS = 60;
const MIN_CALLS = 2;
const IMMEDIATE = ['physical_harm', 'confinement'];
const PATTERN_KINDS = ['unpaid_pay', 'excess_hours', 'confinement',
  'verbal_abuse', 'physical_harm', 'denied_food_or_rest'];
Enter fullscreen mode Exit fullscreen mode

A concern has to come up on two different calls inside 60 days before anyone is flagged. I count distinct calls, not mentions, so saying the same thing five times in one conversation is still one piece of evidence. Physical harm and confinement skip the repetition rule and raise a check after a single call.

A signal is something she said. A flag means a coordinator should look at those signals. Neither one is a finding about the household, and the code and the UI both say so.

Keeping her words instead of a score

The most important decision was to store what the helper actually said, not a model's summary of it. Every signal keeps a short quote, the call it came from and the date:

function recordSignals({ helperId, householdId = null, callId, concerns, saidAt = dayOf(Date.now()) }) {
  let n = 0;
  for (const c of Array.isArray(concerns) ? concerns : []) {
    const evidence = String((c && c.evidence) || '').trim().slice(0, 300);
    if (!c || !SAFETY_KINDS.includes(c.kind) || !evidence) continue;
    db.prepare(`INSERT INTO safety_signals (id, helper_id, household_id, call_id, kind, evidence, said_at, source, created_at)
                VALUES (?, ?, ?, ?, ?, ?, ?, 'call', ?)`)
      .run(newId('sig_'), helperId, householdId, callId, c.kind, evidence, saidAt, nowSql());
    n += 1;
  }
  return n;
}
Enter fullscreen mode Exit fullscreen mode

The concerns come from a structured extraction step after each call, and that step is only allowed to use the helper's lines, never the agent's. If the agent said "it sounds like your salary is late" and she didn't confirm it, nothing is recorded.

There is no anonymous "risk = high" column anywhere. A coordinator who opens a flag sees three short quotes with dates, and can judge for themselves.

Where Hindsight comes in

SQLite holds the state machine: signals, open flags, reviews. What it can't do on its own is remember everything she has said across months of calls, including things that were never tagged as a safety concern at the time. That is Hindsight's job.

Every call is retained to Hindsight with the helper's tag, a real timestamp and a stable document id. When a flag opens, a background step asks Hindsight to look across all of her calls and return what she said, with dates, in a fixed shape:

const out = await hindsight.reflect(
  `Across every call with ${helper.name}, has she herself described being paid late or not paid, ` +
  'working excessive hours, not being allowed to leave, being shouted at or insulted, being hurt, ' +
  'or being denied food or rest? List each mention with its date and her words. Then write two ' +
  'neutral sentences for the agency coordinator. This is a flag for a human to check with her, ' +
  'not a conclusion about the household: do not accuse anyone, and use only what she said herself.',
  { tags, budget: 'low', responseSchema: { /* summary + mentions[{ when, quote }] */ } }
);
Enter fullscreen mode Exit fullscreen mode

Three details in that call matter. The tags scope it to one helper, so nothing from another person leaks in. The response schema forces dated quotes back instead of free prose. And reflect returns the memories it used, which I show under the summary so the coordinator can trace every sentence.

Hindsight's documentation describes recall and reflect in detail, and Vectorize has a good explanation of how agent memory differs from a context window. That difference is the whole point here. A context window holds one call. Durable memory is what lets the system notice that three calls, weeks apart, are about the same thing.

Hindsight Core: retained facts, observations and links across the agency

Before and after

Before this layer, each call stood alone. Lakshmi's first complaint went into the call notes, the second went into different call notes a month later, and nothing connected them unless the same coordinator happened to remember.

After it, the second mention inside the window raises a private Safety check needed card on the coordinator's dashboard. It shows both quotes with their dates, the kind of concern, and the summary Hindsight wrote across her calls. The household is never told. The helper is never shown a label about herself.

The coordinator's dashboard, where the private safety card appears

The bug that taught me about time

My first version had a flaw I only found in testing. A coordinator reviewed a flag, wrote a note, closed it. On the next call Lakshmi said the bus was crowded that morning, the extraction tagged it as a minor concern, and the flag reopened, because the old salary signals were still inside the 60-day window and still being counted.

The fix was to make the review part of the data:

const lastReview = db.prepare("SELECT MAX(reviewed_at) AS t FROM safety_flags WHERE helper_id = ? AND status = 'reviewed'").get(helperId).t;
const recent = signalsFor(helperId, { withinDays: WINDOW_DAYS }).filter(s => !lastReview || s.created_at > lastReview);
Enter fullscreen mode Exit fullscreen mode

Only concerns raised after the last review count now, and a new flag also needs the latest call itself to contain one of the serious kinds. A remark about a crowded bus can no longer bring back a pattern a person has already dealt with.

The tests for this layer turned out to be about time and boundaries rather than model quality. One concern alone does not flag. Two calls inside the window do. An old mention outside the window does not count. Harm or confinement flags at once. A reviewed flag stays closed. Because the inputs are dated statements, all of that can be checked deterministically.

What I would tell anyone building something similar

Detect patterns, not labels. A review should point at the dated events that caused it and say which threshold they crossed.

Keep the source evidence. Three short quotes from the helper are worth more than one confident generated conclusion.

Make time part of the data model. Signals need dates, windows need edges, and reviews need to reset what counts.

Let memory inform, and let people decide. Hindsight makes it practical to see months of context in one place. Whether a flag is open, reviewed or closed is still decided by a coordinator, and I think that is the right split.

The safety layer lives in server/care.js. The hardest part of building it was knowing where to stop: notice the pattern, keep the evidence, and hand it to a human.

Top comments (0)