DEV Community

Sameera
Sameera

Posted on

Call 1 Brief vs Call 5 Brief: Watching Hindsight Memory Compound

Before a sales rep's fifth call with a customer, the useful context is scattered across four transcripts, a CRM note nobody updated, and the rep's own head. I wanted to see what happens when the agent, not the rep, carries that context, so I built DealMind around one idea: every call gets written into memory, and every brief is generated from what memory recalls.
The memory layer is Hindsight, an open-source agent memory system. This post is about how I wired it in, what it changed about the briefs, and the parts that were harder than they looked.

What DealMind does

A rep pastes a call transcript. DealMind extracts objections, competitors, stakeholders, commitments, next steps and what worked, stores them in SQLite for the dashboard, and retains the whole call in Hindsight. Before the next call, a "Prepare Me" button recalls the deal's history and generates a brief. There is also an "Ask DealMind" box for free-text questions on a deal.

The stack is React with Vite, Express, SQLite, and Groq running openai/gpt-oss-120b for extraction and brief generation. Hindsight sits between them as the long-term memory. SQLite answers "what is the current state of this deal." Hindsight answers "what do we know, and when did we learn it."

Keeping those two jobs separate turned out to be the most useful decision in the project.

The through-line: one bank per deal

The first design question was isolation. A deal's memory should never leak into another deal's brief. I gave every deal its own Hindsight memory bank and also tagged every write, so a recall can be scoped even inside a bank:
export const bankIdForDeal = (dealId) => ${config.hindsight.bankPrefix}-${dealId}
export const dealTag = (dealId) => deal:${dealId}

When I log a call, the retain looks like this:
const memory = await retainFacts(dealId, [{
content: [
Sales call between DealMind rep and ${deal.company_name}...,
Call outcome: ${extraction.sentiment} sentiment. ${extraction.summary},
extraction.objections.length ? Objections raised: ... : '',
extraction.resolutions?.length ? Resolved earlier objections: ... : '',
Raw transcript:\n${transcript},
].filter(Boolean).join('\n'),
context: sales call — ${deal.company_name},
timestamp,
document_id: deal-${dealId}-call-${callId},
tags: [dealTag(dealId), call:${callId}, ...activeTags(extraction)],
metadata: { deal_id: String(dealId), company: deal.company_name, call_id: String(callId) },
}])
Three details here earned their place.

The timestamp is the call date, not the time I clicked save. Memory that knows when something happened, rather than when it was entered, is what makes the rest of this work.

The document_id is deterministic. If I re-log or edit a call, retaining it again replaces the earlier version instead of stacking a duplicate. I got this wrong at first and the brief started repeating the same objection twice.

The category tags (objection, competitor, commitment, and so on) mean I can ask for a slice of the history without writing a new prompt for each.

Recall: four questions beat one

My first brief used a single query, something like "what should I know about this deal." The results were fine but shallow, dominated by whatever was most semantically central. A brief needs breadth: open objections, competitors, who decides, what was promised.
So recallDealContext fans out into four focused queries, runs them in parallel, and merges the results:
const queries = [
'objections, concerns and unresolved pushback the customer raised',
'competitors, alternatives and vendors the customer is evaluating',
'stakeholders, decision makers and their positions',
'commitments, promises, next steps and agreed dates',
]

const results = await Promise.all(
queries.map((query) =>
recallFacts(dealId, query, {
budget, maxTokens: Math.round(maxTokens / queries.length),
tags: [dealTag(dealId)], tagsMatch: 'any_strict',
temporalWindow, limit: 8,
}).catch(() => ({ results: [], mode: 'unavailable' })),
),
)

After that I deduplicate by the first 120 characters, sort by score, and keep the top 20. Each query gets a quarter of the token budget, so one noisy topic can't crowd out the others. The .catch matters too: if one recall fails, the brief still gets the other three.

Time travel is the feature I like most

Because every retain carries the real call date, recall accepts a temporal window. The brief endpoint takes a through date and turns it into one:
const temporalWindow = throughDate
? { start: '1970-01-01T00:00:00Z',
end: new Date(Date.parse(toIso(throughDate)) + 86_400_000).toISOString() }
: null
That lets me ask: what would the brief have said the day after Call 2? The same deal, the same bank, only a different cutoff. I use it for a replay view that steps through five sequential calls for one account and regenerates the brief at each step. It is also a good way to test whether memory is actually doing anything.

What the difference looks like

The seeded account is a mid-size operations customer with 14 sites and 300 technicians.

In Call 1 the customer says they are not shopping, that a previous field-scheduling rollout died at roughly 15 percent adoption, and that multi-site scheduling is a hard requirement. They name the people involved, including a director who will run it and a CFO who will only show up once there are numbers, and they define success as 80 percent of technicians active within the first quarter.

A brief generated after Call 1 has little to work with beyond those facts. That is expected.

In Call 2 the price lands. I quote $68k for the multi-site tier, and the customer reacts because they had told their CFO it would come in under fifty. They also say they can't explain a premium over list price and can't yet see ROI against the free tier.

A brief generated before Call 3 should now carry both calls: the hard requirement and the adoption target from Call 1, plus the price gap and the missing ROI sentence from Call 2. A rep reading it walks in knowing the customer's board-level fear and the specific number that caused friction, without opening a single transcript.
[HERE ADD YOUR CALL 3/4 DETAILS AFTER YOU GIVE THEM TO ME]
By Call 5, the brief also knows which objections were raised, and which were closed.

Objections that get resolved

Memory that only accumulates gets noisy. A brief that keeps warning about a pricing objection the customer already settled is worse than no brief. So when a call resolves something, the extraction returns a resolutions list, and I match each claim against the open objections using a token-overlap similarity with a 0.3 floor:
That last part is worth explaining.
return bestScore >= 0.3 ? { objection: best, score: bestScore } : null
Matches get marked resolved in SQLite, and the retained memory includes a line like Resolved earlier objections: -> . So Hindsight holds the story, including how each problem ended, and SQLite holds the current status for the dashboard.
[SC-5 HERE — Screenshot showing resolved objection/current status, if available]
Asking questions with reflect

For the "Ask DealMind" box I use Hindsight's reflect endpoint, which reasons over a bank rather than just returning matches:
const result = await postJson(bankPath(bankId, '/reflect'), {
body: { query: String(query || '').slice(0, 4000), budget: 'mid', max_tokens: 1200 },
...
})
The split I settled on: recall for structured, predictable inputs to a brief, and reflect for open questions like "why did Call 3 go badly." Recall gives me control, and reflect gives me synthesis.

DealMind keeps the current deal state, long-term memory, and AI processing as separate responsibilities.

[SC-7 HERE — Architecture diagram]
The flow is:

SQLite answers what the current state of the deal is. Hindsight stores the history of what was learned and when. Groq handles extraction and brief generation.

What was harder than expected

Fallbacks can hide bugs. I wrapped every Hindsight call so that if the API is unreachable, it falls back to a local file-backed store with keyword scoring. It keeps the app usable offline. But early on, a malformed request quietly triggered the fallback and everything looked fine while nothing reached the live API. I now log a loud warning when a 4xx comes back, saying it looks like a request-shape bug rather than an outage. If you build a fallback, make it noisy.

The cross-deal playbook is not memory. The "Playbook Insights" page surfaces patterns across deals, like pricing pressure and security reviews. Today that is computed from extracted facts in SQLite with pattern matching, not from Hindsight. It works, but it isn't learning in any deep sense, and I'd rather say so than dress it up.

Deal health is deliberately dumb. The 0–100 score is a deterministic function of stage, call cadence, sentiment trend, open objections and stakeholder count. The same history always gives the same number. I did not want a language model deciding a number a rep might act on.

Lessons I'd reuse

Give each customer or user their own bank, then tag anyway. Isolation plus tags means you can never accidentally recall the wrong account.
Store the real event time. Timestamps at the moment of the event, not the moment of the write, unlock time-scoped recall.
Use deterministic document IDs. They make re-ingestion safe.
Fan out recall into focused queries. Several narrow questions beat one broad one, as long as you split the token budget and deduplicate.
Keep structured state outside memory. Let a normal database own status and let the memory layer own history.
Make fallbacks loud. A silent fallback hides the bug you most need to see.

Top comments (2)

Collapse
 
sameera_001 profile image
Sameera •

Very good