DEV Community

Nikith Chowdary Vadde
Nikith Chowdary Vadde

Posted on

I Gave My AI Sales Assistant a Memory Bank for Every Deal

An AI sales assistant can answer a question perfectly and still be useless three weeks later.

The problem isn't necessarily the model. It's memory.

I built DealMind around a simple idea: a sales deal should have its own long-term memory boundary. Every deal gets its own Hindsight memory bank, so the assistant can retrieve relevant information from weeks of conversations without mixing one customer's context with another's.

That decision ended up shaping almost the entire architecture.

What I Built

DealMind is a deal workspace with an AI assistant.

Each deal has its own:

  • Customer information
  • Conversation thread
  • Activity timeline
  • Insights
  • Stored memories
  • AI assistant

The assistant supports several actions:

  • Free-form chat
  • Meeting preparation
  • Risk analysis
  • Insight generation

The frontend is built with React, Vite, and TypeScript.

Supabase handles authentication, Postgres storage, and row-level security. All AI work happens inside a single Supabase Edge Function called deal-ai, running on Deno.

Gemini 1.5 Flash handles generation through REST.

Hindsight handles long-term memory.

The browser never calls Gemini or Hindsight directly.

Instead, it sends the Edge Function the action, deal ID, and question together with the user's JWT:

{ action, dealId, question }
Enter fullscreen mode Exit fullscreen mode

The Edge Function authenticates the request, verifies that the user owns the deal, retrieves the relevant context, calls the model, saves the result, and then updates long-term memory.

That gives me one controlled boundary around the entire AI workflow.

The Problem With "Just Give the Model the Chat History"

My first instinct could have been to treat the conversation itself as memory.

That works until the conversation gets long.

DealMind intentionally sends only the 10 most recent messages as conversational history.

That is useful for keeping the current conversation coherent, but it creates an obvious problem:

recency isn't importance.

Imagine a customer says during the first meeting:

"Our security team has to approve the integration before procurement can sign."

Three weeks later, that message isn't anywhere near the most recent ten messages.

But if the salesperson asks:

"What could block this deal?"

that old security requirement may be one of the most important pieces of information in the entire deal.

I didn't want the assistant to depend on how recently a fact happened to be mentioned.

I wanted it to depend on whether that fact was relevant.

That's where Hindsight's long-term memory system became useful.

Hindsight Github Repo : https://github.com/vectorize-io/hindsight

Hindsight Documentation : https://hindsight.vectorize.io/

One Deal, One Memory Bank

The most important design decision in DealMind is also the simplest one:

one deal gets one Hindsight memory bank.

The bank is named:

deal-${dealId}
Enter fullscreen mode Exit fullscreen mode

So if a deal has the ID abc123, its memory boundary is effectively:

deal-abc123
Enter fullscreen mode Exit fullscreen mode

Another deal gets another bank.

I deliberately didn't create one global memory store containing every conversation.

Sales data already has a natural isolation boundary: the deal.

I wanted the memory architecture to follow the same boundary.

This makes the system easier to reason about.

If I am processing Deal A, I know which Hindsight bank I am allowed to recall from.

The question determines what information is relevant.

The deal ID determines where that information is allowed to come from.

That distinction is important.

How a Request Flows Through DealMind

Every AI request follows roughly the same pipeline.

User request
     ↓
Authenticate user
     ↓
Verify deal ownership
     ↓
Recall from deal-specific Hindsight memory
     ↓
Load deal context
     ↓
Build grounded model prompt
     ↓
Generate response
     ↓
Save response
     ↓
Retain exchange in Hindsight
Enter fullscreen mode Exit fullscreen mode

The context assembled for the model contains several different sources:

  1. The deal record
  2. Customers associated with the deal
  3. Recent conversation messages
  4. Stored memories
  5. Recent activities
  6. Memories recalled from Hindsight

The assistant then receives that context through a strict grounding prompt.

The basic instruction is intentionally boring:

Use only the supplied context. Do not invent customer statements or memories.

That's exactly what I want.

The model runs with a temperature of 0.4.

I'm not trying to make the assistant creative when it is answering questions about a customer's requirements.

I'm trying to make it useful without inventing things.

Hindsight Is the Recall Layer

For normal chat, the user's question becomes the recall query.

If the user asks:

What security concerns did the customer raise?
Enter fullscreen mode Exit fullscreen mode

that question is used to retrieve relevant information from the current deal's memory bank.

For structured actions such as meeting preparation or risk analysis, there isn't always a natural conversational question.

In those cases, DealMind uses a broader fallback query:

core deal preferences, timeline, and updates
Enter fullscreen mode Exit fullscreen mode

The important thing is that both paths remain inside the current deal's memory boundary.

The architecture therefore separates two concerns:

Question → determines relevance

Deal ID → determines memory scope
Enter fullscreen mode Exit fullscreen mode

That's much easier to reason about than giving an AI assistant access to a giant pool of historical sales information.

The Hindsight documentation goes deeper into the underlying memory system and its recall/retain model.

I Keep Two Kinds of Memory

One of the design decisions I didn't expect at the beginning was keeping two different representations of memory.

Hindsight is the memory system the AI reads from.

Postgres has a separate memories table containing durable facts extracted from conversations.

Why keep both?

Because they serve different purposes.

Hindsight is optimized around the assistant's need to retrieve relevant historical context.

The Postgres memory table is part of the application itself. It gives the user a visible, human-auditable representation of important facts about a deal.

Conceptually:

                    Conversation
                         │
              ┌──────────┴──────────┐
              │                     │
              ▼                     ▼
         Hindsight              Postgres
              │                     │
       Recall-oriented        Human-auditable
       long-term memory       durable facts
Enter fullscreen mode Exit fullscreen mode

I don't think every internal memory representation needs to be exposed directly to users.

And I don't think every user-visible memory needs to be the retrieval representation.

Keeping those responsibilities separate makes the architecture clearer.

Extracting Durable Facts

The Postgres memory table is populated using a separate LLM extraction prompt.

The extraction prompt is intentionally narrower than the main assistant.

It looks for durable facts.

It doesn't need to preserve small talk.

The expected output is a JSON array, which is then parsed by the application.

I also don't blindly trust the model's JSON.

The parser looks for the first [ and last ], extracts the section between them, and attempts to parse it.

If parsing fails, the application returns an empty array instead of turning a formatting mistake into a failed request.

That may sound like a small implementation detail.

It isn't.

When an LLM produces data that my application is going to store, I treat that output as untrusted input.

The same principle applies to the main response.

The model is given explicit grounding instructions because I don't want a generated assumption to silently become a customer fact.

Memory Shouldn't Be Able to Take Down the Application

Long-term memory is useful, but it is still a dependency.

That means it can fail.

A new deal may not have a memory bank yet.

A recall request can fail.

A retain request can fail.

I didn't want any of those situations to automatically turn the entire AI request into a 500.

Recall and retain are therefore isolated with error handling.

For example, the retain operation is treated as a separate step after the response has already been generated and saved.

try {
  await hindsight.retain(
    `deal-${dealId}`,
    memoryContent
  );
} catch (error) {
  console.error("Hindsight retain failed:", error);
}
Enter fullscreen mode Exit fullscreen mode

If memory retention fails, the assistant's answer shouldn't disappear.

The application can still work with the deal record, recent activities, recent messages, and other context.

The answer might be less informed.

That's a much better failure mode than making the entire workspace unusable.

The More Interesting Failure: The AI Can Remember Itself

This is the part of the system I'm least comfortable with.

The current retain flow stores the assistant's answer together with the user's question.

That creates a subtle problem.

Suppose the model makes a mistake.

If that response is retained as memory, the mistake could appear as historical context in a future request.

Now the system has a feedback loop:

User asks X
      ↓
Model incorrectly answers Y
      ↓
Y becomes memory
      ↓
Future request recalls Y
      ↓
Model sees Y as historical context
Enter fullscreen mode Exit fullscreen mode

That's exactly the kind of failure I don't want long-term memory to create.

The next step is to separate user-stated facts from model-generated output.

The safer flow is:

Customer/user states X
        ↓
Extract durable fact
        ↓
Store X as memory
Enter fullscreen mode Exit fullscreen mode

rather than:

User asks X
        ↓
Model answers Y
        ↓
Store Y as memory
Enter fullscreen mode Exit fullscreen mode

Those are fundamentally different things.

A model's interpretation shouldn't automatically become a historical fact.

That's one of the most important lessons I've taken from building this system.

Illustrative Example

Imagine a deal has been active for several weeks.

During an early conversation, the customer says that their security team must approve the integration before procurement can proceed.

The conversation continues for weeks.

Eventually, that message is no longer inside the most recent ten messages.

Before tomorrow's meeting, the salesperson asks:

Prepare me for tomorrow's meeting.
What could block this deal?
Enter fullscreen mode Exit fullscreen mode

DealMind retrieves relevant memories from the deal's Hindsight bank.

The assistant now has access to:

  • The current deal information
  • Recent activities
  • Recent conversation
  • Stored durable memories
  • Relevant long-term Hindsight memories

The historical security requirement can therefore become part of the meeting-preparation context even though it isn't in the most recent conversation messages.

That is the behavior I wanted from long-term memory.

Not "remember everything."

Remember what matters to this deal when it becomes relevant again.

Why the Deal Is the Memory Boundary

There is a broader lesson here that goes beyond sales.

Memory needs boundaries.

If an application has multiple independent entities, those entities can often provide a natural partition for memory.

For DealMind, the partition is obvious:

Deal A
 └── deal-A memory

Deal B
 └── deal-B memory

Deal C
 └── deal-C memory
Enter fullscreen mode Exit fullscreen mode

The application doesn't need to invent an abstract concept of memory ownership.

The business object already provides one.

This is one reason I like the way Hindsight fits into DealMind. The memory layer doesn't need to define what a deal is.

DealMind already knows.

Hindsight gives me the mechanism for retaining and recalling information within that boundary.

For more background on why persistent memory matters for agent-style systems, Vectorize also has a useful agent memory overview.

Security Is Part of the Memory Design

The database uses row-level security on every table, scoped to the deal owner.

The Edge Function also authenticates the user and verifies ownership before it builds the AI context or accesses the deal's memory.

That means a dealId coming from the browser is not treated as authorization.

It's just an identifier.

The application still has to establish:

Authenticated user
        ↓
Owns requested deal?
        ↓
Allowed to access deal context
        ↓
Allowed to access deal memory
Enter fullscreen mode Exit fullscreen mode

The browser also never receives the Gemini or Hindsight provider keys.

Those credentials live in Supabase secrets.

For me, this is part of the AI architecture, not a separate security checklist.

If memory can expose historical customer information, then memory access is an authorization problem.

What I Learned

1. Memory boundaries matter more than memory volume

The interesting decision wasn't simply adding long-term memory.

It was deciding that one deal equals one memory boundary.

That made the system easier to isolate, reason about, and secure.

2. Recency isn't memory

Keeping the latest ten messages is useful for conversational context.

It isn't enough for long-running workflows.

Important information doesn't necessarily remain important because it is recent.

Sometimes the oldest fact is the one that matters most.

3. Retrieval and human visibility are different problems

Hindsight gives the assistant a mechanism for recalling relevant historical information.

Postgres gives the application a human-readable memory surface.

Separating those responsibilities made the system easier to work with.

4. Memory failures should degrade gracefully

A memory service shouldn't automatically become a single point of failure for the entire application.

Recall and retain can fail independently.

The rest of the deal context should still be useful.

5. AI-generated memory needs stricter boundaries

This is the biggest remaining issue in my implementation.

If an assistant can store its own answers as memories, an incorrect answer can become future context.

The next version should retain explicitly extracted facts separately from model-generated interpretations.

Long-term memory is powerful precisely because it persists.

That also means mistakes can persist.

What's Next

The next changes are fairly concrete.

First, I want to separate user- and customer-stated facts from assistant-generated responses during retention.

Second, I want to populate the hindsight_memory_id field so application-level audit records can map directly to individual Hindsight records.

Neither requires making the system more complicated for the sake of complexity.

The goal is the opposite.

I want the memory lifecycle to be obvious:

Deal
 ↓
Relevant historical information
 ↓
Hindsight recall
 ↓
Grounded context
 ↓
Model response
 ↓
Durable facts
 ↓
Hindsight retain
Enter fullscreen mode Exit fullscreen mode

The interesting part of DealMind isn't that an LLM can answer questions about a sales deal.

That's table stakes.

The interesting part is what happens when the assistant needs to remember something from three weeks ago, while being absolutely sure that the memory belongs to the deal currently being worked.

That's where most of the engineering work ended up.

Top comments (0)