DEV Community

Praseedha Bairy
Praseedha Bairy

Posted on

Why I Put Hindsight Behind One Adapter File

The first thing I wrote in this project wasn't an agent. It was a 90-line file called hindsight_adapter.py. Every retain, recall, and reflect call in the codebase goes through it, and it's the decision I'd defend first.

LoopCall is a voice agent that handles support, sales, and follow-up calls with a shared memory in Hindsight. This post is about the interface around that memory and what it bought me.

Why an adapter at all

I'd never used Hindsight before. Before writing code I read the docs and wrote my findings into a notes file: real method names, parameters, limits, and how tags and metadata behave. Then I wrote the adapter from those notes.

Three reasons:

  1. The API will change or I'll misread it. One file means one place to fix.
  2. I need a fair baseline. "Memory off" should be the same code path minus memory, not a separate branch of the app.
  3. Live calls can't wait. Every memory operation needs a timeout and a safe failure mode.

The interface

class Memory(Protocol):
    async def retain(self, bank: str, content: str,
                     tags: list[str], metadata: dict) -> None: ...
    async def recall(self, bank: str, query: str,
                     tags: list[str] | None = None,
                     timeout: float | None = None) -> list[MemoryItem]: ...
    async def reflect(self, bank: str, query: str) -> str: ...
Enter fullscreen mode Exit fullscreen mode

Three methods. Nothing else in the app imports the Hindsight client.

Memory off is a null object

This is the trick I like most. The A/B switch isn't an if memory_on: scattered across the code. It's dependency injection:

class NullMemory:
    async def retain(self, *args, **kwargs) -> None:
        return None

    async def recall(self, *args, **kwargs) -> list[MemoryItem]:
        return []

    async def reflect(self, *args, **kwargs) -> str:
        return ""
Enter fullscreen mode Exit fullscreen mode

A session starts with either HindsightMemory or NullMemory. The brief builder, the agent, and the post-call pipeline behave identically otherwise. When I show a before/after comparison, the only difference is the object passed in, so nobody can say I gave the baseline a worse prompt.

Timeouts and degraded recall

The real implementation wraps every call:

class HindsightMemory:
    def __init__(self, client, default_timeout: float = 2.0):
        self._client = client
        self._timeout = default_timeout

    async def recall(self, bank, query, tags=None, timeout=None):
        try:
            async with asyncio.timeout(timeout or self._timeout):
                raw = await self._client.recall(bank_id=bank, query=query)
        except (TimeoutError, HindsightError) as err:
            log.warning("recall degraded", extra={"bank": bank, "error": str(err)})
            return []
        return [MemoryItem.from_raw(r) for r in raw]
Enter fullscreen mode Exit fullscreen mode

If recall fails mid-call, the agent continues without that context instead of freezing. In-call recalls get an even shorter timeout than the pre-call ones. A late memory is worse than no memory when a person is waiting for the next sentence.

Conventions the adapter enforces

The adapter also owns naming, so the rest of the code can't invent its own:

  • Banks: cust_{id}, org_playbook, org_product.
  • Tags: tech, emotion, commitment, tactic, feedback, release, pref, conflict.
  • Metadata on every retain: call_id, timestamp, mode, customer_id, outcome.

Retained content is always a standalone sentence. "Acme runs v2.1 on Postgres 14. Restarting the worker didn't fix export timeouts; raising the batch size to 500 did" is more useful at recall time than a pile of fields.

The write path: validate, repair, fall back

Post-call extraction turns a transcript into structured JSON, then the adapter retains the pieces. The fast Groq-hosted models sometimes fail on structured output, so the pipeline assumes they will:

async def extract_call(transcript: str) -> CallExtraction:
    for model in (PRIMARY_MODEL, FALLBACK_MODEL):
        raw = await llm.json(model, EXTRACT_PROMPT, transcript)
        for attempt in range(2):
            try:
                return CallExtraction.model_validate_json(raw)
            except ValidationError as err:
                raw = await llm.json(model, REPAIR_PROMPT, raw, errors=str(err))
    return CallExtraction.degraded(transcript)
Enter fullscreen mode Exit fullscreen mode

Valid output goes to SQLite and Hindsight. Broken output gets one repair pass, then a second model, then a degraded record that keeps the raw transcript so nothing is lost.

Testing with a fake

Because the app only sees the Memory protocol, tests use an in-memory fake. The test I trust most checks privacy:

def test_pii_is_masked_before_retain(fake_memory):
    transcript = "My email is dana@acme.io and my card is 4111 1111 1111 1111"
    asyncio.run(retain_call(transcript, memory=fake_memory))

    stored = " ".join(item.content for item in fake_memory.items)
    assert "dana@acme.io" not in stored
    assert "4111" not in stored
Enter fullscreen mode Exit fullscreen mode

Masking runs through Presidio before anything reaches the adapter. The test exists because I don't want that guarantee to depend on me remembering it.

Where the abstraction leaks

It isn't free. Three things got awkward:

  • Tag filtering. The interface takes tags, but how much filtering happens server-side versus after recall depends on the Hindsight features I'm using. The adapter hides that, which can hide a performance problem too.
  • Reflect output is text. I turn weekly reflections into stored "lessons" and append them to the agent's instructions. That step is mine, not the adapter's, and it's the least tidy part of the system.
  • A fake is not the real thing. The fake passes tests that real recall ranking might not. I keep a small smoke test against a live Hindsight instance for that reason.

Takeaways

  1. Put your memory system behind a small interface on day one.
  2. Make "memory off" a null implementation, not a code branch.
  3. Give every memory call a timeout and a defined failure result.
  4. Let the adapter own bank names, tags, and metadata.
  5. Mask sensitive data before retain, and test that it happens.

For background on what agent memory should and shouldn't hold, the Vectorize agent memory guide is a good read. The code for everything above is at [GITHUB_URL].

Top comments (0)