DEV Community

Cover image for I gave my assistant a memory that survives the server it runs on
OKAFOR KOSISOCHUKWU
OKAFOR KOSISOCHUKWU

Posted on

I gave my assistant a memory that survives the server it runs on

I built an assistant that remembers durable facts about the person using it, and
the interesting part turned out not to be the storing.

A quick summary of what it is. Cheta is a personal assistant you can reach from
Telegram, a web page, a command line client, or a browser extension, and all four
share one memory space per person. It picks between thirteen tools on its own,
transcribes voice notes, reads PDF, Word, PowerPoint with the speaker notes,
Excel and any plain text or source file, and can crawl a site you point it at.
The browser extension reads the page you are currently on and can act on it.

Now the failure that taught me the most.

My first version kept its index in a SQLite file on the server. Every redeploy
erased it, so the assistant told a user with twenty stored facts that it had
nothing about them. The facts were safe the whole time. They live on Walrus
Memory, which stores data as blobs on the Walrus network and is reached through
an HTTP relayer, and the local index was only a cache. I had forgotten that, and
the result was an assistant confidently denying it knew someone it had been
talking to for weeks.

It now rebuilds its index from a snapshot held in Walrus itself, so a fresh
server recovers what it lost.

The second lesson was about deciding what to keep. My first consolidation pass
compared raw text, so two notes saying the same thing in different words never
matched and an exact duplicate could be stored twice. Comparing a normalised form
fixed that. A changed preference now retires the older note rather than sitting
beside it, and two facts that conflict get flagged so the assistant can say so
instead of contradicting itself later.

The third lesson was about latency, and the cause surprised me. Turns were taking
up to forty seconds. It was not the model. The provider was rate limiting, and
the client library's automatic retry was sleeping for up to fifty seconds before
trying again, against a timeout of thirty. Every one of those retries was wasted
work. Turning the SDK retry off and rotating across several API keys took a rate
limited turn from about four seconds of dead waiting to effectively nothing.

The thing I would tell anyone adding memory to a bot: do not trust your own
summary of whether it works. I compared two answers to the same question, one
with memory switched on and one with it switched off, same model and same minute,
and put them side by side. That comparison caught more problems than any test I
wrote.

Code and setup instructions are here, and it runs without any paid service:
https://github.com/Ksschkw/

Longer write-up of the build:
https://medium.com/@kookafor893/i-gave-my-chatbot-a-memory-that-outlived-its-own-database-449ed737d0ec

Top comments (0)