DEV Community

Cover image for I Gave a Chatbot Memory It Can Prove : Here's the Architecture
Alex Miano
Alex Miano

Posted on

I Gave a Chatbot Memory It Can Prove : Here's the Architecture

Mnemo AI is a chatbot that remembers you across sessions — and every memory is encrypted, stored on Walrus mainnet, and provable. Postgres is only a mirror; the source of truth lives on-chain, and a daily reconciler proves it. Built with Next.js 16 + Gemini Flash + Sui/Walrus. Live at mnemoai.xyz, source at github.com/mianohh/mnemo.

The problem: every chatbot starts from zero

Language models are stateless. They see the current conversation and nothing else. You tell a bot "never ship on Fridays," come back the next day, ask "can we ship this Friday?" and it answers: "Sure, what time works?"

Chat history files and bigger context windows are the usual patch. I wanted something better: memory that survives sessions and devices, is readable only by its owner, and whose existence you can verify — not trust.

The loop around every reply

Every chat turn runs a fixed pipeline:

flowchart TD
    U["User message"] --> R["Recall in namespace<br/>relevance floor 0.2"]
    R -->|"hits + per-request nonce"| G["Gemini streams the answer"]
    G --> CHIP["UI chip: [N] memories applied<br/>expand to see fact + score"]
    G --> E["Extract pass<br/>keep at most 3 durable facts"]
    E --> L["Relayer: embed, SEAL-encrypt,<br/>upload to Walrus, index"]
    L --> W[("Walrus mainnet<br/>source of truth")]
    L --> M[("Postgres mirror")]
    M -.->|"daily parity check"| W

Three details do the heavy lifting:

Receipts. Every contextual reply carries a [N] memories applied chip. Expand it and you see each stored fact and its cosine relevance score — the model's memory, auditable by anyone reading the screen.

Memories can't become instructions. Recalled facts are injected behind a nonce boundary (BEGIN_UNTRUSTED_WALRUS_MEMORY_…) in their own message, with a fixed untrusted-data policy in the system prompt. No memory byte can ever hold system priority so a poisoned blob can't prompt-inject the bot.

Write, encrypt, then forget locally. The extract pass keeps at most three durable facts per turn (Zod-validated), and the relayer encrypts them with SEAL before they ever hit storage.

Your address is the namespace

Multi-tenancy isn't a user_id column. it's cryptographic. Both sign-in paths resolve to a verified Sui address:

  • Wallet: personal-message signature challenge.
  • Email: Google OIDC → zkLogin address derived server-side from the verified JWT (no proving service).

The namespace is then deterministically mnemo-user-{address}. Clients cannot forge, switch, or inspect someone else's namespace the server derives it from the signer, every request.

The mirror is not the truth

The design rule: PostgreSQL serves exclusively as an ephemeral read-through mirror, while Walrus Memory is the decentralized source of truth. So the dashboard doesn't display its own row count — it cross-checks it against the relayer's on-chain metric:

{
  "mirrorCount": 146,
  "chain": { "count": 146 }
}
Enter fullscreen mode Exit fullscreen mode

The dashboard renders those two numbers side by side (146 ↔ 146) and only claims parity when they're equal. Two independent sources agreeing is the proof the memories really exist on-chain, not just in one database.

Proving what's on-chain (the part nobody else has)

Walrus Memory's SDK exposes no "what is the status of blob X?" endpoint, so I built a reconciler against Sui's GraphQL API:

  • Owned Blob objects are the truth. The relayer transfers every blob it writes to the account owner, so the owner's object set on Sui is what actually exists.
  • Exact epoch math. The current Walrus epoch is read from the shared system object's future_accounting ring buffer — not guessed from timestamps. Every blob has a mandatory end epoch, and a lapsed blob cannot be renewed or recovered.
  • A cron that fails loudly. A daily GitHub Actions job reports chainBlobs, per-epoch expiry buckets, and the object IDs needed to renew — and fails the run within 3 epochs of any risk.
flowchart LR
    GQL["Sui GraphQL API"] --> REC["Reconciler"]
    SYS["System object<br/>future_accounting ring"] --> REC
    REC --> SET["Owned Blob set<br/>start + end epochs"]
    SET --> CRON["Daily cron<br/>fails within 3 epochs of risk"]

Honest wart, because engineering posts should have them: 21 mirror rows still point at blob IDs the chain doesn't recognize. The account-level report is chain-driven and correct; the mismatch is tracked as a known issue.

Does it actually remember?

The chat UI has a Memory ON/OFF toggle, which makes the before/after a one-button demo. Memory off: I told Mnemo "Never ship on Fridays — that's our release-freeze window," then asked in a later session "Can we ship this Friday?" → "Sure, what time works?" Memory on, same question: "You've said Friday is your release-freeze window — want Thursday or Monday?" — with the stored constraint and its score sitting right there in the chip. One bot, two very different Fridays.

More of the build story — what broke along the way, the prompt-injection boundary, and the honest version of the A/B — is in the long-form article.

Try it

git clone https://github.com/mianohh/mnemo
npm install && cp .env.example .env
npm run dev
Enter fullscreen mode Exit fullscreen mode

Node ≥ 24, any Postgres, a Gemini key — and it remembers you tomorrow. Or talk to it first at mnemoai.xyz.

Top comments (0)