DEV Community

Cover image for Sleeptime memory consolidation: an agent that edits its own memory and proves it still remembers
Harish Kotra (he/him)
Harish Kotra (he/him)

Posted on AI-assisted

Sleeptime memory consolidation: an agent that edits its own memory and proves it still remembers

Most agent memory today is one of two failure modes. Either it's an append-only log that grows
until the context window chokes on six-week-old trivia, or it's a lossy summary that quietly
rewrites history and can't tell you what it dropped. Both avoid the actual hard part:

Deliberate, auditable forgetting under a stated retention policy — plus proof that recall still
works afterwards.

So I built Nocturne: a chat agent with a Sleep button. You chat, mixing signal and noise; you
press Sleep; four passes run over its memory blocks — Dedupe → Merge → Promote → Forget — each one
an LLM call whose output is refused or accepted by code, with a snapshot before and after. A diff
view shows every kept / changed / forgotten item with its reason. Then a Wake panel quizzes the
agent: five questions only the consolidated store can answer, plus one about something it was
supposed to have forgotten — where "I don't have that" is the correct answer.

Everything in this post is from the run in the repo: 18/18 verification checks, exit 0.


1. The rule that makes the whole thing defensible

One sentence, and it decides every design choice:

The model proposes. The code applies.

GUARDRAIL = (
    "You may only reference itemIds that appear in the input. You may not invent facts, "
    "names, numbers or dates. If you are unsure, omit the item from all lists."
)
Enter fullscreen mode Exit fullscreen mode

That constant is appended to every pass prompt. It is not the safety mechanism — it's a polite
request. The safety mechanism is that each pass's JSON is re-derived from the data before anything
touches the store:

LLM proposal ──► gate (pure function over store state) ──► apply ──► snapshot
                         │
                         └── refuse ──► log with reason ──► surface in the UI
Enter fullscreen mode Exit fullscreen mode

A rejected proposal is never swallowed. It shows up in the per-pass counts, in the diff panel under
Rejected mutations, and in the mutations table with its rejection_reason. In the run I
documented, the verifier caught the model proposing a drop for an item that had already been dropped
earlier in the same pass:

- drop   m_17   itemId m_17 was already removed by an earlier mutation in this pass
Enter fullscreen mode Exit fullscreen mode

That line is the entire argument for the design: an unverified "consolidation" would have applied a
mutation against a phantom row.


2. Memory model

Four blocks — the shape Letta popularised:

block holds
persona durable facts about the assistant itself
human stable facts about the user: identity, preferences, commitments
episodes events anchored in time or place
scratch working memory: unconsolidated facts pulled from recent conversation

Facts are rows, not prose. That's what makes a per-item diff, a per-item policy and a per-item
undo possible at all.

CREATE TABLE items(
    item_id TEXT PRIMARY KEY,          -- m_17
    block   TEXT NOT NULL REFERENCES blocks(name),
    text    TEXT NOT NULL,
    category TEXT NOT NULL DEFAULT 'context',   -- identity|preferences|commitments|event|...
    source_turn INTEGER,                          -- which transcript turn produced it
    created_at REAL NOT NULL,                     -- what the retention policy measures against
    updated_at REAL NOT NULL,
    active  INTEGER NOT NULL DEFAULT 1            -- forgetting flips this; rows are never deleted
);
CREATE TABLE snapshots(
    snapshot_id TEXT PRIMARY KEY, seq INTEGER NOT NULL UNIQUE,
    label TEXT NOT NULL, created_at REAL NOT NULL, active_count INTEGER NOT NULL
);
CREATE TABLE snapshot_items(            -- a full immutable copy of the state, every version
    snapshot_id TEXT NOT NULL, item_id TEXT NOT NULL, block TEXT NOT NULL, text TEXT NOT NULL,
    category TEXT NOT NULL, source_turn INTEGER, created_at REAL NOT NULL,
    updated_at REAL NOT NULL, active INTEGER NOT NULL
);
CREATE TABLE mutations(                 -- the audit trail, accepted AND refused
    mutation_id INTEGER PRIMARY KEY AUTOINCREMENT, snapshot_id TEXT, pass TEXT NOT NULL,
    kind TEXT NOT NULL, item_id TEXT NOT NULL, refs TEXT NOT NULL DEFAULT '{}',
    accepted INTEGER NOT NULL, reason TEXT, rejection_reason TEXT, policy TEXT, created_at REAL NOT NULL
);
Enter fullscreen mode Exit fullscreen mode

The store exposes six methods, and that's the only interface the rest of the app uses:

list_blocks() · get_block(name) · append(block, item) · replace(block, items) · snapshot() · restore(snapshot_id)
Enter fullscreen mode Exit fullscreen mode
flowchart LR
  UI[React :5173] -->|fetch + SSE over /api proxy| API[FastAPI :3001]
  API --> ENC[write-time encoder] --> G0[traceability gate]
  API --> RC[run_consolidation] --> P[Dedupe → Merge → Promote → Forget] --> G[ gates ]
  API --> RQ[run_quiz] -->|scored against store text| S[(SQLite<br/>items · snapshots · mutations)]
  G --> S
  API --> RS[restore] --> S
  API -.->|LETTA_MODE=letta| L[(Letta memory blocks)]

Backend that actually ran: SQLite. pip install letta today resolves to Letta Code (0.34.x), a
cloud-agent CLI — it no longer ships a self-hosted memory-block server, and there's no Docker daemon
on this machine, so nothing could bind :8283. I implemented the same block model over SQLite behind
the identical six-method interface, and kept letta-client installed so the swap is a config change.
The consolidation, diff, restore and quiz code never learns which backend is live.


3. The write-time encoder (where invention gets cut off first)

Before any sleep, each chat turn is encoded into atomic facts in scratch. Same guardrail, same kind
of gate — every content word of a proposed fact must exist in the transcript, and it may not
introduce a name the transcript doesn't contain:

cov, missing = tm.coverage(text, sources)
if cov < 1.0:
    rejected.append({"proposal": raw, "rejectionReason":
        f"fact is not traceable to the transcript: {len(missing)} content token(s) invented ({', '.join(missing[:8])})"})
    continue
new_names = tm.proper_nouns(text) - transcript_names
if new_names:
    rejected.append({..., "rejectionReason": f"fact introduces proper nouns absent from the transcript: {sorted(new_names)}"})
Enter fullscreen mode Exit fullscreen mode

Real output from the verifier's seed step — the model tried to store something the user never said:

encoder rejected: fact is not traceable to the transcript: 1 content token(s) invented (prefer)
encoder rejected: fact is not traceable to the transcript: 1 content token(s) invented (run)
Enter fullscreen mode Exit fullscreen mode

Duplicates are deliberately not filtered here. Append-only bloat is precisely the condition the
Dedupe pass exists to fix, and if the encoder silently collapsed repeats there would be nothing left
to consolidate.

Encoding is also incremental — it tracks an extracted_through cursor and only sends new turns.
Without that, every message re-encodes the entire transcript: prompt growth was quadratic and a
single chat turn took ~200 s before the fix.


4. The four passes

pass contract the model must satisfy what the code checks before applying
Dedupe {"drop":[{"itemId","duplicateOf","reason"}]} both ids exist in the pre-pass store and are active; not self-pairs; no duplicate chains; normalised token-set similarity ≥ 0.9; reason non-empty
Merge {"merge":[{"intoId","fromIds","mergedText","reason"}]} ids exist, unconsumed, same block; coverage(mergedText, sources) == 1.0; no new proper nouns; reason non-empty
Promote {"promote":[{"itemId","toBlock","reason"}]} toBlock ∈ {persona, human, episodes}; item active and elsewhere; reason non-empty; identity/preference/commitment facts applied first
Forget {"forget":[{"itemId","reason","policy"}]} item active; category not protected; and it qualifies: ageDays > STALE_AFTER_DAYS (stale) or itemId ∈ overflowIds (capacity). The code records the policy it verified, not the one the model claimed

Dedupe gets a little extra engineering: the code computes candidatePairs with its own comparison
and hands them to the model to adjudicate, so a duplicate the model failed to spot is still put in
front of it. The model still decides; the code still checks.

payload["candidatePairs"] = [
    {"dropId": loser, "keeperId": keeper, "similarity": round(score, 3)}
    for loser, keeper, score in tm.find_duplicate_pairs(state.active(), policy.similarity_threshold)
]
Enter fullscreen mode Exit fullscreen mode

Merge in code, because it's the pass where a plausible-sounding rewrite is the whole danger:

cov, missing_tokens = tm.coverage(merged_text, sources_text)
if cov < 1.0:
    reject(raw, f"mergedText is not traceable to its sources: {len(missing_tokens)} content token(s) "
                f"appear in none of them ({', '.join(missing_tokens[:8])})")
    continue
new_names = tm.proper_nouns(merged_text) - {n for s in sources for n in tm.proper_nouns(s["text"])}
if new_names:
    reject(raw, f"mergedText introduces proper nouns absent from every source: {sorted(new_names)}")
    continue
Enter fullscreen mode Exit fullscreen mode

And Forget, where the gate is a permission check, not a vibe check:

qualifies_stale = item["ageDays"] > policy.stale_after_days
qualifies_capacity = item_id in overflow
if not qualifies_stale and not qualifies_capacity:
    reject(raw, f"item is {item['ageDays']:.1f}d old (staleAfterDays={policy.stale_after_days:g}) "
                f"and not over capacity, so no retention rule authorises forgetting it")
    continue
if policy.is_protected(item):
    reject(raw, f"category {item['category']!r} is protected by the retention policy "
                f"({', '.join(policy.protected_categories)})")
Enter fullscreen mode Exit fullscreen mode

5. The measures, with real numbers

Everything above rests on four tiny functions, which is what makes the rule set auditable in one
screen. Values below are computed by the shipped code, not written by hand:

similarity(a, b) = |A ∩ B| / min(|A|, |B|) over lowercased, punctuation-stripped,
inflection-folded, stopword-free token sets.

a b similarity gate at 0.9
I live in Lisbon and work best in the morning. User lives in Lisbon and works best in the morning. 1.000 duplicate
The user keeps a bearded dragon. The user keeps a bearded dragon called Saffron. 1.000 duplicate (subset)
I live in Lisbon and work best in the morning. I live in Porto and work best in the morning. 0.800 refused — different fact
My stack is Airflow and Postgres. My stack is Airflow and Kubernetes. 0.667 refused

Containment rather than Jaccard is a deliberate choice: a restatement that adds words is the same
fact, and "lives in Lisbon" ⊂ "lives in Lisbon, Portugal" should collapse. The cost is that the two
sentences above score identically to the merge case, so near-duplicate in this project means
"same content, different wording", not "sits in a 0.9–0.99 band".

coverage(claim, sources) — fraction of the claim's content tokens present in the source
vocabulary. Sources keep their stopwords in the vocabulary, because a rewrite legitimately promotes a
function word into content ("use Python" → "uses Python").

claim (against two sources) coverage verdict
The user installed postgres 16 on the backup box and the Kitchen leaks. 1.00 tokens all traceable — but stopped by the proper-noun gate: introduces proper nouns absent from every source: ['Kitchen']
The user installed postgres 16 on the backup box near the kettle. 0.86 refused: near appears in no source
The user migrated to Kubernetes on the backup box. 0.50 refused: migrated, kubernete invented

stem_token is intentionally not a real stemmer — possessive stripped, then a trailing plural
-s unless the word ends in -ss/-us or is ≤ 3 chars: Lisbon's → lisbon, camps → camp,
business → business, gas → gas. Bounded rules mean every rejection can be explained in one
sentence, which is the difference between a guardrail and a mystery.

proper_nouns = capitalised-initial words that are not at a sentence start, so
"Marco moved to Lisbon in March, said Ana." → {Marco, Lisbon, March, Ana} while
"Lisbon is nice. The weather is warm." → {}.


6. Three deterministic backstops (and why they aren't cheating)

A rule the model can simply forget to apply is not a rule. Where the rule is derivable from the
data
, the engine applies it itself and labels it:

enforcedBy: engine-duplicate-rule    # a >=0.9 pair the model left un-adjudicated
enforcedBy: engine-promotion-rule    # identity/preferences/commitments still in scratch → human; events → episodes
enforcedBy: policy-engine            # active set above MAX_ACTIVE_FACTS → lowest-salience unprotected facts
Enter fullscreen mode Exit fullscreen mode

This is not the model being overruled by taste. Each backstop runs the same published measurement
the gate runs, and every such mutation is counted separately in the UI and the mutation log
(enforced=3 in the run below, distinct from the 20 the model proposed). Nothing is ever invented by
a backstop: they only move or retire text already in the store.

Salience, for the capacity rule, is explicit arithmetic rather than a judgement:

score = block_weight(persona 60 | human 50 | episodes 30 | scratch 10)
      + 40 if the category is protected
      + max(0, 30 - 30 * ageDays / STALE_AFTER_DAYS)          # recency, decaying to zero
Enter fullscreen mode Exit fullscreen mode

7. Versioning: forgetting is a view change, not a delete

A snapshot is a full copy of the state, written before and after every pass. One Sleep looks like
this on the version line (real labels from nocturne.db):

 #10 consolidation:start  8 facts
 #11 pass:dedupe:before   8        → #12 pass:dedupe:after    7   drop m_6 (similarity 1.00)
 #13 pass:merge:before    7        → #14 pass:merge:after     7   REJECT 1 merge
 #15 pass:promote:before  7        → #16 pass:promote:after   7   promote ×5
 #17 pass:forget:before   7        → #18 pass:forget:after    6   forget m_8 (stale, 70d)
 #19 restore:s0011        8        ← "undo this sleep" (the restore is itself a version)
Enter fullscreen mode Exit fullscreen mode

restore() deletes the live rows and re-inserts them from snapshot_items, then takes a new
snapshot — so undoing is undoable, and the diff between any two versions is exact because both
versions are fully materialised. Diffing is by item_id:

kept      present in both, same text and block
changed   text differs (merged) or block differs (promoted/moved) — with the sources it came from
forgotten active in A, absent or inactive in B — with the pass, the policy and the reason
added     appears only in B
Enter fullscreen mode Exit fullscreen mode

8. Wake: proving recall, and proving forgetting

Five questions are generated against surviving facts, one against a fact the Forget pass removed. Two
details make this a real test rather than a demo.

(a) The answer key is the store, not the model. A reply is correct when it reproduces ≥ 0.9 of
the fact's content tokens. The question itself is also gated: a generated question that leaks half of
its own answer's content words is refused and recorded, because "Does the user work at Northwind
Logistics?" proves nothing about memory.

(b) The forgotten probe needs three independent conditions to pass, and it reports each:

passed = absent_from_store and said_absent and not leaked
Enter fullscreen mode Exit fullscreen mode
absentFromStore=True      no active fact scores >= 0.9 against the forgotten text (checked in code)
saidAbsent=True           the agent's reply matches an explicit not-in-memory pattern
leakedForgottenContent=False   the reply does not reproduce the fact's content tokens
Enter fullscreen mode Exit fullscreen mode

That third one is the interesting one: an agent that says "I don't have that" while blurting the
detail anyway has not forgotten anything. The measured run:

forgotten probe [OK]: "I don't have that in memory."
     absentFromStore=True saidAbsent=True leaked=False
Enter fullscreen mode Exit fullscreen mode

Merged items are excluded from the question pool on purpose — their text is a concatenation of
several facts, so a faithful answer becomes a paragraph and the test stops measuring recall of a fact.


9. Streaming a four-minute operation without lying about progress

EventSource can't POST a body, so the client reads the response stream itself; the server runs the
consolidation on a worker thread and relays queue events:

const reader = res.body.getReader();
buffer += decoder.decode(value, { stream: true });
while ((cut = buffer.indexOf("\n\n")) >= 0) { /* parse event: / data: */ }
Enter fullscreen mode Exit fullscreen mode
item = events.get(timeout=15.0)
if not thread.is_alive(): break
yield ": keepalive\n\n"        # a pass can be silent for minutes; keep proxies awake
Enter fullscreen mode Exit fullscreen mode

Events are named pass:<name>:start|done|error, which is what lets the Sleep panel show live counts
per pass — input, proposed, accepted, rejected, policy-enforced, active-after, and the snapshot ids on
both sides — while it's still running. Dry run runs the identical four passes but applies nothing:
no live mutation and no snapshot, so the forgetting curve doesn't gain a fake point. The verifier
asserts exactly that (active stays 22->22).

Provider plumbing is one plain POST {baseUrl}/chat/completions call with
response_format={"type":"json_object"}. LM Studio answers that with
400 "'response_format.type' must be 'json_schema' or 'text'", so the client probes once per base
URL, remembers the rejection, retries without the field, and Settings reports json mode unsupported
rather than hiding it.


10. What the verifier proves

npm run verify (→ python -m agent.verify) seeds a scripted 8-turn conversation through the real
chat path — the assistant replies are generated by the model — and the real encoder, then runs a
dry run, a real consolidation, the quiz, and a policy comparison. 18/18, exit 0. From one run:

seeded in 92.4s; encoder proposed 21 facts, accepted 20, rejected 1, backdated 2 episodic fact(s)
active facts=22  exact-dup pairs=7  near-dup pairs=3  stale=2  protected=8  durable=14
3 pair(s) sit below the 0.9 gate (0.7-0.9): m_4/m_16 0.857, m_7/m_18 0.875, m_9/m_18 0.875

dry-run plan: proposed=19 accepted(if applied)=21 rejected=1 active stays 22->22
pass dedupe  in=22 proposed=10 accepted= 9 rejected=1 enforced=0 activeAfter=13
pass merge   in=13 proposed= 2 accepted= 2 rejected=0 enforced=0 activeAfter=11
pass promote in=11 proposed= 7 accepted=10 rejected=0 enforced=3 activeAfter=11
pass forget  in=11 proposed= 2 accepted= 2 rejected=0 enforced=0 activeAfter= 9
totals: {'proposed': 21, 'accepted': 23, 'rejected': 1, 'enforced': 3}

23 accepted, 1 rejected; 24 re-verified, 0 problems
9/9 active facts fully traceable to the transcript
MAX_ACTIVE_FACTS=120: forgotten=2,  active left=20
MAX_ACTIVE_FACTS=6:   forgotten=14, active left=8  (floor 8 = protected facts the policy refuses to forget)
Enter fullscreen mode Exit fullscreen mode

The assertions that matter most are the paranoid ones:

  • Every accepted mutation is re-audited from the database, independently of the consolidation code: dropped pairs are re-measured for similarity, merged text is re-checked for coverage and new proper nouns, promote targets are re-checked against the whitelist, forgettings are re-checked against the rules. "The model said it merged" is never evidence.
  • A global no-invention sweep: every surviving active fact must be 100 % traceable to the raw transcript and must introduce no name absent from it. 9/9.
  • The gates are tested without any model at all — 16 adversarial and legal proposals pushed straight through them (ghost itemIds, self-pairs, dissimilar "duplicates", invented tokens, TitleCased names, illegal blocks, a stale-but-protected commitment, an ineligible forgetting), each required to refuse for the stated reason. This is the part that must not depend on how cooperative the model is today:
dedupe: refused as required -- duplicateOf 'm_99' does not exist in the pre-pass store
merge:  refused as required -- mergedText introduces proper nouns absent from every source: ['Kitchen']
promote: refused as required -- toBlock 'longterm' is not one of ('persona', 'human', 'episodes')
forget: refused as required -- category 'commitments' is protected by the retention policy
capacity enforcement dropped exactly the 4 unprotected facts
promotion enforcement moved m_6 scratch -> human
Enter fullscreen mode Exit fullscreen mode
  • Changing the policy changes the outcome measurably: identical seeded facts, two caps, 2 vs 14 forgotten.

11. Bugs that the design surfaced

Worth writing down, because each one is a general trap in this class of system.

  1. KeyError: 'refs' killed the whole Sleep run. Rejected mutations don't carry a refs dict, and the mutation-logging loop indexed it unconditionally. It only fired when a pass had a rejection and was not dry-running — a path the demo never hit until a merge was refused for spanning two blocks. Fix: .get("refs") or {"proposal": ...}, and the verifier now forces the rejection path with model-free gate cases.
  2. The encoder was quadratic. Re-extracting the full transcript on every turn took the chat response for the memory-write event from ~15 s to ~200 s. Fix: an extracted_through cursor.
  3. Apostrophes produced phantom tokens. normalize("user's") → user s, and s isn't a stopword, so perfectly good facts were refused for "inventing" a token. Fix: delete apostrophes rather than mapping them to spaces.
  4. Inflection broke traceability. "only use Python" → a fact saying "uses Python" scored as invented. Fix: fold both sides with the same bounded stem_token, and keep stopwords in the source vocabulary while excluding them from the claim's content tokens.
  5. A dry run that wrote a snapshot polluted the forgetting curve with a no-op point. Fix: a dry run writes nothing at all, and the assertion got stronger as a result.
  6. Rejections vanished from the diff view. The range filter compared a mutation's snapshot to the range's endpoint only; a merge rejected mid-range never appeared. Fix: lo < seq <= hi.
  7. Misleading rejection text. An item dropped earlier in the same pass was reported as "does not exist in the pre-pass store" — false, and exactly the kind of audit-log lie that makes an auditable system unauditable. Fix: _lookup() distinguishes never existed from consumed earlier in this pass.
  8. The quiz failed on a synonym. Store: …sending emails. Agent: …sending messages. Coverage 0.875 → a strict-but-flaky failure. Fix: rank merged items out of the question pool and allow one stricter verbatim re-ask, recorded as attempts: 2 in the report — visible, not hidden.
  9. Local-model ghosting. Killing a client mid-request does not cancel LM Studio's generation; the orphan kept the queue, and the next request looked like a 15-minute hang with an idle server. Diagnosis saved: if a local run stalls with no CPU on either side, suspect an abandoned request.

12. Trade-offs I'd argue about

  • Containment similarity treats subset restatements as duplicates. Right for memory collapse, wrong if you consider "lives in Lisbon" and "lives in Lisbon with a view" different facts.
  • 0.9 is strict for short sentences: one swapped word in an 8-token fact scores 0.875 and is refused. I kept the spec's threshold and let those pairs surface as rejections rather than quietly lowering the bar — the Merge pass is where they legitimately belong.
  • No embeddings. Intentional: similarity must be explicable in a rejection message. A cosine threshold you can't narrate is a guardrail in name only.
  • The answer key is token overlap, so a genuinely correct answer phrased differently fails. That is a cost I accept for a test that can't be satisfied by confident nonsense.
  • Small model, slow loop. A 4B model makes a full Sleep ~2–4 minutes. The UI is honest about it with live per-pass counts, which turns out to be a feature: watching a pass stall for 90 s is information about your provider.

13. Run it, or fork it

git clone https://github.com/<you>/nocturne && cd nocturne
uv venv --python 3.12 .venv && uv pip install -r requirements.txt
cp .env.example .env            # OPENAI_BASE_URL / OPENAI_API_KEY / MODEL
npm install --prefix client
npm run dev                      # API :3001, client :5173
npm run verify                   # 18 checks; ~13 min on a local model, exit 0 on success
Enter fullscreen mode Exit fullscreen mode

Stack: FastAPI 0.142 + Uvicorn + httpx on Python 3.12; SQLite (WAL) with no ORM; React 18.3 +
TypeScript 5.9 (strict) + Vite 5.4 + Tailwind 3.4, with the forgetting curve as hand-written SVG —
no chart dependency. The whole client bundle is 190 KB (59.6 KB gzipped).

If you want to extend it, the highest-value next pieces are a contradiction pass (conflicting
facts should supersede with a validity interval, not coexist), a lineage view (the ancestry data
already lives in mutations.refs), and a cross-model consolidation benchmark — the verifier
already accepts --provider/--model, so "which model edits memory most honestly" is one loop away.

The repo's README.md documents every gate, rejection rule and measure in full, plus the verifier
output. The most instructive file is agent/memory/textmath.py: 136 lines that decide whether your
agent's memory is auditable or vibes.

Code & more: https://www.dailybuild.xyz/project/272-nocturne

Top comments (0)