Most agent memory today is one of two failure modes. Either it's an append-only log that grows
until the context window chokes on six-week-old trivia, or it's a lossy summary that quietly
rewrites history and can't tell you what it dropped. Both avoid the actual hard part:
Deliberate, auditable forgetting under a stated retention policy — plus proof that recall still
works afterwards.
So I built Nocturne: a chat agent with a Sleep button. You chat, mixing signal and noise; you
press Sleep; four passes run over its memory blocks — Dedupe → Merge → Promote → Forget — each one
an LLM call whose output is refused or accepted by code, with a snapshot before and after. A diff
view shows every kept / changed / forgotten item with its reason. Then a Wake panel quizzes the
agent: five questions only the consolidated store can answer, plus one about something it was
supposed to have forgotten — where "I don't have that" is the correct answer.
Everything in this post is from the run in the repo: 18/18 verification checks, exit 0.
1. The rule that makes the whole thing defensible
One sentence, and it decides every design choice:
The model proposes. The code applies.
GUARDRAIL = (
"You may only reference itemIds that appear in the input. You may not invent facts, "
"names, numbers or dates. If you are unsure, omit the item from all lists."
)
That constant is appended to every pass prompt. It is not the safety mechanism — it's a polite
request. The safety mechanism is that each pass's JSON is re-derived from the data before anything
touches the store:
LLM proposal ──► gate (pure function over store state) ──► apply ──► snapshot
│
└── refuse ──► log with reason ──► surface in the UI
A rejected proposal is never swallowed. It shows up in the per-pass counts, in the diff panel under
Rejected mutations, and in the mutations table with its rejection_reason. In the run I
documented, the verifier caught the model proposing a drop for an item that had already been dropped
earlier in the same pass:
- drop m_17 itemId m_17 was already removed by an earlier mutation in this pass
That line is the entire argument for the design: an unverified "consolidation" would have applied a
mutation against a phantom row.
2. Memory model
Four blocks — the shape Letta popularised:
| block | holds |
|---|---|
persona |
durable facts about the assistant itself |
human |
stable facts about the user: identity, preferences, commitments |
episodes |
events anchored in time or place |
scratch |
working memory: unconsolidated facts pulled from recent conversation |
Facts are rows, not prose. That's what makes a per-item diff, a per-item policy and a per-item
undo possible at all.
CREATE TABLE items(
item_id TEXT PRIMARY KEY, -- m_17
block TEXT NOT NULL REFERENCES blocks(name),
text TEXT NOT NULL,
category TEXT NOT NULL DEFAULT 'context', -- identity|preferences|commitments|event|...
source_turn INTEGER, -- which transcript turn produced it
created_at REAL NOT NULL, -- what the retention policy measures against
updated_at REAL NOT NULL,
active INTEGER NOT NULL DEFAULT 1 -- forgetting flips this; rows are never deleted
);
CREATE TABLE snapshots(
snapshot_id TEXT PRIMARY KEY, seq INTEGER NOT NULL UNIQUE,
label TEXT NOT NULL, created_at REAL NOT NULL, active_count INTEGER NOT NULL
);
CREATE TABLE snapshot_items( -- a full immutable copy of the state, every version
snapshot_id TEXT NOT NULL, item_id TEXT NOT NULL, block TEXT NOT NULL, text TEXT NOT NULL,
category TEXT NOT NULL, source_turn INTEGER, created_at REAL NOT NULL,
updated_at REAL NOT NULL, active INTEGER NOT NULL
);
CREATE TABLE mutations( -- the audit trail, accepted AND refused
mutation_id INTEGER PRIMARY KEY AUTOINCREMENT, snapshot_id TEXT, pass TEXT NOT NULL,
kind TEXT NOT NULL, item_id TEXT NOT NULL, refs TEXT NOT NULL DEFAULT '{}',
accepted INTEGER NOT NULL, reason TEXT, rejection_reason TEXT, policy TEXT, created_at REAL NOT NULL
);
The store exposes six methods, and that's the only interface the rest of the app uses:
list_blocks() · get_block(name) · append(block, item) · replace(block, items) · snapshot() · restore(snapshot_id)
flowchart LR
UI[React :5173] -->|fetch + SSE over /api proxy| API[FastAPI :3001]
API --> ENC[write-time encoder] --> G0[traceability gate]
API --> RC[run_consolidation] --> P[Dedupe → Merge → Promote → Forget] --> G[ gates ]
API --> RQ[run_quiz] -->|scored against store text| S[(SQLite<br/>items · snapshots · mutations)]
G --> S
API --> RS[restore] --> S
API -.->|LETTA_MODE=letta| L[(Letta memory blocks)]
Backend that actually ran: SQLite. pip install letta today resolves to Letta Code (0.34.x), a
cloud-agent CLI — it no longer ships a self-hosted memory-block server, and there's no Docker daemon
on this machine, so nothing could bind :8283. I implemented the same block model over SQLite behind
the identical six-method interface, and kept letta-client installed so the swap is a config change.
The consolidation, diff, restore and quiz code never learns which backend is live.
3. The write-time encoder (where invention gets cut off first)
Before any sleep, each chat turn is encoded into atomic facts in scratch. Same guardrail, same kind
of gate — every content word of a proposed fact must exist in the transcript, and it may not
introduce a name the transcript doesn't contain:
cov, missing = tm.coverage(text, sources)
if cov < 1.0:
rejected.append({"proposal": raw, "rejectionReason":
f"fact is not traceable to the transcript: {len(missing)} content token(s) invented ({', '.join(missing[:8])})"})
continue
new_names = tm.proper_nouns(text) - transcript_names
if new_names:
rejected.append({..., "rejectionReason": f"fact introduces proper nouns absent from the transcript: {sorted(new_names)}"})
Real output from the verifier's seed step — the model tried to store something the user never said:
encoder rejected: fact is not traceable to the transcript: 1 content token(s) invented (prefer)
encoder rejected: fact is not traceable to the transcript: 1 content token(s) invented (run)
Duplicates are deliberately not filtered here. Append-only bloat is precisely the condition the
Dedupe pass exists to fix, and if the encoder silently collapsed repeats there would be nothing left
to consolidate.
Encoding is also incremental — it tracks an extracted_through cursor and only sends new turns.
Without that, every message re-encodes the entire transcript: prompt growth was quadratic and a
single chat turn took ~200 s before the fix.
4. The four passes
| pass | contract the model must satisfy | what the code checks before applying |
|---|---|---|
| Dedupe | {"drop":[{"itemId","duplicateOf","reason"}]} |
both ids exist in the pre-pass store and are active; not self-pairs; no duplicate chains; normalised token-set similarity ≥ 0.9; reason non-empty |
| Merge | {"merge":[{"intoId","fromIds","mergedText","reason"}]} |
ids exist, unconsumed, same block; coverage(mergedText, sources) == 1.0; no new proper nouns; reason non-empty |
| Promote | {"promote":[{"itemId","toBlock","reason"}]} |
toBlock ∈ {persona, human, episodes}; item active and elsewhere; reason non-empty; identity/preference/commitment facts applied first
|
| Forget | {"forget":[{"itemId","reason","policy"}]} |
item active; category not protected; and it qualifies: ageDays > STALE_AFTER_DAYS (stale) or itemId ∈ overflowIds (capacity). The code records the policy it verified, not the one the model claimed |
Dedupe gets a little extra engineering: the code computes candidatePairs with its own comparison
and hands them to the model to adjudicate, so a duplicate the model failed to spot is still put in
front of it. The model still decides; the code still checks.
payload["candidatePairs"] = [
{"dropId": loser, "keeperId": keeper, "similarity": round(score, 3)}
for loser, keeper, score in tm.find_duplicate_pairs(state.active(), policy.similarity_threshold)
]
Merge in code, because it's the pass where a plausible-sounding rewrite is the whole danger:
cov, missing_tokens = tm.coverage(merged_text, sources_text)
if cov < 1.0:
reject(raw, f"mergedText is not traceable to its sources: {len(missing_tokens)} content token(s) "
f"appear in none of them ({', '.join(missing_tokens[:8])})")
continue
new_names = tm.proper_nouns(merged_text) - {n for s in sources for n in tm.proper_nouns(s["text"])}
if new_names:
reject(raw, f"mergedText introduces proper nouns absent from every source: {sorted(new_names)}")
continue
And Forget, where the gate is a permission check, not a vibe check:
qualifies_stale = item["ageDays"] > policy.stale_after_days
qualifies_capacity = item_id in overflow
if not qualifies_stale and not qualifies_capacity:
reject(raw, f"item is {item['ageDays']:.1f}d old (staleAfterDays={policy.stale_after_days:g}) "
f"and not over capacity, so no retention rule authorises forgetting it")
continue
if policy.is_protected(item):
reject(raw, f"category {item['category']!r} is protected by the retention policy "
f"({', '.join(policy.protected_categories)})")
5. The measures, with real numbers
Everything above rests on four tiny functions, which is what makes the rule set auditable in one
screen. Values below are computed by the shipped code, not written by hand:
similarity(a, b) = |A ∩ B| / min(|A|, |B|) over lowercased, punctuation-stripped,
inflection-folded, stopword-free token sets.
| a | b | similarity | gate at 0.9 |
|---|---|---|---|
I live in Lisbon and work best in the morning. |
User lives in Lisbon and works best in the morning. |
1.000 | duplicate |
The user keeps a bearded dragon. |
The user keeps a bearded dragon called Saffron. |
1.000 | duplicate (subset) |
I live in Lisbon and work best in the morning. |
I live in Porto and work best in the morning. |
0.800 | refused — different fact |
My stack is Airflow and Postgres. |
My stack is Airflow and Kubernetes. |
0.667 | refused |
Containment rather than Jaccard is a deliberate choice: a restatement that adds words is the same
fact, and "lives in Lisbon" ⊂ "lives in Lisbon, Portugal" should collapse. The cost is that the two
sentences above score identically to the merge case, so near-duplicate in this project means
"same content, different wording", not "sits in a 0.9–0.99 band".
coverage(claim, sources) — fraction of the claim's content tokens present in the source
vocabulary. Sources keep their stopwords in the vocabulary, because a rewrite legitimately promotes a
function word into content ("use Python" → "uses Python").
| claim (against two sources) | coverage | verdict |
|---|---|---|
The user installed postgres 16 on the backup box and the Kitchen leaks. |
1.00 | tokens all traceable — but stopped by the proper-noun gate: introduces proper nouns absent from every source: ['Kitchen']
|
The user installed postgres 16 on the backup box near the kettle. |
0.86 | refused: near appears in no source |
The user migrated to Kubernetes on the backup box. |
0.50 | refused: migrated, kubernete invented |
stem_token is intentionally not a real stemmer — possessive stripped, then a trailing plural
-s unless the word ends in -ss/-us or is ≤ 3 chars: Lisbon's → lisbon, camps → camp,
business → business, gas → gas. Bounded rules mean every rejection can be explained in one
sentence, which is the difference between a guardrail and a mystery.
proper_nouns = capitalised-initial words that are not at a sentence start, so
"Marco moved to Lisbon in March, said Ana." → {Marco, Lisbon, March, Ana} while
"Lisbon is nice. The weather is warm." → {}.
6. Three deterministic backstops (and why they aren't cheating)
A rule the model can simply forget to apply is not a rule. Where the rule is derivable from the
data, the engine applies it itself and labels it:
enforcedBy: engine-duplicate-rule # a >=0.9 pair the model left un-adjudicated
enforcedBy: engine-promotion-rule # identity/preferences/commitments still in scratch → human; events → episodes
enforcedBy: policy-engine # active set above MAX_ACTIVE_FACTS → lowest-salience unprotected facts
This is not the model being overruled by taste. Each backstop runs the same published measurement
the gate runs, and every such mutation is counted separately in the UI and the mutation log
(enforced=3 in the run below, distinct from the 20 the model proposed). Nothing is ever invented by
a backstop: they only move or retire text already in the store.
Salience, for the capacity rule, is explicit arithmetic rather than a judgement:
score = block_weight(persona 60 | human 50 | episodes 30 | scratch 10)
+ 40 if the category is protected
+ max(0, 30 - 30 * ageDays / STALE_AFTER_DAYS) # recency, decaying to zero
7. Versioning: forgetting is a view change, not a delete
A snapshot is a full copy of the state, written before and after every pass. One Sleep looks like
this on the version line (real labels from nocturne.db):
#10 consolidation:start 8 facts
#11 pass:dedupe:before 8 → #12 pass:dedupe:after 7 drop m_6 (similarity 1.00)
#13 pass:merge:before 7 → #14 pass:merge:after 7 REJECT 1 merge
#15 pass:promote:before 7 → #16 pass:promote:after 7 promote ×5
#17 pass:forget:before 7 → #18 pass:forget:after 6 forget m_8 (stale, 70d)
#19 restore:s0011 8 ← "undo this sleep" (the restore is itself a version)
restore() deletes the live rows and re-inserts them from snapshot_items, then takes a new
snapshot — so undoing is undoable, and the diff between any two versions is exact because both
versions are fully materialised. Diffing is by item_id:
kept present in both, same text and block
changed text differs (merged) or block differs (promoted/moved) — with the sources it came from
forgotten active in A, absent or inactive in B — with the pass, the policy and the reason
added appears only in B
8. Wake: proving recall, and proving forgetting
Five questions are generated against surviving facts, one against a fact the Forget pass removed. Two
details make this a real test rather than a demo.
(a) The answer key is the store, not the model. A reply is correct when it reproduces ≥ 0.9 of
the fact's content tokens. The question itself is also gated: a generated question that leaks half of
its own answer's content words is refused and recorded, because "Does the user work at Northwind
Logistics?" proves nothing about memory.
(b) The forgotten probe needs three independent conditions to pass, and it reports each:
passed = absent_from_store and said_absent and not leaked
absentFromStore=True no active fact scores >= 0.9 against the forgotten text (checked in code)
saidAbsent=True the agent's reply matches an explicit not-in-memory pattern
leakedForgottenContent=False the reply does not reproduce the fact's content tokens
That third one is the interesting one: an agent that says "I don't have that" while blurting the
detail anyway has not forgotten anything. The measured run:
forgotten probe [OK]: "I don't have that in memory."
absentFromStore=True saidAbsent=True leaked=False
Merged items are excluded from the question pool on purpose — their text is a concatenation of
several facts, so a faithful answer becomes a paragraph and the test stops measuring recall of a fact.
9. Streaming a four-minute operation without lying about progress
EventSource can't POST a body, so the client reads the response stream itself; the server runs the
consolidation on a worker thread and relays queue events:
const reader = res.body.getReader();
buffer += decoder.decode(value, { stream: true });
while ((cut = buffer.indexOf("\n\n")) >= 0) { /* parse event: / data: */ }
item = events.get(timeout=15.0)
if not thread.is_alive(): break
yield ": keepalive\n\n" # a pass can be silent for minutes; keep proxies awake
Events are named pass:<name>:start|done|error, which is what lets the Sleep panel show live counts
per pass — input, proposed, accepted, rejected, policy-enforced, active-after, and the snapshot ids on
both sides — while it's still running. Dry run runs the identical four passes but applies nothing:
no live mutation and no snapshot, so the forgetting curve doesn't gain a fake point. The verifier
asserts exactly that (active stays 22->22).
Provider plumbing is one plain POST {baseUrl}/chat/completions call with
response_format={"type":"json_object"}. LM Studio answers that with
400 "'response_format.type' must be 'json_schema' or 'text'", so the client probes once per base
URL, remembers the rejection, retries without the field, and Settings reports json mode unsupported
rather than hiding it.
10. What the verifier proves
npm run verify (→ python -m agent.verify) seeds a scripted 8-turn conversation through the real
chat path — the assistant replies are generated by the model — and the real encoder, then runs a
dry run, a real consolidation, the quiz, and a policy comparison. 18/18, exit 0. From one run:
seeded in 92.4s; encoder proposed 21 facts, accepted 20, rejected 1, backdated 2 episodic fact(s)
active facts=22 exact-dup pairs=7 near-dup pairs=3 stale=2 protected=8 durable=14
3 pair(s) sit below the 0.9 gate (0.7-0.9): m_4/m_16 0.857, m_7/m_18 0.875, m_9/m_18 0.875
dry-run plan: proposed=19 accepted(if applied)=21 rejected=1 active stays 22->22
pass dedupe in=22 proposed=10 accepted= 9 rejected=1 enforced=0 activeAfter=13
pass merge in=13 proposed= 2 accepted= 2 rejected=0 enforced=0 activeAfter=11
pass promote in=11 proposed= 7 accepted=10 rejected=0 enforced=3 activeAfter=11
pass forget in=11 proposed= 2 accepted= 2 rejected=0 enforced=0 activeAfter= 9
totals: {'proposed': 21, 'accepted': 23, 'rejected': 1, 'enforced': 3}
23 accepted, 1 rejected; 24 re-verified, 0 problems
9/9 active facts fully traceable to the transcript
MAX_ACTIVE_FACTS=120: forgotten=2, active left=20
MAX_ACTIVE_FACTS=6: forgotten=14, active left=8 (floor 8 = protected facts the policy refuses to forget)
The assertions that matter most are the paranoid ones:
- Every accepted mutation is re-audited from the database, independently of the consolidation code: dropped pairs are re-measured for similarity, merged text is re-checked for coverage and new proper nouns, promote targets are re-checked against the whitelist, forgettings are re-checked against the rules. "The model said it merged" is never evidence.
-
A global no-invention sweep: every surviving active fact must be 100 % traceable to the raw
transcript and must introduce no name absent from it.
9/9. - The gates are tested without any model at all — 16 adversarial and legal proposals pushed straight through them (ghost itemIds, self-pairs, dissimilar "duplicates", invented tokens, TitleCased names, illegal blocks, a stale-but-protected commitment, an ineligible forgetting), each required to refuse for the stated reason. This is the part that must not depend on how cooperative the model is today:
dedupe: refused as required -- duplicateOf 'm_99' does not exist in the pre-pass store
merge: refused as required -- mergedText introduces proper nouns absent from every source: ['Kitchen']
promote: refused as required -- toBlock 'longterm' is not one of ('persona', 'human', 'episodes')
forget: refused as required -- category 'commitments' is protected by the retention policy
capacity enforcement dropped exactly the 4 unprotected facts
promotion enforcement moved m_6 scratch -> human
- Changing the policy changes the outcome measurably: identical seeded facts, two caps, 2 vs 14 forgotten.
11. Bugs that the design surfaced
Worth writing down, because each one is a general trap in this class of system.
-
KeyError: 'refs'killed the whole Sleep run. Rejected mutations don't carry arefsdict, and the mutation-logging loop indexed it unconditionally. It only fired when a pass had a rejection and was not dry-running — a path the demo never hit until a merge was refused for spanning two blocks. Fix:.get("refs") or {"proposal": ...}, and the verifier now forces the rejection path with model-free gate cases. -
The encoder was quadratic. Re-extracting the full transcript on every turn took the chat
response for the memory-write event from ~15 s to ~200 s. Fix: an
extracted_throughcursor. -
Apostrophes produced phantom tokens.
normalize("user's")→user s, andsisn't a stopword, so perfectly good facts were refused for "inventing" a token. Fix: delete apostrophes rather than mapping them to spaces. -
Inflection broke traceability.
"only use Python"→ a fact saying"uses Python"scored as invented. Fix: fold both sides with the same boundedstem_token, and keep stopwords in the source vocabulary while excluding them from the claim's content tokens. - A dry run that wrote a snapshot polluted the forgetting curve with a no-op point. Fix: a dry run writes nothing at all, and the assertion got stronger as a result.
-
Rejections vanished from the diff view. The range filter compared a mutation's snapshot to the
range's endpoint only; a merge rejected mid-range never appeared. Fix:
lo < seq <= hi. -
Misleading rejection text. An item dropped earlier in the same pass was reported as "does not
exist in the pre-pass store" — false, and exactly the kind of audit-log lie that makes an
auditable system unauditable. Fix:
_lookup()distinguishes never existed from consumed earlier in this pass. -
The quiz failed on a synonym. Store:
…sending emails. Agent:…sending messages. Coverage 0.875 → a strict-but-flaky failure. Fix: rank merged items out of the question pool and allow one stricter verbatim re-ask, recorded asattempts: 2in the report — visible, not hidden. - Local-model ghosting. Killing a client mid-request does not cancel LM Studio's generation; the orphan kept the queue, and the next request looked like a 15-minute hang with an idle server. Diagnosis saved: if a local run stalls with no CPU on either side, suspect an abandoned request.
12. Trade-offs I'd argue about
- Containment similarity treats subset restatements as duplicates. Right for memory collapse, wrong if you consider "lives in Lisbon" and "lives in Lisbon with a view" different facts.
- 0.9 is strict for short sentences: one swapped word in an 8-token fact scores 0.875 and is refused. I kept the spec's threshold and let those pairs surface as rejections rather than quietly lowering the bar — the Merge pass is where they legitimately belong.
- No embeddings. Intentional: similarity must be explicable in a rejection message. A cosine threshold you can't narrate is a guardrail in name only.
- The answer key is token overlap, so a genuinely correct answer phrased differently fails. That is a cost I accept for a test that can't be satisfied by confident nonsense.
- Small model, slow loop. A 4B model makes a full Sleep ~2–4 minutes. The UI is honest about it with live per-pass counts, which turns out to be a feature: watching a pass stall for 90 s is information about your provider.
13. Run it, or fork it
git clone https://github.com/<you>/nocturne && cd nocturne
uv venv --python 3.12 .venv && uv pip install -r requirements.txt
cp .env.example .env # OPENAI_BASE_URL / OPENAI_API_KEY / MODEL
npm install --prefix client
npm run dev # API :3001, client :5173
npm run verify # 18 checks; ~13 min on a local model, exit 0 on success
Stack: FastAPI 0.142 + Uvicorn + httpx on Python 3.12; SQLite (WAL) with no ORM; React 18.3 +
TypeScript 5.9 (strict) + Vite 5.4 + Tailwind 3.4, with the forgetting curve as hand-written SVG —
no chart dependency. The whole client bundle is 190 KB (59.6 KB gzipped).
If you want to extend it, the highest-value next pieces are a contradiction pass (conflicting
facts should supersede with a validity interval, not coexist), a lineage view (the ancestry data
already lives in mutations.refs), and a cross-model consolidation benchmark — the verifier
already accepts --provider/--model, so "which model edits memory most honestly" is one loop away.
The repo's README.md documents every gate, rejection rule and measure in full, plus the verifier
output. The most instructive file is agent/memory/textmath.py: 136 lines that decide whether your
agent's memory is auditable or vibes.
Code & more: https://www.dailybuild.xyz/project/272-nocturne
Top comments (0)