How Sentinel's memory works, why it can't grow forever, and why forgetting turned out to be the hardest part to design.
Part 1 — For everyone
The problem nobody talks about
Every AI agent demo shows you what the agent remembers. Nobody shows you what happens after three months — when the agent's memory file is 400 MB of stale decisions, half-true observations and resolved incidents it keeps re-acting on.
Human memory solved this a long time ago. You don't remember every day of your life. You remember what was important, you forget what was routine, and a few bad experiences shape your caution for years. The rest quietly fades.
We built our autonomous agent — Sentinel — around the same idea. Its memory is not a database that grows forever. It is a layered system that forgets by design, keeps what matters, and archives the rest instead of losing it completely.
The four ideas, in plain language
1. Everything has a hard limit. Each kind of memory — decisions, audit trail, rollbacks, observations — has a fixed maximum number of entries. The memory is one JSON file, so the total size is capped by the sum of those limits. It physically cannot swell to infinity. Disk stays quiet.
2. It forgets the least important, not the oldest. When a list is full, Sentinel doesn't drop the oldest entry — it scores every entry for importance and drops the weakest. A decision to actually change code is more memorable than a routine skip. A serious rejection is memorable. A "nothing happened" day is not.
3. Some things are never forgotten. The agent's identity — its mission, its rules, its hard "never do this" list — cannot be wiped by compaction. Forgetting is for experience, not for DNA.
4. Forgotten ≠ deleted. Dropped entries go to a compressed, checksummed archive (gzip + SHA-256) written in the background. If the agent ever needs the old detail, it can be restored. Losing an archive write can never crash the agent — it fails open and moves on.
A concrete picture
Think of Sentinel's memory like a desk:
- The desktop (hot memory): small, capped, only the most important working items.
- The filing cabinet (the archive): everything pushed off the desk, compressed and checksummed, retrievable but never in the way.
- The spine (identity): never leaves the body, never ages out.
There's even a "trauma center" — a counter of things that went badly (rollbacks, security blocks, failed simulations), categorized so the agent knows what it is cautious about, not just that it is cautious. Same file judged the same way twice is collapsed into one entry with a repeats count — repeated evidence, not new facts.
Why this matters
An agent that never forgets becomes slow, expensive, and eventually wrong — acting on context that stopped being true months ago. An agent that forgets everything can't learn. Sentinel tries to sit in between: bounded, importance-weighted, restorable. And because every forgetting event is counted (memory_compaction counters), the forgetting itself is auditable — it's a designed behavior, not silent data loss.
Part 2 — For developers
One file, layered, hard-capped
All of Sentinel's working memory lives in a single JSON state file (data/sentinel_memory.json, gitignored, mirrored to GitLab on a schedule). The caps are constants in libs/core/memory-manager.js:
// libs/core/memory-manager.js:9-24
const LIMITS = {
decisionStream: 10,
evolutionHistory: 50,
auditLog: 50,
externalSignals: 10,
rollbackHistory: 20,
maxStringLength: 180,
observationSamples: 30,
semanticModules: 200,
reflectionLog: 30,
intentEntries: 100,
intentExamples: 5,
decisionAnalytics: 40,
// Same horizon DecisionQuality writes with (HISTORY_LIMIT), one row per day.
decisionQualityHistory: 90
};
Two TTLs sit on top: CACHE_TTL_MS = 5 min for ephemeral scratch (L0) and WORKING_MEMORY_TTL_MS = 1 h for stale current-task cleanup (L1) — memory-manager.js:26-27. Individual strings are clipped at 180 chars, so a single oversized payload can't bloat a capped list.
Importance-scored compaction — the "human-like" part
This is the mechanism that makes it more than a ring buffer:
// libs/core/memory-manager.js:301-327 (condensed)
// Compaction forgets the LEAST important memory, not merely the oldest —
// closer to how human memory decays.
function decisionImportance(entry) {
const action = String(decision.action || '').toUpperCase();
let importance = 0.3;
if (action === 'EVOLVE') importance = 0.8;
else if (action === 'REJECT') importance = 0.7;
else if (action === 'SKIP') importance = 0.2;
// ... reason keywords (security, rollback, blocked) push it higher
return importance;
}
Real changes (EVOLVE) and refusals (REJECT) outrank routine skips. Within the same class, a security/rollback-flavored reason adds weight. Age is not the primary axis — a six-week-old rejection that mattered still beats yesterday's noise.
Deduplication: same fact, counted not stored
compactDecisionAnalytics folds entries with the same fingerprint into one row with a repeats counter — the comment in memory-manager.js:836 puts it as "the same file, judged the same way, is one fact repeated — not new evidence."
Trauma center: categorized, not just counted
// libs/core/memory-manager.js:67-112
const TRAUMA_CATEGORIES = ['rollback', 'security', 'performance',
'hallucination', 'architecture'];
trauma_center tracks shadow_guard_blocks, simulation_failures, rollback_events, plus per-category counters — so downstream policy can be cautious about something rather than generically timid.
Identity: the never-forgotten layer
defaultIdentity() (memory-manager.js:34-60) holds the mission, four operating rules, a five-item never list (never send raw repo dumps or secrets to external agents, never persist foreign source in memory, never bypass governance or the daily budget, never commit output that failed the trust boundary, never move funds without an approved decision), an always list, and the budget philosophy. It's versioned via IDENTITY_VERSION — bumping it refreshes stale snapshots from code, so identity lives in the code, not in whatever an old file contains. Compaction never touches it.
L6 archive: forget, but don't lose
libs/knowledge/l6-archive.js is the cold-storage layer:
-
Background only —
enqueue()returns immediately; gzip + disk write happens onsetImmediate, never blocking the runtime hot path. -
Restorable — each snapshot is gzipped JSON plus a sidecar with a SHA-256 checksum;
restore()verifies integrity. - Fail-open — an archive error is logged and swallowed; a failed archive write must never crash the agent or corrupt live memory.
- Stored under
data/archive/l6(gitignored, never committed).
Auditable forgetting
Every compaction run increments counters (memory-manager.js:243-256): runs, duplicates removed, entries trimmed per stream (decision, analytics, quality, evolution, audit, external, rollback), plus last_compaction_ts. You can diff two snapshots and see exactly what the agent chose to forget.
What this gets you — and what it costs
Gets you:
- Memory file size is bounded by the sum of LIMITS — no unbounded growth, no disk pressure.
- Retention is semantic: important events outlive routine ones regardless of age.
- Cold detail is never truly gone — it's an
l6restore away, checksum-verified. - The forgetting policy itself is inspectable and testable (three dedicated test files:
tests/unit/memory-brain-layers.test.js,memory-layers.test.js,memory-sync.test.js; full suite currently 575 tests).
Costs you:
- Importance is heuristic — a fixed weight table, not learned. A false SKIP can outrank a true REJECT less often than it should.
- The archive trades disk for recall latency: restoring means decompression and integrity checks, done offline.
- Bounded analytics windows mean "history" is the last N events, not all history —
decisionAnalytics: 40is a rolling window, which is why quality snapshots (decisionQualityHistory: 90, one row/day) exist as the long-horizon view.
The honest summary
This is engineered memory inspired by human decay — importance scoring, TTLs, categorized trauma, restorable archives — not neuroscience, and not a vector DB. It works for an agent whose job is bounded-state decision-making, and its limits are deliberate, documented, and measured. If you're building an agent that runs unattended for months, "how does it forget?" is probably a better first question than "how much can it store?"

Top comments (0)