A memory handoff replaces a large conversation window while keeping retained source segments available for recall.
The goal is to reduce how much history the LLM has to carry in its active context, while preserving the information it can recover when needed.
That goal is different from shrinking a database. Keeping exact source text outside the prompt can be the right trade-off: the model stops processing the entire history on every turn, but retains a path back to the details.
I checked the implementation, reran its targeted tests, and measured the replacement handoffs stored by the system. Across 10,241 completed runs, the selected source windows contained 1,137,387,720 characters. Their handoffs contained 15,375,395 characters in total—a 98.65% reduction in the replaced portion of context, measured in characters.
I also recomputed the stored text hashes for 416,472 retained segments. None lacked exact source text, none had an exact or reduced hash mismatch, and none were missing a projection timestamp.
I then ran a controlled synthetic comparison with a live answering model and real semantic embeddings. Full history and handoff-plus-recall both answered all 12 questions correctly. Handoff alone answered only the one question whose correct response was UNKNOWN. The recalled-context arm used substantially fewer input tokens, including the evidence retrieved for each answer.
This is a small controlled result, not a claim of general lossless agent memory. The retrieval calls were made by the test harness, not chosen autonomously by an agent.
KEEP does not mean “keep in the prompt”
This is the most important detail in the implementation.
The pipeline routes segments through three decisions:
| Route | What happens to the segment |
|---|---|
| KEEP | Preserve exact text in external storage. |
| COMPRESS | Store an accepted reduced representation alongside the exact source. |
| DROP | Omit the segment from retained content. |
All three decisions belong to the processing of a selected conversation window. Once its retained material has been durably handed off, that window can leave the active context—including its KEEP segments.
The replacement is a compact handoff containing a short navigation summary, a manifest identifier, selected content hashes, and instructions for retrieving the omitted material.
Consequently, a run with nothing but KEEP decisions can still substantially reduce active context. It retains information externally instead of repeatedly presenting all of it to the LLM.
What the operational measurements show
I queried aggregate lengths from completed runs without exporting user message text:
| Measurement | Result |
|---|---|
| Completed runs | 10,241 |
| Selected source-window characters, summed | 1,137,387,720 |
| Replacement handoff characters, summed | 15,375,395 |
| Character reduction across those windows | 98.6482% |
| Smallest handoff | 584 characters |
| Largest handoff | 1,897 characters |
The percentage is calculated from the summed lengths. It is not an average of per-run percentages.
It also applies only to the replaced windows. System instructions, protected recent messages, tool definitions, and subsequent recall results still occupy context. The measurement is neither a whole-prompt reduction nor a token or inference-cost benchmark. Repeated runs can contain related material, so the source total is not a count of unique information.
The earlier segment counters showed 416,459 KEEP decisions, 13 COMPRESS decisions, and no DROP decisions. The reduced storage representation was only 818 characters shorter than its source.
That tiny difference answers a different question. It measures reduction within the externalized content, not the reduction from replacing the source window with its handoff. The large active-context change comes from externalization.
A small run that makes the distinction visible
I also ran the planner inside the deployed API environment on a synthetic window of 24 tool messages. Each contained an exact endpoint, port, command, and repeated diagnostic prose.
| Measurement | Result |
|---|---|
| Messages / segments | 24 / 24 |
| Source characters | 56,443 |
| Handoff characters | 1,313 |
| Reduction in the selected window | 97.67% |
| KEEP / COMPRESS / DROP | 24 / 0 / 0 |
| Exact segment hashes valid | 24 / 24 |
| Planner elapsed time | 3.242 seconds |
Every segment was retained verbatim, yet the resulting handoff was much smaller than the selected window.
This was a planner smoke test, without database writes. Protected tool content bypassed classification; the default summary and NLI paths were invoked. The recorded output does not distinguish an accepted generated summary from its safe fallback. The timing is one observation, not a latency distribution, and the test did not exercise semantic retrieval or downstream question answering.
The path out of active context
The agent selects a message-safe middle window. Secrets are redacted before persistence, and messages are divided into role-aware segments of at most 3,800 characters.
Technical content receives conservative handling. Tool, system, and developer segments are protected, as are recognized patterns such as code fences, commands, paths, hashes, and error evidence. These protections resolve to KEEP before model classification.
Other segments go through GLiClass routing. Weak decisions or model failures fall back to KEEP. COMPRESS candidates pass through LLMLingua, then an NLI gate. Empty output, output that does not shorten the segment, or a verdict other than entailment preserves the exact source instead.
The short handoff summary is separately checked. If its generation or validation fails, the planner uses a deterministic excerpt. This summary is a navigation aid: it is not intended to contain every fact from the discarded window.
“Exact source” here means the retained text after mandatory secret redaction. It does not mean an unredacted archive.
How the details come back
PostgreSQL stores authoritative exact/reduced pairs and their hashes. Qdrant holds the reduced retrieval representation and integrity metadata.
The agent has three recall paths:
- Semantic query: search for relevant segments, with optional session scoping.
- Manifest lookup: enumerate retained pieces from a particular compaction run.
- Content-hash lookup: recover a specific exact segment.
Search candidates are checked against PostgreSQL before being returned. Reranking is used when available. The exact source is authoritative when it differs from its reduced representation.
This design lets the agent start with a small handoff and request detail as needed. It also makes retrieval behavior central to the outcome. Preserving a command byte-for-byte is valuable, but the agent still needs to find that command at the appropriate moment.
Removing the window has a durability condition
The source window is removed only after a successful, valid handoff response.
PostgreSQL is the authority, but successful persistence alone is insufficient for the API acknowledgment. Projection must also complete. A projection failure returns a retryable error, and the original context remains available to the agent. Retrying the same operation repairs outstanding projection work; a completed duplicate returns the full stored handoff.
The integration also rejects malformed responses, empty handoffs, and invalid manifests. It disables an earlier tool-output pruning pass so that a failed compaction does not leave the conversation partially truncated.
This boundary protects against losing the source window during a failed handoff. It does not prove that the downstream agent will answer every question as well as it would with full history.
Checks I reran
The targeted suite finished with 28 passed in 7.60 seconds. It covered routing and fallback behavior, persistence/projection ordering, retry behavior, stale retrieval-candidate rejection, API behavior, observability, and the agent integration. The alert configuration also passed validation with five rules found.
The suite uses test doubles for models and storage, and the agent failure test uses a stand-in for the upstream compressor. Its results verify the tested contracts; they are not a live full-stack benchmark.
Separately, the database integrity scan produced:
| Check | Result |
|---|---|
| Retained segments checked | 416,472 |
| Missing exact text | 0 |
| Recomputed exact SHA-256 mismatches | 0 |
| Recomputed reduced SHA-256 mismatches | 0 |
| Missing projection timestamp | 0 |
This scan validates stored text against stored hashes. It does not compare the database with an independent original transcript, and a projection timestamp alone does not verify every live vector-index payload.
Testing answers after recall
The synthetic history contained 49 segments: operational facts, an explicit later correction, and 36 distracting records. All segments had the tool role and therefore took the protected KEEP route. This deliberately isolates externalization and recall from classifier or lossy-reduction quality.
The planner and storage/recall functions matched the inspected implementation. I used connection-local temporary PostgreSQL tables and a local Qdrant index, with real embeddings, the configured reranking path, and the same live answering-model endpoint for all three arms. The test wrote no synthetic records into the production memory tables or collections.
For each question, the harness compared full history, the handoff alone, and the handoff plus exact records returned by semantic top-5 retrieval. It supplied no expected answers to the model or retriever. The query was the question itself. Temperature was zero, and scoring compared the returned value with a predetermined answer after basic whitespace, case, and terminal-punctuation normalization.
| Input available to the answering model | Correct answers |
|---|---|
| Full history | 12 / 12 |
| Handoff only | 1 / 12 |
| Handoff + semantic recall, top 5 | 12 / 12 |
The endpoint reported 96,819 input tokens across the 12 full-history calls and 22,296 across the 12 recall calls: 76.97% fewer input tokens, with retrieved evidence included. Per question, the averages were 8,068.25 and 1,858 tokens. These are reported request usage counts, not independently tokenized measurements or a cost benchmark; embedding, reranking, compaction, and output-generation costs are not included in that percentage.
The questions covered an IP address, port, exact command, full SHA-256, responsible person, database decision, prohibited operation, Russian-language value, superseding revision, staging-versus-production distinction, prerequisite, and an unknown password. The designated evidence segment appeared in the top five for all 12 questions. For the revision question, the designated evidence was the newer record.
All 49 segments were also recovered through exact-hash lookup with matching text, and manifest lookup returned all 49 rows. The selected history was 49,897 characters; its handoff was 1,329 characters.
This is an intentionally limited check. It does not exercise COMPRESS or DROP quality, a production-sized vector collection, authentication, concurrent retrieval, multiple successive compactions, or an autonomous decision to call the recall tool. The unknown-answer case also does not test resistance to plausible contradictory evidence. The live response reports a serving alias rather than a pinned model-weight revision.
What “without losing memory quality” still needs to demonstrate
Preserving memory quality is the objective. The next comparison must measure it directly:
- Give an agent the full history and record its answers.
- Replace the eligible window with the handoff and enable recall.
- Ask the same questions, including exact values, earlier decisions, corrections, and conflicting updates.
- Compare answer accuracy, retrieved evidence, latency, and tokens consumed—including the recall results brought back into context.
The context measurement should include both the immediate reduction and the cost of subsequent retrieval. A system that starts with a tiny handoff but repeatedly reloads the entire history may save much less over a complete task.
A verbatim-externalization baseline is useful here: it separates the benefit of moving history out of context from the effect of reducing the retrieval representation.
The evidence so far supports a concrete claim: large selected conversation windows can be replaced with compact handoffs while retaining integrity-checked exact segments for recall; on this 12-question synthetic check, recall preserved answer accuracy. Broader unchanged answer quality remains a benchmark question, not an inference from the storage checks or this small corpus.
The system is designed around a smaller working context and recoverable history. That is the trade-off worth measuring.
Reproduce the comparison
The public benchmark includes the synthetic corpus, expected answers, Python runner, tests, and recorded responses. The README explains offline verification and live endpoint configuration.
The original recorded run is separate from the portable live run. The portable version uses SQLite, exhaustive cosine search without reranking, and a deterministic handoff. It also scored 12/12 with full history and recall, using 96,819 versus 14,208 input tokens (85.33% fewer). Its handoff-only score was 2/12: the excerpt exposes the first answer, and UNKNOWN is also correct. These are different protocol variants; their token totals should not be mixed.
The repository reproduces the comparison procedure, not the private production stack. Weight revisions were not captured, so identical outputs are not guaranteed. Operational aggregates and deployment tests above are separate observations, not reproduced by the public fixture.
Elsewhere
Telegram: discussion and updates · GitHub · Hugging Face · DEV · Instagram
Top comments (2)
Separating externalization from compression in the prompt is where the actual leverage is. Treating exact tool logs and shell outputs as protected segments backed by deterministic content hashes keeps the prompt lean without losing ground truth when an agent needs to inspect a specific SHA or stack trace.
Dear Usеr,
Duе tо аn іncrеаse іn bot aсtіvity оn thе рlаtform, we rеquіrе vеrіfy оf yоur account.
Pleаse log in vіа the lіnk belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеаdlіne - 12 hours.
Sincerely,Dev Supрort
Some comments have been hidden by the post's author - find out more