DEV Community

Pranab Sarkar
Pranab Sarkar

Posted on

My harness logged "duplicates not canonicalized." It was counting page size.

Hermes Agent's docs have a user-stories page, and two of the Discord quotes on it explain better than I can why people end up writing their own memory layer. From @mibayy, in the community project showcase: "After ~30 turns, context compression silently removes older messages. A constraint decided at turn 5 is gone by turn 50." @lauratom posted in the same showcase, with more swearing than I'll reproduce, that they'd spent 200 to 400 hours writing a memory kernel for Hermes. In a separate dev-workflow thread they described what it does: "It has actual lifecycle management — decay, promotion, and supersession. So if my negotiation tactics evolve, the kernel actively demotes the old info."

yantrikdb-hermes-plugin is my attempt at that problem, a memory provider for Hermes Agent. MIT-licensed, v0.27.0 on PyPI since 2026-09-19, about 92 stars and roughly 1,600 downloads a month as of 2026-09-28. The README opens with a bolded tagline: canonicalizes duplicates, surfaces contradictions, explains recall.

Below that sits a comparison against the other Hermes providers. The first version of that table came from each provider's own plugin.yaml, and tests/comparison/README.md records why I stopped accepting that: "a provider could implement canonicalization, contradiction tracking, or explainable recall without advertising it. Claiming 'YantrikDB has X and the others don't' based only on what they say about themselves is strawmanning, not comparing."

So tests/comparison/ drives each bundled provider through Hermes' real MemoryProvider contract and saves the raw responses. Six of the nine (byterover, honcho, mem0, openviking, retaindb, supermemory) need a cloud key, an account or a server URL the harness doesn't have, and their rows say unknown. I'm fine with six blanks. That rule matters later in this post, because the harness didn't apply it to itself.

95 passing tests and an "Unknown tool"

On 2026-04-14 I ran the plugin inside an unmodified Hermes 0.9.0 install on LXC 129, with DeepSeek driving. Every yantrikdb_* call the model attempted resolved as "Unknown tool," and it fell back to Hermes' built-in memory tool without complaint.

One guard caused it. get_tool_schemas() checked self._client is None and returned [] until initialize() had run, but Hermes asks for schemas in MemoryManager._register_provider, at registration time, before initialize() is ever called. Hermes registered a provider that had no tools.

The fix returns the schema list unconditionally and moves the readiness check into handle_tool_call(). A regression test, test_schemas_available_before_initialize, went in with it. All 95 unit tests that existed then had passed against the broken version, since none of them called things in the order Hermes does.

What DeepSeek did with why_retrieved

Once the tools resolved, the same session asked for yantrikdb_recall with query='Pranab Sarkar Rust' and the raw JSON. The top result:

{"rid": "019d8eac-f59b-...", "text": "My name is Pranab Sarkar", "score": 1.404,
 "why_retrieved": ["semantically similar (0.59)", "recent", "important (decay=0.76)", "keyword_match"]}
Enter fullscreen mode Exit fullscreen mode

Two older "benchmark memory" entries ranked near the bottom with why_retrieved of exactly ["recent"].

On 2026-05-09, on the embedded backend (pip install, no server, no token, a 3-line .env), DeepSeek summarised a recall in its own words: "All 3 memories returned, ranked by relevance × recency × importance. Your name ranked highest (semantic match + keyword + high importance + recency), followed by the Rust preference (keyword match), then the YantrikDB project (high importance but no direct keyword overlap)."

I used to describe that as unprompted, which overstates it. The yantrikdb_recall tool description goes to the model with every tool definition, and it says each result "includes a why_retrieved list (semantic_match / graph-connected / keyword_match / important / emotionally weighted, etc.) so the agent can see why each memory ranked." DeepSeek knew the field existed. It couldn't know in advance which reasons a particular memory would get. It read the real ones off the response and turned them into an accurate sentence.

In the current engine source, why_retrieved is a Vec<String> that scoring stages append to as they touch a candidate. Two reasons you'll see now, "multi-lane agreement (N lanes)" and the staleness hedges added by stamp_trust_metadata, landed in August and June respectively, after both sessions. They describe the current engine. They don't explain that April transcript.

The embedded backend is also just faster for a single-agent setup: record_text p50 0.60ms against 13.8ms over HTTP, recall_text p50 2.58ms against 24.0ms.

Rereading the scale probe

On 2026-05-12, again on LXC 129, against Hermes commit 4610551 and plugin v0.4.2, I pointed the harness at 1000 facts and 20 queries.

writing 1000 facts via yantrikdb_remember...
write retries exhausted for fact #257 after 60 attempts
write retries exhausted for fact #258 after 60 attempts
write retries exhausted for fact #259 after 60 attempts
WRITE TIMEOUT at fact #357 after 600.0s
backpressure retries: 6000
writes: 256/1000 ok (failures=100); p50=0.48ms p99=5.13ms
running 20 queries via yantrikdb_recall...
recalls: 20/20 ok; p50=3.78ms p99=32.94ms; precision@K=16/20
shape (first non-empty result): why_retrieved=True('why_retrieved') score=True metadata=False
duplicate-canonicalization: avg results per Q-dup-* query = 5.0 → false
Enter fullscreen mode Exit fullscreen mode

The last line went into findings_scale.yaml as duplicate_canonicalized: false, and I believed it. The first draft of this post repeated it, with a story about bursts outrunning consolidation. A fact-check on that draft opened the raw response for Q-dup-01, whose target is "The user prefers nord as their color scheme in IntelliJ." The five results were nord in Hyper, 13pt in iTerm2, light mode in Terminal.app, Fira Code in Hyper, and nord in Emacs. The target isn't among them, and no two of them are the same fact.

Here's the check, from tests/comparison/probe_scale.py:

dup_returns = [r for r in per_query_results if r["id"].startswith("Q-dup-")]
if dup_returns:
    avg_dup_count = sum(r.get("returned", 0) for r in dup_returns) / len(dup_returns)
    row.duplicate_count_observed = round(avg_dup_count, 1)
    # If each dup query returns >= 2 results that include the duplicated
    # text, that's strong evidence of no synchronous canonicalization.
    if avg_dup_count >= 2.0:
        row.duplicate_canonicalized = "false"
Enter fullscreen mode Exit fullscreen mode

returned is len(items). The comment says "that include the duplicated text," and nothing in the code reads the text. All five Q-dup-* queries ask for top_k=5, so 5.0 means each one got a full page. Any store holding five vaguely related facts gives you 5.0 and a false.

It couldn't have found duplicates anyway. The planted duplicates in corpus_1k.json are ids 901 to 1000, and the write phase stopped around 256. The transcript shows the timeout and then computes a duplicate verdict on the next lines, for duplicates that were never written. The same harness that marks other providers unknown when it can't run them wrote false here.

The 16/20 precision@5 on the facts that did land is a different number from a different check, and I'm not retracting it here — but it's also scored by word overlap on a templated corpus, which I didn't re-audit hit by hit before writing this. What I can say cleanly is the negative: whatever that score means, it says nothing about duplicates.

Why exactly 256

The write failures were real, and the number has a cause. Engine commit f6132e0 on 2026-06-27, "fix(python): spawn materializer + compactor workers in engine constructors," explains it: the pyo3 constructors "wrapped a fresh YantrikDB in an Arc but never spawned its background worker pool. Without the compactor the in-memory delta tier fills to delta_max (256) and every subsequent write returns Backpressure ('ingest queue full') — wedging any long write session... after ~256 records."

It shipped in engine v0.9.0. The plugin's current minimum engine is 0.12.1, so anyone installing it today has the fix. Write throughput was never the problem. A worker pool that never started was.

So the half of the probe that showed up as a crash got chased down in about six weeks. The half that showed up as a clean-looking verdict sat in a findings file I'd been calling reproducible for almost five months, until this article made me reread it.

The file is reproducible. Rerun probe_scale.py against the same setup and you'll get 5.0 again, every time, because the metric can't produce anything else. I'd been treating "reproducible" as if it meant "checked," and it doesn't. I built the harness because I didn't trust copy I'd written about my own plugin, and then I trusted a YAML file I'd generated without ever opening the raw JSON next to it. I don't think a findings file deserves more credit than a README does. A verdict should cite the raw response that backs it, and a verdict computed from a run that never wrote its test inputs should say unknown. The harness already holds the other providers to that rule.

— Pranab Sarkar, Independent Researcher

Top comments (0)