DEV Community

Cover image for Why does my vector DB retrieval keep returning duplicate noise for my chat agent, and how to fix it?
Edward Izgorodin
Edward Izgorodin

Posted on Originally published at mnemoverse.com

Why does my vector DB retrieval keep returning duplicate noise for my chat agent, and how to fix it?

A user tells your chat agent they prefer dark mode, and it saves that. A week later they say they like the dark theme, and it saves that too. Now every question about settings pulls back "The user prefers dark mode.", "User likes the dark theme." and three chunks of the same onboarding guide: five slots holding one fact and one document, while whatever else was relevant never reached the prompt.

The short answer: duplicates come back for three different reasons, each with its own setting. A near-copy of a good match is close to the query too, so turn on maximal marginal relevance (MMR), which picks results that are relevant to the query and different from those already picked. One document's overlapping chunks fill the list, so group or collapse results by a document field, in the store. One fact stored twice, by a re-ingest or by the agent writing back what it read, is stopped on the write: stable IDs, and capture that stores each turn once; an old and a new version of one fact need lifecycle rules. A higher similarity threshold fixes none of the three.

A threshold cuts weak matches, not copies of each other

A similarity threshold is a floor against the query: it keeps or drops each result on its own score for the question, and it never compares one result with another. Qdrant says of its score_threshold: "It will exclude all results with a score worse than the given." LangChain's similarity_score_threshold retriever takes a "Minimum relevance threshold", the threshold on Mem0's search API is a "Minimum semantic relevance score", and the threshold on Supermemory's search is a "Similarity cutoff".

Two copies of one fact are close to the question for the same reason they are close to each other, so they score almost the same and tend to pass or fail together. In the library page's toy script, on hand-set vectors, "The user prefers dark mode." and "User likes the dark theme." score 0.89 and 0.90 in cosine similarity against one query: a floor placed between them is tuning to one pair, not a rule.

Thresholds remain the right tool for weak or irrelevant matches. Removing a copy asks a different question of each result: is it too close to something already picked? A threshold cuts weak matches, not copies of each other.

Three symptoms, three settings

Qdrant's beginner course pairs the first two the same way; the third row is mine.

What you see Where it comes from Setting
The top results are near-identical A near-copy of a good match is almost as close to the query as the match Maximal marginal relevance, with more candidates than you keep
One document fills the list Its overlapping chunks all match the same query Group or collapse by a document field, in the store
One fact comes back twice, or an old and a new version both come back A re-ingest with fresh IDs, or an agent writing back what it read Stable IDs on re-ingest; store each turn once; lifecycle rules for versions

They can look identical in a prompt, but their causes differ, and so do their settings.

Near-copies: MMR needs spare candidates

Carbonell and Goldstein defined maximal marginal relevance in 1998: "a document has high marginal relevance if it is both relevant to the query and contains minimal similarity to previously selected documents" (SIGIR 1998, authors' copy). Each candidate is scored against the question and against the picks already made, which is the comparison a threshold never makes.

So MMR is a selection from a pool, and it needs more candidates than it returns. Qdrant's course names the trap in its own MMR, where the candidate limit defaults to the query's limit, "which leaves MMR nothing spare to choose from, so all it can do is reorder the results it was already given. This is the most common reason MMR looks like it did nothing." Behind LangChain it is as_retriever(search_type="mmr", search_kwargs={"k": 4, "fetch_k": 20, "lambda_mult": 0.5}), with the defaults from LangChain's reference: k results picked from fetch_k candidates.

The trade-off knob runs in different directions, so check before you tune it:

Setting Toward relevance Toward diversity
Qdrant diversity 0.0 1.0
Weaviate balance 1.0 0.0
LangChain lambda_mult 1 0
LlamaIndex mmr_threshold close to 1 close to 0
Zep mmr_lambda, with reranker="mmr" 1.0 0.0
Mnemoverse diversity, on a REST read 0 1

Mnemoverse is the author's company; its diversity applies when top_k is below 200 and picks from up to 200 ranked results that clear the relevance floor.

In langchain-qdrant, LangChain's lambda_mult reaches Qdrant as its diversity, so the direction flips, as an open LangChain bug report records. At 0.5 the two readings coincide; the trap opens when you move it.

One document's chunks: group in the store

Grouping collapses the results that share a document field, so one source takes one slot. Each store names it differently: the Grouping API with group_by in Qdrant, GroupBy in Weaviate, group_by_field in Milvus, GroupBy in the Search API of Chroma Cloud (Chroma Cloud only for now, its overview says), and the collapse parameter in Elasticsearch and OpenSearch.

Do it in the store, not after the fetch. Qdrant's course says why: "Deduplicating the results yourself after the search does not fill the page." Ask for ten, receive ten chunks of one document, drop nine yourself, and the model sees one result where you budgeted ten. Upstream, the troubleshooting list in Chroma's chunking guide answers duplicate results with one instruction: "decrease chunk overlap".

Copies on the write: stable IDs, and each memory once

Qdrant's clean-up post names the cause: "Duplicates accumulate when each ingest assigns fresh IDs to content the collection already holds." With stable IDs, written with an upsert rather than a plain insert, a re-ingest does not add a second copy: Qdrant says "points with the same id will be overwritten when re-uploaded", Pinecone says "If a record ID already exists, upserting overwrites the entire record.", and Weaviate says "Use deterministic IDs to avoid inserting duplicate objects". On a memory layer, Zep's ingestion guide warns: "Reingesting the same file creates new episodes, even when its contents are unchanged." It asks for an import manifest, with source IDs and content hashes, before the first write. On Mnemoverse a write that is too similar to the nearest memory in the same domain is not stored, and a write can carry external_ref, a client reference that makes it idempotent: reuse one on POST /memory/write and the write returns 409 (API reference).

A chat agent adds two sources of its own: writing back what it just read, and injecting the same memory every turn. The first is fixed in the wiring: capture each turn once, and do not store the memories you injected. Plugins answer the second on the read. Mem0's DeepSeek Harness plugin page says recall "adds unseen results to the model context", and Supermemory's Cursor plugin page says recall "deduplicates results" and then injects them. For its Context Block, Zep's placement guidance says to replace the previous turn's block "instead of appending a second one", for its own reason: the order preserves the cacheable prefix that prompt caching needs. The Mnemoverse agent memory rules ask the same of the client, because its hosted connector is stateless and does not know what the conversation already holds.

Stable IDs catch the same content loaded twice. They cannot tell which of two contradicting facts is current, and neither can a similarity score: that is lifecycle, the third row of the table.

Check it with a control

  1. Store one fact twice in different words, or index one long document, and pick a query that hits it. Run your current search and count how many of the top results are the same fact or the same document.
  2. Turn on the setting for that symptom: MMR with a candidate pool larger than the result count, or grouping by the document field. Copies should give way to distinct, still relevant results.
  3. Control: run MMR again with the candidate pool equal to the result count, in LangChain fetch_k equal to k. The set should match the plain search, only reordered. If it does not, something other than MMR changed between the runs.

The library page has every store's setting with its source, the default window bounds, and a short step that compares results with each other when your store has no diversity setting. It also carries the film of the same map. Which of the three was filling your agent's context?

Top comments (1)

Collapse
 
agentsearchhq profile image
AgentSearch •

"A threshold cuts weak matches, not copies of each other" is the clearest way I've seen this put.

One more source of your third row when the corpus comes from the web: the same page reachable under several URLs (tracking parameters, redirects, mobile or AMP versions, a trailing slash). Each one gets a fresh ID at ingest, so stable IDs only help if the ID comes from the canonical or final URL after redirects, plus a hash of the cleaned text to catch true mirrors.

Cleaning before hashing matters too. If nav and footers stay in, two copies of the same article can differ only by a date in the footer and never match.