DEV Community

Lohith J
Lohith J

Posted on

How I Designed a Two-Stage Recall to Stop Other Vendors Crowding Out Exact Matches

How I Designed a Two-Stage Recall to Stop Other Vendors Crowding Out Exact Matches
Here's the failure mode I was trying to fix: the agent is processing a price variance exception for Vendor A. It recalls 5 records from Hindsight. Four of them are about Vendor B, which has a richer history. The one record about Vendor A gets buried at position 5 — or knocked out entirely.
The agent sees mostly Vendor B precedents, concludes the case has strong precedent for approval, and recommends approval. The Vendor A record that said "always reject this vendor's price variances" never made it into the top 5.
This is a real problem when one vendor has 10 decisions in memory and another has 2. The semantic similarity of the query pulls in whoever has the most similar-sounding history, regardless of whether it's the right vendor.
The Fix: Two Recall Calls
The solution I landed on was running two separate recall passes against Hindsight for every exception, using different tag-matching strategies.
Pass 1 — all_strict: Tags must ALL match. For a price variance exception from Vendor V06, this means only records tagged vendor:V06 AND type:price_variance come back. Exact matches only.
Pass 2 — any_strict: Tags match if ANY matches. This pulls in records tagged with either the vendor OR the exception type — broader context, other vendors' price variance history, the same vendor's other exception types.
The combination: exact matches fill the top slots. Broader matches fill whatever's left up to the limit of 5. If exact matches already fill all 5 slots, pass 2 is skipped entirely.

# memory/hindsight_store.py
def recall(self, query: RecallQuery, k: int = 5) -> list[RecalledMemory]:
    """Exact matches first (vendor AND type), so other vendors can never crowd
    relevant precedents out of the top k."""
    exact = self._recall_once(query, "all_strict")
    if len(exact) >= k:
        return exact[:k]
    seen = {r.id for r in exact}
    broader = [r for r in self._recall_once(query, "any_strict") if r.id not in seen]
    return (exact + broader)[:k]
Enter fullscreen mode Exit fullscreen mode

Three lines of logic, but it required understanding how Hindsight's tag-matching modes work before I could write it.
How Hindsight Tags Work
Every record retained into Hindsight gets two tags when we store it:

# memory/hindsight_store.py — tags applied at retain time
def _item(self, record: MemoryRecord) -> dict:
    return {
        "content": record.content(),
        "document_id": record.id,
        "tags": record.tags(),  # ["vendor:V06", "type:price_variance"]
        "update_mode": "replace",
    }
Enter fullscreen mode Exit fullscreen mode

And at recall time, we pass those same tags back with the tags_match parameter:

def _recall_once(self, query: RecallQuery, tags_match: str) -> list[RecalledMemory]:
    resp = self._call(
        lambda: self.client.recall(
            self.bank_id,
            query.to_text(),
            tags=[f"vendor:{query.vendor_code}", f"type:{query.exception_type}"],
            tags_match=tags_match,  # "all_strict" or "any_strict"
        )
    )
Enter fullscreen mode Exit fullscreen mode

all_strict means the result must have ALL the provided tags. any_strict means it must have at least ONE. The difference sounds minor. In a memory bank with dozens of vendors and multiple exception types, it's the difference between getting the right vendor's records and getting swamped by whoever has the most entries.
What the Recommender Does With This
The recommender receives the combined list and filters it into two groups: relevant (same vendor AND same exception type) and not relevant (everything else from pass 2).

# agent/recommender.py
for r in recalled:
    r["relevant"] = r["vendor_code"] == vendor_code and r["exception_type"] == exc_type
relevant = [r for r in recalled if r["relevant"]]
if not relevant:
    return result("escalate", 0.0,
        f"No relevant precedent: {len(recalled)} memory(ies) were recalled but "
        f"none concern {vendor_name} with a {exc_type} exception.", [], recalled, "guard")
Enter fullscreen mode Exit fullscreen mode

Only relevant records are passed to the LLM for decision-making. The broader pass 2 records are shown in the UI for transparency ("evidence recalled") but the agent can't cite them and they don't influence the recommendation.
This means the agent is allowed to say "I recalled 5 records but none of them are for this vendor, so I'm escalating" — which is more honest and more useful than citing a record from a different vendor.
The Edge Case That Motivated This
The two-stage approach came directly from a test I wrote after finding a real failure:

# tests/test_memory.py
def test_hindsight_exact_matches_are_never_crowded_out():
    """A vendor with 1 decision must not be beaten by another vendor with 4 decisions."""
    store = LocalMemoryStore()
    # Retain 4 records for vendor B
    for i in range(4):
        store.retain(MemoryRecord(id=f"PREC-B-{i}", vendor_code="VB", exception_type="price_variance", ...))
    # Retain 1 record for vendor A
    store.retain(MemoryRecord(id="PREC-A-0", vendor_code="VA", exception_type="price_variance", ...))

    results = store.recall(RecallQuery(vendor_code="VA", exception_type="price_variance"), k=5)
    # The vendor A record must appear, regardless of how many vendor B records exist
    assert any(r.id == "PREC-A-0" for r in results)
Enter fullscreen mode Exit fullscreen mode

With a single-pass any_strict recall, this test fails when Vendor B has enough records to fill the top 5 slots by semantic similarity. With two-stage recall, it passes every time because pass 1 guarantees the Vendor A record appears first.
What I Would Do Differently
The two-pass approach doubles the number of Hindsight API calls. For a batch of 17 exceptions, that's potentially 34 recall calls instead of 17. In practice, pass 2 is skipped whenever pass 1 returns a full 5 results, which reduces the overhead. But it's still worth noting.
A cleaner solution might be a single recall call with a "prefer these tags, fall back to any" mode, if Hindsight adds that. For now, two calls with deduplication is the right tradeoff: correctness over API efficiency.
The Broader Point
Agent memory is only as good as the retrieval strategy. Storing decisions in Hindsight was the easy part. Getting the right decisions back when it matters — without being drowned out by a noisier part of the memory bank — required deliberate design.
If you're building something similar and using tag-based recall, test the crowding-out scenario early. Write the test I described above. If it fails, you have a recall problem that will only get worse as your memory bank grows.
The two-stage approach solved it cleanly. The exact matches always win their slots first. Everything else fills in around them.

Top comments (0)