Most voice agents that "forget" things or reference stale info aren't failing at the LLM layer. They're failing at retrieval. The usual pattern is: embed the query, grab top-50 chunks by cosine similarity, concatenate, ship it to the model. That works in a demo. It falls apart when you have 500K+ context documents across thousands of leads and dozens of campaigns, and the agent has a sub-200ms telephony budget to decide what it knows before it says "Hello".
At Alchemyst AI we model this as a pipeline of set operations. We call it context arithmetic:
Final Context = (SemanticMatches ∩ ScopeMatches ∩ MetadataMatches) − Superseded
→ rank(score) → top_k
Each stage is a monotonic narrowing of the candidate set. Here's how each one works, what it costs, and the bugs it prevents.
The funnel at a glance
| Stage | Op | Candidates | Failure it prevents |
|---|---|---|---|
| 1. Semantic search | ANN over embeddings | 500,000 → ~2,000 | Irrelevant topics |
| 2. Scope filter | Hierarchical prefix match | ~2,000 → ~200 | Cross-tenant / cross-campaign leakage |
| 3. Metadata filter | Exact / range predicates | ~200 → ~50 | Wrong lead, wrong language, stale window |
| 4. Semantic dedup | Supersession detection | ~50 → ~15 | Agent quoting outdated facts |
| 5. Rank + top-K | Weighted composite score | ~15 → 5 (~400 tokens) | Context window dilution |
Implementation note: the code below is a reference sketch in Python to illustrate the logic, not our production code or SDK surface. For the actual API, see the Alchemyst docs.
Stage 1: Semantic recall
The query isn't the user's utterance (there isn't one yet for outbound calls). It's composed from the campaign objective, the lead profile, and per-call instructions.
def build_query(campaign, lead, instructions):
return " | ".join([
f"objective: {campaign.objective}",
f"lead: {lead.persona} in {lead.region}, interested in {lead.product}",
f"instructions: {instructions}",
])
q_vec = embed(build_query(campaign, lead, instructions))
candidates = vector_index.search(q_vec, k=2000) # recall-oriented, deliberately wide
Keep k wide here. This stage is about recall; precision comes later. Optimizing for precision at the ANN step is the classic mistake: you throw away the right document before the filters that would have surfaced it get a chance to run.
Stage 2: Hierarchical scoping with groupName
Every document carries a groupName path like jkshah/ca-foundation/gujarat. Visibility is ancestor-based: a query scoped at jkshah/ca-foundation sees that doc, but jkshah/cs-executive doesn't.
def in_scope(doc_group: str, query_scope: str) -> bool:
d, q = doc_group.split("/"), query_scope.split("/")
return d[:len(q)] == q
scoped = [d for d in candidates if in_scope(d.group_name, "jkshah/ca-foundation")]
This gives you shared context flowing down the tree (org-wide FAQs, campaign-wide fee structures) without sibling campaigns bleeding into each other. In production you push this into the index as a pre-filter (prefix on a keyword field) rather than post-filtering in Python. Post-filtering 2,000 rows is fine; post-filtering with a tight k silently returns empty sets.
Stage 3: Metadata predicates
Structured, deterministic constraints. No embeddings involved:
filters = {
"lead_id": "GJ-4521",
"language": {"$in": ["gu", "hi", "en"]},
"created_at": {"$gt": "2026-01-01"},
"interaction_type": {"$in": ["call_summary", "objection", "email"]},
}
filtered = apply_filters(scoped, filters) # ~200 -> ~50
Rule of thumb: if a constraint can be expressed as an exact or range match, don't make the embedding model infer it. Vectors are bad at "only lead GJ-4521" and great at "about fee objections".
Stage 4: Semantic deduplication (the stage everyone skips)
Context mutates. In December the lead says "I want the January batch." In February: "Actually, March." Both are in the store, and both are semantically near-identical to a query about batch preference. Top-K by similarity alone happily returns both, and the LLM now has a coin flip.
Timestamp dedup doesn't work, because the two docs aren't duplicates by ID or text. You need to detect same topic, newer assertion:
def drop_superseded(docs, sim_threshold=0.85):
docs = sorted(docs, key=lambda d: d.created_at, reverse=True) # newest first
kept = []
for d in docs:
if any(cosine(d.vec, k.vec) > sim_threshold and same_slot(d, k) for k in kept):
continue # an equal-or-newer doc on this topic already survived
kept.append(d)
return kept
same_slot is where the real work is: checking that two docs talk about the same attribute (batch preference, fee quote, enrollment status) and that the newer one updates it, not just mentions it. You can do this with extracted slot keys at ingest time (cheap, deterministic) or an LLM contradiction check (expensive, keep it off the hot path). ~50 → ~15.
Stage 5: Composite ranking
Similarity alone over-weights verbose, generic notes. We score on three signals:
| Signal | Weight | Why |
|---|---|---|
| Semantic relevance | 0.40 | Match to call objective |
| Recency | 0.35 | Newer = more likely current |
| Information density | 0.25 | Specific, actionable facts beat general notes |
import math, time
def score(doc, q_vec, half_life_days=30):
relevance = cosine(doc.vec, q_vec)
age_days = (time.time() - doc.created_at_ts) / 86400
recency = math.exp(-math.log(2) * age_days / half_life_days)
density = doc.entity_count / max(doc.token_count, 1) # normalize per corpus
return 0.40 * relevance + 0.35 * recency + 0.25 * normalize(density)
final = sorted(deduped, key=lambda d: score(d, q_vec), reverse=True)[:5]
Exponential decay with a half-life is easier to tune than linear recency, and it degrades gracefully for leads you haven't contacted in months.
Worked trace: one outbound call
Lead GJ-4521, a parent in Gujarat, third call across two campaigns, interested in CA Foundation for her son.
| Stage | Docs | What happened |
|---|---|---|
| Semantic | 500,000 → 2,000 | Everything about CA Foundation, Gujarat, parent inquiries |
| Scope | 2,000 → 200 |
jkshah/ca-foundation, including the retarget campaign |
| Metadata | 200 → 50 | This lead, preferred languages, last 90 days |
| Dedup | 50 → 12 | Dropped Jan batch preference (superseded by Mar), old fee quote |
| Rank | 12 → 5 | Last call summary, current objection, batch pref, language pref, enrollment status |
What the agent actually receives (~400 tokens):
Lead: Priya Mehta (GJ-4521) | Parent | Gujarat | Preferred language: Gujarati
Prior interactions: 2 calls across "CA Foundation Gujarat" and "CA Foundation Retarget Q1"
...
Recommended approach: Open in Gujarati. Reference fee breakdown email.
Address fee objection with installment option. Confirm March batch interest.
The naive alternative, dumping all 50 post-filter docs, is ~4,000 tokens with the 400 that matter buried in the middle (where long-context models are measurably worst at attending). Same model, 10x the noise.
Same pipeline, different workload: NPS calls
The stages are use-case agnostic; only the parameters change. For an Unacademy NPS feedback campaign:
| Stage | Docs | Notes |
|---|---|---|
| Semantic | 80,000 → 3,500 | Course, engagement, support history |
| Scope | 3,500 → 400 | unacademy/nps-feedback |
| Metadata | 400 → 80 |
learner_id, course_id, 60 days, interaction_type IN (support_ticket, course_milestone)
|
| Dedup | 80 → 20 | Resolved tickets, stale progress reports |
| Rank | 20 → 6 | Progress, open tickets, milestones, last NPS, negative flags |
Result: ~450 tokens, and the agent knows to probe on session timings and doubt-clearing if the score comes back low, instead of reading a survey script.
What you need to build this yourself
-
Indexed interaction storage, not logs. Every call transcript goes through extraction → embedding → storage with metadata and
groupNameat write time. If you're enriching at read time, you've already blown the latency budget. - Hierarchical scoping in the index. Prefix-filterable keyword field, configurable per deployment.
- Slot extraction at ingest. Makes Stage 4 a cheap lookup instead of an LLM call.
- A hard latency budget. All five stages together need to land under ~200ms. That means pre-filters pushed into the vector store, no LLM calls on the hot path, and caching the lead-level metadata.
Takeaways
- Treat retrieval as composable set operations, not one top-K call.
- Wide recall first, then deterministic filters, then precision.
- Supersession is a correctness bug, not an optimization. Handle it explicitly.
- Rank on more than cosine similarity.
- Measure tokens delivered, not documents retrieved.
This is a developer-focused cut of Context Arithmetic for Voice: A Technical Primer on the Alchemyst blog, which goes deeper on how Kathan applies this at 500K+ calls/day.
Top comments (1)
Bài viết rất chi tiết về pipeline 5 stage — phần hierarchical scoping + metadata filters trước khi dedup là điểm thú vị nhất. Mình hay gặp trường hợp metadata filter loại bỏ sớm quá nhiều candidate khiến semantic search sau đó thiếu context để rank chính xác. Bạn có thử điều chỉnh thứ tự: để semantic search chạy rộng trước, rồi apply metadata filter + dedup ở stage sau không? Còn composite ranking — weight giữa semantic similarity vs metadata match vs recency bạn set như nào? Thử grid search hay dùng learning-to-rank? (site: labagent .tech)