DEV Community

Cover image for Context Arithmetic: A 5-Stage Retrieval Pipeline for Voice Agents (500K docs to 400 tokens in <200ms)
Debayan Pradhan
Debayan Pradhan

Posted on Originally published at getalchemystai.com

Context Arithmetic: A 5-Stage Retrieval Pipeline for Voice Agents (500K docs to 400 tokens in <200ms)

Most voice agents that "forget" things or reference stale info aren't failing at the LLM layer. They're failing at retrieval. The usual pattern is: embed the query, grab top-50 chunks by cosine similarity, concatenate, ship it to the model. That works in a demo. It falls apart when you have 500K+ context documents across thousands of leads and dozens of campaigns, and the agent has a sub-200ms telephony budget to decide what it knows before it says "Hello".

At Alchemyst AI we model this as a pipeline of set operations. We call it context arithmetic:

Final Context = (SemanticMatches ∩ ScopeMatches ∩ MetadataMatches) − Superseded
                → rank(score) → top_k
Enter fullscreen mode Exit fullscreen mode

Each stage is a monotonic narrowing of the candidate set. Here's how each one works, what it costs, and the bugs it prevents.

The funnel at a glance

Stage Op Candidates Failure it prevents
1. Semantic search ANN over embeddings 500,000 → ~2,000 Irrelevant topics
2. Scope filter Hierarchical prefix match ~2,000 → ~200 Cross-tenant / cross-campaign leakage
3. Metadata filter Exact / range predicates ~200 → ~50 Wrong lead, wrong language, stale window
4. Semantic dedup Supersession detection ~50 → ~15 Agent quoting outdated facts
5. Rank + top-K Weighted composite score ~15 → 5 (~400 tokens) Context window dilution

Implementation note: the code below is a reference sketch in Python to illustrate the logic, not our production code or SDK surface. For the actual API, see the Alchemyst docs.

Stage 1: Semantic recall

The query isn't the user's utterance (there isn't one yet for outbound calls). It's composed from the campaign objective, the lead profile, and per-call instructions.

def build_query(campaign, lead, instructions):
    return " | ".join([
        f"objective: {campaign.objective}",
        f"lead: {lead.persona} in {lead.region}, interested in {lead.product}",
        f"instructions: {instructions}",
    ])

q_vec = embed(build_query(campaign, lead, instructions))
candidates = vector_index.search(q_vec, k=2000)   # recall-oriented, deliberately wide
Enter fullscreen mode Exit fullscreen mode

Keep k wide here. This stage is about recall; precision comes later. Optimizing for precision at the ANN step is the classic mistake: you throw away the right document before the filters that would have surfaced it get a chance to run.

Stage 2: Hierarchical scoping with groupName

Every document carries a groupName path like jkshah/ca-foundation/gujarat. Visibility is ancestor-based: a query scoped at jkshah/ca-foundation sees that doc, but jkshah/cs-executive doesn't.

def in_scope(doc_group: str, query_scope: str) -> bool:
    d, q = doc_group.split("/"), query_scope.split("/")
    return d[:len(q)] == q

scoped = [d for d in candidates if in_scope(d.group_name, "jkshah/ca-foundation")]
Enter fullscreen mode Exit fullscreen mode

This gives you shared context flowing down the tree (org-wide FAQs, campaign-wide fee structures) without sibling campaigns bleeding into each other. In production you push this into the index as a pre-filter (prefix on a keyword field) rather than post-filtering in Python. Post-filtering 2,000 rows is fine; post-filtering with a tight k silently returns empty sets.

Stage 3: Metadata predicates

Structured, deterministic constraints. No embeddings involved:

filters = {
    "lead_id": "GJ-4521",
    "language": {"$in": ["gu", "hi", "en"]},
    "created_at": {"$gt": "2026-01-01"},
    "interaction_type": {"$in": ["call_summary", "objection", "email"]},
}
filtered = apply_filters(scoped, filters)   # ~200 -> ~50
Enter fullscreen mode Exit fullscreen mode

Rule of thumb: if a constraint can be expressed as an exact or range match, don't make the embedding model infer it. Vectors are bad at "only lead GJ-4521" and great at "about fee objections".

Stage 4: Semantic deduplication (the stage everyone skips)

Context mutates. In December the lead says "I want the January batch." In February: "Actually, March." Both are in the store, and both are semantically near-identical to a query about batch preference. Top-K by similarity alone happily returns both, and the LLM now has a coin flip.

Timestamp dedup doesn't work, because the two docs aren't duplicates by ID or text. You need to detect same topic, newer assertion:

def drop_superseded(docs, sim_threshold=0.85):
    docs = sorted(docs, key=lambda d: d.created_at, reverse=True)  # newest first
    kept = []
    for d in docs:
        if any(cosine(d.vec, k.vec) > sim_threshold and same_slot(d, k) for k in kept):
            continue   # an equal-or-newer doc on this topic already survived
        kept.append(d)
    return kept
Enter fullscreen mode Exit fullscreen mode

same_slot is where the real work is: checking that two docs talk about the same attribute (batch preference, fee quote, enrollment status) and that the newer one updates it, not just mentions it. You can do this with extracted slot keys at ingest time (cheap, deterministic) or an LLM contradiction check (expensive, keep it off the hot path). ~50 → ~15.

Stage 5: Composite ranking

Similarity alone over-weights verbose, generic notes. We score on three signals:

Signal Weight Why
Semantic relevance 0.40 Match to call objective
Recency 0.35 Newer = more likely current
Information density 0.25 Specific, actionable facts beat general notes
import math, time

def score(doc, q_vec, half_life_days=30):
    relevance = cosine(doc.vec, q_vec)
    age_days = (time.time() - doc.created_at_ts) / 86400
    recency = math.exp(-math.log(2) * age_days / half_life_days)
    density = doc.entity_count / max(doc.token_count, 1)   # normalize per corpus
    return 0.40 * relevance + 0.35 * recency + 0.25 * normalize(density)

final = sorted(deduped, key=lambda d: score(d, q_vec), reverse=True)[:5]
Enter fullscreen mode Exit fullscreen mode

Exponential decay with a half-life is easier to tune than linear recency, and it degrades gracefully for leads you haven't contacted in months.

Worked trace: one outbound call

Lead GJ-4521, a parent in Gujarat, third call across two campaigns, interested in CA Foundation for her son.

Stage Docs What happened
Semantic 500,000 → 2,000 Everything about CA Foundation, Gujarat, parent inquiries
Scope 2,000 → 200 jkshah/ca-foundation, including the retarget campaign
Metadata 200 → 50 This lead, preferred languages, last 90 days
Dedup 50 → 12 Dropped Jan batch preference (superseded by Mar), old fee quote
Rank 12 → 5 Last call summary, current objection, batch pref, language pref, enrollment status

What the agent actually receives (~400 tokens):

Lead: Priya Mehta (GJ-4521) | Parent | Gujarat | Preferred language: Gujarati
Prior interactions: 2 calls across "CA Foundation Gujarat" and "CA Foundation Retarget Q1"
...
Recommended approach: Open in Gujarati. Reference fee breakdown email.
Address fee objection with installment option. Confirm March batch interest.
Enter fullscreen mode Exit fullscreen mode

The naive alternative, dumping all 50 post-filter docs, is ~4,000 tokens with the 400 that matter buried in the middle (where long-context models are measurably worst at attending). Same model, 10x the noise.

Same pipeline, different workload: NPS calls

The stages are use-case agnostic; only the parameters change. For an Unacademy NPS feedback campaign:

Stage Docs Notes
Semantic 80,000 → 3,500 Course, engagement, support history
Scope 3,500 → 400 unacademy/nps-feedback
Metadata 400 → 80 learner_id, course_id, 60 days, interaction_type IN (support_ticket, course_milestone)
Dedup 80 → 20 Resolved tickets, stale progress reports
Rank 20 → 6 Progress, open tickets, milestones, last NPS, negative flags

Result: ~450 tokens, and the agent knows to probe on session timings and doubt-clearing if the score comes back low, instead of reading a survey script.

What you need to build this yourself

  1. Indexed interaction storage, not logs. Every call transcript goes through extraction → embedding → storage with metadata and groupName at write time. If you're enriching at read time, you've already blown the latency budget.
  2. Hierarchical scoping in the index. Prefix-filterable keyword field, configurable per deployment.
  3. Slot extraction at ingest. Makes Stage 4 a cheap lookup instead of an LLM call.
  4. A hard latency budget. All five stages together need to land under ~200ms. That means pre-filters pushed into the vector store, no LLM calls on the hot path, and caching the lead-level metadata.

Takeaways

  • Treat retrieval as composable set operations, not one top-K call.
  • Wide recall first, then deterministic filters, then precision.
  • Supersession is a correctness bug, not an optimization. Handle it explicitly.
  • Rank on more than cosine similarity.
  • Measure tokens delivered, not documents retrieved.

This is a developer-focused cut of Context Arithmetic for Voice: A Technical Primer on the Alchemyst blog, which goes deeper on how Kathan applies this at 500K+ calls/day.

Top comments (1)

Collapse
 
koev3kcjausd profile image
koev3kcjausd •

Bài viết rất chi tiết về pipeline 5 stage — phần hierarchical scoping + metadata filters trước khi dedup là điểm thú vị nhất. Mình hay gặp trường hợp metadata filter loại bỏ sớm quá nhiều candidate khiến semantic search sau đó thiếu context để rank chính xác. Bạn có thử điều chỉnh thứ tự: để semantic search chạy rộng trước, rồi apply metadata filter + dedup ở stage sau không? Còn composite ranking — weight giữa semantic similarity vs metadata match vs recency bạn set như nào? Thử grid search hay dùng learning-to-rank? (site: labagent .tech)