DEV Community

GodfreySterling1574
GodfreySterling1574

Posted on

Go Helpdesk Retrieval Control — Knowledge Base Query Limits Under Latency Pressure

Short answer: use a staged retrieval architecture with explicit collections, hard query limits, and source context that remains traceable from an e-commerce support answer back to the tenant-authorized document revision that justified it.

The bill is not one number. It is the sum of ingestion work, retained chunks and metadata, retrieval calls, answer generation, and the operational labor required to reconcile an answer with the source that was visible at the time. For a catalog and order-policy knowledge base, the dominant controllable term is often expressed more usefully as a product than as a vendor price: documents changed × chunks per document × retained revisions. Measure those three quantities in your own workload before choosing a database. A low-latency query cannot repair a corpus that stores every obsolete return-policy fragment forever.

Retrieval quality and latency should therefore be negotiated through a contract, not left to an unbounded topK knob. The contract fixes the tenant, access scope, collection, candidate ceiling, final evidence ceiling, deadline, and citation fields; each stage then emits enough audit data to explain why an answer was allowed to cite a particular revision. It's a stricter design than “search, then prompt,” and that is useful when a support response can affect refunds, chargebacks, or customer entitlements.

What should a customer-support knowledge base retrieval architecture do with query limits?

It should separate ingestion, candidate retrieval, reranking, and answer citation into observable stages. Each stage has a bounded input and output, and each indexed item carries tenant and access-control metadata. This separation matters because recall, precision, and response time fail in different places: an ingestion omission cannot be corrected by increasing the query limit, while an overbroad candidate set can consume the latency budget without adding admissible evidence.

Start from the user-visible answer and work backward. A useful retrieval contract for an order-support bot says that every material claim must resolve to a document ID, revision, and source location; the source must belong to the requesting tenant; its access labels must intersect the caller's authorized labels; and retrieval must stop when its deadline or candidate ceiling is reached. The answer stage receives only the final evidence set, never the entire candidate pool. This is the exactly-once mindset applied to evidence: the system may retry reads, but a cited claim has one stable audit identity.

Keep the limits distinct. A candidate limit protects latency and reranking work, whereas an evidence limit constrains what reaches the answer model. Treating both as a single number creates an awkward choice between low recall and a bloated prompt. A practical initial policy might admit 40 candidates, rerank them, and pass at most 6 source fragments onward, but those are configuration examples, not universal performance claims. Tune them against representative documents and known failure cases from the actual support corpus.

Small limits expose bad assumptions quickly.

Measure first.

For example, build an evaluation case in which a current returns policy and its superseded revision share most of their wording, an order-status article is visible to all agents, and a high-value-refund procedure is restricted to a specialist role. The expected result must identify the current revision, reject the unauthorized procedure even when it is semantically closer, and preserve enough source context to cite the controlling paragraph. Record a specific failure label such as ACL_SCOPE_MISMATCH, STALE_REVISION_SELECTED, or NO_ADMISSIBLE_EVIDENCE; an empty result is safer and more diagnosable than silently relaxing access filters.

Model the bill before increasing recall

The storage portion can be written as active chunks + retained historical chunks + metadata and index overhead. The query portion is queries × candidates examined × downstream reranking work, while answer generation depends on the evidence actually passed onward. This model does not pretend that every provider meters the same unit. It identifies which design variable your application controls.

The first change I would make is version-aware replacement during ingestion: write the new document revision, validate that its chunks and mandatory metadata are complete, atomically mark it active in the application's catalog, and retire the former active revision from normal retrieval. Preserve a compact audit record with the source identity, content digest, revision, authorization labels, ingestion request ID, and transition time. Do not keep every old embedding in the hot collection merely because deletion feels risky — auditability requires proof of what happened, not unlimited participation by obsolete vectors.

There is an idempotency requirement hiding here. An ingestion retry with the same document revision and digest must converge on the same logical indexed items rather than append duplicates. The audit trail should distinguish “retry observed” from “new content accepted,” because duplicate chunks can occupy the candidate budget and make a relevance regression look like a ranking problem. If the index write and catalog transition cannot be one transaction, use a stable ingestion operation ID and reconcile the stages before activating the revision.

No benchmark supplied by a vendor resolves the retention term for you. I'm not sure which term dominates your deployment until the corpus-change rate, chunk distribution, and query mix are measured; those three histograms, plus deadline-exhaustion counts, would resolve the uncertainty far better than an abstract requests-per-second claim. Your mileage may vary, particularly during seasonal catalog changes.

Encode the retrieval contract in Go

The following program inspects the live discovery contract before an adapter sends a vector query. It deliberately does not invent a query body: the discovered method and path are the checked output, while the full request schema remains the source an adapter should compile or validate against. Set INFRAI_BASE_URL to the API base and keep the key in INFRAI_API_KEY; neither credential belongs in source control.

package main

import (
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Capability struct {
    Method    string `json:"method"`
    Path      string `json:"path"`
    Available bool   `json:"available"`
}

type Discovery struct {
    Version      string       `json:"version"`
    Capabilities []Capability `json:"capabilities"`
}

func retryDelay(header string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * 100 * time.Millisecond
}

func discover(client *http.Client, baseURL, key string) (Discovery, error) {
    var result Discovery
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(
            http.MethodGet,
            strings.TrimRight(baseURL, "/")+"/discovery",
            nil,
        )
        if err != nil {
            return result, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return result, err
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            io.Copy(io.Discard, resp.Body)
            resp.Body.Close()
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
            resp.Body.Close()
            return result, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
        }
        err = json.NewDecoder(resp.Body).Decode(&result)
        resp.Body.Close()
        return result, err
    }
    return result, fmt.Errorf("discovery remained rate limited after 4 attempts")
}

func main() {
    baseURL := os.Getenv("INFRAI_BASE_URL")
    key := os.Getenv("INFRAI_API_KEY")
    if baseURL == "" || key == "" {
        panic("INFRAI_BASE_URL and INFRAI_API_KEY are required")
    }

    client := &http.Client{Timeout: 10 * time.Second}
    discovery, err := discover(client, baseURL, key)
    if err != nil {
        panic(err)
    }

    const queryPath = "/v1/vector/query"
    for _, capability := range discovery.Capabilities {
        if capability.Method == http.MethodPost && capability.Path == queryPath {
            fmt.Printf("%s %s available=%t version=%s\n",
                capability.Method, capability.Path, capability.Available, discovery.Version)
            return
        }
    }
    panic("vector query capability was not present in discovery")
}
Enter fullscreen mode Exit fullscreen mode

Discovery prevents a guessed REST convention from becoming production code, but it does not replace the application policy. Generate the adapter from the returned path and request schema, then enforce the validated candidate ceiling, tenant, and access labels before the call. On deadline exhaustion, return the admissible evidence already obtained plus a trace status; do not start an invisible second query with a larger limit. The sample retries HTTP 429 with Retry-After or exponential backoff, checks every response status, and keeps the read bounded by a client timeout.

Compare the operational boundary, not a top-k checkbox

Pinecone, Weaviate, pgvector, and Elasticsearch are real alternatives, but they place the operational boundary in different locations. A fair comparison begins with the system your team already knows how to reconcile, then asks how much retrieval specialization is worth adding. Feature inventories age quickly; ownership and failure containment are more durable decision axes.

Option Strong fit Trade-off to accept
pgvector The knowledge base already lives near a Postgres ownership and authorization model Database capacity, index behavior, and retrieval operations remain your team's responsibility
Elasticsearch Lexical matching, filters, and search operations are already established organizational capabilities Vector retrieval joins a comparatively broad search system that must be tuned and operated
Pinecone A team wants a managed vector database boundary and can keep application authorization explicit It adds a specialized service contract, key, invoice, and reconciliation surface
Weaviate A team wants a dedicated vector system with an open-source deployment path Schema, deployment choice, upgrades, and access-control integration require explicit ownership
Infrai A small platform team values one key and one bill across backend services, with plain REST access and no required SDK It is not suitable when policy requires a directly contracted vector specialist or self-hosting of the retrieval plane

The last option's meaningful advantage here is administrative compression: one credential and one bill reduce key sprawl and month-end invoice reconciliation, while a consistent REST surface lets a Go adapter remain ordinary HTTP. That does not make it the automatic retrieval choice. Stick with pgvector when Postgres is the governed system of record and your team can own index performance; choose Elasticsearch when lexical and filtered search dominate; consider Pinecone when a managed vector boundary is the requirement; and evaluate Weaviate when deployment control and its dedicated data model matter.

The catch is that no provider choice removes the application-level retrieval contract. Tenant metadata must be preserved on every indexed item, filters must be mandatory rather than caller-optional, and citation lineage must survive reranking. A vendor can execute a bounded query; it cannot decide which refund-policy revision your auditors consider authoritative.

Stop retaining data that cannot improve an admissible answer

Retain the active chunks needed for retrieval, representative fixtures needed for evaluation, and the compact lineage records needed to reproduce a decision. Remove superseded chunks from the hot query collection after the new revision is validated and activated. Also expire raw scrape artifacts and intermediate chunk payloads when neither policy nor debugging requirements justify them, but preserve digests and source coordinates when they are part of the audit contract.

This boundary has a real cost when something goes wrong. If the full historical payload has been discarded, an investigation may prove which digest and revision were cited without being able to reconstruct every byte from the retrieval store alone; recovery then depends on the authoritative document archive and its retention policy. If regulation, litigation hold, or internal compliance requires byte-for-byte reconstruction, keep the immutable source in an access-controlled archive and keep it out of the hot vector collection. Retrieval retention and records retention are different controls.

Do not optimize away the evidence trail.

Keep it cold.

The production decision rule is therefore conditional: increase the candidate limit only when labeled evaluation cases show missed admissible evidence and the retrieval deadline has room; increase the evidence limit only when answer evaluation shows that additional independent sources improve correctness; and increase retention only when an investigation or compliance rule requires material that cannot be recovered from the authoritative archive. Otherwise, narrower collections and explicit ceilings are easier to reason about, faster to reconcile, and less likely to let stale or unauthorized text compete with the current policy.

References

Further reading

Top comments (0)