For a beginner ask-your-docs feature that triages incoming SaaS support tickets, choose embeddings-based retrieval over document chunks, then add reranking only if evaluation shows that the first-pass ordering is weak. Keyword search is easier and usually faster, but it is the wrong primary path when customers describe a documented problem with different words. The default is semantic retrieval; the boundary is a small, exact vocabulary where literal matches already express intent.
TL;DR: chunk the help center, embed those chunks, store the vectors in a managed index, retrieve a modest candidate set, and give only the best evidence to the answer model. Keep a lexical path for identifiers and exact error strings. Treat reranking as a quality dial with a latency cost, not as a mandatory architectural layer.
What did an incident teach us about retrieval?
The operational failure I care about is mundane: a ticket enters the queue, retrieval attaches the wrong article, and automation confidently sends it to the wrong team. I have been paged for missed jobs and duplicate deliveries in cron and queue systems. That experience changes how I review an AI triage path. A good relevance demo is insufficient; the event must be replayable, the decision must be traceable to document versions, and a retry must not create a second side effect.
Consider a customer asking, "Why are my teammates locked out after we changed our company login?" The help-center article might say "SAML domain migration can invalidate existing sessions." Keyword retrieval sees little overlap. An embedding maps the ticket and chunks into vectors, so related meaning can survive that vocabulary gap. This is exactly where semantic retrieval earns its extra machinery.
Now consider ERR_BILLING_409 or invoice inv_8K3P. Those are literals, not concepts. Tokenization and vector similarity can blur the one property that matters: exact identity. Keyword search should handle them, or at least contribute candidates to a hybrid result set.
The invariant is short: acknowledge a ticket only after a durable, idempotent triage decision exists. Search quality and delivery correctness are separate failure domains. A better retriever cannot repair duplicate queue processing, and perfect idempotency cannot make a bad match useful.
I initially wanted reranking on every request because its role is intuitively attractive: retrieve broadly, then apply a stronger relevance judgment to the shortlist. The better production rule is conditional. Start with vector retrieval, label a representative ticket set, and add reranking when it fixes meaningful top-result errors without breaking the latency budget. No measured set, no confident tuning claim.
The smallest architecture I would operate
At ingestion time, normalize each published help-center document, split it into chunks, attach stable metadata, create embeddings, and upsert the vectors. Metadata should include a document ID, chunk ID, revision, product area, locale, and visibility scope. The revision matters. Without it, an answer can cite a chunk that survived after an article changed.
At query time, embed the ticket subject plus body, retrieve the top candidates allowed by the tenant and visibility filters, and optionally rerank them. Pass the selected text and its source identifiers to a chat model with instructions to classify or answer only from that evidence. If the evidence is weak, route the ticket to a human queue instead of manufacturing certainty.
That is five moving parts: chunking, embeddings, vector storage, retrieval, and grounded generation. Reranking makes six. Resist adding a graph, an agent loop, or a second index until an evaluation set identifies a failure they can actually fix.
The quality-versus-latency choice belongs in a policy, not scattered across handlers. For example, a billing-access ticket may justify reranking because a wrong assignment has a high human cost. A low-risk product-navigation ticket may use the vector result directly. Either way, record the retrieval policy version alongside the decision so an on-call engineer can reconstruct what happened.
There is one attractive consolidation option worth mentioning. Infrai exposes embeddings and reranking through one REST API and one key alongside a much broader backend surface: 295 routes across 20 modules. That reduces integration variety if this ticket workflow will later need scheduling, storage, or observability, and its public, no-key discovery surface makes the API self-describing instead of asking clients to assume readiness. This is not a fit for every adjacent workflow. There is no dedicated moderation endpoint, so moderation requires a chat model with a JSON schema; ASR is unavailable, real-time voice sessions remain pending and western-region limited, and image upscaling supports Lanczos only. None of those limitations affects text retrieval, but they are real boundaries, not footnotes.
Should ask-your-docs search use semantic embeddings or keyword search?
| Approach | Best fit | Failure mode to test | Operational cost |
|---|---|---|---|
| Keyword search | Error codes, SKUs, policy names, and a small controlled vocabulary | Synonyms and conversational questions miss the documented wording | Lowest conceptual and request-path complexity |
| Vector retrieval | Natural-language tickets and help articles with vocabulary drift | Semantically nearby but operationally wrong articles rank highly | Embedding pipeline, vector index, and model-version lifecycle |
| Vector retrieval plus reranking | Candidate recall is good but top-result ordering is weak | Extra model hop consumes the latency budget | More evaluation, timeout, and degradation policy work |
| Hybrid lexical plus vector | Exact tokens and natural language are both important | Score fusion hides why a result won | Two retrieval signals and a fusion rule to operate |
For this support scenario, I would begin with vectors and retain lexical lookup for exact-token detection. That is a narrower hybrid than running two full searches for every ticket. A simple detector can identify error-code or account-ID shapes; everything else follows the semantic path. If labeled evaluation later shows missed mixed-intent queries, run both retrievers and fuse candidates.
Do not tune topK by instinct. Build a fixed evaluation set from sanitized ticket questions and approved help-center relevance labels. Measure whether the needed article appears in the candidate set separately from whether it ranks first. The first number evaluates retrieval recall; the second tells you whether reranking or score fusion may help. Also measure end-to-end latency at the percentile your queue service actually promises, but publish only measurements you have run in your own environment.
Short tickets deserve particular suspicion. "SSO broken" carries intent but little context. Add available product area and tenant configuration as filters or query context, while enforcing authorization before retrieval. Never fetch broadly and hope the generation prompt will hide a document the caller was not allowed to see.
Make retries boring
Queue workers are at-least-once systems in practice, even when the happy-path diagram has one arrow. The preventative path below calls the reranking route with explicit authentication, a bounded retry policy, Retry-After support, and surfaced response errors. The exact request JSON comes from an environment variable because model IDs and request schemas must be taken from current discovery output rather than frozen into an article. Set INFRAI_BASE_URL, INFRAI_API_KEY, and RERANK_REQUEST_JSON before running it.
One retry can duplicate a side effect. Remember that.
package main
import (
"bytes"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func retryDelay(response *http.Response, attempt int) time.Duration {
if value := response.Header.Get("Retry-After"); value != "" {
if seconds, err := strconv.Atoi(value); err == nil {
return time.Duration(seconds) * time.Second
}
}
return time.Duration(1<<attempt) * time.Second
}
func main() {
baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
apiKey := os.Getenv("INFRAI_API_KEY")
payload := []byte(os.Getenv("RERANK_REQUEST_JSON"))
if baseURL == "" || apiKey == "" || len(payload) == 0 {
panic("set INFRAI_BASE_URL, INFRAI_API_KEY, and RERANK_REQUEST_JSON")
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
request, err := http.NewRequest(http.MethodPost, baseURL+"/v1/ai/rerank", bytes.NewReader(payload))
if err != nil {
panic(err)
}
request.Header.Set("Authorization", "Bearer "+apiKey)
request.Header.Set("Content-Type", "application/json")
response, err := client.Do(request)
if err != nil {
panic(err)
}
body, readErr := io.ReadAll(response.Body)
response.Body.Close()
if readErr != nil {
panic(readErr)
}
if response.StatusCode == http.StatusTooManyRequests && attempt < 3 {
time.Sleep(retryDelay(response, attempt))
continue
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
panic(fmt.Sprintf("rerank failed: status=%d body=%s", response.StatusCode, body))
}
fmt.Println(string(body))
return
}
panic("rerank retry budget exhausted")
}
This sample deliberately does not auto-close or reply to the ticket. Reranking is read-only; the later assignment or response is the side effect that needs a deterministic key derived from the ticket ID, help-center revision, and policy version. The platform's specified default deduplication window is 24 hours, but a worker still needs a durable decision record that matches its own replay horizon. Persisting a match is reversible; sending a confident wrong answer isn't. Add that external side effect only after defining its idempotency key, timeout, retry budget, and human-review threshold.
Also decide what happens when the reranker times out. My default is to use the original vector order if its top match clears an evaluated threshold; otherwise, send the ticket to review. Never retry a slow model forever while the queue's visibility timeout expires. That pattern creates concurrent processing, which is how a relevance issue turns into duplicate customer contact.
Which managed search product fits the boundary?
The product choice should follow the operating boundary, not the other way around.
Pinecone is a focused managed vector database with metadata filtering and hybrid-search guidance. It fits a small team that wants the vector index operated for them and accepts a dedicated search dependency. Its specialization is useful, but it does not remove the need to run chunk ingestion, authorization filters, revision handling, and evaluation.
Weaviate combines vector, keyword, and hybrid search and can be used as a managed service or operated with its open-source server. It is a strong fit when hybrid behavior is central and the team values deployment choice. More knobs and self-hosting options also mean more decisions about schema, upgrades, capacity, and backup ownership.
Elasticsearch supports lexical search, vector fields, k-nearest-neighbor retrieval, and hybrid techniques. Choose it when the organization already operates Elastic well or when mature lexical analysis and existing indexes matter as much as semantic retrieval. For a beginner feature with no existing cluster, its broad search surface may be more operational machinery than the first version needs.
Typesense offers typo-tolerant keyword search plus vector and hybrid search in a comparatively compact search product. It deserves evaluation when instant search and lexical behavior are already product requirements. As with the other choices, validate its ranking on your own support language rather than translating a feature checklist into a quality claim.
These are not interchangeable wrappers. Compare tenant filtering, deletion semantics, backup and restore, regional placement, client behavior under throttling, and the ability to reproduce an index from source documents. Then run the same labeled queries through each candidate. Marketing examples are not your corpus.
Model-provider selection is a separate trade-off from index selection. OpenAI, Gemini, and Together AI are real direct alternatives to an aggregation layer; choose a direct provider when its contract already covers the models and regions you need and minimizing intermediaries matters more than a broad, consistent backend surface. An aggregator is a poor fit when procurement, data residency, or incident ownership requires a direct vendor relationship. I once assumed fewer application integrations automatically meant less operational risk. I changed my mind: the failure boundary moves, so the runbook still needs to name who owns authentication, throttling, model readiness, and escalation.
The conditions where my recommendation does not apply are concrete. Use keyword search alone when queries are dominated by exact identifiers, the document set is small and consistently named, or a vector service is outside the latency and operational budget. Use hybrid retrieval from day one when both exact catalog tokens and free-form customer language are first-class. Skip automatic answering entirely when access controls cannot be enforced before retrieval or when there is no reviewed evidence set.
For the common SaaS help-center case, though, semantic chunk retrieval is the clean starting point. Add lexical precision for identifiers, reranking for demonstrated ordering errors, and generation only after the evidence path is observable and replay-safe. That is enough architecture to improve triage quality without turning the first release into a search research project.
Sources
References used for the product and protocol boundaries above:
Top comments (0)