Choose reranking with a strict evidence gate when RAG hallucination makes an ask-your-docs chatbot return wrong answers; keep embedding-only retrieval for small, stable invoice collections where latency matters more than handling ambiguous matches.
TL;DR: Wrong answers from an ask-your-docs system usually start before generation. The relevant invoice rule was missed, split badly, crowded out, or silently truncated, and the model was then allowed to fill the gap from general knowledge. Retrieve candidates, rerank them, fit only the strongest evidence into a measured token budget, and require not found when the context does not support a field.
This is an operational choice, not a model beauty contest. For a media company extracting supplier name, invoice number, purchase-order number, currency, and total, a plausible invented value is worse than an explicit miss. The miss enters a review queue. The plausible value enters accounting.
Why Does RAG Hallucination Make Ask Your Docs Chatbot Answers Wrong?
An embedding proves that two pieces of text are close in a vector space. It does not prove that a retrieved passage contains the evidence needed for a particular field. An invoice mentioning "campaign total" may sit near a query for "invoice total," while the actual payable total appears in a footer chunk with little semantic context. Top-k retrieval can confidently return the wrong neighborhood.
Chunking creates a second failure boundary. Large chunks carry unrelated line items and boilerplate into the prompt. Tiny chunks detach a value from its label, supplier, page, or currency. There is no universal chunk size that repairs both cases. Preserve document structure and keep stable identifiers such as document ID, page, section, and chunk ID so an answer can point back to its evidence.
Then comes prompt construction. If the assembled context exceeds the model limit, an important passage may be dropped. If the instruction merely says "answer the question," the model still has permission to use general knowledge. Counting tokens before generation and reserving room for the answer makes truncation an explicit admission decision instead of a surprise.
The practical failure chain is short:
- Retrieval misses or dilutes the relevant invoice fragment.
- Context assembly admits weak chunks or exceeds its token budget.
- Generation is allowed to infer an unsupported value.
Fix it in that order. Swapping the chat model first leaves the bad evidence path intact.
Embedding-only retrieval versus reranked evidence
Pure vector retrieval has the shorter path: embed the query, fetch the nearest chunks, and generate. Use it when the corpus is narrow, terminology is consistent, and a missed field can safely become not found. It also remains a useful baseline because every extra stage adds latency and another dependency to observe.
Reranked retrieval fetches a wider candidate set, then applies a second relevance judgment before prompt assembly. That costs time, but it can remove semantically nearby passages that do not answer the field-level question. For supplier invoices with repeated totals, tax lines, credits, and purchase-order references, reranked evidence is the safer default. The quality gain comes from improving what the model sees, not from asking generation to reason around noisy context.
The products make different architectural bets:
| Product | Relevant approach | Operational boundary |
|---|---|---|
| Pinecone | Managed vector search with metadata filtering and reranking options | A focused fit when retrieval infrastructure should be managed; generation and evidence policy remain application concerns. |
| Weaviate | Vector, keyword, and hybrid search with reranking integrations | Useful when hybrid retrieval is central; schema and search tuning still need ownership. |
| Elasticsearch | BM25, vector, and hybrid retrieval in an established search platform | Strong when the team already operates Elastic and needs lexical controls; cluster and relevance work are not removed. |
| OpenAI | Hosted embeddings and generation APIs | Convenient model primitives, but document structure, retrieval policy, and citations stay in the application. |
| Anthropic Claude | Hosted generation with a focus on supplying source context to the model | A reasonable generator choice when Claude already fits the stack; retrieval and evidence admission remain separate work. |
| Google Gemini | Hosted embedding and generation capabilities in Google's AI platform | Fits teams standardized on Google Cloud, while grounding quality still depends on the application retrieval path. |
| OpenRouter | One API for accessing multiple model providers | Useful for model routing experiments; it does not replace document retrieval, reranking, or field-level evidence rules. |
| Infrai | A self-describing API surface with discovery schemas and runnable examples, plus token counting and reranking capabilities | Useful when one key and one bill reduce credential and vendor-account sprawl; readiness should be checked per capability, and the application still owns grounding. |
None of these products can define what counts as sufficient invoice evidence for you. That policy belongs beside the field schema and must survive a provider change.
Put the evidence gate before generation
The safe implementation separates candidate retrieval, evidence admission, and answer generation. Do not let the generator see every candidate. Admit only chunks that pass a relevance threshold, fit within the token budget, and retain a source locator. If no admitted chunk supports the requested field, return not found without calling generation.
Here is a small, runnable Go client for the token-count step. COUNT_REQUEST_JSON must contain a request validated against the public discovery schema; keeping that JSON outside the example avoids freezing an undeclared field shape into application code. The client uses the required bearer credential, an explicit method, bounded 429 retries, Retry-After, and real error bodies.
package main
import (
"bytes"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func retryDelay(header string, attempt int) time.Duration {
if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func countTokens(client *http.Client, key string, body []byte) ([]byte, error) {
baseURL := os.Getenv("INFRAI_BASE_URL")
if baseURL == "" {
return nil, errors.New("INFRAI_BASE_URL is required")
}
endpoint := baseURL + "/ai/tokens/count"
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodPost, endpoint, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
resp, err := client.Do(req)
if err != nil {
return nil, err
}
responseBody, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("token count failed: status=%d body=%s", resp.StatusCode, responseBody)
}
return responseBody, nil
}
return nil, errors.New("token count failed after rate-limit retries")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
body := []byte(os.Getenv("COUNT_REQUEST_JSON"))
if key == "" || len(body) == 0 {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and COUNT_REQUEST_JSON are required")
os.Exit(2)
}
result, err := countTokens(&http.Client{Timeout: 15 * time.Second}, key, body)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(result))
}
The generation instruction should be equally plain: extract only requested fields supported by the supplied chunks, attach the supporting chunk ID to each value, and emit not found when evidence is absent or conflicting. Do not ask the model to choose between two invoice totals without a rule that distinguishes payable total from campaign subtotal.
There are two deliberate trade-offs here. A wider first-stage fetch improves recall but raises reranking latency. A high admission threshold reduces unsupported answers but increases review volume. Start conservative. False negatives are visible and recoverable; confident false positives tend to travel downstream. My decision rule is to spend the extra retrieval time on payable amounts and purchase-order IDs, then accept the review queue as the cost of refusing weak evidence.
The same integration also reduces operational sprawl: one credential covers a discovery surface of 295 routes across 20 modules, and documented capabilities include runnable examples in 10 languages. In this workflow, that means token counting and reranking follow one set of API conventions instead of adding another SDK and credential lifecycle. It does not make the evidence policy automatic.
Verification before traffic moves
Build an evaluation set from representative document shapes, not only clean single-page invoices. Include multi-page invoices, credit notes, repeated currencies, scanned text, missing purchase orders, and a document where the requested value truly does not exist. Label the expected field, acceptable evidence chunk, and expected abstention.
Track retrieval separately from generation. For each field, record whether the labeled evidence appeared in the candidate set, survived reranking, fit the prompt budget, and supported the final value. A single end-to-end accuracy number cannot tell an on-call engineer which stage failed.
Use at least these release checks:
- Retrieval recall: did the candidate set contain the labeled source chunk?
- Rerank retention: did the correct chunk remain after the second-stage ranking?
- Grounding: does every emitted value cite an admitted chunk that contains it?
- Abstention: does missing or conflicting evidence produce
not found? - Budget: were any admitted chunks dropped during prompt construction?
- Idempotency: does replaying the same invoice produce one durable result rather than duplicate downstream actions?
No silent drops. Log document ID, field name, candidate chunk IDs, admitted chunk IDs, token count, model identifier, and request ID. Avoid logging raw invoice text unless the data policy explicitly permits it; invoices commonly contain sensitive commercial details.
Run the embedding-only and reranked paths against the same frozen set. Compare field-level outcomes and end-to-end latency distributions. The decision boundary is concrete: choose the simpler path only if it meets the agreed retrieval, grounding, and abstention targets within the latency objective. Do not claim a quality win from a handful of visually pleasing answers.
Rollout and rollback
Ship reranking behind a per-document decision flag. Shadow it first, storing rankings and token budgets without changing production answers. Next, canary a small deterministic slice of invoice IDs so retries stay on the same path. Watch abstention rate, unsupported-field rate, review-queue depth, and latency together; improving one while overwhelming another is not a successful rollout.
Rollback should disable the reranker and restore the previous candidate limit and prompt policy as one versioned configuration change. Keep the strict source-only instruction and not found behavior in place. Those are correctness controls, not experiment knobs.
If the reranker is unavailable, fail closed for fields whose wrong value can trigger payment or reconciliation. Queue the invoice for retry or review. For lower-risk descriptive fields, an explicitly configured embedding-only fallback may be acceptable, provided it retains citations and the same evidence threshold.
That is the durable recommendation: improve retrieval before generation, measure the context before sending it, and make abstention a normal result. Reranking is justified when invoice ambiguity makes quality the dominant constraint. Embedding-only retrieval wins when the corpus is predictable, evidence is easy to retrieve, and the measured latency budget is tight.
Top comments (1)
The main claim feels under-supported.
You recommends reranking as the "safer default," but provides no actual comparison results. There's no benchmark for embedding-only vs reranked retrieval on recall, grounding, hallucination, or latency.