TL;DR
- On 50k documents, retrieval quality matters more than your choice of vector database. At this size you can run exact search and skip the infrastructure.
- BM25 is great at exact terms (error codes, names, IDs) and weak at paraphrases. Dense retrieval is the reverse.
- Hybrid (BM25 + dense, merged with Reciprocal Rank Fusion) is the best default. It's cheap to build and covers both failure modes.
- Add a reranker only after you've measured that hybrid isn't enough. Build a small eval set first, because without one you're guessing.
The setup nobody tells you about
Say you've built a RAG app over 50,000 documents: support tickets, internal docs, maybe PDFs. The demo worked. Then a real user asked a real question, and the model confidently answered from the wrong paragraph.
Nine times out of ten, the LLM isn't the problem. Retrieval is. If the right chunk never reaches the prompt, no model can save you.
One clarification before we start. "50k documents" usually means far more chunks. If each document splits into about 10 chunks, you're searching roughly 500k pieces of text. That number matters for memory and latency, so keep it in mind below.
I haven't run a fresh benchmark on your data, and I won't pretend to. This article uses published research and documented numbers to show what each approach does, where it breaks, and how to decide. If you want the background on what embeddings actually are, my post on the anatomy of AI covers vectors from the ground up.
The three contenders at a glance
| BM25 | Dense | Hybrid (RRF) | |
|---|---|---|---|
| Matches on | Exact words | Meaning | Both |
| Strong at | IDs, codes, names, jargon | Paraphrases, natural questions | Mixed real-world queries |
| Weak at | Synonyms, vocabulary mismatch | Rare tokens, unseen domains | Slightly more moving parts |
| Needs a model | No | Yes (embeddings) | Yes |
| Build effort | Low | Low to medium | Low, once the other two exist |
BM25: the librarian with perfect index cards
BM25 is a keyword-ranking function from the 1990s. It scores a chunk higher when:
- it contains your query words (term frequency),
- those words are rare across the corpus (inverse document frequency), and
- the chunk isn't just long and rambling (length normalization).
Analogy: imagine a librarian who never reads the books. She only keeps meticulous index cards of which words appear where. Ask for "ERR-4012" and she hands you the exact page in seconds. Ask "why does my app crash on startup" and she's stuck unless the page literally says "crash" and "startup."
Here's a minimal version using rank_bm25:
import re
from rank_bm25 import BM25Okapi
def tokenize(text: str) -> list[str]:
# keep dots, dashes and underscores *inside* a token (ERR-4012, user_id,
# v2.4.1) but not at its edges, so "startup." still matches "startup"
return re.findall(r"[a-z0-9]+(?:[._\-][a-z0-9]+)*", text.lower())
tokenized_chunks = [tokenize(c) for c in chunks]
bm25 = BM25Okapi(tokenized_chunks)
def bm25_search(query: str, k: int = 50):
scores = bm25.get_scores(tokenize(query))
top = scores.argsort()[::-1][:k]
return [(int(i), float(scores[i])) for i in top]
Two honest notes. First, rank_bm25 is fine for prototypes, but it scores every document per query. For production, use something with an inverted index like Elasticsearch, OpenSearch, Tantivy, or Postgres full-text search. Second, your tokenizer is half the battle. The regex above keeps ERR-4012 as one token on purpose. A default tokenizer that splits on punctuation can quietly wreck exact-match queries.
Where BM25 wins: product codes, function names, legal clause numbers, rare jargon, anything where the user types the exact words.
Where it loses: vocabulary mismatch. The user says "refund," the doc says "reimbursement." BM25 sees two unrelated strings.
Dense retrieval: the friend who gets what you mean
Dense retrieval turns text into vectors (embeddings) using a neural model. Chunks with similar meaning land close together in vector space, and you search by finding the nearest neighbors to your query vector.
Analogy: instead of index cards, you have a well-read friend. You describe a vague idea ("that thing where the server keeps retrying and makes it worse") and they say "oh, you want the doc on retry storms." They understand meaning, but they might blank on the exact serial number you asked for.
import numpy as np
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-en-v1.5")
# 50k chunks x 384 float32 is about 77 MB. 500k chunks is about 770 MB.
doc_vecs = model.encode(chunks, normalize_embeddings=True, batch_size=64)
def dense_search(query: str, k: int = 50):
q = model.encode([query], normalize_embeddings=True)[0]
scores = doc_vecs @ q # cosine similarity (vectors are normalized)
top = np.argpartition(-scores, k)[:k]
top = top[np.argsort(-scores[top])]
return [(int(i), float(scores[i])) for i in top]
Notice what's missing: no vector database, no ANN index. That's deliberate.
You probably don't need an ANN index yet
Approximate nearest neighbor indexes like HNSW trade a bit of recall for speed. At small scale, you may be paying that trade for nothing. pgvector, for example, does exact nearest neighbor search by default, which gives perfect recall. Plenty of teams run tens of thousands of vectors on a plain sequential scan, and skipping the index avoids build time, maintenance cost, and the recall loss of approximate search.
If your 50k documents become 500k chunks, revisit that, but measure first. A brute-force matrix multiply over a few hundred thousand vectors is still quick on a decent machine; it's the memory (see the comment in the code above) you'll feel first. When you do need HNSW, here's the standard pgvector shape:
CREATE INDEX ON chunks
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
-- higher ef_search = better recall, slower queries
SET hnsw.ef_search = 100;
Always compare ANN results against exact results on your own queries before trusting the index.
Where dense wins: paraphrases, natural-language questions, cross-vocabulary matches, multilingual queries.
Where it loses: exact identifiers, rare tokens, and anything where the embedding model has never seen your domain's vocabulary.
What the research says (and what it doesn't)
The most cited evidence here is BEIR, a benchmark covering zero-shot retrieval across a wide range of settings. Its headline finding surprised people at the time. Dense models beat BM25 by a clear margin on MS MARCO, the data they were trained on, yet BM25 held up as a robust baseline almost everywhere else. The authors found dense models do well when training and target data overlap heavily, but fall short on datasets with large domain shift.
That's the key insight for your 50k-document corpus. Your data is probably not MS MARCO. It has your company's acronyms, your product names, your weird internal vocabulary. A dense model trained on generic web text is, in effect, out-of-domain. That's exactly where BM25 refuses to embarrass itself.
Two honest caveats:
- BEIR is from 2021. Embedding models have improved a lot since. Don't read it as "BM25 beats dense." Read it as "dense models can fail unpredictably on unfamiliar domains, and lexical search is a safety net."
- Benchmark averages hide your queries. A model can win on average and still lose on the 20% of queries your users care about most.
BEIR also showed the cost side. BM25 answers in milliseconds on a CPU, while a cross-encoder reranker on top of it was orders of magnitude slower per query in the paper's setup. Reranking helps, but it isn't free, and we'll come back to that.
Hybrid: two witnesses are better than one
Analogy: picture two witnesses to a crime. One remembers exact details (the license plate), and the other remembers the overall impression (a dark sedan, driving erratically). Either alone can mislead you. Together, they converge on the right car.
That's hybrid retrieval: run BM25 and dense search in parallel, then merge the two ranked lists.
The tricky part is the merge. BM25 scores and cosine similarities live on completely different scales, so you can't just add them. You can normalize and weight them, but the weights drift as your data changes.
Reciprocal Rank Fusion (RRF)
RRF sidesteps the scale problem by ignoring scores entirely and using only rank positions. It comes from a 2009 SIGIR paper by Cormack, Clarke, and Büttcher. The formula is one line:
RRF(d) = Σ over rankers 1 / (k + rank(d))
with k = 60 as the usual default. The paper found that value near-optimal in a pilot study, and it has stuck around ever since. Because RRF is score-independent, using only ranks, one retriever with weird score ranges can't dominate the result.
from collections import defaultdict
def rrf(ranked_lists: list[list[int]], k: int = 60) -> list[tuple[int, float]]:
fused = defaultdict(float)
for ranking in ranked_lists:
for rank, doc_id in enumerate(ranking, start=1):
fused[doc_id] += 1.0 / (k + rank)
return sorted(fused.items(), key=lambda x: x[1], reverse=True)
def hybrid_search(query: str, k: int = 10, pool: int = 50):
bm25_ids = [i for i, _ in bm25_search(query, pool)]
dense_ids = [i for i, _ in dense_search(query, pool)]
return rrf([bm25_ids, dense_ids])[:k]
Analogy: RRF is like a talent show with two judges, where each judge only gives rankings, never scores. A contestant who is 2nd on both lists beats one who is 1st on one list and 40th on the other. Consensus beats a single loud opinion.
Two practical tips: pull a generous candidate pool (50 or so) from each retriever before fusing, and treat k = 60 as a starting point rather than a law.
A real-world data point on hybrid
Anthropic published numbers on this in their Contextual Retrieval write-up. The headline result: combining contextual embeddings with contextual BM25 cut the top-20-chunk retrieval failure rate by 49%, from 5.7% to 2.9%. They also explain why BM25 earns its place: embedding models can miss exact-match queries like unique identifiers, which BM25 handles well.
A fair caution comes from critics of that post. The best number, a 67% reduction, comes from stacking several techniques, including reranking, not from any single trick. One analysis pointed out that the headline figures measure only top-k retrieval failure rate, which may not match what you care about. So don't copy their pipeline wholesale. Take the lesson: hybrid beats either alone, and each added layer needs its own measurement.
How to actually decide: build a tiny eval
This is the part most tutorials skip, and it's the part that matters. Before you pick a strategy, write down 50 to 100 real questions with the chunk IDs that should answer them. Then measure recall@k: what fraction of the time the right chunk shows up in the top k results.
def recall_at_k(search_fn, eval_set, k=10):
hits = 0
for query, relevant_ids in eval_set:
retrieved = {i for i, _ in search_fn(query)[:k]}
hits += bool(retrieved & set(relevant_ids))
return hits / len(eval_set)
for name, fn in [("bm25", bm25_search),
("dense", dense_search),
("hybrid", hybrid_search)]:
print(name, recall_at_k(fn, eval_set, k=10))
Pull your eval questions from real user logs if you have them, and include the ugly ones: typos, error codes, vague questions. Hand-label them yourself. It's boring, and it's the highest-leverage hour you'll spend on this project.
Look at the failures, not just the number. If BM25 misses are paraphrase problems and dense misses are ID problems, you've just proven hybrid is justified for your data, not someone else's.
What I'd do on a 50k-document corpus
If I were handed this project tomorrow, here's the order I'd work in:
- Start with BM25 alone. It's fast to build and a strong baseline. If it already hits your recall target, you may be done.
- Add dense retrieval with a decent open embedding model, brute-force, no vector DB. Compare the two on your eval set.
- Fuse with RRF. In most real corpora with mixed vocabulary, this is where recall jumps.
- Fix chunking before adding fancier models. Chunks that lose their context (like a paragraph saying "it increased 12%" with no idea what "it" is) hurt both retrievers. This is the problem Anthropic's contextual approach targets.
- Add a reranker last, and only if hybrid recall@50 is good but precision@5 is poor. That pattern means the right chunk is in the pool but ranked too low, which is exactly the reranker's job. Budget for the latency, since reranking can cost far more per query than the retrieval itself.
- Reach for HNSW only when measurements say so, such as high query volume or a corpus that grew far past your current size.
Common mistakes I see
- Skipping evaluation and trusting vibes from five hand-picked demo queries.
- Tokenizing BM25 carelessly, so identifiers and code get shredded.
- Normalizing and summing raw scores from different retrievers instead of fusing ranks.
- Over-engineering infrastructure (a managed vector DB, three services) for a dataset whose vectors fit in under a gigabyte of RAM.
- Blaming the LLM for answers that were doomed by what retrieval handed it.
- Ignoring metadata filters. If users usually ask about one product or one date range, filtering before search often beats any clever ranking.
Final thoughts
There's no universally best retriever. BM25 is the reliable, literal-minded librarian. Dense retrieval is the intuitive friend who sometimes gets overconfident. Hybrid puts them in the same room and lets them check each other's work.
At 50k documents, you have the luxury of simplicity: brute-force vectors, a plain keyword index, and fusion code you can read in 15 seconds. Spend your effort on a good eval set and good chunking. Those two things will move your answer quality more than any model swap.
The snippets above are illustrative, so test them on your own data before relying on them.
References
- Thakur et al. (2021). BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. NeurIPS Datasets and Benchmarks.
- Cormack, Clarke, Büttcher (2009). Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods. SIGIR '09.
- Anthropic. Introducing Contextual Retrieval.
- Almond AI. Responding to Anthropic's Contextual Retrieval: Why Context is NOT All You Need.
- pgvector. pgvector on GitHub.
Find me across the web:
Portfolio: ahmershah.dev
Crunchbase: @syed-ahmer-shah
Crunchbase Company: @syedahmershah
Clutch: @syed-ahmer-shah
Tech Behemoth: @syed-ahmer-shah
Design Rush: @syed-ahmer-shah
Edverise: @syed-ahmer-shah
Trust Pilot: ahmershah.dev
LinkedIn: Syed Ahmer Shah
GitHub: @ahmershahdev
AWS Builder Profile: @syedahmershah
DEV: @syedahmershah
Medium: @syedahmershah
Hashnode: @syedahmershah
Substack: @syedahmershah
HackerNoon: @syedahmershah
Substack: @syedahmershah
Facebook: @ahmershahdev
Linkedin Page: @syedahmershah
YouTube: @ahmershahdev
Instagram: @ahmershahdev
TikTok: @ahmershahdev
Top comments (0)