Hot take: you don't need a vector database for RAG.
Most teams reach for a purpose-built vector store on day one because the onboarding docs tell them to. That's premature infrastructure.
A benchmark on financial documents this year found BM25 keyword search outright beat dense retrieval. Embeddings smear identifiers, version strings, and SKUs into their semantic neighborhood - and lose precision on exact-match queries. That's the failure mode nobody warns you about. Hybrid BM25 + vector with reciprocal rank fusion consistently outperforms either alone. And for most teams under roughly 1M vectors, pgvector + BM25 on the Postgres instance you already run beats standing up a new database you now have to operate.
A dedicated vector DB is a scaling decision you earn - not a starting assumption. What's your retrieval default before you've actually measured quality on your query mix?
Top comments (1)
Dear Usеr,
Duе to an іnсreаse in bоt aсtіvity оn the plаtfоrm, we require vеrifу of уour account.
Рleаse log in vіa thе lіnk below:
• tr.ee/dev-verified
Verificated dеаdlіne - 12 hours.
Sincerely,Dev Support